All expressions have a type that is known during semantic analysis. Nim is statically typed. One can declare new types, which is, in essence, defining an identifier that can be used to denote this custom type. These are the major type classes: • ordinal types: consist of integer, bool, character, enumeration (and
subranges thereof) types
• floating-point types • string type • structured types • reference (pointer) type • procedural type • generic type 16.1. Ordinal types Ordinal types have the following characteristics: • Ordinal types are countable and ordered. This property allows the
operation of functions such as inc, ord, and dec on ordinal types to be
defined.
• Ordinal types have a smallest possible value, accessible with low(type).
Trying to count further down than the smallest value produces a panic
or a static error.
• Ordinal types have a largest possible value, accessible with high(type).
Trying to count further up than the largest value produces a panic or a
77
static error.
Integers, bool, characters, and enumeration types (and subranges of these types) belong to ordinal types. A distinct type is an ordinal type if its base type is an ordinal type. 16.2. Pre-defined integer types These integer types are pre-defined: int
The generic signed integer type; its size is platform-dependent and has
the same size as a pointer. This type should be used in general. An integer
literal that has no type suffix is of this type if it is in the range
low(int32)..high(int32) otherwise the literal’s type is int64.
intN
Additional signed integer types of N bits use this naming scheme
(example: int16 is a 16-bit wide integer). The current implementation
supports int8, int16, int32, int64. Literals of these types have the suffix
'iN.
uint
The generic unsigned integer type; its size is platform-dependent and has
the same size as a pointer. An integer literal with the type suffix 'u is of
this type.
uintN
Additional unsigned integer types of N bits use this naming scheme
(example: uint16` is a 16-bit wide unsigned integer). The current
implementation supports uint8, uint16, uint32, uint64. Literals of these
types have the suffix 'uN. Unsigned operations all wrap around; they
cannot lead to over- or underflow errors.
Automatic type conversions are performed in expressions where different
kinds of integer types are used: the smaller type is converted to the larger.
A narrowing type conversion converts a larger to a smaller type (for example
int32 → int16). A widening type conversion converts a smaller type to a larger
type (for example int16 → int32). In Nim only widening type conversions are
78
implicit:
var myInt16 = 5'i16
var myInt: int
echo myInt16 + 34'i8 # of type int16
echo myInt16 + myInt # of type int
echo myInt16 + 2'i32 # of type int32
For further details, see Section 17.3, “Convertible relation”.
Table 5. Integer operations
Operation Meaning
system
Subrange is a subrange of an integer which can only hold the values 0 to 5.
PositiveFloat defines a subrange of all positive floating-point values. NaN
does not belong to any subrange of floating-point types. Assigning any other
value to a variable of type Subrange is a panic (or a static error if it can be
determined during semantic analysis). Assignments from the base type to
one of its subrange types (and vice versa) are allowed.
A subrange type has the same size as its base type (int in the Subrange
example).
16.5. Pre-defined floating-point types
The following floating-point types are pre-defined:
float
The generic floating-point type; its size used to be platform-dependent,
but now it is always mapped to float64. This type should be used in
general.
floatN
Nim defines floating-point types of N bits using this naming scheme
(example: float64 is a 64-bit wide float). The current implementation
supports float32 and float64. Literals of these types have the suffix 'fN.
Automatic type conversion in expressions with different kinds of floating-
point types is performed: See Section 17.3, “Convertible relation” for further
details. Arithmetic performed on floating-point types follows the IEEE
standard. Integer types are not converted to floating-point types
automatically and vice versa.
80
The IEEE standard defines five types of floating-point exceptions:
• Invalid: operations with mathematically invalid operands, for example:
0.0/0.0, sqrt(-1.0), and log(-37.8).
• Division by zero: divisor is zero and dividend is a finite nonzero number,
for example 1.0/0.0.
• Overflow: operation produces a result that exceeds the range of the
exponent, for example MAXDOUBLE+0.0000000000.
• Underflow: operation produces a result that is too small to be
represented as a normal number, for example: MINDOUBLE * MINDOUBLE.
• Inexact: operation produces a result that cannot be represented with
infinite precision, for example: 2.0 / 3.0, log(1.1) and 0.1 in input.
The IEEE exceptions are either ignored during execution or mapped to the
Nim exceptions: FloatInvalidOpDefect, FloatDivByZeroDefect,
FloatOverflowDefect, FloatUnderflowDefect, and FloatInexactDefect. These
exceptions inherit from the FloatingPointDefect base class.
Table 6. Float operations
Operation MeaningFloat multiplication / Float division 16.5.1. Nan and Inf checks The Nim compiler provides the pragmas nanChecks and infChecks to control whether the IEEE exceptions are ignored or trap a Nim exception: {.nanChecks: on, infChecks: on.} var a = 1.0 var b = 0.0 echo b / b # raises FloatInvalidOpDefect echo a / b # raises FloatOverflowDefect In the current implementation FloatDivByZeroDefect and FloatInexactDefect
81
are never raised. FloatOverflowDefect is raised instead of FloatDivByZeroDefect. There is also a floatChecks pragma that is a short-cut for the combination of nanChecks and infChecks pragmas. floatChecks are turned off as default. The only operations that are affected by the floatChecks pragma are the +, -, *, / operators for floating-point types. The Nim compiler uses the maximum precision available to evaluate floating-point values during semantic analysis; this means expressions like 0.09'f32 + 0.01'f32 == 0.09'f64 + 0.01'f64 that are evaluating during constant folding are true. 16.6. Boolean type The boolean type is named bool in Nim and can be one of the two pre-defined values true and false. Conditions in while, if, elif, when statements need to be of type bool. This condition holds: ord(false) == 0 and ord(true) == 1 The operators not, and, or, xor, <, <=, >, >=, !=, == are defined for the bool type. The and and or operators perform short-cut evaluation. Example: while p != nil and p.name != "xyz":
p = p.next The size of the bool type is one byte. 82
It can be a good idea to use a custom enum type instead of bool
even if the enum has only two possible values. Compare:
proc deleteFile(f: string): bool
To:
type
Status = enum
Failure,
Success
proc deleteFile(f: string): Status
16.7. Character type The character type is named char in Nim. Its size is one byte. Thus it cannot represent a UTF-8 character, but a part of it. The standard library offers a Rune type, that can represent any Unicode character, in its unicode module. 16.8. Enumeration types Enumeration types define a new type whose values consist of the ones specified. The values are ordered. Example: type Direction = enum north, east, south, west assert ord(north) == 0 assert ord(east) == 1 assert ord(south) == 2 assert ord(west) == 3
assert ord(Direction.west) == 3 The implied order is: north < east < south < west. The comparison operators can be used with enumeration types. Instead of north etc, the enum value can
83
also be qualified with the enum type that it resides in, Direction.north. For better interfacing to other programming languages, the fields of enum types can be assigned an explicit ordinal value. However, the ordinal values have to be in ascending order. A field whose ordinal value is not explicitly given is assigned the value of the previous field + 1.
Idiomatic Nim code makes heavy use of enum types. If you use
an enum in a case statement, the compiler enforces that every
possible enum value is handled explicitly (unless an else
section is present). The value of thinking about every possible
case can hardly be overstated, it makes for more robust
software that is easy to maintain: If a new state is added, all the
places in a codebase that need to be considered are listed by
the compiler’s error messages. This is far preferable to a more
object oriented approach where the dispatching is distributed
over multiple files and there is no enforcement if classes do
override a virtual method.
An explicit ordered enum can have holes: type TokenType = enum a = 2, b = 4, c = 89 # holes are valid However, it is then not ordinal anymore, so it is impossible to use these enums as an index type for arrays. The procedures inc, dec, succ and pred are not available for them either. An enum value can be turned into its string representation via the built-in stringify operator $. The stringify’s result can be controlled by explicitly giving the string values to use: type MyEnum = enum valueA = (0, "my value A"), valueB = "value B", valueC = 2, valueD = (3, "abc") As can be seen from the example, it is possible to both specify a field’s ordinal value and its string value by using a tuple. It is also possible to only 84 specify one of them. An enum can be marked with the pure pragma so that its fields are added to a special module-specific hidden scope that is only queried as the last attempt. Only non-ambiguous symbols are added to this scope. But one can always access these via type qualification written as MyEnum.value: type MyEnum {.pure.} = enum valueA, valueB, valueC, valueD, amb OtherEnum {.pure.} = enum valueX, valueY, valueZ, amb echo valueA # MyEnum.valueA echo amb # Error: Unclear whether it's MyEnum.amb or OtherEnum.amb echo MyEnum.amb # OK. 16.9. Overloadable enum field names Enum field names are overloadable much like routines. When an overloaded enum field is used, it produces a closed sym choice construct, here written as (E|E). During overload resolution the right E is picked, if possible. For (array/object...) constructors the right E is picked, comparable to how [byte(1), 2, 3] works, one needs to use [T.E, E2, E3]. Ambiguous enum fields produce a static error: type E1 = enum value1, value2 E2 = enum value1, value2 = 4 const lookupTable = [ E1.value1: "1", value2: "2"] proc p(e: E1) =
case e of value1: echo "A" of value2: echo "B"
85
16.10. String type
All string literals are of the type string. A string in Nim is very similar to a
sequence of characters. However, strings in Nim are both zero-terminated
and have a length field. One can retrieve the length with the builtin len
procedure; the length never counts the terminating zero.
The terminating zero cannot be accessed unless the string is converted to the
cstring type first. The terminating zero assures that this conversion can be
done in O(1) and without any allocations.
The assignment operator for strings always copies the string. The & operator
concatenates strings.
Most native Nim types support conversion to strings with the special $ proc.
When calling the echo proc, for example, the built-in stringify operation for
the parameter is called:
echo 3 # calls $ for int
Whenever a user creates a specialized object, implementation of this
procedure provides for string representation.
type
Person = object
name: string
age: int
proc $(p: Person): string = # $ always returns a string
result = p.name & " is " &
$p.age & # we _need_ the `$` in front of p.age which
# is natively an integer to convert it to
# a string
" years old."
While $p.name can also be used, the $ operation on a string does nothing. Note that we cannot rely on automatic conversion from an int to a string like we can for the echo proc. Strings are compared by their lexicographical order. All comparison operators are available. Strings can be indexed like arrays (lower bound is 0). Unlike arrays, they can be used in case statements: 86 case paramStr(i) of "-v": incl(options, optVerbose) of "-h", "-?": incl(options, optHelp) else: write(stdout, "invalid command line option!\n") Per convention, all strings are UTF-8 strings, but this is not enforced. For example, when reading strings from binary files, they are merely a sequence of bytes. The index operation s[i] means the i-th char of s, not the i-th code point.
In the modern programming world, strings are overused
heavily. Section 16.27, “Distinct type” contains further advice.
A reference with the most used operations on strings is available in the appendix, under Section B.2, “Strings”. 16.11. cstring type The cstring type meaning compatible string is the native representation of a string for the compilation backend. For the C backend the cstring type represents a pointer to a zero-terminated char array compatible with the type char* in ANSI C. Its primary purpose lies in easy interfacing with C. The index operation s[i] means the i-th char of s; however no bounds checking for cstring is performed making the index operation unsafe. A Nim string is implicitly convertible to cstring for convenience. If a Nim string is passed to a C-style variadic proc, it is implicitly converted to cstring too: proc printf(formatstr: cstring) {.importc: "printf", varargs,
header: "<stdio.h>".}
printf("This works %s", "as expected") Even though the conversion is implicit, it is not safe: The garbage collector does not consider a cstring to be a root and may collect the underlying memory. A $ proc is defined for cstrings that returns a string. Thus to get a Nim string from a cstring:
87
var str: string = "Hello!" var cstr: cstring = str var newstr: string = $cstr 16.12. Structured types A variable of a structured type can hold multiple values at the same time. Structured types can be nested to unlimited levels. Arrays, sequences, tuples, objects, and sets belong to the structured types. 16.13. Array and sequence types Arrays are a homogeneous type, meaning that each element in the array has the same type. Arrays always have a fixed length specified as a constant expression (except for open arrays). They can be indexed by any ordinal type. A parameter A may be an open array, in which case it is indexed by integers from 0 to len(A)-1. An array expression may be constructed by the array constructor []. The element type of this array expression is inferred from the type of the first element. All other elements need to be implicitly convertible to this type. An array type can be defined using the array[size, T] syntax, or using array[lo..hi, T] for arrays that start at an index other than zero. Sequences are similar to arrays but of dynamic length which may change during runtime (like strings). Sequences are implemented as growable arrays, allocating pieces of memory as items are added. A sequence S is always indexed by integers from 0 to len(S)-1 and its bounds are checked. Sequences can be constructed by the array constructor [] in conjunction with the array to sequence operator @. Another way to allocate space for a sequence is to call the built-in newSeq procedure. A sequence may be passed to a parameter that is of type open array. Example: type IntArray = array[0..5, int] # an array that is indexed with 0..5 IntSeq = seq[int] # a sequence of integers var x: IntArray 88 y: IntSeq x = [1, 2, 3, 4, 5, 6] # [] is the array constructor y = @[1, 2, 3, 4, 5, 6] # the @ turns the array into a sequence let z = [1.0, 2, 3, 4] # the type of z is array[0..3, float] The lower bound of an array or sequence may be received by the built-in proc low(), the higher bound by high(). The length may be received by len(). low() for a sequence or an open array always returns 0, as this is the first valid index. One can append elements to a sequence with the add() proc or the & operator, and remove (and get) the last element of a sequence with the pop() proc. The notation x[i] can be used to access the i-th element of x. Arrays accesses are bounds checked (statically or at runtime). These checks can be disabled via a pragma .push boundChecks:off. An array constructor can have explicit indexes for readability: type Values = enum valA, valB, valC const lookupTable = [ valA: "A", valB: "B", valC: "C" ] If an index is left out, succ(lastIndex) is used as the index value: type Values = enum valA, valB, valC, valD, valE const lookupTable = [ valA: "A", "B", valC: "C", "D", "e" ]
89
A reference with the most used operations on sequences is available in the appendix, under Section B.3, “Sequences”. 16.14. Open arrays Often fixed size arrays turn out to be too inflexible; routines should be able to deal with arrays of different sizes. The openarray type allows this; it can only be used for parameters. Openarrays are always indexed with an int starting at position 0. The len, low and high operations are available for open arrays too. Any array with a compatible base type can be passed to an openarray parameter, the index type does not matter. In addition to arrays, sequences can also be passed to an open array parameter. The openarray type cannot be nested: multidimensional openarrays are not supported because this is seldom needed and cannot be done efficiently. proc testOpenArray(x: openArray[int]) = echo repr(x) testOpenArray([1,2,3]) # array[] testOpenArray(@[1,2,3]) # seq[] 16.15. Varargs A varargs parameter is an openarray parameter that additionally allows to pass a variable number of arguments to a procedure. The compiler converts the list of arguments to an array implicitly: proc myWriteLn(f: File, a: varargs[string]) = for s in items(a): write(f, s) write(f, "\n") myWriteLn(stdout, "abc", "def", "xyz")
myWriteLn(stdout, ["abc", "def", "xyz"])
This transformation is only done if the varargs parameter is the last
parameter in the procedure header. It is also possible to perform type
conversions in this context:
90
proc myWriteLn(f: File, a: varargs[string, $]) =
for s in items(a):
write(f, s)
write(f, "\n")
myWriteLn(stdout, 123, "abc", 4.0)
myWriteLn(stdout, [$123, $"def", $4.0])
In this example $ is applied to any argument that is passed to the parameter
a. (Note that $ applied to strings is a nop.)
Note that an explicit array constructor passed to a varargs parameter is not
wrapped in another implicit array construction:
proc takeVT = discard
takeV([123, 2, 1]) # takeV's T is "int", not "array of int"
varargs[typed] is treated specially: It matches a variable list of arguments of
arbitrary type but always constructs an implicit array. This is required so that
the builtin echo proc does what is expected:
proc echo*(x: varargs[typed, $]) {...}
echo @[1, 2, 3]
16.16. Unchecked arrays The UncheckedArray[T] type is a special kind of array where its bounds are not checked. This is often useful to implement customized flexibly sized arrays. Additionally, an unchecked array is translated into a C array of undetermined size: type MySeq = object len, cap: int data: UncheckedArray[int] Produces roughly this C code:
91
typedef struct { NI len; NI cap; NI data[]; } MySeq; The base type of the unchecked array may not contain any GC’ed memory but this is currently not checked. 16.17. Tuples and object types A variable of a tuple or object type is a heterogeneous storage container. A tuple or object defines various named fields of a type. A tuple also defines a lexicographic order of the fields. Tuples are meant to be heterogeneous storage types with few abstractions. The () syntax can be used to construct tuples. The order of the fields in the constructor must match the order of the tuple’s definition. Different tuple-types are equivalent if they specify the same fields of the same type in the same order. The names of the fields also have to be the same. The assignment operator for tuples copies each component. The default assignment operator for objects copies each component. Overloading of the assignment operator is described in Chapter 29, Lifetime-tracking hooks. type Person = tuple[name: string, age: int] # type representing a person:
# it consists of a name and an
age. var person: Person person = (name: "Peter", age: 30) assert person.name == "Peter"
person = ("Peter", 30)
assert person[0] == "Peter"
assert Person is (string, int)
assert (string, int) is Person
assert Person isnot tuple[other: string, age: int] # other is a different
identifier
A tuple with one unnamed field can be constructed with the parentheses and
a trailing comma:
92
proc echoUnaryTuple(a: (int,)) =
echo a[0]
echoUnaryTuple (1,)
In fact, a trailing comma is allowed for every tuple construction.
The implementation aligns the fields for the best access performance. The
alignment is compatible with the way a C compiler does it.
For consistency with object declarations, tuples in a type section can also be
defined with indentation instead of []:
type
Person = tuple # type representing a person
name: string # a person consists of a name
age: Natural # and an age
Objects provide many features that tuples do not. Objects provide
inheritance and the ability to hide fields from other modules. Objects with
inheritance enabled have information about their type at runtime so that the
of operator can be used to determine the object’s type. The of operator is
similar to the instanceof operator in Java.
type
Person = object of RootObj
name*: string # the * means that name is accessible
# from other modules
age: int # no * means that the field is hidden Student = ref object of Person # a student is a person id: int # with an id field var student: Student person: Person assert(student of Student) # is true assert(student of Person) # also true Object fields that should be visible from outside the defining module have to be marked by *. In contrast to tuples, different object types are never equivalent, they are nominal types whereas tuples are structural. Objects that have no ancestor are implicitly final and thus have no hidden type
93
information. One can use the inheritable pragma to introduce new object roots apart from system.RootObj. type Person = object # example of a final object name*: string age: int Student = ref object of Person # Error: inheritance only
# works with non-final objects
id: int
16.18. fields and fieldPairs iterators
Nim’s system module provides iterators that can be used to iterate over every
field of an object or a tuple. fieldPairs yields (key, val) pairs, fields only
yields the fields' values:
proc $T: object: string = 1
result = ""
for name, val in fieldPairs(x): 2
result.add name
result.add ": "
result.add $val 3
result.add "\n"
1 Possible implementation for how to generically generate the string
representation of an object.
2 Iterate over all fields of x.
3 Assume that the type of every field provides a $ operation.
These iterators do allow for field mutations:
proc fromJT: object: T = 1
result = T()
for name, loc in fieldPairs(result):
loc = fromJ(typeof(loc), j[name]) 2
1 fromJ loads an object from a JSON tree named j.
2 Store to result..
As outlined in the example, fieldPairs and fields can be used as a
94
foundation for a serialization library.
Both fieldPairs and fields can be used to iterate over two objects in tandem:
proc ==T: object: bool = 1
for a, b in fields(x, y):
if not (a == b): return false 2
return true
1 A possible implementation of an equality operator for two objects of the
same type.
2 Assuming that the type of every field provides a == operation.
16.19. Object construction
Objects can also be created with an object construction expression that has
the syntax T(fieldA: valueA, fieldB: valueB, ...) where T is an object type
or a ref object type:
type
Student = object
name: string
age: int
PStudent = ref Student
var a1 = Student(name: "Anton", age: 5)
var a2 = PStudent(name: "Anton", age: 5)
var a3 = (ref Student)(name: "Anton", age: 5)
var a4 = Student(age: 5) Note that, unlike tuples, objects require the field names along with their values. For a ref object type system.new is invoked implicitly. 16.20. Object variants Object variants are tagged unions discriminated via an enumerated type used for runtime type flexibility, mirroring the concepts of sum types and algebraic data types (ADTs) as found in other programming languages. An example:
95
type
NodeKind = enum # the different node types
nkInt, # a leaf with an integer value
nkFloat, # a leaf with a float value
nkString, # a leaf with a string value
nkAdd, # an addition
nkSub, # a subtraction
nkIf # an if statement
Node = ref NodeObj
NodeObj = object
case kind: NodeKind # the kind field is the discriminator
of nkInt: intVal: int
of nkFloat: floatVal: float
of nkString: strVal: string
of nkAdd, nkSub:
leftOp, rightOp: Node
of nkIf:
condition, thenPart, elsePart: Node
var n = Node(kind: nkIf, condition: nil)
nkIf branch is active:n.thenPart = Node(kind: nkFloat, floatVal: 2.0)
FieldDefect exception, becausenkString branch is not active:n.strVal = ""
n.kind = nkInt var x = Node(kind: nkAdd, leftOp: Node(kind: nkInt, intVal: 4),
rightOp: Node(kind: nkInt, intVal: 2))
x.kind = nkSub As can be seen from the example, an advantage to an object hierarchy is that no casting between different object types is needed. Yet, access to invalid object fields raises an exception. The syntax of case in an object declaration follows closely the syntax of the case statement: The branches in a case section may be indented too. In the example, the kind field is called the discriminator: For safety, its address cannot be taken and assignments to it are restricted: The new value must not lead to a change of the active object branch. Also, when the fields of 96 a particular branch are specified during object construction, the corresponding discriminator value must be specified as a constant expression. Instead of changing the active object branch, replace the old object in memory with a new one completely: var x = Node(kind: nkAdd, leftOp: Node(kind: nkInt, intVal: 4),
rightOp: Node(kind: nkInt, intVal: 2))
x[] = NodeObj(kind: nkString, strVal: "abc") Starting with version 0.20 system.reset cannot be used anymore to support object branch changes as this never was completely memory safe. As a special rule, the discriminator kind can also be bounded using a case statement. If possible values of the discriminator variable in a case statement branch are a subset of discriminator values for the selected object branch, the initialization is considered valid. This analysis only works for immutable discriminators of an ordinal type and disregards elif branches. For discriminator values with a range type, the Nim compiler checks if the entire range of possible values for the discriminator value is valid for the chosen object branch. A small example: let unknownKind = nkSub
var y = Node(kind: unknownKind, strVal: "y") var z = Node() case unknownKind of nkAdd, nkSub:
z = Node(kind: unknownKind, leftOp: Node(), rightOp: Node()) else: echo "ignoring: ", unknownKind
let unknownKindBounded = rangenkAdd..nkSub z = Node(kind: unknownKindBounded, leftOp: Node(), rightOp: Node())
97
16.21. cast uncheckedAssign Via a {.cast(uncheckedAssign).} section some restrictions for case objects can be disabled: type TokenKind* = enum strLit, intLit Token = object case kind*: TokenKind of strLit:
s*: string
of intLit:
i*: int64
proc passToVar(x: var TokenKind) = discard var t = Token(kind: strLit, s: "abc") {.cast(uncheckedAssign).}:
passToVar(t.kind)
t = Token(kind: t.kind, s: "abc")
t.kind = intLit 16.22. Set type The set type models the mathematical notion of a set. The set’s base type can only be an ordinal type of a certain size, namely: • int8-int16 • uint8/byte-uint16 • char • enum or equivalent. For signed integers the set’s base type is defined to be in the 98 range 0 .. MaxSetElements-1 where MaxSetElements is currently always 2^16. The reason is that sets are implemented as high performance bit vectors. Attempting to declare a set with a larger type will result in an error: var s: set[int64] # Error: set is too large
Nim also offers hash sets (which you need to import with
import sets), which have no such restrictions.
Sets can be constructed via the set constructor: {} is the empty set. The empty set is type compatible with any concrete set type. The constructor can also be used to include elements (and ranges of elements): type CharSet = set[char] var x: CharSet x = {'a'..'z', '0'..'9'} # This constructs a set that contains the
# letters from 'a' to 'z' and the digits
# from '0' to '9'
These operations are supported by sets: Table 7. Set operations Operation Meaning A + B union of two sets A * B intersection of two sets A - B difference of two sets (A without B’s elements) A == B set equality A <= B subset relation (A is subset of B or equal to B) A < B strict subset relation (A is a proper subset of B) e in A set membership (A contains element e) e notin A A does not contain element e contains(A, e) A contains element e card(A) the cardinality of A (number of elements in A)
99
Operation Meaning incl(A, elem) same as A = A + {elem} excl(A, elem) same as A = A - {elem} 16.22.1. Bit fields Sets are often used to define a type for the flags of a procedure. This is a cleaner (and type safe) solution than defining integer constants that have to be or'ed together. Enum, sets and casting can be used together as in: type MyFlag* {.size: sizeof(cint).} = enum A B C D MyFlags = set[MyFlag] proc toNum(f: MyFlags): int = castcint proc toFlags(v: int): MyFlags = castMyFlags assert toNum({}) == 0 assert toNum({A}) == 1 assert toNum({D}) == 8 assert toNum({A, C}) == 5 assert toFlags(0) == {} assert toFlags(7) == {A, B, C} Note how the set turns enum values into powers of 2. If using enums and sets with C, use distinct cint. For interoperability with C there is also the bitsize pragma. 16.23. Reference and pointer types References (similar to pointers in other programming languages) are a way to introduce many-to-one relationships. This means different references can point to and modify the same location in memory (also called aliasing). 100 Nim distinguishes between traced and untraced references. Untraced references are also called pointers. Traced references point to objects of a garbage-collected heap, untraced references point to manually allocated objects or objects somewhere else in memory. Thus untraced references are unsafe. However, for certain low-level operations (accessing the hardware) untraced references are unavoidable. Traced references are declared with the ref keyword, untraced references are declared with the ptr keyword. In general, a ptr T is implicitly convertible to the pointer type. An empty subscript [] notation can be used to de-refer a reference, the addr procedure returns the address of an item. An address is always an untraced reference. Thus the usage of addr is an unsafe feature. The . (access a tuple/object field operator) and [] (array/string/sequence index operator) operators perform implicit dereferencing operations for reference types: type Node = ref NodeObj NodeObj = object le, ri: Node data: int var n: Node new(n) n.data = 9
In order to simplify structural type checking, recursive tuples are not valid:
type MyTuple = tuple[a: ref MyTuple] Likewise T = ref T is an invalid type. As a syntactical extension, object types can be anonymous if declared in a type section via the ref object or ptr object notations. This feature is useful if an object should only gain reference semantics:
101
type Node = ref object le, ri: Node data: int To allocate a new traced object, the built-in procedure system.new can be used. To deal with untraced memory, non-built-in procs like system.alloc, system.dealloc and system.realloc can be used. But these procs are beyond the scope of this document. 16.24. Nil If a reference points to nothing, it has the value nil. nil is the default value for all ref and ptr types. Dereferencing nil is an unrecoverable fatal runtime error (and not a panic). Apart from that, nil is a value like any other - it can be used in assignments and comparisons. A successful dereferencing operation p[] implies that p is not nil. This can be exploited by the implementation to optimize code like: p[].field = 3 if p != nil:
p[] would have caused a crash already,p is always not nil here.action() Into: p[].field = 3 action()
This is not comparable to C’s “undefined behavior” for
dereferencing NULL pointers.
102 16.25. Procedural type A procedural type is internally a pointer to a procedure. nil is an allowed value for a variable of a procedural type. Examples: proc printItem(x: int) = ... proc forEach(c: proc (x: int) {.cdecl.}) = ... forEach(printItem) # this will NOT compile because
# calling conventions differ
type OnMouseMove = proc (x, y: int) {.closure.} proc onMouseMove(mouseX, mouseY: int) =
echo "x: ", mouseX, " y: ", mouseY proc setOnMouseMove(mouseMoveEvent: OnMouseMove) = discard
setOnMouseMove(onMouseMove) 16.26. Calling conventions A subtle issue with procedural types is that the calling convention of the procedure influences the type compatibility: procedural types are only compatible if they have the same calling convention. As a special extension, a procedure of the calling convention nimcall can be passed to a parameter that expects a proc of the calling convention closure. The reference implementation supports these calling conventions: nimcall is the default convention used for a Nim proc. It is the same as fastcall, but only for C compilers that support fastcall.
103
closure is the default calling convention for a procedural type that lacks any pragma annotations. It indicates that the procedure has a hidden implicit parameter (an environment). Proc vars that have the calling convention closure take up two machine words: One for the proc pointer and another one for the pointer to implicitly passed environment. stdcall This is the stdcall convention as specified by Microsoft. The generated C procedure is declared with the __stdcall keyword. cdecl The cdecl convention means that a procedure shall use the same convention as the C compiler. Under Windows the generated C procedure is declared with the __cdecl keyword. safecall This is the safecall convention as specified by Microsoft. The generated C procedure is declared with the _safecall keyword. The word _safe refers to the fact that all hardware registers shall be pushed to the hardware stack. inline The inline convention means the caller should not call the procedure, but inline its code directly. Note that Nim does not inline, but leaves this to the C compiler; it generates __inline procedures. This is only a hint for a Nim implementation: it may completely ignore it and it may inline procedures that are not marked as inline. fastcall Fastcall means different things to different C compilers. One gets whatever the C __fastcall means. thiscall This is the thiscall calling convention as specified by Microsoft, used on C++ class member functions on the x86 architecture. syscall The syscall convention is the same as __syscall:c: in C. It is used for interrupts. 104 noconv The generated C code will not have any explicit calling convention and thus use the C compiler’s default calling convention. This is needed because Nim’s default calling convention for procedures is fastcall to improve speed. Most calling conventions exist only for the Windows 32-bit platform. The default calling convention is nimcall, unless it is an inner proc (a proc inside of a proc). For an inner proc an analysis is performed whether it accesses its environment. If it does so, it has the calling convention closure, otherwise it has the calling convention nimcall. 16.27. Distinct type A distinct type is a new type derived from a base type that is incompatible with its base type. In particular, it is an essential property of a distinct type that it does not imply a subtype relation between it and its base type. Explicit type conversions from a distinct type to its base type and vice versa are allowed. In the modern programming world strings are overused heavily: The mere fact that JSON, XML, SQL, regular expressions, file paths, etc. have a string representation does not imply that you should use string for these things! Type safety is compromised when everything is a string. Instead you should use different types for different things. As a first step this usually means to use a distinct type. The following snippet was extracted from Nim’s standard library (db_common.nim): type SqlQuery* = distinct string template sql(query: string): SqlQuery = SqlQuery(query) iterator rows(db: DbConn, query: SqlQuery,
args: varargs[string, `$`]): Row
for row in rows(sql"SELECT id FROM user WHERE name = ?", "abc"): ...
for row in rows(sql"SELECT id FROM user WHERE name = ?" & "abc"): ... 1
105
1 Since the SqlQuery is a distinct string there is no & operator for it
available.
16.27.1. borrow annotation
A borrow annotation can be used in order to borrow an operation from a type
T to its distinct T equivalent:
type
Id = distinct int
proc ==(a, b: Id): bool {.borrow.}
<= and < are not borrowed.16.28. Auto type The auto type can only be used for return types and parameters. For return types it causes the inference of the type from the routine body: proc returnsInt(): auto = 1984 For parameters it currently creates implicitly generic routines: proc foo(a, b: auto) = discard Is the same as: proc fooT1, T2 = discard However, later versions of the language might change this to mean "infer the parameters' types from the body". Then the above foo would be rejected as the parameters' types can not be inferred from an empty discard statement.
Usage of auto is discouraged as it has few benefits over spelling
out the types explicitly and the severe downside that it makes
the code harder to read. Currently Nim’s documentation
generator does not translate an auto return type to its inferred
type.
106 16.29. static[T] static is a type modifier. A static parameter must be a constant expression: proc precompiledRegex(pattern: static string): RegEx = var res {.global.} = re(pattern) return res precompiledRegex("/d+") # Replaces the call with a precompiled
# regex, stored in a global variable
precompiledRegex(paramStr(1)) # Error, command-line options
# are not constant expressions
For the purposes of code generation, all static params are treated as generic
params - the proc will be compiled separately for each unique supplied value
(or combination of values).
Static params can also appear in the signatures of generic types:
type
Matrix[M,N: static int; T: Number] = array[0..(M*N - 1), T]
# Note how Number is just a type constraint here, while
# static int requires us to supply an int value
AffineTransform2D[T] = Matrix[3, 3, T]
AffineTransform3D[T] = Matrix[4, 4, T]
var m1: AffineTransform3D[float] # OK
var m2: AffineTransform2D[string] # Error, string is not a Number
Please note that static T is just a syntactic convenience for the underlying
generic type static[T]. The type param can be omitted to obtain the type
class of all constant expressions. A more specific type class can be created by
instantiating static with another type class.
One can force an expression to be evaluated at compile time as a constant
expression by coercing it to a corresponding static type:
import std/math
echo static(fac(5)), " ", staticbool
107
The Nim compiler should report any failure to evaluate the expression or a possible type mismatch error. In future versions of the Nim programming language the static metatype might not be required at all. It could delay the reporting of an error until the generic type is instantiated incorrectly: type Matrix[M, N, T] = array[0..(M*N - 1), T] var a, b: int var m: Matrix[a, b, int] # Error: the array size must be provided at compile-time. 16.30. typedesc[T] In many contexts, Nim treats the names of types as regular values. These values exist only during the compilation phase, but since all values must have a type, typedesc is considered their special type. typedesc acts as a generic type. For instance, the type of the symbol int is typedesc[int]. Just like with regular generic types, when the generic param is omitted, typedesc denotes the type class of all types. As a syntactic convenience, one can also use typedesc as a modifier. Procs featuring typedesc params are considered implicitly generic. They will be instantiated for each unique combination of supplied types, and within the body of the proc, the name of each param will refer to the bound concrete type: proc new(T: typedesc): ref T = echo "allocating ", T.name new(result) var n = Node.new var tree = new(BinaryTree[int]) When multiple type params are present, they will bind freely to different types. To force a bind-once behavior, one can use an explicit generic param: proc acceptOnlyTypePairsT, U 108 Once bound, type params can appear in the rest of the proc signature: template declareVariableWithType(T: typedesc, value: T) = var x: T = value declareVariableWithType int, 42 Overload resolution can be further influenced by constraining the set of types that will match the type param: template maxval(T: typedesc[int]): int = high(int) template maxval(T: typedesc[float]): float = Inf var i = int.maxval var f = float.maxval when false: var s = string.maxval # error, maxval is not implemented for string
109
16.31. typeof
typeof(x) can for historical reasons also be written as type(x)
but type(x) is discouraged.
One can obtain the type of a given expression by constructing a typeof value from it (in many other languages this is known as the typeof operator): var x = 0 var y: typeof(x) # y has type int If typeof is used to determine the result type of a routine call c(X) (where X stands for a possibly empty list of arguments), the interpretation where c is an iterator is preferred over the other interpretations, but this behavior can be changed by passing typeOfProc as the second argument to typeof: iterator split(s: string): string = discard proc split(s: string): seq[string] = discard
y has the typestring:
assert typeof("a b c".split) is string
assert typeof("a b c".split, typeOfProc) is seq[string]
typedesc[T] provides a mechanism for inferring the return type which
cannot be overloaded. The interaction between typeof, overloading, iterators,
typedesc[T] and generics allows for idioms that are not obvious to the casual
user of the language. Here is an example showing how to map JSON data to a
generic object type.
import std / json
proc fromJT: enum: T {.inline.} = T(j.getInt)
1
proc fromJ(t: typedesc[string]; j: JsonNode): string {.inline.} = j.getStr
proc fromJ(t: typedesc[bool]; j: JsonNode): bool {.inline.} = j.getBool
proc fromJ(t: typedesc[int]; j: JsonNode): int {.inline.} = int(j.getInt)
proc fromJ(t: typedesc[float]; j: JsonNode): float {.inline.} = j.getFloat
proc fromJT: seq: T = 2
result = newSeq[typeof(result[0])]()
assert j.kind == JArray
for elem in items(j):
result.add fromJ(typeof(result[0]), elem) 3
110
proc fromJT: object: T = 4
result = T()
assert j.kind == JObject
for name, loc in fieldPairs(result): 5
if j.hasKey(name):
loc = fromJ(typeof(loc), j[name]) 6
1 The fromJ family of procs supports the loading of enums, int, bool, float, string from JSON. 2 A seq can also be loaded from JSON. 3 Depending on the sequence element’s type call the correct overloaded fromJ proc. typeof(result[0]) is passed to the typedesc[T] parameter enabling static dispatching. 4 An object can also be loaded from JSON. 5 Iterate over every field of the object via fieldPairs. 6 fieldPairs allows for the mutation of the object that is iterated over so that loc = ... can be read as result. = .... Note that this solution allows to load complex objects which have fields that themselves are objects or primitives or sequences thereof. The compiler will produce type specialized code with no runtime overhead because the dispatching is resolved at compile-time.