EK9 Built in Types

EK9 naturally comes with a number of different built-in types, but also a set of collection types; these have been designed in from the outset and form the core of the language and its features.

There is a blog post here on strong typing and why it is important. In addition this is another post on modern types for languages. This makes the case for adding more sophisticated types to a programming language.

The Types

There are several predefined classes and types available; these are detailed in sections collection types, standard types and network types as they are used for specific purposes; these are more like a standard library/API. Whereas the types above are the basic building block types you will need. Most of the names above will be quite obvious and a number of examples of their use are give below.

Both String and Bits are also a sort of collection, String can also be viewed as an ordered list of Characters and Bits and order list of Booleans.

All of the types above are Objects, they are not primitive types, and they are all passed by reference and not by value. Please also note that as discussed at length in the operators and basics section a variable of any type can also be allocated space to hold a value (or not), but can also be un set. This means it does have space to hold a value but the value it holds is meaningless. Please also read the declarations section as this section will be using that terminology.

EK9 does not have static methods that are attached to classes; but has a number of mechanisms that facilitate that sort of functionality. The first is the use of functions to return an Object in a specific state, the second is the use of methods on classes that do not alter/mutate the Object they are called on but do return some value. This might seem a little strange, but it removes the fixed and binding nature of static methods on a fixed class, it reduces the need for specific Factory classes to some degree.

Important: No 'null' Literal

EK9 does not have a null literal. If you try to use null, you will receive error E01073.

Instead, EK9 uses tri-state semantics where values can be:

  • Absent - The variable doesn't exist (e.g., missing Dict key)
  • Unset - The variable exists but has no meaningful value
  • Set - The variable exists with a valid, usable value
// Creating an unset value (NOT null)
name <- String()      //Unset String
assert not name?      //_isSet() returns false

// Checking if set before use
if name?
  process(name)       //Safe - guaranteed set

// Guard for safe access
if validName <- getName()
  process(validName)  //Safe - guaranteed set

This design is intended to eliminate NullPointerException entirely: the compiler checks that every variable, property and return value holds an object before it is read, and the work to make that a proven guarantee (closing the known gaps in those checks) is ongoing (see Excluded Features for the evidence-based rationale).

Any

Any is the implicit universal supertype of all EK9 constructs. This includes classes, records, components, traits, functions, enumerations, constrained types, and generic instantiations.

Unlike Java's Object

In Java and similar languages, Object is a class and everything extends it as a class. EK9's Any is fundamentally different: it is a universal supertype that preserves the 'genus' (kind) of whatever it holds:

This makes Any somewhat "strange" compared to other languages - it adapts to hold any construct while preserving what that construct actually is.

Limited Operations

Any only provides one operator: the ? (is-set) operator. There are no other methods, operators, or call semantics available on Any directly. To work with the actual value, you need to know its real type.

Type Detection with Dispatchers

When working with Any, there comes a time when you need to "know" the actual type of the object. EK9 does not provide casting or instanceof - instead, use a dispatcher method that routes to the appropriate overload based on the runtime type. See Advanced Class Methods for details on dispatchers.

Use Cases for Any

  • //A heterogeneous list - declare List of Any (an undeclared mix sharing only Any is E50030)
  • mixedList as List of Any: [ 1, "hello", 2024-01-15, CardSuit.Hearts ]
  • //To process these, use a dispatcher - see Advanced Class Methods

Boolean

Variables of type Boolean can have the values of:

This is very different from most other languages in the sense that most other languages only support true and false.

  • //Declare a Boolean - using type inference but with no value yet known.
  • reviewAccess ← Boolean()
  • //Declare a Boolean - with type inference with a known initial value
  • allowModification ← false
  • //Declare a Boolean - but with no space allocated to hold the value.
  • provideAccess as Boolean
  • //Declare a Boolean - but with no value yet known.
  • denyAccess as Boolean: Boolean()
  • //Declare a Boolean - without type inference
  • allowViewing as Boolean := true

To understand what you can do with the boolean values; review the operators section. In general, it is best to employ type inference whenever possible. The Boolean is the simplest type.

Character

Variables of type Character can have the values of:

The above represents a single character, to represent a word or a sentence then use the String type.

  • //Some examples of using Character, some with type inference
  • c1 ← Character()
  • c2 ← 's'
  • c3 as Character := '\u00E9'
  •  
  • //Comparisons
  • c1 < c2
  • c1 <= c2
  • c1 > c2
  • c1 >= c2
  • c1 == c2
  • c1 != c2
  • c1 <> c2
  •  
  • //Other operators
  • length c1 //results in not set
  • length c2 //results in 1
  • length c3 //results in 1
  •  
  • c3.length() //same result (1) - but written in object form
  • #? c3 //results in 233 (hashcode method)
  • c3.#?() //same result (233) - but written in object form
  •  
  • //Promotion to String
  • str as String := #^c2 //results String of "s"
  • str ← #^c2 //also results String of "s" - uses type inference
  • str ← c2.#^() //also results String of "s" - uses type inference and object form
  •  
  • //Standard methods
  • str ← c2.upperCase() //results in 'S'
  • str ← c2.upperCase().lowerCase() //results in 's'

String

Variables of type String can have the values of:

The String is used to represent an order sequence of characters. It can be altered and can grow and unlike some languages does not have 'termination' character. It also implements the comparison and functional operators.

But note the important distinction between not set and "" an empty String. It also implements the following methods and operators.

It also adds a few additional methods.

String Literals

The String literal can take two forms in EK9; this is to support fixed Strings and also Interpolated Strings.

  • //Declare a fixed String literal
  • language ← "EK9"
  •  
  • //Declare a variable to be used in an interpolated String
  • age ← 30
  • //Use the variable in an interpolated String
  • message ← `Your age is ${age}`

The interpolated String starts and end with a back tick '`'. The sequence '${' indicates that the compiler should now look for a variable/function/method/expression in the scope to execute. This is the same sort of syntax as used in Javascript. This must return a variable of a type that either is a String, could be coerced to a String or has a '$' String operator on it.

Escapes in String

In a normal String i.e. "A Test" the " character must be escaped: i.e. "A \"Test\"", but back tick '`' and '$' do not need to be escaped.

A normal String is never interpolated, so it cannot hold '${': "${count} items" would print those characters rather than the count, and the compiler rejects it (E07901). Double quotes do interpolate in Kotlin, Groovy, Dart, PHP and the shells, which is why it is so easy to write by habit - and before the error existed it compiled silently and showed up only as wrong output. To interpolate, use a back tick String: `${count} items`. To keep the characters ${ as text, escape the dollar in a back tick String: `\${count} items`. A backslash is not an escape for '$' in a normal String ("\${count}" keeps the backslash), so the back tick form is the way to write it. The error message spells out both rewrites of the String it is on.

If you have a handful of Strings you need to use in code; then directly defining them in the source code might be done. But for very large text elements or a large number of text elements in different spoken languages consider using the text construct. This enables you to separate and format plain text Strings and interpolated Strings in a specific construct away from source code.

In an interpolated String the converse of escaping characters is true: i.e. `Your \`salary\` is \$${salary}`. Every '$' that does not open an interpolation is written \$ - including a regular expression replacement that names a group, `\${day}.\${month}`.

Interpolation Example

See text properties for defining a large number of String text items.

#!ek9
defines module introduction

  defines record
    Person
      firstName String: String()
      lastName String: String()
      
      Person()
        ->
          firstName String
          lastName String
        require firstName? and lastName?
        
        this.firstName :=: firstName
        this.lastName :=: lastName
      
      operator $ as pure
        <- rtn String: `${firstName} ${lastName}`
  
  defines function
    getAuthor()
      <- rtn Person: Person("Steve", "Limb")
                      
  defines program

    ShowStringInterpolation()
      stdout <- Stdout()
            
      normal <- "With normal double quote \" text"
      stdout.println("[" + normal + "]")

      var <- 90.9
      interpolated <- `A tab\t ${var} a " but a back tick \` the dollar have to be escaped \$\nOK?`
      stdout.println("[" + interpolated + "]")            

      me <- Person("Steve", "Limb")
      stdout.println(`Author: ${me} or via function: ${getAuthor()}, direct access "${me.firstName}"`)
//EOF        
        

This above program produces the following output:

[With normal double quote " text]
[A tab   90.9 a " but a back tick ` the dollar have to be escaped $
OK?]
Author: Steve Limb or via function: Steve Limb, direct access "Steve"
    

The two specific types of String literal enable a choice of functionality for creating output Strings. They have been designed to be explicitly different: using either " or `. More importantly the variables in interpolated string must always be used with ${...} and never just $ by itself. Because the two are explicitly different the compiler holds the line between them: ${ inside a " String is an error (E07901), never quietly treated as text.

Accessing Characters in a String

If you want to access specific a specific Character or a range of Characters in a String, pipeline processing as shown below.

#!ek9
defines module introduction

  defines function

    isO() as pure
      -> c as Character
      <- o as Integer: c == 'o' or c == 'O' <- 1 else 0
        
  defines program
      
    ShowStringType()
    
      i1 <- "The Quick Brown Fox"
      i2 <- "Jumps Over The Lazy Dog"
      space <- " "      
      i3 <- String()
                   
      hashCode <- #? i1
      //Hashcode is -732416445      
      
      sentence1 <- i1 + " " + i2      
      //sentence1 is "The Quick Brown Fox Jumps Over The Lazy Dog"      
      l1 <- length sentence1 //will be 43 
      l2 <- sentence1.length() //will also be 43 but in object form.
      
      oCount <- sentence1.count('o') //will be three as three 'o' in sentence
      oOCount <- cat sentence1 | map with isO | collect as Integer //will be four as 1 'O' and 3 'o'
      
      //Alternative way of joining values
      
      parts <- [i1, space, i2]
      sentence2 <- cat parts | collect as String
      //sentence2 is "The Quick Brown Fox Jumps Over The Lazy Dog"      
      
      //OR
      sentence3 <- cat [i1, space, i2] | collect as String 
      //sentence3 is "The Quick Brown Fox Jumps Over The Lazy Dog"
      
      //Mechanism to extract parts of a String
      jumpingBrownFox <- cat sentence2 | skip 10 | head 15 | collect as String

      //jumpingBrownFox is "Brown Fox Jumps"
      
      //Note that there is no issue with extending beyond the length of the sentence
      dog <-  cat sentence2 | skip 40 | head 15 | collect as String
      //dog is "Dog"
      
      //Even this is 'safe' and the result is just not set      
      nonSuch <-  cat sentence2 | skip 50 | head 5 | collect as String
      //nonSuch is not set
//EOF

There are few new and different ideas going on in the example above, the first thing to note are the standard operators and the type inference and also the length operator in 'object form'.

But importantly the two mechanisms that can be used to join String together; the standard '+' addition operator, but also the use of a List of String being collected into a single String (with the '|' operator).

There is the use of the Stream/Pipeline commands map, skip and head. These are used as there is no way to index Characters on a String. This is by design, this approach is safe in terms of indexing past the end of the actual length as shown in the example above nonSuch just is not set to any value (ie un set).

As an aside most Object Oriented languages tend to add more methods to Object types for tasks like getting substrings. EK9 does provide some methods on Objects, but mainly focuses on providing tooling to enable you to write your own functions to provide this type of functionality.

Clearly very frequent operations such as converting to lower case or upper case should be provided via the Object; but other operations that are less frequent should be developed as a standard library of functions by the developer. This keeps the EK9 standard library smaller, over time the EK9 standard library will increase but in a controlled manner.

Integer

Variables of type Integer can have the values of:

The comparison, mathematical , modification and ternary operators have been covered in other sections. But here are a couple of examples of operations on Integers for clarity. Note that idea of plus/minus infinity (or NaN) does not exist in EK9, you can make the argument infinity is not known; in EK9 that is represented by un set.

  • //Some examples of using Integer
  • i1 ← -9223372036854775808
  • hashCode ← #? i1 //results in -2147483648
  • asFloat ← #^ i1 //promotion to Float results in -9.223372036854776E18
  •  
  • //Division operations with zero
  • nonResult ← 0 / 90 //result is 0 (zero)
  • nonResult := 90 / 0 //result is un set i.e not known
  • nonResult := 0 / 0 //result is un set

There are no unsigned, short, long or other types; Integer is the only type that holds whole numbers.

Pipeline with Integer result

The accumulation of amounts (Integer/Float values etc.) is quite common; many languages have very distinct syntax for this specific task. EK9 does not have a specific syntax like list comprehensions, but it does have a much more general syntax that is covered in depth in the Streams/Pipelines section.

But as this section describes the use of the Integer type, a short example of how the Integer type can be used in a pipeline is given below. The objective of the code below is two-fold, first to get a list of Integers starting at 1 and incrementing by two up to and including 11; secondly to get the total of those values. The example below could have been done with a simple for loop, but it is shown here as a pipeline. The key point in the example is the collect as Integer syntax, you can use your own types here if you wish.

  • theValues as List of Integer := List() //create a list to capture values in
  •  
  • //Simple pipeline processing
  • intSum ← for i in 1 ... 11 by 2 | tee in theValues | collect as Integer
  •  
  • //theValues has the following contents [1, 3, 5, 7, 9, 11]
  • //intSum has the value 36

The first couple of examples here and all the standard operators are probably what you would normally expect with and Integer type. But the last example is probably something very different from what you've seen before. This is discussed in a little more detail here but is much more detail in the Streams/Pipelines section. In general this pipeline approach enables reuse and collection of objects during the pipeline processing - but in a standard syntax.

As an example, later sections will discuss Durations as part of date/time processing, it is possible to use a pipeline like the one above but with durations rather than integers. But the concept of the processing pipeline and accumulations, mapping and teeing is the same irrespective of type.

It is accepted that the syntax is not as terse as list comprehensions.

Float

Variables of type Float can have the values of:

As with the Integer the comparison, mathematical , modification and ternary operators have been covered in other sections.

  • //Some examples of using Float
  • i1 ← -4.9E-324
  • hashCode ← #? i1 //results in -2147483647
  •  
  • //Division operations with zero
  • nonResult ← 0.0 / 90.0 //result is 0.0 (zero)
  • nonResult := 90.0 / 0.0 //result is un set i.e not known
  • nonResult := 0.0 / 0.0 //result is un set

As you can see the Float has most of the same operators as Integer (except promotion '#^') but has variable precision (minimum and maximum values). Note it too can be used in pipeline processing and can receive Float values.

Float also provides trigonometric and transcendental methods that operate in radians and delegate to the platform maths library: sin, cos, tan, asin, acos, atan, atan2, log (natural), log10 and exp, plus round(places). Any result that is not a finite number (for example asin of a value outside -1..1) is returned un set. The constants pi and e are provided by Maths.

BigInteger

BigInteger is an arbitrary-precision integer. Unlike Integer - which is 64-bit and becomes un set on overflow - a BigInteger grows as large as needed, so '+', '-', '*' and '^' (power) never overflow. Use it for very large whole numbers such as big factorials, unbounded Fibonacci or large combinatorics.

It has the same comparison and mathematical operators as Integer, plus sqrt and abs. The mod and rem operators return an Integer (as EK9 requires); when a result will not fit in 64 bits it is returned un set rather than silently losing data. There is deliberately no promotion ('#^'): a BigInteger is wider than a Float, so an implicit conversion would lose precision.

#!ek9
defines module introduction

  defines program

    BigIntegerExample()
      stdout <- Stdout()

      //A value far beyond the range of a 64-bit Integer
      big <- BigInteger("123456789012345678901234567890")

      //Multiplication never overflows
      squared <- big * big
      stdout.println("Squared is " + $squared)

BigDecimal

BigDecimal is an arbitrary-precision decimal - the exact-decimal counterpart of Float. Where Float uses 64-bit binary and so cannot represent values like 0.1 exactly, a BigDecimal is exact: 0.1 + 0.2 really is 0.3. Use it for high-precision or exact calculations (for money specifically use Money).

Equality is by value, not scale, so 1.0 equals 1.00 ('==', '<=>' and '#?' all use value comparison). Division and sqrt keep 34 significant digits (so non-terminating results such as 1/3 are well defined); division by zero is un set. It also provides round(places), abs and '^' (power). Like BigInteger it has no promotion operator.

#!ek9
defines module introduction

  defines program

    BigDecimalExample()
      stdout <- Stdout()

      total <- BigDecimal("0.1") + BigDecimal("0.2")
      stdout.println("0.1 + 0.2 = " + $total)

Maths

Maths is a small, stateless utility type that exposes the mathematical constants as Float values: pi() and e(). Create one with Maths() and read the constants; the trigonometric and transcendental functions themselves live on Float.

#!ek9
defines module introduction

  defines program

    MathsExample()
      stdout <- Stdout()

      maths <- Maths()

      //sin of pi/2 is 1.0
      halfPi <- maths.pi() / 2.0
      stdout.println("sin(pi/2) = " + $halfPi.sin())

Bits

Variables of type Bits can have the values of:

The Bitwise operators have been discussed in the operators section. The '+' and '+=' operators are provided to join Bits together as is the '|' operator when used with Streams/Pipelines.

  • //Bits values
  • a ← 0b010011
  • b ← 0b101010
  • c ← 0b010011
  •  
  • //Equality Operators
  • result ← a == b //result would be false
  • result ← a <> b //result would be true
  • result ← a < b //result would be true
  • result ← a <= b //result would be true
  • result ← a > b //result would be false
  • result ← a >= b //result would be false
  • result ← a == c //result would be true
  • result ← a != c //result would be false
  • result ← a <> c //result would be false (alternate syntax)
  •  
  • //Joining/Adding Bits
  • set6 ← 0b010011
  • set7 ← 0b101010
  •  
  • set6set7 ← set6 + set7 //result would be 010011101010
  • set6true ← set6 + true //result would be 0100111
  • set6false ← set6 + false //result would be 0100110
  • set6false += true //set6false would now be 01001101
  •  
  • //Bitwise operations
  • set6 ← 0b010011
  • set7 ← 0b101010
  •  
  • ored ← set6 or set7 //result would be 0b111011
  • xored ← set6 xor set7 //result would be 0b111001
  • anded ← set6 and set7 //result would be 0b000010
  • notted ← ~set6 //result would be 0b101100
  • alsonotted ← not set6 //result would be 0b101100

As you can see in the addition example above, Bits are not treated like numbers.

A Stream/Pipeline approach should be taken to access ranges of specific bits from a set of Bits as shown below.

  • //Extracting ranges of Bits
  • //assumes function booleanToBits has been defined
  • partial ← set6false | skip 3 | map booleanToBits | collect as Bits
  • //partial has the value of 0b01001

Importantly Bits are read right to left as you would expect the least significant bit is to the right and the most significant bit is to the left. So the streaming of the bits (Boolean values) starts with the least significant bit. Hence, in the example above the "101" is skipped this just leaves the "01001" left. You can use head and tail to further refine your sub-selection.

Hopefully you can now start to see that rather than defining an Object specific set of methods on classes to get the result of a sub-selection; the functional methodology here really is repeatable and consistent. So given a page it is feasible to filter and sub-select a number of paragraphs using various criteria.

This point is quite important and is in no way aimed at denigrating and Object-Oriented approach, but rather splits development between Object Orientation and a Functional Programming. You may find this difficult at first, but when you think "hey why is there no method on this class to do X", it's probably time to consider a more functional approach to the problem.

Over time you will find your classes become smaller/concise, but your library of functions grows and is much more reusable in different contexts.

Byte

Byte is an unsigned 8-bit value - 0 to 255 - and it is the scalar of the byte family. It exists so that raw binary data has an element type with real EK9 semantics (value semantics, immutability and tri-state) rather than a naked small integer.

It is deliberately unsigned, not signed -128..127. Raw binary data is the use case; arithmetic is not. There are therefore no arithmetic operators ('+', '-', '*', '/') and no '++' or '--' - so the question "what is 0xFF + 1?" cannot be asked of a Byte at all. Construction outside the range gives an un set value rather than a wrapped one, so a bad octet is visibly absent instead of quietly wrong.

To do arithmetic, convert deliberately with toInteger(), or let the '#^' promotion widen it inside an Integer expression. Note the asymmetry: 1 + aByte resolves, because promotion applies to the argument of Integer's '+'; aByte + 1 does not resolve, because Byte has no '+' of its own. Arithmetic over a Byte therefore always widens to Integer and never wraps.

The String form is hex in both directions. The '$' operator emits exactly two upper-case hex characters, and the String constructor parses hex with an optional 0x prefix, so Byte($b) == b holds for all 256 values. Decimal text is deliberately rejected - Byte("255") is un set - because "10" cannot mean sixteen to '$' and ten to the constructor. Decimal arrives through Byte(Integer).

The bitwise operators 'and', 'or', 'xor' and '~' work as they do on Bits and Boolean. The shift operators '<<' and '>>' stay inside eight bits - bits shifted out are lost and the value never widens - and '>>' is always a logical shift because there is no sign bit to propagate.

#!ek9
defines module introduction

  defines program

    ByteExample()
      stdout <- Stdout()

      //Hex text, an Integer, or a copy
      fromHex <- Byte("0xFF")
      fromInteger <- Byte(255)

      //$ always emits two upper-case hex characters, so the round trip holds
      stdout.println($fromHex)                  //FF
      stdout.println($Byte(15))                 //0F

      //Out of range is un set, never wrapped
      stdout.println(`Byte(256) is set: ${Byte(256)?}`)

      //Bitwise, and shifts that stay inside eight bits
      masked <- Byte("AA") and Byte("0F")       //0A
      stdout.println($masked)
      stdout.println($(Byte("0F") << 4))        //F0
      stdout.println($(Byte("0F") << 8))        //00 - shifted out, not widened

      //Arithmetic is deliberate - and widens rather than wrapping
      stdout.println(`1 + 0xFF = ${1 + #^fromInteger}`)   //256

      //Conversions
      stdout.println($fromInteger.toInteger())  //255
      stdout.println($Byte("AA").toBits())      //10101010

Bytes

Bytes is an immutable sequence of unsigned 8-bit values - the public face of binary data. It follows the sequence semantics of Bits rather than those of a number: '+' concatenates, '|' accumulates, and there is no arithmetic at all.

🔑 The literal is 0y followed by hex digits in PAIRS - 0yDEADBEEF. 0x remains an Integer literal and 0b remains a Bits literal, so the three never collide. The even-digit rule is enforced by the lexer, which is why 0yF does not parse at all rather than becoming an un set Bytes to be diagnosed later - half a byte is not a byte, and the earliest place to say so is the one that cannot be bypassed. 0Y works too, as with 0X/0B.

Empty is not un set. Bytes() is un set; a zero-length Bytes is set and empty. Binary protocols depend on that difference - a field that was absent and a field that was present and empty are not the same fact.

The String form is hex both ways, as it is on Byte: '$' emits upper-case hex and the String constructor parses hex, so Bytes($b) == b always holds. An odd number of hex digits is un set - half a byte is not a byte. Base64 is deliberately not guessed at, because "DEADBEEF" is simultaneously valid hex and valid Base64; name the encoding instead with the two-argument constructor, Bytes(text, "base64").

Buffers come from the withLength factories, following the 'with' convention used across the standard library: Bytes().withLength(16) is sixteen zero bytes, and Bytes().withLength(16, Byte("FF")) fills them with a value.

🔑 'and', 'or' and 'xor' require both operands to be the same length. A mismatch is un set, never zero-extended. This is the one place Bytes departs from Bits, and the reason is a bug category rather than a preference: zero-extending a short key leaves the tail of the longer operand unchanged, so the "ciphertext" would carry plaintext while the operation reported success. Length is semantic for binary data - a 16-byte key is 16 bytes - so a mismatch is a defect at the call site and is made visible.

'<<' and '>>' shift the whole sequence by a number of bits, carrying them across byte boundaries, and preserve the length - bits shifted off the end are lost, exactly as on Byte. That is deliberately not the Bits behaviour, where a left shift lengthens the sequence. To change width, say so: growLeft(count) and growRight(count), each taking an optional fill Byte, add space at one end or the other. The two pair naturally - Bytes("DEAD").growLeft(2) is 0000DEAD, and a following '<< 16' moves the content along into the room just made.

Ordering ('<', '<=', '>', '>=', '<=>') is lexicographic over unsigned octets, so 0xFF sorts above 0x7F. slice, first and last return values rather than views, so a slice can never mutate its source, and an out-of-range request is un set rather than clamped. indexOf is un set when the needle is absent - there is no -1 to be mistaken for a real index. Use constantTimeEquals for MACs, tokens and digests: it examines every byte regardless of where the first difference is, whereas '==' is free to short-circuit.

#!ek9
defines module introduction

  defines program

    BytesExample()
      stdout <- Stdout()

      //Hex in, hex out - the round trip always holds
      payload <- Bytes("DEADBEEF")
      stdout.println(`${payload} has ${length payload} bytes`)
      stdout.println(payload.toBase64())

      //Buffers via the 'with' factory convention
      mask <- Bytes().withLength(4, Byte("FF"))

      //xor is its own inverse - the workhorse of low level code
      masked <- payload xor mask
      stdout.println(`masked ${masked}, recovered ${(masked xor mask) == payload}`)

      //Unequal lengths are un set, never zero extended
      stdout.println(`short key set: ${(payload xor Bytes("FF"))?}`)

      //'+' concatenates and slices are values, not views
      stdout.println($(Bytes("DEAD") + Bytes("BEEF")))
      stdout.println($payload.first(2))

      //Shifts cross byte boundaries and keep the length
      stdout.println($(payload << 4))

      //Grow to make room, then shift the content into it
      stdout.println($(Bytes("DEAD").growLeft(2) << 16))

BytesReader

BytesReader is a forward-only cursor over a Bytes, for binary protocol and format parsing. Obtain one with aBytes.reader() rather than constructing it - a reader with no source has nothing to read.

🔑 Malformed input is a value, not an exception. Every read is tri-state, and a read that would run past the end returns un set and leaves the cursor where it was. So parsing is a sequence of ordinary reads followed by a set-check, with no exception handling and no bounds test before every field. Because a failed read does not move the cursor, a smaller read still succeeds afterwards, and a wide read can never part-consume.

Cursor state is position(), remaining() and atEnd(), with peek() to look at the next byte without consuming it. skip(count) advances, answering false and not moving if it would run past the end - it never clamps.

🔑 Endianness is always explicit. There is deliberately no readUInt32() with a default - only readUInt16BE/LE, readUInt32BE/LE, readUInt64BE/LE and the signed readInt16/32/64BE/LE. A silent endianness default is a known bug category: it is correct on the machine it was written on and wrong in production. EK9 eliminates bug categories by making them unrepresentable, so the defaulted spelling does not exist. Signed reads sign-extend; unsigned ones do not.

⚠️ readUInt64BE/LE is un set when the top bit is set. EK9's Integer is 64-bit signed, so an unsigned 64-bit value that large has no correct representation - returning it would hand back a negative number for a large positive one. It is also not consumed, so readInt64BE can reinterpret the same bytes.

slice(count) is what makes nested formats safe: a bounded sub-reader over the next n bytes which consumes them from the parent. The sub-reader cannot read past its own boundary however malformed its content turns out to be, and the parent carries on after it - exactly what a length-prefixed structure needs.

#!ek9
defines module introduction

  defines program

    BytesReaderExample()
      stdout <- Stdout()

      //A TLS-record shape: type, version, body length, then a length-prefixed body
      record <- Bytes("1603030008034142430358595a")
      cursor <- record.reader()

      stdout.println(`type ${cursor.readByte()}, version ${cursor.readUInt16BE()}`)
      bodyLength <- cursor.readUInt16BE()

      //A bounded view - it cannot read past its own boundary
      body <- cursor.slice(bodyLength)
      fieldLength <- body.readByte().toInteger()
      stdout.println(`first field: ${body.readBytes(fieldLength)}`)

      //A short read is un set and does NOT move the cursor
      short <- Bytes("010203").reader()
      stdout.println(`wide read set: ${short.readUInt32BE()?}`)
      stdout.println(`still at ${short.position()}, smaller read gives ${short.readUInt16BE()}`)

BytesBuilder

BytesBuilder is the mutable working buffer - the write side of the family and the counterpart to BytesReader. It is the only type here that permits an indexed byte write, and it exists so that that capability does not leak into Bytes or general application code.

🔑 Speed is not why it exists. Bytes is a value, so every + and | copies the whole accumulator and building byte-by-byte is quadratic - but that path is correct, merely slow, and fine at protocol sizes. Three capabilities are inexpressible on Bytes at any speed, and those are the justification.

1. Endian-aware writes. appendUInt16BE/LE, appendUInt32BE/LE, appendUInt64BE/LE and the signed appendInt16/32/64BE/LE - twelve writes mirroring BytesReader's twelve reads one for one. Without them there is no endian-aware write anywhere in the language, so callers hand-assemble bytes and reintroduce exactly the endianness bug class the reader eliminated. As on the read side, there is no defaulted appendUInt32() spelling.

2. In-place mutation - setByte(index, aByte), xorInto(offset, source) and copyInto(offset, source). A value type cannot offer these, because every operation on Bytes yields a new array. xorInto is one primitive with three uses: the AES AddRoundKey, the Argon2 block XOR, and the ChaCha20 keystream application. All three write within the current length and never extend - use append to grow.

3. wipe() - a Bytes cannot be securely erased, because + and | have already scattered copies of it. A builder is one buffer that can actually be zeroed, which is what key material and password buffers need. Best-effort by nature: the JVM may have moved it during garbage collection and the OS may have paged it out.

⚠️ Range is enforced, not truncated. appendUInt16BE(0x10000) makes the builder un set rather than writing 0x0000, and appendUInt64BE(-1) is refused rather than written as two's complement - the mirror image of readUInt64BE refusing to hand back a negative for a large positive. Use the signed spelling when a signed value is meant. A write that would run past the end is refused outright, never part-applied, because half an applied round key is worse than none.

An empty builder is set - empty is not un set. capacity() is backing storage and is not length; reserve(n) grows it and never shrinks. Growth is geometric, so appending is amortised rather than quadratic. toBytes() is the single visible boundary back to an immutable value and always copies, so a snapshot cannot change underneath whoever holds it.

⚠️ $ reports shape, never content - BytesBuilder(4/16), the length and capacity, not the bytes, so a buffer holding a key cannot leak through a debug print. For the same reason there is deliberately no #?: a hash that changes as the buffer is appended to makes a mutable key look usable, and that is a bug rather than a feature.

#!ek9
defines module introduction

  defines constant

    tlsVersionTls12 <- 0x0303
    tlsBodyLength <- 8

  defines program

    BytesBuilderExample()
      stdout <- Stdout()

      //Building the TLS-record shape that the BytesReader example parses. The endianness is in
      //the method name, so it is visible on the line doing the writing.
      record <- BytesBuilder()
      record.append(Byte("16"))
      record.appendUInt16BE(tlsVersionTls12)
      record.appendUInt16BE(tlsBodyLength)
      record.append(Bytes("034142430358595a"))
      stdout.println(`framed: ${record.toBytes()}`)

      //Big and little endian genuinely differ - there is no defaulted spelling
      bigEndian <- BytesBuilder()
      bigEndian.appendUInt32BE(0x01020304)
      littleEndian <- BytesBuilder()
      littleEndian.appendUInt32LE(0x01020304)
      stdout.println(`big ${bigEndian.toBytes()}, little ${littleEndian.toBytes()}`)

      //A number too wide for its field is rejected, never truncated into a quiet lie
      tooWide <- BytesBuilder()
      tooWide.appendUInt16BE(0x10000)
      stdout.println(`too wide accepted: ${tooWide?}`)

      //xorInto is AddRoundKey, the Argon2 block XOR and the ChaCha20 keystream application,
      //and it is its own inverse
      block <- BytesBuilder(Bytes("ffffffff"))
      block.xorInto(0, Bytes("0f0f0f0f"))
      stdout.println(`after xor: ${block.toBytes()}`)

      //wipe zeroes the backing store and leaves the builder usable - something an immutable
      //Bytes fundamentally cannot offer
      secret <- BytesBuilder(Bytes("cafebabedeadbeef"))
      secret.wipe()
      stdout.println(`wiped to length ${length secret}, still usable ${secret?}`)

Digest

Digest produces unkeyed cryptographic hashes. It answers one question - has this content changed? - and it is the honest name for what this capability has always been. A digest proves integrity and nothing else: anyone at all can compute one, so anyone who can replace your content can replace the digest beside it too.

Three overloads of SHA256, differing in what they take and what they hand back: SHA256(String) and SHA256(GUID) return a String of hexadecimal, while SHA256(Bytes) takes and returns Bytes. The last is the one to reach for when the input is a file, a request body, or anything else that is not text - it hashes the OCTETS with no decode and no encoding assumption, and it chains directly into the next step (an HMAC key, a signature) without a hex round-trip.

SHA512 has exactly the same three overloads, so the two algorithms are learned as one form: SHA512(String), SHA512(GUID) and SHA512(Bytes), the last returning 64 octets where SHA-256 returns 32. SHA-1 and MD5 are deliberately not offered here - both are broken for security use.

🔑 Hex is UPPER-case across the whole family. SHA256(String), Bytes' $ and toText("hex") all render the same way, so a String digest and a Bytes digest of the same input compare equal with no normalising. That was not true before 2026-09-25: the String form emitted lower-case, inherited from the type formerly misnamed HMAC, and the two spellings of one SHA-256 were silently unequal.

⚠️ Published vectors are conventionally written lower-case - shasum, git and openssl all print that way - so a literal copied from a specification needs its case folded before comparison. The value is identical; only the rendering differs.

🔑 If you need to prove WHO produced something, this is the wrong type - reach for HMAC. A hash over attacker-supplied content is exactly as valid as a hash over yours, which is the mistake this split of the two types exists to prevent.

HMAC

HMAC is a keyed message authentication code. It proves integrity AND authenticity, because producing one requires the key: SHA256(key, message), where both are Bytes and the result is Bytes.

🔑 The key is not optional, and that is the entire point. Until the Bytes family existed there was no honest way to express a keyed MAC, and this type carried the unkeyed hashes instead - so code reaching for HMAC got a bare digest and the authenticity it assumed was imaginary. Those hashes now live on Digest, and this type finally means what its name says.

⚠️ Verify a received tag with constantTimeEquals, never with ==. Both answer the same question, but == is free to stop at the first differing byte, and how long it took then tells an attacker how many leading bytes they guessed correctly - which forges a tag one byte at a time. Bytes.constantTimeEquals folds in the XOR at every position, and the length difference too, so there is nothing to time.

Unset in, unset out: an unset key or an unset message yields an unset result, never a MAC computed over a substitute value. An empty key is a different thing from an unset one and is accepted, because empty is set.

The example below uses RFC 4231 test case 2 - a published vector, so it agrees with every conforming implementation or it is wrong. A value captured from our own output would pass just as happily over a subtly broken MAC.

#!ek9
defines module introduction

  defines constant

    sha256HexLength <- 64

  defines program

    DigestAndHmacExample()
      stdout <- Stdout()

      //Digest is UNKEYED - anyone can compute it, so it proves only integrity
      fingerprint <- Digest().SHA256("the quick brown fox")
      stdout.println(`fingerprint length ${length fingerprint}, expected ${sha256HexLength}`)

      //The Bytes overload hashes OCTETS and answers in octets - no encoding assumption
      //anywhere, so it works on a binary payload as readily as on text
      payload <- Bytes().withUtf8("the quick brown fox")
      asOctets <- Digest().SHA256(payload)
      stdout.println(`same value both ways: ${asOctets.toText("hex") == fingerprint}`)

      //HMAC is KEYED - producing one requires the secret, so it proves WHO
      key <- Bytes().withUtf8("Jefe")
      message <- Bytes().withUtf8("what do ya want for nothing?")
      mac <- HMAC().SHA256(key, message)
      stdout.println(`mac: ${mac}`)

      //🔑 Change the key and the MAC changes completely. That is the whole difference:
      //a Digest anyone can forge over new content, an HMAC they cannot without the key.
      forged <- HMAC().SHA256(Bytes().withUtf8("wrong-key"), message)
      stdout.println(`a wrong key cannot reproduce it: ${mac <> forged}`)

      //⚠️ VERIFY WITH constantTimeEquals, never with ==. Both answer the same question, but
      //== is free to stop at the first differing byte, and how long it took then tells an
      //attacker how many leading bytes they guessed right - a tag forged one byte at a time.
      stdout.println(`verified: ${mac.constantTimeEquals(HMAC().SHA256(key, message))}`)
      stdout.println(`forgery rejected: ${mac.constantTimeEquals(forged)}`)

      //Unset in, unset out - never a MAC computed over a substitute value
      stdout.println(`unset key gives unset mac: ${HMAC().SHA256(Bytes(), message)?}`)

Which produces:

fingerprint length 64, expected 64
same value both ways: true
mac: 5BDCC146BF60754E6A042426089575C75A003F089D2739839DEC58B964EC3843
a wrong key cannot reproduce it: true
verified: true
forgery rejected: false
unset key gives unset mac: false

SecureRandom

SecureRandom produces cryptographically secure random octets - for nonces, salts, keys and tokens. It has one operation: nextBytes(count), which returns count octets as Bytes. A token is those octets rendered as text - nextBytes(32).toBase64() or .toText("hex"). There is deliberately no next() returning an Integer: reducing random numbers into a range by hand is how modulo bias creeps into tokens.

🔑 It is deliberately not a Random. SystemRandom and SeededRandom implement the Random trait; SecureRandom does not. If it did, anything accepting a Random would accept a SeededRandom just as happily - and a seeded generator gives the same sequence every run, so a nonce or salt made from it is reproducible by anyone who knows the seed. Because it stands outside the trait, a function that needs unpredictable octets declares a SecureRandom parameter, and handing it a SeededRandom is a compile error rather than a code-review catch. For the same reason there is no seeded constructor.

Tri-state: an unset or negative count gives unset Bytes, never a substitute length. A count of zero gives empty Bytes, which are set - asking for nothing and receiving nothing is a well-defined answer.

#!ek9
defines module introduction

  defines function

    //Declaring SecureRandom is the safety property - a SeededRandom will not compile here
    newSalt() as pure
      ->
        generator as SecureRandom
      <-
        rtn as Bytes: generator.nextBytes(16)

  defines program

    SecureRandomExample()
      stdout <- Stdout()
      generator <- SecureRandom()

      //Exactly the number of octets asked for
      salt <- newSalt(generator)
      stdout.println(`salt length: ${length salt}`)

      //A token is octets rendered as text - there is no next() Integer on purpose
      token <- generator.nextBytes(32).toBase64()
      stdout.println(`token length: ${length token}`)

      //Two generators never agree - there is no seed for them to share
      stdout.println(`generators differ: ${SecureRandom().nextBytes(32) <> SecureRandom().nextBytes(32)}`)

      //Zero is EMPTY and SET; an unset or negative count is UNSET
      stdout.println(`empty is set: ${generator.nextBytes(0)?}`)
      stdout.println(`negative count: ${generator.nextBytes(-1)?}`)

Which produces:

salt length: 16
token length: 44
generators differ: true
empty is set: true
negative count: false

Time

Variables of type Time can have the values of:

Time represents the concept of the time of day HH:MM or HH:MM:SS. This is in isolation to any particular date. There is a separate type that represents the time of day on a particular date in a particular time zone This is a DateTime.

The comparison and ternary operators have been covered in other sections. But here are a couple of examples of operations on Time for clarity. Time can also be used with Duration. Indeed, it only makes sense to use the '+', '-', '+=' and '-=' operators on Time with a Duration. Moreover, it makes no sense to be able to add, multiply or divide by a Time (though with a Duration this does make sense). But it is logical to be able to subtract one Time from another to get a Duration.

Some examples and also a demonstration of the require key word, very useful for pre-conditions / post-conditions.

#!ek9
defines module introduction

  defines program
  
    ShowTimeType()
    
      //These declarations and operations should be obvious by now if you've read the other sections.
      
      t1 <- 12:00
      t2 <- 12:00:01
      t3 <- 12:00:01
      t4 <- Time()
      t5 <- Time(12, 00)
      t6 <- Time(12, 00, 01)
            
      require t1 <> t2
      require t2 == t3     

      require t1 < t2
      require t1 <= t2
      require t2 <= t3
      
      require t2 > t1
      require t2 >= t1
      require t3 >= t2
      
      //Check not set
      require ~t4? 
          
      require t1 == t5      
      require t2 == t6

      //Just create a Time object to get the current time
      t4a <- SystemClock().time() //t4a will be set to current time
      t4b <- t4.startOfDay() //t4b will be set to 00:00:00      
      t4c <- t4.endOfDay() //t4c will be set to 23:59:59
      
      //Now actually alter t4 itself
      t4.set(SystemClock().time()) //t4 will now be set to the current time
      require t4?      
      t4.setStartOfDay() //Will be set to 00:00:00
      require t4?      
      t4.setEndOfDay() //Will be set to 23:59:59
      require t4?
      
      //Durations      
      d1 <- PT1H2M3S //one hour, two minutes and three seconds
      d2 <- P3DT2H59M18S //three days, two hours, fifty nine minutes and eighteen seconds
      
      //Note how only the hours minutes and seconds of 'd2' is used.
      
      //Addition operations
      t4 := t5 + d1
      require t4 == 13:02:03      
      t4 += d2
      require t4 == 16:01:21
      
      //Subtraction operations
      t4 := t5 - d1
      require t4 == 10:57:57
      t4 -= d2
      require t4 == 07:58:39
      
      //Get the Duration between two times
      d3 <- 12:04:09 - 06:15:12
      require d3 == PT5H48M57S      
      
      //Note the different views of a Duration
      d4 <- 06:15:12 - 12:04:09
      require d4 == PT-6H11M3S
      require d4 == PT-5H-48M-57S
      
      //Example of building a time from a number of durations
      durations as List of Duration := [d1, d2]
      
      t7 <- cat durations | collect as Time
      require t7 == 04:01:21
      
      //t7 will default to start of the day plus each of the durations piped in
      //Note that time will just roll over just after 23:59:59
      
      //It is also possible to get the hour, minute and second from a Time
      second <- t7.second()
      minute <- t7.minute()
      hour <- t7.hour()
      
      require 4 == hour
      require 1 == minute
      require 21 == second
      
//EOF

Duration

Variables of type Duration can have the values of:

The Duration has been shown in the previous example using Time. With Time the duration values of hour, minute and second are really the only relevant ones. But a Duration can also hold years, months and days. Which will be important when the types Date and DateTime are covered.

The format of the duration is as per ISO 8601 Durations, though fractional values are not supported see Millisecond on how to work with fractions of a second. Durations take the following format: P[n]Y[n]M[n]W[n]DT[n]H[n]M[n]S.

The 'P' denotes a period, then each '[n]' value is then followed by what the field is to be applied. So for example, "P3Y6M4DT12H30M5S" represents a duration of "three years, six months, four days, twelve hours, thirty minutes, and five seconds".

It is also possible to use 'W' for weeks in addition to months and days.

Durations are designed to be used without milliseconds and in some ways anything less than a second is really in a different category of time (as a human would understand). But clearly it is possible to multiply up milliseconds so they become significant (to a human). The Millisecond type and the Duration are designed to work together (but are different).

  • //Some examples of using Duration
  • d1 ← PT1H2M3S //results in 1 hour, 2 minutes and 3 seconds
  • d2 ← P3DT2H59M18S //results in 3 days, 2 hours, 59 minutes and 18 seconds
  • d3 ← P2Y6W3DT8H5M8S //results in 2 years, 45 days, 8 hours, 5 minutes and 8 seconds
  • //Note that 6 weeks was converted to 42 days and then the additional 3 days are added to that
  • d4 ← P2Y2M1W3DT8H5M8S //results in 2 years, 2 months, 10 days, 8 hours, 5 minutes and 8 seconds
  • d5 ← P2W //results in 14 days
  •  
  • //Mathematics with durations
  • d6 ← PT1H2M1S
  • d6a ← d6 * 2 //results in PT2H4M2S i.e. twice the duration in d6
  • d6b ← d6 * 8.25 //results in PT8H31M38S
  • d6c ← d6b / 3 //results in PT2H50M32S
  • d6d ← d6 + PT20M //results in PT1H22M1S
  • d6e ← d6 - PT2M6S //results in PT59M55S
  • d6e += PT2H5S //results in PT3H

As you can see from the examples above; once you have a Duration of any period of time you can use normal mathematics to manipulate that Duration. The Duration has logical and rational mathematics operators; for example it makes no sense to add an Integer, Float or String to a Duration. However, it does make sense to be able to add and subtract a Duration, but also multiply and divide a Duration by an Integer or Float.

It is these specific types and your own defined types like records and classes where adding specific operators and overloading those operators with varying (but suitable) types that enable EK9 to remain readable and easy to understand.

Millisecond

Variables of type Millisecond can have the values of:

The format of the Millisecond literal is digits'ms' for example 250ms is two hundred and fifty milliseconds. The compiler will check the validity of millisecond literals. You can have any valid integer (positive or negative). So 6000ms is actually 6 seconds. Naturally you can add these types to Durations and if there values result in seconds, minutes or hours etc. then those will be added to the Duration. You can also convert them to a Duration by using the duration() method. The value of the milliseconds will be rounded up to the nearest second where necessary, for example 1501ms will round up to two seconds.

Naturally the Millisecond type supports many of the normal mathematical and comparison operators and can be promoted to a Duration and also constructed with the Duration type.

Date

Variables of type Date can have the values of:

The format of the date literal is YYYY-MM-DD for example 2020-11-18 is the Eighteenth of November 2020. The compiler will check the validity of date literals, so for example declaring a date of '2020-11-31' would not even compile.

It is important to note that the Date is not associated with any sort of time zone. It is just a Date so this could be used to hold a date of birth for example. If you need to be exact on the Date then using a DateTime would be more appropriate.

The range of operations on Date are very similar to those offered on the Time type; comparison and ternary for example. As with Time Duration can be used to do calculations, it only makes sense to use the '+', '-', '+=' and '-=' operators on Date with a Duration. Moreover, it makes no sense to be able to add, multiply or divide a Date (though with a Duration this does make sense). But it is logical to be able to subtract one Date from another to get a Duration.

#!ek9
defines module introduction

  defines program
      
    ShowDateType()
            
      //3rd of October 2020
      date1 <- 2020-10-03
      date2 <- 2020-10-04
      date3 <- 2020-10-04
      date4 <- Date()
      date5 <- Date(2020, 10, 03)      
      date6 <- Date(2020, 10, 04)
           
      require date1 <> date2
      require date2 == date3      

      require date1 < date2
      require date1 <= date2
      require date2 <= date3
      
      require date2 > date1
      require date2 >= date1
      require date3 >= date2
      
      require ~date4?
          
      require date1 == date5      
      require date2 == date6
                  
      date4a <- Date().today()
      require date4a?
      
      date4.setToday() //Now set to the current date
      
      //Durations
      d1 <- P1Y1M4D //one year, one month and four days
      d2 <- P4W1DT2H59M18S //twenty nine days, two hours, fifty nine minutes and eighteen seconds
      
      //Note how only the days part of 'd2' is used
      date4 := date5 + d1
      require date4 == 2021-11-07
      
      date4 += d2
      require date4 == 2021-12-06
      
      d3 <- date4 - date5
      require d3 == P1Y2M3D
      
      //Pipeline with durations and dates.
      durations as List of Duration := [d1, d2]
      
      //Because the epoch is 1970-01-01 accumulation starts from that date            
      date7 <- cat durations | collect as Date
      require date7 == 1971-03-06
      
      //If you want to accumulate from a specific date you can do this instead
      date8 <- 0000-01-01
      cat durations >> date8
      require date8 == 0001-03-06
      
      dayOfMonth <- date8.day()
      monthOfYear <- date8.month()
      year <- date8.year()      
      dayOfTheWeek <- date8.dayOfWeek()      
      
      require dayOfMonth == 6
      require monthOfYear == 3
      require year == 1
      require dayOfTheWeek == 2

//EOF

When using Dates and Durations you must take care because transitioning and adding durations to specific points in time and subtracting points in time to recover Durations can lead to some strange effects. For example adding a Duration of a few months and days to date where the span will cover a leap year will result in a new date that takes that particular year (leap) into account.

Also the addition of durations uses 30 days per month to when summing up days. This means that adding a number of durations together and then applying that to a date will give one result. Whereas, taking an initial date and adding each duration in turn will give a different date.

If you are looking for a final date from a base date when using a number of durations, apply each duration in turn to the base date, rather than summing all the durations and applying that sum of durations to the base date.

  • //Adding Durations to a Base Date
  • date ← 2019-06-01
  • //Definition of durations omitted
  • cat durations >> date
  • //The Easiest way to move a date on is by a number of durations

By using the approach above and number of positive, negative or even empty durations can be applied to a date. The date will just increment by the appropriate amount but take into account the base date it started from and include leap years and the normal variation in the days per month.

DateTime

This is the first built in type that is a sort of aggregate. Whilst this type really holds and instant in time it also has the concept of a timezone. So it relates Date, Time and a timezone together all in one type.

Variables of type DateTime can have the values of:

The format of the DateTime literal is YYYY-MM-DDTHH:MM:SS+/-HH:MM for example 2020-11-18T13:45:21-01:00 is the eighteenth of November 2020 @ thirteen forty-five and twenty-one seconds in a time zone offset by minus one hour from UTC. But also note that 'Z' Zulu or GMT/UTC can also be used as a shortcut for +/-00:00 as follows 2020-11-18T14:45:21Z.

Working with Time and Date as separate concepts is quite natural for most developers, but DateTimewith timezones sometimes take a bit more thought.

For example is 2020-11-18T13:45:21-01:00 less than or greater than 2020-11-18T13:45:21Z? Seeing the "-01:00" may confuse you.

It is greater; if you adjust them both to the same time zone by adding/removing the hour and making the same adjustment to the hour of the day is becomes obvious.

The range of operations on DateTime are very similar to those offered on the Date and Time types; comparison and ternary for example. As with Date; Duration can be used to do calculations.

The following example with DateTime uses timezones - doing mathematics with dates, times and timezone combinations can be a bit confusing. The main thing to think about here is the same instant in time. i.e. if you are in London what time (and date) will it be right now in New York.

#!ek9
defines module introduction

  defines program
    
    ShowDateTimeType()
      
      //3rd of october 2020 @ 12:00 UTC
      dateTime1 <- 2020-10-03T12:00:00Z
      dateTime2 <- 2020-10-04T12:15:00-05:00
      dateTime3 <- 2020-10-04T12:15:00-05:00
      dateTime4 <- DateTime()
      dateTime5 <- DateTime(year: 2020, month: 10, dayOfMonth: 3, hour: 12)      
      dateTime6 <- DateTime(year: 2020, month: 10, dayOfMonth: 4, hour: 12, minute: 15)
                 
      require dateTime1 <> dateTime2
      require dateTime2 == dateTime3

      require dateTime1 < dateTime2
      require dateTime1 <= dateTime2
      require dateTime2 <= dateTime3
      
      require dateTime2 > dateTime1
      require dateTime2 >= dateTime1
      require dateTime3 >= dateTime2
      
      require ~dateTime4?
          
      require dateTime1 == dateTime5         
      
      dateTime4a <- DateTime().today()
      require dateTime4a?
     
      dateTime4.setToday() //Now set to the current dateTime
      
      //Durations
      d1 <- P1Y1M4D //one year, one month and 4 days
      d2 <- P4W1DT2H59M18S //twenty nine days, two hours, fifty nine minutes and eighteen seconds
      
      dateTime4 := dateTime5 + d1
      require dateTime4 == 2021-11-07T12:00:00Z
      
      dateTime4 += d2
      require dateTime4 == 2021-12-06T14:59:18Z
      
      d3 <- dateTime4 - dateTime5
      require d3 == P1Y2M3DT2H59M18S
      
      d4 <- dateTime2 - dateTime3
      require d4 == PT0S
      
      //Think carefully about this (focus on the instant)
      //At ZULU/GMT/UTC what time will it be in the -05:00 time zone?
      //2020-10-04T12:15:00-05:00
      //2020-10-04T12:15:00Z
      d5 <- dateTime3 - dateTime6
      require d5 == PT5H
      
      d6 <- dateTime3.offSetFromUTC()
      require d6 == PT-5H
      
      //processing of durations
      durations as List of Duration := [d1, d2]
      dateTime7 <- cat durations | collect as DateTime
      //Because the epoch is 1970-01-01
      require dateTime7 == 1971-03-06T02:59:18Z
      
      //From a specific date time add on the durations in turn
      dateTime8 <- 0000-01-01T00:00:00Z
      cat durations >> dateTime8
      require dateTime8 == 0001-03-06T02:59:18Z
      
      dayOfMonth <- dateTime8.day()
      monthOfYear <- dateTime8.month()
      year <- dateTime8.year()      
      dayOfTheWeek <- dateTime8.dayOfWeek()      
      hourOfTheDay <- dateTime8.hour()
      minuteOfTheHour <- dateTime8.minute()
      secondOfTheMinute <- dateTime8.second()
      timeZone <- dateTime8.zone()
      offset <- dateTime8.offSetFromUTC()
      
      require dayOfMonth == 6
      require monthOfYear == 3
      require year == 1
      require dayOfTheWeek == 2
      require hourOfTheDay == 2
      require minuteOfTheHour == 59
      require secondOfTheMinute == 18
      require timeZone == "Z"
      require offset == PT0S

//EOF

The date time example is quite long, but as you can see the ideas are the same; in the sense of using normal mathematical operations as with just Date and Time with Duration. The idea is to try and keep the same semantics irrespective of type, this reduces the surface area of method names and ideas. It also reduces the number of classes involved.

So while the EK9 language does have a lot of types, constructs and uses a wide range of operators it does reduce the need for a wide range of method names, helper classes and third party API's.

This is a key point if you decide to adopt EK9 - 'go with the grain and flow'. Don't fight against it, it would be better to adopt a different language like C or C# for example.

Money

This is the first built-in type that models an important modern real world concept. You can make the argument that dates, times and timezones are abstract concepts, but Money really is a modern abstract business concept. Many programming languages do not have a built-in type for money, but as EK9 is aimed at a wide range of uses incorporating Money from the outset provides a more rounded language. There are some particularly tricky things to deal with when it comes to money, these are the number of fractional parts (cents, etc.) and the rounding methodology. It is this aspect that means Money is in general more difficult to work with than a DateTime as there is additional variation in Money.

Variables of type Money can have the values of:

The format of the Money literal is D+(.D+)?#SSS where D is any number, it is possible to have an optional point then the fractional part of the amount; this is then always followed by a '#' and then the three digit currency code.

So for example the Chilean Unidad de Fomento is a bit special as it has four fractional places i.e. 45.9999#CLF and the Iraqi Dinar has three places (45.000#IQD). When displaying these values it is normal to use a Locale. But when coding them use the format above. Currencies like GBP and USD only use two decimal places as do many other currencies, however some like Guinea Franc have none at all.

Please note that the mechanism of specifying and checking these values is code and it is not intended to be viewed in this way by end users. The presentation of Money is subject to the Locale you may for example decide that while a currency does have 2 decimal places because of your business application it is not appropriate to display them (for example transaction charges could be a fixed amount of £50 or $75 - so it is not necessary to display the pennies/cents).

Examples of the use of Money:

#!ek9
defines module introduction

  defines program
  
    ShowMoneyType()
      
      tenPounds <- 10#GBP
      require tenPounds == 10.00#GBP
      
      thirtyDollarsTwentyCents <- 30.2#USD
      require thirtyDollarsTwentyCents == 30.20#USD
            
      //Mathematical operators
      nintyNinePoundsFiftyOnePence <- tenPounds + 89.51#GBP
      require nintyNinePoundsFiftyOnePence == 99.51#GBP
      
      require tenPounds < nintyNinePoundsFiftyOnePence
      require nintyNinePoundsFiftyOnePence > tenPounds
      require tenPounds <> nintyNinePoundsFiftyOnePence
      require tenPounds <= 10.00#GBP
      require tenPounds >= 10.00#GBP
      
      //As they are different currencies - operation like this are meaningless
      //You will have to convert to the same currency to compare or combine them.
      require ~(tenPounds != thirtyDollarsTwentyCents)?
      require ~(tenPounds == thirtyDollarsTwentyCents)?
      //The above requirements are checking if the equality test is a valid test, not the result of the equality test.
      
      //rounding up for money values 49.755 is 49.76
      fourtyNinePoundsSeventySixPence <- nintyNinePoundsFiftyOnePence/2
      require fourtyNinePoundsSeventySixPence == 49.76#GBP
      
      minusFourHundredAndThirtyFivePoundsSixtyPence <- fourtyNinePoundsSeventySixPence * -8.754
      require minusFourHundredAndThirtyFivePoundsSixtyPence == -435.60#GBP
      
      nineHundredAndThirtyFivePoundsSixtyPence <- 500#GBP - minusFourHundredAndThirtyFivePoundsSixtyPence
      require nineHundredAndThirtyFivePoundsSixtyPence == 935.60#GBP
      
      variableAmount <- (nineHundredAndThirtyFivePoundsSixtyPence * 3) / 18
      require variableAmount == 155.93#GBP
      
      multiplier <- nineHundredAndThirtyFivePoundsSixtyPence / variableAmount
      //A Money / Money ratio is a Float - 6.0001282627 - so bound it, never == it
      require multiplier > 6.0001 and multiplier < 6.0002
      
      variableAmount += 4.07#GBP
      require variableAmount == 160.00#GBP
      
      variableAmount *= 0.666
      require variableAmount == 106.56#GBP
      
      variableAmount -= 0.56#GBP
      require variableAmount == 106.00#GBP
      
      variableAmount /= 4
      require variableAmount == 26.50#GBP
      
      variableAmount := -variableAmount
      require variableAmount == -26.50#GBP
      
      variableAmount := abs variableAmount
      require variableAmount == 26.50#GBP
      
      variableAmount := sqrt variableAmount
      require variableAmount == 5.15#GBP
      
      variableAmount := variableAmount ^ 6
      require variableAmount == 18657.07#GBP
      
      //Division by zero result is not set
      variableAmount := variableAmount/0
      require ~variableAmount?
      
      //Use of two difference currencies together - result is not set
      variableAmount := tenPounds + thirtyDollarsTwentyCents
      require ~variableAmount?
      
      //Example if pipeline to add amounts together
      amounts as List of Money := [tenPounds, nintyNinePoundsFiftyOnePence, fourtyNinePoundsSeventySixPence]
      total <- cat amounts | collect as Money
      require total == 159.27#GBP
      
      //Provide the exchange rate and convert to USD
      totalInUSD <- total.convert(1.32845, "USD")
      require totalInUSD == 211.58#USD

//EOF

Hopefully you can see from the above example, working with Money in EK9 is easy and natural. For example divide one amount of money by another, and you get a Float as a result; i.e. how many times does one amount go into another. But divide an amount of money by an Integer or Float and you get Money as a result. If you want to convert an amount of Money into another currency - then just provide the exchange rate you wish to use and the new currency code.

But importantly 'Money' is type safe, you cannot add USD values to GBP values for example; the result would be un set!

  • //Pipeline could have been written like this
  • total ← cat tenPounds, nintyNinePoundsFiftyOnePence, fourtyNinePoundsSeventySixPence | collect as Money
  •  
  • //Rather than as shown in the example like this
  • amounts as List of Money := [tenPounds, nintyNinePoundsFiftyOnePence, fourtyNinePoundsSeventySixPence]
  • total ← cat amounts | collect as Money

Money like quite a few other types has an iterator() method on it, this allows it to iterate its own value (if it is set). This means that is can be used as a source in a pipeline. The cat pipeline command accepts multiple parameters (comma separated) and so can trigger iteration over all the parameters. String and Bits provide iteration over Characters and Booleans respectively and so cannot do self iteration, to accomplish that with those typesuse a List or for single values an Optional.

Locale

EK9 has a mechanism for enabling the developer to output a range of types in localised form. The general form of a locale is {language}_{country}, for example en_GB or de_DE. EK9 supports localisation of the following types out of the box. In some cases (Time, Date, DateTime and Money) there are short, medium, long and full formats available for:

For additional localisation of 'Text' suitable for output to end users see Text Properties. The Locale together with the Text construct provide a very good basis to produce textual output suitable for users in a range of different languages.

An example of the formats below is given with a number of sample locales and formats so that difference in output for the same data can be clearly seen.

#!ek9
defines module introduction

  defines program
  
    ShowLocale()
      
      enGB <- Locale("en_GB")
      //Note you can use underscore or dash as the separator.
      enUS <- Locale("en-US")
      deutsch <- Locale("de_DE")
      skSK <- Locale("sk", "SK")
      
      i1 <- -92208
      i2 <- 675807
         
      presentation <- enGB.format(i1)
      require presentation == "-92,208"
      presentation := enGB.format(i2)
      require presentation == "675,807"
      
      presentation := deutsch.format(i1)
      require presentation == "-92.208"
      presentation := deutsch.format(i2)
      require presentation == "675.807"
      
      presentation := skSK.format(i1)
      require presentation == "-92 208"
      presentation := skSK.format(i2)
      require presentation == "675 807"
      
      f1 <- -4.9E-22
      f2 <- -1.797693134862395E12
      
      presentation := enGB.format(f1)
      require presentation == "-0.00000000000000000000049"
      presentation := enGB.format(f2)
      require presentation == "-1,797,693,134,862.395"
      
      //With control over number of decimal places displayed
      presentation := enGB.format(f1, 22)
      require presentation == "-0.0000000000000000000005"
      presentation := enGB.format(f2, 1)
      require presentation == "-1,797,693,134,862.4"
      
      presentation := deutsch.format(f1)
      require presentation == "-0,00000000000000000000049"      
      presentation := deutsch.format(f2)
      require presentation == "-1.797.693.134.862,395"
      
      presentation := skSK.format(f1)
      require presentation == "-0,00000000000000000000049"      
      presentation := skSK.format(f2)
      require presentation == "-1 797 693 134 862,395"
      
      time1 <- 12:00:01
      
      presentation := enGB.shortFormat(time1)
      require presentation == "12:00"
      
      presentation := deutsch.mediumFormat(time1)
      require presentation == "12:00:01"
      
      date1 <- 2020-10-03
      
      presentation := enGB.shortFormat(date1)
      require presentation == "03/10/2020"
      
      presentation := enUS.mediumFormat(date1)
      require presentation == "Oct 3, 2020"
      
      presentation := skSK.longFormat(date1)
      require presentation == "3. októbra 2020"
      
      presentation := deutsch.fullFormat(date1)
      require presentation == "Samstag, 3. Oktober 2020"
      
      dateTime1 <- 2020-10-03T12:00:00Z
      
      presentation := enUS.shortFormat(dateTime1)
      require presentation == "10/3/20, 12:00 PM"
      
      presentation := enGB.mediumFormat(dateTime1)
      require presentation == "3 Oct 2020, 12:00:00"
      
      presentation := deutsch.longFormat(dateTime1)
      require presentation == "3. Oktober 2020, 12:00:00 Z"
      
      presentation := skSK.fullFormat(dateTime1)
      require presentation == "sobota 3. októbra 2020, 12:00:00 Z"
      
      thousandsOfChileanCurrency <- 6798.9288#CLF
      tenPounds <- 10#GBP
      thirtyDollarsEightyNineCents <- 30.89#USD
      
      presentation := enGB.format(tenPounds)
      require presentation == "£10.00"
      presentation := enGB.format(thirtyDollarsEightyNineCents)
      require presentation == "US$30.89"
      presentation := enGB.format(thousandsOfChileanCurrency)
      require presentation == "CLF6,798.9288"
      
      presentation := deutsch.format(tenPounds)
      require presentation == "10,00 £"
      presentation := deutsch.format(thirtyDollarsEightyNineCents)
      require presentation == "30,89 $"
      presentation := deutsch.format(thousandsOfChileanCurrency)
      require presentation == "6.798,9288 CLF"
      
      //Without the currency symbol
      presentation := deutsch.longFormat(tenPounds)
      require presentation == "10,00"
      
      presentation := deutsch.longFormat(thirtyDollarsEightyNineCents)
      require presentation == "30,89"
      
      presentation := deutsch.longFormat(thousandsOfChileanCurrency)
      require presentation == "6.798,9288"
      
      presentation := skSK.format(tenPounds)
      require presentation == "10,00 GBP"
      presentation := skSK.format(thirtyDollarsEightyNineCents)
      require presentation == "30,89 USD"
      presentation := skSK.format(thousandsOfChileanCurrency)
      require presentation == "6 798,9288 CLF"
      
      presentation := enGB.mediumFormat(tenPounds)
      require presentation == "£10"
      presentation := enGB.mediumFormat(thirtyDollarsEightyNineCents)
      require presentation == "US$31"
      presentation := enGB.format(thousandsOfChileanCurrency, true, false)
      require presentation == "CLF6,799"
      
      presentation := enGB.shortFormat(tenPounds)
      require presentation == "10"
      presentation := enGB.shortFormat(thirtyDollarsEightyNineCents)
      require presentation == "31"
      presentation := enGB.format(thousandsOfChileanCurrency, false, false)
      require presentation == "6,799"

//EOF

The example above; albeit quite long should give you a good idea of the range and simplicity of how to localise the output of a range of types. It also shows how the format can be altered for specific types being short, medium, long of full.

Colour

The Colour type has a number of features that add quite a bit of value for any developer needing to manipulate colour. These are built right into the type itself as shown in the following example. It is possible to work with the Bits of the Colour alter them and create new colours. But also add/subtract and change transparency/saturation and lightness.

#!ek9
defines module introduction

  defines program
    
    ShowSimpleColour()
      
      //Colour can hold RGB or ARGB (i.e with alpha channel)
      //Here shown with alpha channel fully opaque
      testColour <- #FF186276
      testColourAsBits <- testColour.bits()
      require testColourAsBits == 0b11111111000110000110001001110110      
      
      //Bitwise manipulation and colour recreation from bits
      modifiedBits <- testColourAsBits and 0b10110111000100000110001010110110
      modifiedColour <- Colour(modifiedBits)
      require modifiedColour == #B7106236
      require modifiedBits == 0b10110111000100000110001000110110
      
      //Note it is also possible to access the HSL values of the colour
      require modifiedColour.hue() == 148
      require modifiedColour.saturation() == 71.9298245614035
      require modifiedColour.lightness() == 22.35294117647059
      
      //Alter the alpha channel to control how transparent/opaque
      moreOpaqueColour <- modifiedColour.withOpaque(80)
      require moreOpaqueColour == #CC106236
      
      //It can be made lighter - much lighter in this case
      lighterColour <- moreOpaqueColour.withLightness(80)
      require lighterColour == #CCA7F1CA
      
      //Then we can saturate it more
      moreSaturatedColour <- lighterColour.withSaturation(90)
      require moreSaturatedColour == #CC9EFAC9
            
      //Remove the Red from this colour
      lessRedColour <- moreSaturatedColour - #3A0000
      require lessRedColour == #CC04FAC9
      
      //Add in some blue
      moreBlueColour <- lessRedColour + #00001D
      require moreBlueColour == #CC04FADD
      
      //Available in different formats as Strings
      require moreBlueColour.RGB() == $#04FADD
      require moreBlueColour.RGBA() == $#04FADDCC
      require moreBlueColour.ARGB() == $#CC04FADD

//EOF

  • //The actual colours from example above, shown with vertical bar background
  • //To show opaque and transparency
  • #186276FF - testColour
  •  
  • #106236B7 - modifiedColour
  •  
  • #106236CC - moreOpaqueColour
  •  
  • #A7F1CACC - lighterColour
  •  
  • #9EFAC9CC - moreSaturatedColour
  •  
  • #04FAC9CC - lessRedColour
  •  
  • #04FADDCC - moreBlueColour
  •  

As you can see from the example above EK9 makes working with colours easy and in general developers always need to be able to alter shades and hues of colours programmatically; including lightening and darkening. This is straight forward with the Colour type, get the current Lightness and just multiply it up or down by a factor and make the alteration to get the new Colour.

It's quite easy to imagine a pipeline process taking a stream of Colours and using a range of functions to manipulate those Colours to produce a new set that are darker, lighter or more muted.

Dimension

The Dimension type is a simple aggregate type that enforces the concept of some sort of size in combination with a unit of measurement. The main focus of this type is to make explicit the use and combinations of mathematical calculations. This could have been done by end user developers with a class hierarchy, but EK9 provides a simpler mechanism in the language itself.

Inspired by the CSS idea of dimensions this type can employ any of the following units for dimensions.

Absolute Lengths (m, km and mile added - not suitable in CSS context)

Relative Lengths (really only applicable for screen rendering)

Example of using the Dimension type.

#!ek9
defines module introduction

  defines program
  
    ShowDimensionType()
 
      dimension1 <- 1cm
      dimension2 <- 10px
      dimension3 <- 4.5em
      dimension4 <- 1.5em

      calc1 <- dimension1 * 2
      calc2 <- dimension2 / 5
      calc3 <- dimension3 + 0.6
      calc4 <- dimension3 + dimension4

      require calc1 == 2cm
      require calc2 == 2px
      require calc3 == 5.1em
      require calc4 == 6em

      //But if we divide a dimension by another dimension we get just a number
      calc5 <- dimension3 / dimension4
      require calc5 == 3
      
      //Normal comparison operators as you would expect
      require calc3 < calc4

      //But also the coalescing operators
      lesser <- calc3 <? calc4
      require lesser == calc3      

      calc3 += 0.9
      require calc3 == calc4

      calc3++
      require calc3 > calc4
      
      //Checking the calculation is not actually valid
      //different types of dimension calc1 is cm and calc2 is px
      require not (calc1 <> calc2)?

      //This is a test that the comparison was valid, not the result of the comparison
      require (calc3 < calc4)?

      calc6 <- sqrt calc4
      require calc6 == 2.449489742783178em

      //Conversion is table-driven: name the target UNIT (1 mile is 1609.344 m by definition)
      eightMiles <- 8mile
      inKM <- eightMiles.convert("km")
      require inKM == 12.874752km

      require length inKM == 12.874752
      returnJourney <- -eightMiles
      require abs returnJourney == eightMiles

      squared <- inKM ^ 2
      require squared == 165.75923906150402km
      
      result <- eightMiles + inKM
      require not result?

//EOF

The Dimension type is designed to help with strong typing and to ensure that only compatible types are used together or are converted through an explicit method call. See Mars Climate Orbiter as to why this is 'quite important'.

Path

Like JSON (next) the concept of an object/array path is built into the EK9 language. This is because that concept of an object graph with a mix of Maps and Arrays is a very common one.

Here is an example of how to define a path to be used with an object graph. The literal definition of a path starts with $?, it is then followed by combinations of property or array addressing.

#!ek9
defines module com.customer.just

  defines program

    ASimpleSetOfPaths()
      path1 <- $?.some.path.inc[0].array
      path2 <- $?.another[2][1].multi-dimensional.array.access_value

      simplePath <- $?.aKey
      firstArrayElement <- $?[0]
      propertyFromFirstElement <- $?[0].a-field

      //Not to be confused with string conversion
      anInt <- 90
      stringRepresentation <- $anInt

      //Or json conversion of a record.
      me <- CustomerDetail(firstName: "Steve", lastName: "Limb")
      jsonOfMe <- $$me

  defines record

    CustomerDetail
      firstName <- String()
      lastName <- String()

//EOF
    

JSON

The JSON type is built right into the EK9 language. Is has been designed to support all the normal types that JSON supports, those being:

Other EK9 types; such as Character, Date, Time, DateTime, Duration, Millisecond, Money, Colour, Dimension and Bits are all converted to the JSON String type.

Importantly developers can convert their own data types of records and classes to JSON using the $$ JSON operator. But EK9 can provide a default $$ operator for records without the developer needing to write any code in most cases.

It is also possible convert your records and classes back from JSON by providing:

A stream can also be poured into a JSON (cat items > target, or collect as JSON). Pouring into JSON follows the target's shape: an unset or array target collects, an object target merges each streamed object in; anything else - a non-object into an object, or anything into a single value - is invalid, leaves the target unset, and later items do not revive it. Each item is one element of an array (an array poured in is not spread); objects are merged by the merge rule, so the first value recorded wins; and a JSON null item is unset, so it is dropped. See pouring into JSON.

JSON is a value. Merging, pouring, combining or converting into JSON copies what it takes: +, :~:, :^:, |, :=: and $$ (of a JSON, or of anything holding one) all copy, and so do iterating a JSON and read(). Changing an operand afterwards never reaches the result, and a JSON can never come to contain itself. The one exception is get(): it is a view into its container, so merging into what get() returns changes the container too.

CSV

CSV is a first-class construct for reading and writing RFC 4180 Comma Separated Values - modelled on JSON. It correctly handles the parts that a naive split(',') gets wrong: quoted fields containing the separator, doubled ("") quotes, and newlines embedded inside quoted fields. The separator is configurable, so tab- and semicolon-separated data work too.

parse returns a Result of (CSV, String): a malformed document carries its reason, and the compiler requires you to check isOk() before accessing the value - you cannot accidentally use a half-parsed document. Rows are read total-style with getOrDefault(index, default) (following List), and iterator() exposes the rows to stream pipelines so filtering, sorting and so on come for free. Stacking CSVs vertically ('+', '+=') is guarded by header alignment - stacking two CSVs whose headers differ yields un set rather than welding mismatched columns together. Merge (':~:') does not stack: it follows the merge rule - a wholly un set CSV takes the source, otherwise only what the target lacks is filled (a missing header, rows past its last, cells that are missing or empty) and nothing it holds is overwritten.

#!ek9
defines module introduction

  defines program

    CsvExample()
      stdout <- Stdout()

      result <- CSV().parse("name,weight\nHydrogen,1.008\nHelium,4.0026")
      if result.isOk()
        table <- result.ok()

        //Read a data row by column name; convert the cell with the target type's constructor
        helium <- table.getOrDefault(1, CSVRow())
        weight <- BigDecimal(helium.getOrDefault("weight", "0"))

        stdout.println(`Helium weighs ${weight}`)
      else
        stdout.println("Bad CSV: " + result.error())

CSVRow

A CSVRow is one row of a CSV - the same type is used for the header and for every data row. It is more than a List of String: a data row is header-aware, so a cell can be read by column name as well as by position. Access is total, following List - getOrDefault(key, default) never goes out of range, it just yields the supplied default:

Cells are text. To get a real type out of a cell, hand it to that type's constructor - which parses it and returns un set if the cell is not that type, e.g. Integer(row.getOrDefault(1, "0")) or BigDecimal(row.getOrDefault("weight", "0")). This is uniform across every type (including your own) and is safer than dedicated typed getters, because a bad cell simply becomes un set rather than throwing.

Regular Expression

The RegEx type is the last of the built-in non-collection types. It is built into the language for the simple reason that regular expressions are so widely used and are so expressive (albeit quite complex).

Standard shorthand sequences

Quantifiers

Boundary Matchers

Here is an example of the use of RegEx with match, split and group. The operations can be RegEx or String focussed.

#!ek9
defines module introduction

  defines program

    ShowRegExType()

      sixEx <- /[a-zA-Z0-9]{6}/

      require sixEx matches "arun32"
      require sixEx not matches "kkvarun32"
      require sixEx matches "JA2Uk2"
      require sixEx not matches "arun$2"

      regEx <- /[S|s]te(?:ven?|phen)/

      Steve <- "Steve"
      steve <- "steve"

      Stephen <- "Stephen"
      stephen <- "stephen"

      Steven <- "Steven"
      steven <- "steven"

      Stephene <- "Stephene"
      stephene <- "stephene"

      require Steve matches regEx and steve matches regEx
      require Stephen matches regEx and stephen matches regEx
      require Steven matches regEx and steven matches regEx

      require Stephene not matches regEx
      require stephene not matches regEx

      //String and Regex can work either way around
      require regEx matches Stephen

      //Examples using escape of '\'

      stockExample <- "This order was placed for QT3000! OK?"
      groupEx <- /(.*?)(\d+)(.*)/
      require groupEx matches stockExample

      fractionExample <- "3/4"
      fractionEx <- /.*\/.*/
      require fractionEx matches fractionExample

      slashEx1 <- /^Some\\Thing$/
      require slashEx1 matches "Some\Thing"

      slashEx2 <- /^Some\/Thing$/
      require slashEx2 matches "Some/Thing"
      
      colonDelimited <- "one:two:three:four:five"      
      colonRegEx <- /:/

      //You can work with RegEx first or String first
      whenSplit <- colonDelimited.split(colonRegEx)
      alsoSplit <- colonRegEx.split(colonDelimited)

      //$ on a List renders the collection with its String elements quoted - it is not a join
      require length whenSplit == 5
      require whenSplit.getOrDefault(0, "?") == "one"
      require whenSplit.getOrDefault(4, "?") == "five"
      require whenSplit == alsoSplit

      //Check on finding groups - the capture groups of the first match, in pattern order
      extractionCheck <- "This is a sample Text 1234 with numbers in between."
      extractNumbersEx <- /(.*?)(\d+)(.*)/
      matchedGroups <- extractionCheck.group(extractNumbersEx)

      //Three groups: the text before, the digits, the text after
      require length matchedGroups == 3
      
      justTheNumber <- cat matchedGroups | skip 1 | head 1 | collect as String
      //Convert to an Integer - if possible
      asAnInteger <- Integer(justTheNumber)
      require $justTheNumber == "1234"
      require asAnInteger == 1234
      
      //Check just the last two (as an example)
      lastTwo <- cat matchedGroups | tail 2 | collect as List of String
      require length lastTwo == 2
      
      nonExtractionCheck <- "Another sample but with no numbers in it."
      matchedGroups := nonExtractionCheck.group(extractNumbersEx)
      //Expecting this to be unset
      require not matchedGroups?
      
      //Safe way of accessing some result that might not be set (in this case it is not set)
      justNotTheNumber <- cat matchedGroups | skip 1 | head 1 | collect as String      
      require not justNotTheNumber?

      //Named groups: ask for one by name - a String, unset when absent
      namedDate <- /(?\d{4})-(?\d{2})-(?\d{2})/
      dueLine <- "Due 2024-06-15 at noon"
      require dueLine.group(namedDate, "year") == "2024"
      require not dueLine.group(namedDate, "hour")?

      //Replace every match, referring to the groups as $n or ${name} - a name as \${name} in backticks
      require dueLine.replace(namedDate, "$3/$2/$1") == "Due 15/06/2024 at noon"
      require dueLine.replace(namedDate, `\${day}.\${month}.\${year}`) == "Due 15.06.2024 at noon"
      stdout.println(`Replaced [${dueLine.replace(namedDate, "")}]`)
//EOF

Another reasonably long example (only touches on how powerful regular expression are though), but hopefully you can see how the direct use of regular expressions in the code reduces the number of escape sequences required.

When using the RegEx literal syntax (e.g. /\|/), the pattern is passed directly to the regular expression engine. This is the recommended approach. However, if a RegEx is constructed from a String (e.g. RegEx("\|")), String escaping rules apply first. A common mistake is to double-escape: RegEx("\\|") produces the two-character pattern \\| (literal backslash followed by pipe) rather than the intended single-character escape \|. Use regex literals where possible to avoid this class of error.

Summary

EK9 provides a wide range of types built right into the language. By providing types like Time, Date, DateTime, Money, Dimension and Colour EK9 enables strong typing and provides consistent and simple semantics.

Next Steps

Many of the examples above have used some form of collection type. These are discussed in more detail in the collection types section itself.