Culebra Language Specification
September 25, 2026 · View on GitHub
This document defines the syntax and runtime semantics of the Culebra
programming language. It is normative: both engines — the bytecode VM's
executor and its LLVM lowering (--jit and AOT) — are required to
implement what this document says, and a divergence between them or from
this text is a bug.
For an introductory tour with runnable examples, see
handbook.md. For API reference of the standard library see
stdlib.md.
Table of contents
- Overview and philosophy
- Lexical structure
- Grammar
- Types
- Values and identity
- Variables and scope
- Expressions
- Strings and interpolation
- Arrays
- Objects
- Functions and closures
- Control flow (
if,cond,while,for) - Pattern matching (
match) - Optional type annotations
- Error handling
- Algebraic effects
- Memory model (RC + cycle collector)
- Built-in type methods (incl. iterator protocol)
- Core built-in functions
- Multimethods
- Decorators
- Command-line interface
- Known limitations
- Modules
- Appendix: VM ↔ JIT divergence
- Appendix: conformance test mapping
1. Overview and philosophy
Culebra is a dynamically-typed scripting language. Its priorities are:
- Orthogonal features. New capability arrives as a feature that composes with the rest, not as a special case in the grammar.
- One front end, two engines. A bytecode VM and an LLVM ORC JIT are two consumers of the same parser, AST, and bytecode compiler, so features must be implementable in both.
- Predictable memory. Reference counting in both engines; a cycle collector reclaims cyclic garbage.
The language is dynamic at its core. Type annotations are optional and enforce their invariants at runtime boundaries only; they do not turn Culebra into a statically-typed language.
2. Lexical structure
Source encoding
Source text is UTF-8. Identifiers use ASCII only.
Line terminators and whitespace
Line terminator: \n. Whitespace: space, tab, and the line terminator.
A file with CRLF (\r\n) endings is accepted: the source is normalized to LF
before it is parsed, so a checkout with CRLF endings behaves exactly like one
with LF — including inside triple-quoted and raw string literals, where a
stray \r would otherwise be part of the value. culebra fmt writes LF.
A lone \r — one not followed by \n — is a syntax error rather than a third
kind of line ending. It is either an old-Mac line ending or a stray byte, and
which one changes what the program means.
Comments
- Line comments:
# ...or// ...until end of line or end of file. - Block comments:
/* ... */, not nestable.
Identifiers
IdentInitChar <- [a-zA-Z_]
IdentChar <- [a-zA-Z0-9_]
IDENTIFIER <- IdentInitChar IdentChar*
Identifiers are case-sensitive.
Keywords
Reserved words that cannot name a variable:
nil true false mut debugger return while for in if unless else
fn match break continue throw try catch defer
class trait enum import export yield
Assigning to one is a SyntaxError (let class = 1). The reservation
covers the assignment target only: a parameter, a for variable, an
object key, and a property name are all positions where no keyword can
start, so they accept any of these words.
Four of them — return, throw, break, continue — are the
exception when read rather than bound. They are expressions (§12), so
an occurrence of one in expression position is the control-transfer
form, never a variable reference. for break in xs { f(break) } binds
break and then leaves the loop instead of passing it. Object keys and
property names are unaffected ({break: 1} and o.break still work),
since neither is an expression position.
The grammar's remaining keywords are contextual — each is recognized only where its own construct can begin, and is an ordinary identifier everywhere else:
cond nobreak from with effect perform handle static get by
static and get are markers inside a class body and by is
contextual inside a range literal (0..10 by 2). The other seven
introduce constructs with a distinctive shape — cond { … } needs at
least one test => body arm, nobreak follows a loop body, from
follows an import — so let with = 1 and if handle { … } both name
a variable. A loop label (outer: for … { break outer }, §12) is
contextual in the same way but reserves no word at all: it is recognized
only directly before a while / for and directly after a break /
continue, so the name it uses stays free for a variable. Type annotation names (Nil, Bool, Long, Float,
String, Array, Object, Function, Any) are not reserved
either; they are contextual and only recognized after : or ->.
The parser also recognizes let as an optional prefix in assignments.
Literals
- Integer: a
NUMBERis a 64-bit signed integer (Long), written in decimal or with a radix prefix — hex0xFF, octal0o755, binary0b1010(prefix letter is case-insensitive; digits match the base). - Float:
FLOAT <- [0-9]+ '.' [0-9]+ ([eE] [-+]? [0-9]+)? / [0-9]+ [eE] [-+]? [0-9]+. A literal with either a decimal point (followed by digits, not a bare trailing dot) or ane/Eexponent is aFloat(IEEE 754 binary64). Examples:1.0,0.5,-2.5(unary minus applied to2.5),1e-5,1.5e3. - String:
'...'is a raw string — every character between the quotes is taken literally, including backslashes. There is no escape syntax, no interpolation, and apostrophes inside are not expressible (use a backtick string`...`for raw content containing'). - Interpolated string:
"...{expr}...".{expr}embeds any expression. Recognized escape sequences:\n\r\t\\\"\{(use\{to embed a literal{without starting an interpolation; raw}is fine since it's only special as the interpolation terminator),\xHH(exactly two hex digits → one raw byte,0x00–0xFF),\uXXXX(exactly four hex digits → a BMP Unicode scalar value, UTF-8 encoded), and\UXXXXXXXX(exactly eight hex digits → a Unicode scalar value, UTF-8 encoded; the only form that reaches beyond the BMP). Both reject values aboveU+10FFFFand the surrogate range. An unknown\Xis preserved as the two literal characters\andX. - Boolean:
true,false. - Nil:
nil.
Trailing comma. A single trailing comma is allowed after the last
element of an array [1, 2,], object {a: 1,}, set {1, 2,}, tuple
(1, 2,), argument list f(1, 2,), parameter list fn g(a, b,),
lambda params |a, b,|, and destructuring pattern let [a, b,] = …
(cleaner diffs, and reordering without touching the neighbours).
A leading or doubled comma is still a syntax error. Note (x,) and
{x,} are not merely cosmetic: the comma marks a one-element tuple
and a one-element set respectively.
Operators and punctuation
== != <= < >= > # comparison
+ - * / % ** # arithmetic (`**` exponentiation); `+` also concatenates strings and Arrays
@ # matmul (user-defined via `__matmul__`)
| & ^ << >> ~ # bitwise or / and / xor / shifts / complement (Long only)
! # logical not
&& || # logical and/or (short-circuit)
?? # nil coalesce (lower precedence than ||)
?. ?[ ] # optional chaining / optional index
? : # ternary conditional `c ? a : b` (right-assoc)
!! # non-null assertion (postfix)
.. ..= # range literals (exclusive / inclusive)
= # assignment
+= -= *= /= %= **= @= # compound assignment
&= |= ^= <<= >>= # compound assignment (bitwise)
??= # nil-coalescing assignment
=> # match arm separator
-> # return type
... # rest pattern
| # or-pattern separator
: # type annotation
( ) [ ] { } # grouping / container delimiters
, ; # separators
. # member access
Statement separators
Statements are separated by ; or by a line terminator. Empty lines
between statements are fine.
3. Grammar
The full PEG grammar is one source of truth, which the parser loads verbatim. Its top-level rules:
PROGRAM <- STATEMENTS
STATEMENTS <- (STATEMENT ((';' / newline) STATEMENT?)*)?
STATEMENT <- DEBUGGER / RETURN / THROW / YIELD_FROM / YIELD / BREAK
/ CONTINUE / DEFER / IMPORT_STMT / EXPORT_STMT
/ EFFECT_FN_DECL / MULTIFN_DECL / ENUM_DECL / CLASS_DECL
/ TRAIT_DECL / LEXICAL_SCOPE / EXPRESSION
EXPRESSION <- DESTRUCTURE_ASSIGN / PLACE_ASSIGN / ASSIGNMENT / TRY
/ CONDITIONAL
ASSIGNMENT <- LET MUTABLE PRIMARY (ARGUMENTS / INDEX / DOT)*
TYPE_ANNOTATION? ASSIGN_OP EXPRESSION
CALL <- PRIMARY (ARGUMENTS / INDEX / SAFE_DOT / SAFE_INDEX
/ DOT / NONNULL)*
PRIMARY <- WHILE / FOR / IF / MATCH / COND / HANDLE / PERFORM
/ RETURN / THROW / BREAK / CONTINUE
/ FUNCTION / LAMBDA / OBJECT / SET / ARRAY / NIL
/ BOOLEAN / FLOAT / NUMBER / REGEX_LIT / IDENTIFIER
/ TRIPLE_STRING / STRING / RAW_STRING
/ INTERPOLATED_STRING / TUPLE / '(' EXPRESSION ')'
Statements are the constructs that may not appear in expression
position: the declaration forms, yield / yield from, defer, and the
module forms. Everything else — including if, while, for, match,
cond, handle and perform — is a PRIMARY, which is what makes
them usable as values (§7).
The four diverging forms — return, throw, break, continue —
are PRIMARY too, so they may sit wherever an expression may. They never
produce a value; reaching one transfers control instead, so there is
nothing for the surrounding expression to consume (§12).
Operator precedence
From lowest to highest. Below item 1 sit assignment (=, lowest), then the
ternary c ? a : b (right-associative — a ? b : c ? d : e is
a ? b : (c ? d : e)), then ??; so a ?? b ? c : d is (a ?? b) ? c : d
and x = c ? a : b is x = (c ? a : b).
||&&==,!=,<,<=,>,>=- Binary
+,- - Binary
*,/,%,@(matmul) - Unary
+,-,!,~(right-associative) **(right-associative). Binds tighter than a unary prefix —-2**2 == -4— but the RHS of**itself accepts one, so2 ** -1 == 0.5and2 ** -3 ** 2 == 0.00195…both parse.- Call / index / dot (
x(...),x[i],x.k) — left-associative
A unary prefix binds tighter than *, so -a * b is (-a) * b. For
numbers the two readings agree, but a type defining __neg__ and
__mul__ sees which one runs first:
class N {
new(v) {
self.v = v
}
__neg__() {
println('neg')
N(-self.v)
}
__mul__(o) {
println('mul')
N(self.v * o.v)
}
}
_ = -N(2) * N(3)
# => |
# neg
# mul
= is right-associative but appears only in ASSIGNMENT, not as a
general expression operator.
Expression-oriented syntax
Most constructs are expressions that yield a value:
if { a } else { b }yields the taken branch's value.match x { ... }yields the matching arm's body value (ornil).cond { test => a, _ => b }yields the body of the first truthytest, ornilif none match (seecond).whilealways yieldsnil.- A block
{ ... }in statement position is aLEXICAL_SCOPEand yieldsnil; in expression position{starts anOBJECTliteral. - The last expression evaluated in a sequence is the sequence's value.
4. Types
Culebra has exactly twelve types — the names type_of can return:
| Type | Description |
|---|---|
Nil | Single value nil |
Bool | true or false |
Long | 64-bit signed integer |
Float | IEEE 754 binary64 (double precision) |
String | Immutable heap-allocated byte string |
StringView | Immutable zero-copy borrow of another string's bytes, produced by slice / split / view (§18.1) |
Array | Mutable ordered collection of values |
Object | Mutable map of hashable keys to values (String keys go through a fast shape path) |
Function | Closure (function pointer + captures) |
Tensor | N-dimensional numeric tensor with a lazy graph evaluated on CPU or GPU (see stdlib §8) |
Tuple | Immutable, hashable sequence ((1, 2)) — usable as an Object/Set key |
Set | Insertion-ordered collection of unique hashable values ({1, 2, 3}) |
Classes, enum variants, modules, iterators, and stdlib namespaces are
all Objects. type_of reports 'Object' for a module, an iterator and
a namespace, the class's own name for a class instance, and the variant's
name for an enum variant (§19).
Arithmetic between Long and Float promotes the Long operand to
Float automatically — see §7. Comparison mixes the two by value
without that promotion, which past would land two different
Longs on one Float. Outside that numeric pair there is no
implicit conversion between types, and
arithmetic, comparison, or boolean operators given the wrong kind of
value raise type error (see §15).
Any is only valid in type annotations and matches any value.
Both engines: Float is fully supported on the VM and the LLVM
ORC JIT. Long-only code paths keep their existing inline integer
codegen under --jit; Float values take a narrow Long↔Float
promotion slow path with bit-for-bit identical semantics.
5. Values and identity
Nil,Bool,Long, andFloatare value types: they are compared and copied by value.==compares by the underlying data.String,Array,Object, andFunctionare reference types: variables hold a reference to a heap-allocated object. Assignment and passing copy the reference, not the object. Mutation through one variable is visible through another that refers to the same object.
== on reference types:
String: compares by contents.Array,Object,Tuple,Set: compare by value (structural, recursing through elements). Each pair of elements is compared by==itself, so an element's own__eq__/eq(Operator overloading) decides for it at any depth —[a] == [b]whenevera == b. What is matched as a key — aSet's members, anObject's non-Stringkeys — is matched as a lookup matches it instead, byhashandeq: that step reads no__eq__and crosses no types, so{1, 7} == {1.0, 7}is false where[1, 7] == [1.0, 7]is true. OnlyFunctionandTensorcompare by reference identity. (Value equality does not make arrays hashable — they still can't be Object/Set keys.)
Cross-type numeric equality: Long and Float compare by numeric
value. 1 == 1.0, 0 == 0.0. NaN compares unequal to everything
including itself.
Ordering (<, <=, >, >=) is defined for:
Long,Float, andBool— numeric ordering (booleans orderfalse<true;LongandFloatmix by value, exactly, so ordering and==agree about every pair).String— lexicographic byte ordering.Nil—nilcompares equal toniland always returnsfalsefor ordering comparisons.
Ordering values of different types (outside the numeric pair) raises
type error.
6. Variables and scope
Declaration and assignment
x = 10 # bare assignment
let y = 20 # let binding (immutable)
mut z = 30 # mut binding (mutable)
let mut w = 40 # equivalent to `mut w = 40`
let {name, age} = person # object destructure (shorthand)
let {name: nm, age: a} = person # rename / value-match per entry
let {user: {name}} = req # nested destructure
let [a, b, ...rest] = xs # array destructure (rest allowed)
let (x, y) = pair # tuple destructure
let mut {x, y} = point # destructure with mutable bindings
Assignment with a simple identifier LHS is handled as follows:
- If
letormutis present, a new binding is created in the current scope. It is a compile-time error to shadow a variable captured from an enclosing function (see "Shadow prohibition" below). - Without
let/mut, barex = vsearches the scope chain:- If
xexists in any visible scope (including outer closures and the global scope), that binding is reassigned. This may update captured variables in outer scopes (see §11). This is the mechanism by which closure-based objects mutate their state. - Otherwise a new (immutable) binding is created in the innermost
scope: the block, loop body or match arm the assignment is written in.
An
ifarm shares its enclosing scope, so a binding made there stays visible after theif.
- If
Compound assignment
Twelve compound-assignment operators rewrite LHS OP= RHS as
LHS = LHS OP RHS, with the side-effect that the LHS is evaluated
exactly once. They cannot be combined with let or mut (compound
assignment only updates an existing binding):
x += 1 # x = x + 1
x -= y # x = x - y
x *= 2
x /= 3
x %= 5
x **= 2
x @= M # matrix multiply (via Tensor / __matmul__)
x &= MASK # the bitwise five are Long-only, like their operators
x |= FLAG
x ^= 1
x <<= 2
x >>= 2
The LHS may be an identifier, an array element (a[i]), or an object
property (o.x). Index expressions and property names are evaluated
exactly once. Compound assignment to an undefined name is an error.
a[next_idx()] += 1 # next_idx() is called once
o.count += delta
bogus += 1 # error: compound assignment on undefined name
Tensor interaction (in-place writes). When the LHS is a Tensor
that owns its storage and the result fits the LHS's shape, +=, -=,
*=, /=, and **= write the result back into the existing buffer
instead of allocating a fresh Tensor. (%= has no Tensor semantics
and @= changes the output shape, so neither is in-place.) Other references to the same
Tensor see the update — this matches NumPy semantics:
mut W = Tensor.randn('f32', 1024, 256)
let alias = W
W -= grad * lr # mutates W's buffer
Tensor.eval(alias) # alias.to_array() observes the new values
W = W - grad * lr allocates a new Tensor every step, so for SGD-style
updates over large weight tensors -= is materially faster (the
per-step weight allocation goes away). When the LHS is a view, an
unevaluated graph node, or has a shape that the RHS does not broadcast
into, the runtime falls back transparently to the regular new-Tensor
path — no observable behaviour change.
The _ sink
_ is a non-binding sink. In any binding form it evaluates the
right-hand side (so side effects still run) but discards the value
instead of introducing a name. The same form may therefore appear
repeatedly within one scope, and the shadow-prohibition rule below
does not apply to _.
let _ = side_effect() # value dropped
let _ = 1; let _ = 2 # repeated _ in one scope is fine
for _ in 0..n { count = count + 1 } # body runs n times, no iter var
fn (_, _, x) { x } # only the third arg is bound
try { ... } catch _ { recover() } # error value dropped
let [first, _, third] = [1, 2, 3] # array slot ignored
let [_, ..._, last] = xs # head and middle dropped
match v { [a, _, c] => a + c } # pattern slot is a sink
Reading _ is an error — the binding never happens, so a later _
reference raises undefined variable '_'. The sink rule applies
uniformly to let/mut declarations, for ... in, function
parameters, try ... catch, and pattern slots inside destructure /
match. (Inside an Object pattern { _, x } is a special case:
_ still requires the object to carry a literal _ key, and is not
a sink there — use a positional pattern if you want to discard.)
Shadow prohibition
Introducing a new binding is an error if a variable with the same name exists in an enclosing function (closure-captured). The rule applies uniformly to three kinds of binding introduction:
let/mutdeclarations- function parameters (
fn (name) { ... }) matchpattern bindings (match v { name => ... })
For example:
make_bumper = fn () {
mut count = 0
bump = fn () {
mut count = 10 # error: cannot shadow outer variable 'count'
}
incr = fn (count) { # error: parameter shadows captured 'count'
count + 1
}
peek = fn (v) {
match v {
count: Long => count # error: pattern shadows captured
}
}
}
This catches a common class of bugs where a nested scope introduces the same name as a captured variable, silently breaking the intent to reassign the outer binding. Renaming the inner variable removes the ambiguity.
Shadowing is allowed in two cases:
-
Globals and builtins (
inspect,min, or any top-level binding) may always be shadowed, so writingmut min = arr[0]in a local function costs nothing. -
Block-scope shadowing within the same function is allowed. A
{ ... }block may introduce a newlet/mutbinding of a name already declared in the enclosing function body:fn () { a = 0 { let a = 1; ... } # OK: new binding scoped to the block }
The restriction applies only at function boundaries: closure-captured
state is mutated via bare x = v, and a new local is declared with
let/mut.
Mutability
Bindings are immutable by default. Reassignment without let to an
immutable binding raises immutable variable 'x'....
mut on the binding allows reassignment, but it does not make the
underlying object immutable or vice versa: you can push to an
Array even if the binding is immutable, since the pointer itself
is not changing.
Scope
- Each function body introduces a new function-level scope. Parameters are bound there.
- A block
{ ... }in statement position (LEXICAL_SCOPE) introduces a nested lexical scope that ends at}. - A
forand awhilebody introduce a fresh scope per iteration. A binding introduced in the body neither leaks out of the loop nor persists across iterations (so a bare immutablex = …re-declares each pass rather than re-assigning), and a bodydeferfires at the end of every iteration. Anifbody shares the enclosing scope (it does not iterate, so nothing collides). matcharms introduce a scope that covers the arm's guard and body; variable bindings from the pattern are visible there.
Assignment targets
Complex assignment targets (LHS) are supported:
arr[0] = x # index assignment
obj.key = v # property assignment (creates key if absent)
Design note: why three-tier shadow rules
Culebra treats the three shadow axes independently:
| Scope relationship | Culebra | Typical convention |
|---|---|---|
| Across a function boundary (closure capture) | Error | Warning or allowed |
| Within the same function (block scope) | Allowed | Warning or allowed |
| A global / builtin name | Allowed | Warning or allowed |
The three positions get different policies because each serves a different purpose in the closure-as-object idiom:
- Captured state is object state. In the closure-based object
pattern (see
handbook.md§9.2), an enclosing function's mutable binding is the object's private field. Accidentally shadowing it — typically by writingmut x = ...intending a new local — silently breaks the object. Making this a compile-time error is worth the small restriction. - Block scope is a computation staging area. Inside a single
function, a
{ ... }block is a local calculation region. Rebinding a name there (let a = transform(a)) is a common, intentional pattern, not a bug. No reason to restrict it. - Globals form a shared vocabulary. Builtins (
inspect,to_string,Math,IO) and top-level names are understood to be ambient. Locals likemut min = arr[0]are an ergonomic idiom, not a confusion risk. Requiring renames would be friction without safety gain.
The effect is that the rule catches the bug class it is designed to catch — confusion between "new local" and "outer reassignment" — without interfering with the two situations where shadowing is natural.
7. Expressions
Arithmetic
+, -, *, /, %, ** take two numeric operands (Long or
Float); a non-numeric operand raises type error.
- Both
Long→Long. Integer arithmetic; division and modulo truncate toward zero (not floored, so the sign of the operands does not change the direction).7 / 2 == 3,-7 % 3 == -1. Overflow wraps (no bignum). For the floored remainder — the one that wraps a negative index into0..ninstead of leaving it negative — useMath.wrap. - Either operand
Float→Float. TheLongoperand is promoted toFloatand the operation runs in IEEE 754 binary64.1 + 2.0 == 3.0,3 / 2.0 == 1.5.
Division or modulo by zero raises divide by 0 error at L:C for both
Long / 0 and Float / 0.0.
** (exponentiation) has a slightly richer rule:
Long ** non-negative Long → Long(integer exponentiation; wraps on overflow).2 ** 10 == 1024.Long ** negative Long → Float.2 ** -1 == 0.5.0 ** -1raisesdivide by 0 error.- Either operand
Float→Float(viastd::pow).2.0 ** 0.5 ≈ 1.4142. **is right-associative and binds tighter than unary minus. Its RHS accepts a unary prefix, so both2 ** 3 ** 4 == 2 ** 81and2 ** -1 == 0.5parse.
Bitwise
| (or), & (and), ^ (xor), << / >> (left / arithmetic-right
shift), and unary ~ (complement) operate on two Long operands
(one for ~); any non-Long operand raises type error. >> is an
arithmetic (sign-preserving) shift, matching Long's signedness. Shifts
wrap like the rest of Long arithmetic (no bignum); the shift count is
taken modulo 64 (its low 6 bits), so 1 << 64 == 1 and a negative count
wraps the same way.
0b1100 & 0b1010 # → 8
12 ^ 10 # → 6
0b1010 | 0b0101 # → 15
1 << 4 # → 16
~0 # → -1
5 & ~1 # → 4 (clear the low bit)
READ | WRITE # combine disjoint flag bits
Precedence: comparison < or | < xor ^ < and & < shift
<</>> < additive +. So 1 << 2 + 1 == 8 (1 << (2+1)),
2 & 3 ^ 1 == 3 ((2 & 3) ^ 1), and 1 | 2 == 3 ((1 | 2) == 3).
~ binds at the unary level.
| and |...| lambdas: a bit-OR | works everywhere a normal
expression is parsed, including a lambda body (|x| x | 1). The one
exception is a parameter default, where a top-level | would be
ambiguous with the closing | of the parameter list — write it with
parentheses there: |x = (A | B)|. || remains logical-or.
Comparison
==,!=: any two values; see §5 for reference-type semantics and for numeric cross-type equality (1 == 1.0).<,<=,>,>=: same-typed operands, with the exception thatLongandFloatcompare by numeric value (1 < 1.5works). Any other cross-type ordering raisestype error.- Chaining:
a < b < cmeans(a < b) && (b < c)— the middle operandbis evaluated once, and the chain short-circuits tofalseat the first failing link. Any mix of comparison operators chains:0 <= i < n,lo < x <= hi,a == b == c.
Logical
!x: requiresxconvertible to bool.x && y: evaluatesx; if falsy returnsx, else evaluates and returnsy. Short-circuit.x || y: evaluatesx; if truthy returnsx, else evaluates and returnsy. Short-circuit.x ?? y: returnsxif it is notnil, elsey. Short-circuit (RHS not evaluated when LHS is non-nil). Lower precedence than||. Chains left-associatively:a ?? b ?? c=(a ?? b) ?? c.
Null-safe access
These operators complement ?? and the T? Optional type (§14) for
working with possibly-nil values.
a?.b/a?.m(...)— optional chaining: ifaisnilthe whole remaining chain short-circuits tonil(the property read or method call is skipped, including its arguments). Otherwise behaves likea.b/a.m(...).a?.b.cshort-circuits the entire.ctoo; usea?.b?.cto guard each link.a?[k]— optional index:nilreceiver short-circuits tonil, otherwise indexes likea[k](Array / Tuple / Object).expr!!— non-null assertion: passesexprthrough unchanged, or raisesNilErrorif it isnil. Postfix, so it chains:a!!.b,a!![0].a ??= b— nil-coalescing assignment: assignsbtoaonly whenacurrently reads asnil, short-circuitingbotherwise (bis not evaluated on the non-nil path). The target may be a plain variable,obj.key,obj[k](including a class instance's__index__/__setindex__fallback), orarr[i]. An absent Object key reads asnilfor this purpose, same as a plainobj.keyread (§10) — soobj[k] ??= vinsertskwhen it is missing, honoring the usual mutable-by-default rule for a runtime-inserted key (§10) and the existing slot'smutflag when the key is already present but nil. An Array index does not auto-extend: an out-of-rangeistill raisesIndexError. Not supported on aFixedArray/SharedBufferelement or a@packablepacked field (nonilsentinel for a packed scalar), or on aShared.newview (unconditionally immutable).
Truthiness
Only Bool, Long, and Float are convertible to bool:
Bool: itself.Long:0is false, all others true.Float:0.0(and-0.0) are false. Every other finite value is true, andNaNis true — truthiness asks "is this a number at all", not "is this a usable number".Nil,String,Array,Object,Function: not convertible — using one in a boolean context (e.g.,if s { ... }) raisestype error.
(This is intentionally strict. Wrap with an explicit check such as
s != nil or !arr.empty() if needed.)
Unary
+x is a no-op (must be numeric). -x negates a Long or Float.
Parentheses
(expr) groups and does not introduce a scope.
Method call
A method call takes the form receiver.name(args). The full
resolution rules — including method/UFCS dispatch order and how
built-ins interact with user-defined properties — are specified in
§10 ("Methods and UFCS"). Operator overloading (+, -, *, ==,
@, …) and __str__ are also defined there.
Evaluation order
Sub-expressions are evaluated left-to-right in source order. This rule is normative on every backend; the JIT may rearrange intermediate IR but must preserve the observable effect order.
- Function call arguments.
f(p1, p2, **splat, k: kv)evaluatesp1, thenp2, thensplat, thenkv— exactly the order they appear at the call site, even when positional,**splat, and keyword arguments are interleaved. - Array, object, tuple, and set literals. Element/property
expressions evaluate in source order. For
{a: v1, b: v2}it isv1thenv2. For[n; default](the size-prefixed form) the count expression evaluates before the default expression. - Binary operators
+,-,*,/,%,==, etc. Left operand first, then right operand, then the operation. ||,&&,??. Short-circuit at the first decisive operand —||stops at the first truthy value,&&at the first falsy,??at the first non-nil. The trailing operands are not evaluated.- Assignment
lval = rhs. RHS evaluates first, then the LHS target chain (and its final subscript key forobj[k] = ...), then the store is performed — the value is computed before the place it lands in is resolved. - Compound assignment
lval op= rhs. RHS evaluates first; the LHS target chain (including any subscript key) is then evaluated exactly once, the implicit read and the store both reuse the same evaluated chain. Subscript keys are not re-evaluated. - Method call. Receiver evaluates first, then the argument list in source order (as above).
When in doubt, write the expression in steps using temporaries — the spec exists so you don't need to.
8. Strings and interpolation
Raw string literals
'hello'
'C:\path\to\file' # backslashes are literal — no escape decoding
'a\nb' # 4 characters: a, \, n, b
Single-quoted strings are raw: every character is taken verbatim
between the quotes. There are no escape sequences and no interpolation.
A literal apostrophe is the one byte a '...' string cannot hold — use
a backtick string (below) when the content contains '. Both raw and
interpolated strings may span multiple lines (a newline in the source is
part of the value), so '\d+' is a ready-made regex literal and
"...\n..." a multi-line template.
Backtick raw string literals
`it's` # holds a single quote
`say "hi"` # holds double quotes — no escaping
`\d{4}-\d{2}` # regex with quantifiers, verbatim
Backtick strings (Go-style) are raw like '...' but may also contain
', ", and { — every byte between the backticks is verbatim, with
no escapes and no interpolation, and they may span multiple lines. The
only byte a backtick string cannot hold is a backtick itself. They are
the cleanest form for regex patterns that mix apostrophes and braces
(e.g. a GPT-2 pre-tokenizer: `'s| ?\p{L}+|\s+(?!\S)`), where neither
'...' (no apostrophes) nor "..." (interpolates {...}) fits.
Regex literals
re'\d+' # a compiled Regex, same as Regex.compile('\d+')
re"\d{4}-\d{2}" # `re` makes the body raw, so {4} is a quantifier
re`["']\w+` # backtick body: holds both ' and "
re"hello"i # trailing flags (i / m / s)
re"^${word}$" # ${expr} interpolation (escaped / spliced, see below)
A re'...' / re"..." / re`...` literal evaluates to a compiled
Regex value — it is exactly Regex.compile(<body>, <flags>) and supports the whole Regex object API, so it chains directly:
re'\w+'.find(" hello world").value # => "hello"
re"hello"i.test("HELLO") # => true
The body is raw regardless of the quote — the re prefix turns off
escape decoding, so \d, \w, and {n} quantifiers pass through verbatim
and the closing quote is the only delimiter.
The one structured form is ${expr} interpolation. It uses $, not
the bare { of a normal "...", precisely so it cannot collide with a
{n} quantifier (a $ anchor is never quantified in practice):
let word = "a.b"
re"^${word}$".test("a.b") # => true
re"^${word}$".test("axb") # => false — the `.` is escaped
Interpolation is type-driven and injection-safe by default:
- a String (or any non-Regex value, stringified) is escaped, so it matches literally — metacharacters in the value lose their meaning;
- a compiled
Regexis spliced as a non-capturing group(?:src), so its quantifiers and alternation compose into the pattern (the engine applies any inline flags it carries to the whole match).
let digits = re"\d+"
re"id=${digits};".test("id=123;") # => true — composed
To write a literal $ immediately before a {, use \$ (the regex
escape for a literal dollar), which also suppresses interpolation:
re"\${2}" matches two dollar signs. Still build a fully dynamic
pattern with Regex.compile(...):
let n = 4
Regex.compile('\d{' + n.to_s() + '}')
Optional trailing flag letters (i case-insensitive, m multi-line, s
dot-matches-newline) follow the closing quote: re"^b"m. There is no
/.../ form: / is division, and disambiguating the two needs a
stateful lexer culebra deliberately avoids. re is only a regex prefix
when a quote immediately follows — otherwise it is an ordinary identifier
(let re = 5 is fine).
Triple-quoted strings
"""multi-line
with "quotes" and 'apostrophes' and {interpolation}"""
"""...""" is interpolated like "..." (the same {expr} / {x:spec}
forms and \n / \{ escapes) but single and double quotes inside need
no escaping — only """ closes it. Handy for embedded LLM prompts and
snippets that contain quotes. Use \{ for a literal open brace.
Block form (dedented)
When the opening """ is immediately followed by a newline, the string is
a block string. The newline after the opening """ and the newline before
the closing """ are not part of the value, and every line is dedented by
the indentation of the closing """, so the literal can be indented to match
the surrounding code without that indentation leaking into the string:
let html = """
<html>
</html>
"""
# => "<html>\n</html>"
The dedent is resolved once, when the literal is built (the JIT folds a
pure-literal block into a single constant) — there is no separate runtime
trimIndent-style pass over an already-allocated string. Relative indentation
beyond the closing delimiter is preserved, and blank lines are normalized to
empty. A non-blank line indented less than the closing """, or a closing
""" that does not sit on its own line, is a SyntaxError.
A triple string whose content begins on the opening line (including the
single-line """...""" form) is not a block string: it is taken raw, with
no dedent, exactly as before.
Interpolated strings
"hello {name}"
"sum = {a + b}"
"nested: {if x > 0 { 'pos' } else { 'neg' }}"
"line one\nline two"
"literal brace: \{ and tab:\there"
A "..." string consists of plain text segments and {expr} segments.
Each expression is evaluated, converted to its display form (see below),
and concatenated with the surrounding text.
Escape sequences (in plain-text segments only):
| escape | byte / codepoint |
|---|---|
\n | newline (0x0A) |
\r | carriage return |
\t | tab |
\\ | backslash |
\" | double quote |
\{ | literal { (does not start an interpolation) |
\xHH | one raw byte from two hex digits (0x00–0xFF) |
\uXXXX | a BMP Unicode scalar value (exactly 4 hex digits), UTF-8 encoded |
\UXXXXXXXX | a Unicode scalar value (exactly 8 hex digits), UTF-8 encoded |
\xHH writes a raw byte — it can produce bytes that aren't valid UTF-8
(e.g. "\xff"), since a String is a byte string. \uXXXX / \UXXXXXXXX
write a codepoint: both reject values above U+10FFFF and the surrogate
range U+D800–U+DFFF (not Unicode scalar values); only \U reaches
beyond the BMP. A } does not need
escaping outside of an interpolation. An unknown \X is preserved
unchanged as two characters (\ and X). An embedded NUL (\x00 /
\u0000) is an ordinary byte: size() counts it and every String
operation preserves it — a String never terminates at a NUL.
Format specs
An interpolation may carry a format spec after a colon — {expr:spec}.
The spec is the C++ std::format mini-language:
[[fill]align][sign][#][0][width][.precision][type].
"{pi:.2f}" # → 3.14 (fixed-point, 2 decimals)
"{5:.2f}" # → 5.00 (a Long is coerced for a float spec)
"{n:05}" # → 00042 (zero-padded width)
"{255:#x}" # → 0xff (hex with prefix)
"{name:>10}" # → right-aligned in 10 columns
"{score:+}" # → +42 (explicit sign)
Numeric values honor the spec's type char: a float type (f/e/g)
formats Longs and Floats as floating-point, an integer type
(d/x/o/b) formats as an integer. Other values format their
display string (so width / alignment apply). An invalid spec for the
value's type raises ValueError. std::format is followed exactly, so
, digit grouping is not supported (the locale L option covers that
case instead).
A spec may itself carry {expr} fields, so a width or a precision can be
computed rather than written out:
"{s:>{w}}" # right-aligned in `w` columns
"{x:.{p}f}" # `p` decimals
"{x:{w}.{p}f}" # both, in one spec
A field is an ordinary expression, evaluated where it stands (left to
right, after the value being formatted). It must be a Long — it is
spliced into the spec as decimal text, which is what a width or precision
is — and anything else is a TypeError at the field's own position. A
field means exactly what the same digits written by hand would mean, so a
width of 0 reads as the zero-pad flag rather than "no width": fine for a
number ("{n:>{0}}"), a ValueError for a string, just as "{s:>0}"
already is. It works anywhere an interpolation does:
let w = 6
inspect(["a", "bb"].map(|s| "{s:>{w}}")) # => [' a', ' bb']
Display conversion
When an expression inside "..." is not a String, its display form
is inserted:
Nil→nilBool→trueorfalseLong→ decimalFloat→ shortest round-trip decimal, always with a.or an exponent so the type is distinguishable fromLong:1.0,0.5,-2.5,1e-05,nan,inf,-infArray→[v1, v2, ...]with each element'sinspectform (strings in brackets are quoted, e.g.['hi'])Object→{key: val, mut key2: val2, ...}in insertion orderFunction→[function]
String values inside "..." are inserted verbatim (no quotes); this
is the one place display differs from inspect.
Cyclic data (a.c = a) displays as {...} / [...] to avoid
infinite recursion.
Concatenation
The + operator joins two strings into a new String:
'foo' + 'bar' # 'foobar'
'a' + 'b' + 'c' # 'abc'
+= appends in place (strings are immutable, so this rebinds the
variable / element / field to the joined result):
let mut s = 'a'
s += 'b' # s is now 'ab'
Both operands must be a String or StringView; the result is always
an owned String. Concatenating a string with a non-string ('n: ' + 1)
is a TypeError — use interpolation ("n: {1}") or to_string to mix
types.
Arrays concatenate the same way (§9); no other type does.
9. Arrays
Construction
[] # empty array
[1, 2, 3] # three elements
[1, 2, 3](5, 0) # 5-element array: [1, 2, 3, 0, 0]
[](5, nil) # 5-element array of nils
[1](10, 0) # [1, 0, 0, 0, 0, 0, 0, 0, 0, 0]
The optional (count, default) tail first fills the array to count
elements with default, then overwrites the first positions with any
literal values. Only the default form is available; omit default to
get nil fill.
Spread. A ...iterable element splices another collection's
elements into the literal:
[0, ...a, 4] # → [0, <a's elements>, 4]
[...a, ...b] # concatenate
Spread sources are Array, Tuple, and Set (a non-iterable raises
TypeError).
Concatenation. + joins two Arrays into a new one, the same way it
joins two Strings (§8). Spread is the more general form — it takes
Tuple / Set sources and can splice into the middle of a literal —
while + reads better for the plain two-Array case and chains:
a + b # new Array; a and b unchanged
a += b # rebinds a to the concatenation
[...a, ...b] # same result as a + b
The copy is shallow: elements are shared with the operands, as with
slice. To append in place instead of building a new Array, use
a.extend(b) (§18.2).
Access and mutation
arr[i] # index access, negative indices count from end
arr[-1] # last element
arr[i] = v # set; element must already exist (no auto-extend)
arr.size() # element count (Long)
arr.push(x) # append to end, returns nil
arr[i] out of range raises index out of range at L:C.
Slicing
A range index (seq[a..b], seq[a..=b]) returns a sub-sequence instead
of a single element. .. is end-exclusive, ..= end-inclusive; either
endpoint may be negative (counted from the end). Either endpoint may also
be omitted — an open start defaults to 0, an open end to the
sequence length: xs[2..] drops the first two, xs[..3] keeps the first
three, xs[..] copies the whole sequence. Out-of-range endpoints
clamp, and a start past the end yields an empty result, so slicing
never raises on bounds. An endpoint is an index, so it must be a Long:
a range with a Float endpoint raises TypeError as a slice.
Arrays return a shallow copy — the slice's spine is independent of
the source, but elements are shared (a reference-semantic array sliced
into a view would alias surprisingly, so copy is the safe default). For
a zero-copy lazy window over a large array, use the iterator instead
(xs.iter().skip(a)). Strings return a byte-unit view (Go-style
byte indexing; a slice that lands mid-codepoint keeps the raw bytes).
Tuples return a tuple.
A range is a first-class value (let r = 1..3) — store it, pass it
to a function, and use it to subscript later (xs[r]). A range bounded
by two Long endpoints is also iterable (for i in 1..4); an unbounded
one (2..) has no iteration end and raises if iterated.
let xs = [10, 20, 30, 40, 50]
inspect(xs[1..3]) # => [20, 30]
inspect(xs[1..=3]) # => [20, 30, 40]
inspect(xs[-3..-1]) # => [30, 40]
inspect(xs[2..]) # => [30, 40, 50]
inspect(xs[..3]) # => [10, 20, 30]
let xs = [10, 20, 30, 40, 50]
inspect(xs[1..100]) # => [20, 30, 40, 50]
inspect(xs[3..1]) # => []
let r = 1..3
inspect(xs[r]) # => [20, 30]
# Shallow copy: mutating the source does not touch the slice.
mut xs = [10, 20, 30]
let s = xs[0..2]
xs[0] = 99
inspect(s) # => [10, 20]
inspect("hello"[1..3]) # => 'el'
inspect("hello"[1..=3]) # => 'ell'
Equality and ordering
Arrays compare by value (structural): [1, 2] == [1, 2] is true,
recursing through nested elements. Tuples, Sets, and
Objects are likewise value-equal; only Function and Tensor keep
reference identity. (Arrays remain unhashable — value equality does
not make them usable as Object/Set keys.) Ordering operators (< etc.)
are not defined on arrays.
10. Objects
Construction
{} # empty object
{name: 'alice', age: 30} # two properties
{mut counter: 0, name: 'x'} # explicit mut on a property
{name, age} # shorthand — same as {name: name, age: age}
Property order in the source is irrelevant for equality or access, but the order is preserved for display and iteration: keys come back in the order they were first written.
The shorthand form {x} is equivalent to {x: x}: it reuses the
identifier as both key and value, looking up x in the current scope.
mut is allowed ({mut n}), which declares the property mutable while
the value still comes from the binding n.
Spread / merge. A ...obj member copies another Object's entries
in; later keys win, so it doubles as config override:
{...defaults, ...overrides} # merge (overrides win)
{...base, port: 8080} # override one field
Merged entries are mutable. Spreading a non-Object raises TypeError.
Display conventions
The default formatter (used when an Object has no __str__) prefixes a
value with the class that built it. So a class-sugar instance with fields
x, y displays as Point {mut x: 3, mut y: 4}. A plain Object has no
class and displays as {...}, whatever its fields are named — a class
field is an ordinary property and appears in the list like any other.
Enumeration
keys(), size(), to_string(), for k, v in obj, spread ({...obj})
and JSON.stringify all enumerate the same set: the object's own
entries, in insertion order. For a class-sugar instance that is the
declared fields in declaration order, then the ones the constructor and
methods set with self.x = y. Neither the class name nor the methods are
own entries — both live on one per-class meta every instance delegates to
(see §25) — so neither appears. A spread of an instance therefore copies
its fields into a plain Object, which is not an instance of anything:
class P { new() { self.a = 1 } m() { 2 } }
let p = P()
p.keys() # ['class', 'a']
p.size() # 2
p.to_string() # 'P {mut a: 1}' (class hoisted to the prefix)
JSON.stringify(p) # '{"class":"P","a":1}'
{...p}.keys() # ['class', 'a'] — a plain Object, no methods
has(key) is the exception: it answers the lookup question, so it also
sees class methods — but not dict builtins:
p.has('m') # true — a class method
p.has('size') # false — a dict builtin, not an entry
The data accessors stay own-entry-only: p.get('m', fallback) returns
the fallback and p['m'] raises KeyError. Assigning to a method name
(p.m = 1) creates an own field that shadows the method.
Access and mutation
obj.key # read, returns nil if absent
obj.key = v # set (creates as mutable if absent; see below)
obj[k] # subscript read with a non-String key (see below)
obj.size() # property count (Long), includes non-String keys
obj.has(key) # Bool — accepts String, Long, Float, Bool, Nil, Tuple
obj.keys() # Array of keys in insertion order (interleaved)
obj.remove(key) # remove the entry for `key` (any hashable key); no-op if absent
Assigning to an existing property that was declared without mut
raises immutable property 'key' at L:C. A property created at
runtime by obj.key = v (or obj[k] = v, get_or_put, to_object,
group_by) is mutable by default — the same rule the {...spread}
merge already uses ("merged entries are mutable"). Only an Object
literal's own keys are immutable by default ({key: v}; mut key: v
opts a literal key in — see Construction, above).
Dot-form names (obj.key) are identifiers ([A-Za-z_][A-Za-z0-9_]*)
that go through the fast shape path. Non-identifier keys reach the
subscript path below.
A key may also be a string literal — '...', a backtick string, or a
"..." with no interpolation — which is how you write keys that aren't
identifiers (hyphens, spaces, etc.):
let headers = {"User-Agent": "curl", "X-Trace-Id": id}
headers["User-Agent"] # read non-identifier keys with [ ]
A key must be a compile-time constant, so an interpolating "...{x}..."
is a SyntaxError; build a dynamic key with o[k] = v instead.
Non-String keys
In addition to String keys ({name: 'alice'}), Object literals accept
Long, Float, Bool, Nil, and Tuple keys:
let grid = {(0, 0): 'origin', (1, 0): 'east'}
let by_id = {1: 'one', 2: 'two', 3.14: 'pi'}
let by_nil = {nil: 'unknown'}
Non-String keys live in a sidecar map and are read with obj[k]:
grid[(0, 0)] # 'origin'
by_id[1] # 'one'
Key identity is type-strict, deliberately finer than the ==
operator. Keys of different types are never the same key, so
1, 1.0, and true are three distinct keys even though 1 == 1.0
holds for the == operator. A literal that repeats a key keeps the
last write —
{1: 'a', 1: 'b'} is {1: 'b'} — the same last-wins rule a literal
always uses for a duplicated key, including String keys.
String and non-String keys share a single insertion-order record, so
to_string renders mixed-key Objects with the keys interleaved in
their actual write order.
Subscript assignment
obj[k] = v writes any hashable key — Long, Float, Bool, Nil,
Tuple, an enum variant, or a runtime String:
mut bag = {}
let k = 'alpha'
bag[k] = 1 # runtime String key
bag[42] = 'long'
bag[(1,2)] = 'tuple'
inspect(bag[k]) # 1
inspect(bag[(1,2)]) # 'tuple'
String keys unify with the shape-based obj.foo path — both forms
reach the same slot:
mut o = {}
o['x'] = 1
inspect(o.x) # 1
o.y = 2
inspect(o['y']) # 2
A key created at runtime is mutable, so writing it again succeeds
(unlike a same-named Object literal key, which stays immutable
without an explicit mut):
mut tally = {}
for w in ['a', 'b', 'a'] {
tally[w] ??= 0 # insert 0 on the first sighting (§7, nil-coalescing assignment)
tally[w] += 1
}
inspect(tally) # {a: 2, b: 1}
Existing slots honor their mut flag, regardless of which form the
write uses; obj[k] = v on an immutable slot raises ImmutableError.
Non-String keys live in a sidecar map. Compound forms (obj[k] += v)
update the slot in place and require the key to already exist
(KeyError otherwise) and the slot to be mutable.
An Object built this way is a dictionary and performs like one: insert, read and delete stay O(1) as it grows, whatever the key type. The fixed-slot layout an Object literal or a class instance gets is a different representation, and an object leaves it once it holds enough keys that the layout no longer pays — a change of representation only, with no change in behaviour.
Methods and UFCS
A method call receiver.name(args) resolves in this order:
- If
receiveris a built-in namespace (IO,Math, …), only its own members resolve — see Namespaces are closed below. UFCS is off for such a receiver. - If
receiverexposes a property or built-in method namedname, invoke it. ForObject/Array, user-defined properties win over built-ins; built-ins fill in otherwise. String methods are the only choice forStringreceivers. A default implementation from a trait the receiver's class conforms to (§14) counts as one of its methods here, so it resolves at this step rather than falling through to UFCS. When the resolved value is aFunction,selfis bound toreceiverfor the duration of the call. - Otherwise, if a free function named
nameis visible in the enclosing scope and its first parameter acceptsreceiver, call it withreceiveras the first argument and the remaining arguments as-is. This is Uniform Function Call Syntax (UFCS), matching the D / Nim convention. See The first parameter decides below. - Otherwise the property lookup returns
nil; the subsequent call fails with a type error.
o = {n: 10, add: fn (x) {
x + self.n
}}
# A receiver call binds self, so `add` reads o.n:
inspect(o.add(5)) # => 15
double = fn (x) {
x * 2
}
42.double() # UFCS → double(42) → 84
word_count = fn (s) {
s.split(' ').size()
}
'hello world'.word_count() # UFCS → 2
# Existing methods always win — a user-defined `size` is shadowed by
# the Array/Object/String built-in `size`.
size = fn (x) {
99
}
[1, 2, 3].size() # 3 (builtin), not 99
UFCS only fires when DOT is immediately followed by an argument list;
bare property access (x.name without ()) never uses UFCS. A UFCS
invocation does not bind self to the receiver — the call is
semantically a free-function call with the receiver in the first
positional slot.
Two names dispatch outside the built-in tables and still follow this
order. A class instance's synthesized parameters() (§10) is a property
for step 1, so a global parameters never claims one. The explicit-drop
form x.drop() (§17) sits below step 2 instead: a receiver carrying
no drop of its own hands the call to a free drop in scope that
accepts it, and only a receiver that resolves the name — a handle with
its own drop, or no such candidate in scope at all — reaches the
at-most-once guard. A class
object, an enum or a namespace never does: C.drop() is an ordinary
call of the static drop.
The first parameter decides
A free function is a UFCS candidate for a receiver only when its first
parameter accepts that receiver. An unannotated first parameter accepts
everything; an annotated one accepts what the annotation would accept as
an ordinary argument (§14) — a union, a trait the receiver's class
conforms to, an enum at any of its three levels. A function with no
leading positional parameter (none at all, or only *args) declares
nothing to fail and stays a candidate. For a multimethod, one overload
whose first parameter accepts the receiver is enough; which overload
runs is then the ordinary dispatch over all the arguments (§20).
A function that does not accept the receiver is not there as far as the call is concerned, so a mistyped method name does not turn into a complaint about some unrelated function's parameter:
# doctest: skip
let g = fn (p: Long) {}
class C {
new() {}
f() { self.g(0) } # C has no `g`, and the free `g` takes a Long
}
C().f() # TypeError: expected Function, got Nil
7.g() # UFCS → g(7)
The question is asked of the receiver alone, before any argument runs — the other arguments are checked where they always were, by the callee's own binder.
Built-in methods bind positionally
A built-in method (Array / String / Set / Tuple / dict /
iterator / Tensor) binds its arguments by position. The one name a
keyword argument may carry is a keyword-only parameter — one with no
positional slot at all, like sorted's reverse:. A keyword naming an
ordinary parameter is a TypeError, even where that parameter exists,
and a ** splat is never accepted:
# doctest: skip
[3, 1, 2].sorted(reverse: true) # → [3, 2, 1] (keyword-only)
'hello world'.truncate(8) # → 'hello...'
'hello world'.truncate(max: 8) # TypeError: built-in method 'truncate'
# does not accept keyword arguments
[3, 1].sorted(**{'reverse': true}) # the same TypeError
The rule is about the callee, not the name: a user-defined method of
the same name is an ordinary function, so doc.truncate(max: 8) binds
its keyword normally. It is also about a receiver that resolves the
name — the property is read before any argument binds, so
[1, 2].take(n: 1) (no take on Array) is the ordinary method miss
rather than the keyword error.
The argument list runs before anything binds
x.m(a, b) reads the property, evaluates every argument left to
right, and only then binds them. So a wrongly typed argument and a
method the receiver does not have are both reported after the whole
list has run:
# doctest: skip
'abc'.truncate(bad(), loud()) # both run, then the `max` type error
'ab'.push(loud()) # loud() runs, then "expected Function, got Nil"
The one thing that fails earlier is a scalar receiver: nil, Bool,
Long and Float carry no members at all, so (5).push(loud()) fails
at the property read and loud() never runs.
Namespaces are closed
A built-in namespace exposes a fixed member set, so an unknown member is
a typo or a removed API rather than an extension point. Reading one
raises AttributeError: namespace 'IO' has no member 'zzz', and UFCS
does not step in — a free function named zzz cannot fill the gap:
# doctest: skip
zzz = fn (ns, v) {
v
}
IO.zzz(9) # AttributeError, not zzz(IO, 9)
Math.to_string() # AttributeError, not to_string(Math)
The dict built-ins are the exception: IO.keys(), Sys.has('script')
and the rest of the Object table still answer, because a namespace is
an Object. Plain objects are unaffected — {a: 1}.zzz(9) is UFCS as
usual.
Writing a namespace member follows the same fixed-set rule: Ns.zzz = v
and Ns.zzz ??= v raise the identical AttributeError rather than
minting a phantom member (Ns.zzz += v already raised on a missing
property under the general compound-assignment rule below). Writing an
existing member follows the ordinary assignment rules for that slot —
most stdlib namespace members are registered immutable, so the common
outcome there is ImmutableError, not a silent overwrite:
# doctest: skip
Math.zzz = 1 # AttributeError: namespace 'Math' has no member 'zzz'
Math.pi = 3.0 # ImmutableError: immutable property 'pi'
A bare read of a built-in method (let m = [1, 2].map, no parens)
is not a first-class value — it raises TypeError: built-in method 'map' cannot be used as a value (call it, or wrap it in a lambda).
The reject applies only when the receiver's own built-in table carries
the name; a built-in-named property the receiver simply lacks reads
as nil like any other miss ({a: 1}.map, or C.join on a class
object whose join is an instance method). User-defined methods read
as bound, first-class values (see self below).
self resolves in two steps. A call that supplied a receiver binds it
for the duration of the call, and that dynamic binding always wins. A
body reached without one — a plain f(x), a UFCS invocation, a
function handed to a built-in like map, a defer block — falls back
to the self of the lexically enclosing function, walking outward as
far as needed, so a nested fn, lambda, or closure returned from a
method keeps seeing the method's receiver. Only when neither a
receiver nor an enclosing self exists does the read raise
NameError: undefined variable 'self', exactly as any other unbound
name does, and at the point of the read rather than on entry to the
body.
Reading a function-valued property as a value (let m = o.f, no
call parens) returns a bound method: a fresh wrapper that carries
o as its self. The binding is permanent — attaching the wrapper to
another object and calling it as a method still runs with the original
receiver — and each read mints a new wrapper, so o.f == o.f is
false. This applies uniformly to object-literal properties, class
methods, statics, and constructors (let mk = C.new). Inside a body,
the implicit recursion handle fn refers to the value the call was
made through, so in a method-invoked frame fn is the bound wrapper —
recursing through fn(...) or returning fn keeps the current
receiver.
JIT: UFCS is supported under --jit. Resolution happens at
runtime: if the receiver carries a property by that name the method
path wins, otherwise the name is looked up as a free function and — if
its first parameter accepts the receiver — invoked with the receiver as
its first argument.
class sugar
Closure-based objects remain the canonical OO idiom. The class form
is a lightweight alternative that desugars to the same runtime shape:
class Car {
new(mpr) { self.miles = 0; self.mpr = mpr }
run(n) { self.miles = self.miles + self.mpr * n }
total() { self.miles }
}
c = Car.new(5); c.run(3); inspect(c.total())
inspect(type_of(c)) # 'Car' — the class that built it
Semantics:
-
The decl binds
Carto anObjectwith a singlenewproperty. -
Car.new(...)returns a freshObjectholding the fields the constructor sets.selfis bound to that object for the duration of the constructor body. Neither the class name nor the methods are copied onto the instance: both live on a single per-class meta each instance delegates to, soc.runresolves through it andtype_of(c)reads the name off it, whilec.keys()reports only the fields (see Enumeration). An Object cannot be given a class by writing a field:type_of,match, a class-typed parameter and==all ask the meta. -
Calling the class calls its constructor:
Car(5)isCar.new(5), keyword arguments and all. The rule is stated over thenewa value carries, not over how the value was written, so it reaches the library's own constructors too —Vector2(1, 2),Scene.Image(4, 4),Channel()— and there is no documented.new(...)it does not cover. What it does not reach is a plainObjectwith anewkey: that stays inert, so a dict cannot be made callable by naming a field. -
Fields created via
self.x = yinside constructors and methods are mutable by default (unlike bareo.x = y, which creates an immutable property). This matches the idiom of classes whose methods routinely mutate instance state. -
The
newmethod is optional; without it the class accepts no arguments and returns an instance with no fields at all. -
selfis immutable inside the constructor body. Attemptingself = newObjraisesImmutableError. The constructor always returns the originally allocated instance — an explicitreturn valuediscardsvalue. Identity-swap factories live asstaticmethods (below) or as plain top-level functions. -
A method's parameters and return value take the same annotations as a
fn(see Return), and they are checked the same way:area() -> Float { ... }checks the returned value on fallthrough and onreturn, andc.area.return_typereports'Float'.staticandgetmembers take them too.newis the exception: its result is always the instance, sonew(...) -> Tis aSyntaxError. -
Methods prefixed with
staticlive on the class object itself (not on instances), providing class-as-namespace for factories, constants-as-functions, and helpers.Shape.create(...)resolves via the usual property-access mechanism; the receiver is the class object, not an instance, soselfis unavailable inside a static body:class Shape { new () { self.kind = 'unknown' } area () { 0 } static circle (r) { let s = Shape.new() s.kind = 'circle'; s.radius = r; s } static square (side) { let s = Shape.new() s.kind = 'square'; s.side = side; s } } let c = Shape.circle(4) # static factory let s = Shape.square(3) inspect(c.kind) # 'circle'Static methods are not visible through instances (
c.circle(...)raisesTypeErrorbecause the instance has nocircleproperty): a static is reached through the class, never through an instance. -
static NAME = EXPRESSIONdeclares a class-level immutable constant (a static field), evaluated eagerly at class declaration time and installed as a property of the class object alongside static methods:class Circle { new (r) { self.radius = r } static PI = 3.14 static MAX = 100 area () { self.radius * self.radius * Circle.PI } } inspect(Circle.PI) # 3.14 inspect(Circle.MAX) # 100The value expression can be arbitrary (
static SUM = [1,2,3].sum()), evaluated in the enclosing scope at class declaration time. Like static methods, static fields are immutable (Circle.PI = 2raisesImmutableError) and not visible through instances. -
NAME = EXPRESSION/NAME: Type/NAME: Type = EXPRESSIONdeclares an instance field — mutable per-instance state initialized before thenewbody runs. The type annotation is optional, as everywhere else in the language:class Player { score = 0 name: String tags = [] best = self.score + 10 new (name) { self.name = name } } let p = Player.new('rocci') inspect(p.score) # 0 inspect(p.best) # 10 p.tags.push('bird') # each instance gets its own ArrayInitializer semantics:
- Per instance: the initializer expression runs once for every
C.new(...)call —tags: Array = []gives each instance an independent Array — nothing is shared between instances. - Declaration order,
selfin scope: an initializer may read fields declared above it and call methods (bestabove). Reading a field declared below yieldsnil— order is meaningful. - After argument binding, before the
newbody: initializers see a fully bound call — an arity/type/default error on the ctor call fires first, with no field side effects — and thenewbody in turn sees every declared field. Constructor parameters are not visible inside initializers (initializers close over the class's defining scope, not over the constructor call); pass ctor args to fields explicitly withself.x = ain the body, or with a field parameter (new(.x), below). - A typed field without an initializer (
name: Stringabove) takes its type's zero value:0/0.0/''/false; reference types (Array,Object, ...) default tonil. The untyped form always carries an initializer (there is no type to infer a zero value from). - Declared fields are mutable instance state, exactly like fields
created via
self.x = yin the constructor. A@valueclass (§21) is the exception: its instances freeze whennewreturns.
A scalar declared type —
Float,Long,Bool, and the fixed-width spellings of §21 (Float32,Int32,Byte, ...), each of which stands for one of those three — is checked on every write, the same runtime-check model parameters andlet x: Tfollow (§14): the field admits only its own values. The check covers every way a value reaches the field: the declaration's own initializer, a literal name, a computed key, the constructor's ownself.x, a native builder. That is what makes such a field's type worth reading: aFloatfield holds aFloat, so code that knows the class knows the field's type without asking.Every other annotation is unchecked. It still chooses the field's zero value above and still records what the code means, but the field accepts whatever it is given:
String, a class, a union, an optional (T?), and a field with no annotation at all. The reasoning is the one@valueuses to narrow its own field types (§21) — a scalar is a value a tag alone decides, with no heap body behind it and no second spelling of the same thing.Stringhas a second spelling:StringView(§18.1) is also a string value, and it is whatslice,split,linesandforover a string hand back. Aname: Stringfield therefore accepts aStringView, and a read of it asks the value what it is, exactly as before.@packableclasses (§21) additionally read a field's declared type to compute their fixed byte layout, and@valueclasses to check that every field is a scalar or another@valueclass, so both require typed fields (x = 7in either is a SyntaxError). - Per instance: the initializer expression runs once for every
-
new(.x)declares a field parameter: anewparameterxthat also stores its argument into the fieldx, so the body carries no assignment for it..x: Tand.x = vgive it a type and a default the way any parameter takes them;.x?makes it optional and takes both from the field's declaration instead:class Cartridge { prg_rom: Array chr_rom: Array chr_ram = self.chr_rom.empty() ? repeat(8192, 0) : [] mapper_id = 0 has_battery = false new(.prg_rom, .chr_rom, *, .mapper_id?, .has_battery?) {} } let c = Cartridge([0x4E, 0x45, 0x53], [], mapper_id: 4) inspect(c.mapper_id) # 4 inspect(c.has_battery) # false — the declaration's default inspect(c.chr_ram.size()) # 8192 — the initializer read chr_rom- It is an ordinary parameter named
x: passed asxby keyword, reported asxin arity and type errors, readable asxin later defaults and in the body. Positional and keyword passing,*, and the ordering rule (a required parameter cannot follow an optional one) apply unchanged;.x?counts as optional. - The argument is stored at the field's own place in the declaration
order, after the arguments bound and before the
newbody, so an initializer declared below the field reads what the argument stored (chr_ramabove) and one declared above it readsnil, as ever. The store is the checked write every other store is: a scalar declared type refuses a value of another type. .x?left out by the caller takes the declaration's own value, the initializer or the type's zero value, and the parameterxreads as that value in the body. Passingnilis passing a value, not leaving it out.- Naming a declared field: the declaration must carry no initializer
unless the parameter is
.x?(the parameter would otherwise always supply the value); a type on the parameter must repeat the declaration's;.x?takes neither a type nor a default of its own. - Naming no declared field:
.xand.x: Tdeclare it, in the class's field list at the firstnew's place (every overload's such fields go there, in the order the overloads name them), with the type given (an untyped one starts asnil, has no zero value and is unchecked). Severalnewoverloads may declare the same field so, spelling its type identically (an untyped.xincluded), and the name is held to the same uniqueness rule as a body declaration. A@valueor@packableclass holds such a field to its field rules (a type is required)..x?has no declaration to default from there and is a SyntaxError. - Only a
newparameter list accepts the form; anywhere else it is a SyntaxError.
- It is an ordinary parameter named
-
get NAME () { ... }declares a getter — a no-parameter method that is invoked on a bare property read, with no call parentheses:class Circle { new (r) { self.radius = r } get area () { self.radius * self.radius * 3.14 } get name () { "circle" } } let c = Circle.new(4) inspect(c.area) # 50.24 — reads like a field inspect(c.area()) # 50.24 — the call spelling also worksA getter reads
selflike any method but presents as a property, so a fluent chain drops its parentheses (p.parent.namerather thanp.parent().name()). Reserve getters for pure, total, O(1) derivations (an inherent quality of the value); anything that does I/O or can fail should stay an ordinary method, so the absence of parentheses signals "no side effects". A getter takes no parameters (get f (x)is a syntax error), andgetis contextual: a member literally namedget(get () { ... }) is still an ordinary method. Both spellings (obj.nameandobj.name()) invoke the getter identically on the VM, the JIT, and AOT. -
Both engines compile classes. Instance construction is a small runtime call —
newitself is a regular JIT closure whose captures are the method closures plus the user'snewbody, and a runtime helper wires them into the fresh object. -
Well-known methods like
drop(§17) can be written as ordinary class methods. Under the JIT, methods are held on a shared per-class meta object via prototype delegation, but the auto-drop lookup walks the proto chain —class C { drop() { ... } }fires as expected withselfbound to the instance.
Operator overloading
Any Object (whether produced by class sugar or a plain literal)
can participate in arithmetic and comparison by defining special
methods. Dispatch happens at runtime: if the left operand is an
Object with the matching special method, it is called with the
right operand as its sole argument; otherwise the built-in numeric
path runs. Both engines route through the same special-method
protocol, so Object arithmetic compiles.
Classes defined via class sugar participate identically since
their instances are plain Objects with methods attached.
| Operator | Special method | Notes |
|---|---|---|
a + b | __add__ | |
a - b | __sub__ | |
a * b | __mul__ | |
a / b | __div__ | |
a % b | __mod__ | |
a ** b | __pow__ | |
a @ b | __matmul__ | Matrix multiply (PEP 465). Same precedence as *. Has no built-in numeric meaning — operand without __matmul__ raises type error. |
-a | __neg__ | 0-arg method on a |
a == b | __eq__ | != derives by negation |
a < b | __lt__ | >= derives by negation |
a <= b | __le__ | If missing and __lt__ exists, derived as a < b or a == b; > derives by negation |
Example:
class Vec {
new(x, y) { self.x = x; self.y = y }
__add__(r) { Vec.new(self.x + r.x, self.y + r.y) }
__mul__(r) { Vec.new(self.x * r, self.y * r) } # scalar
__eq__(r) { self.x == r.x && self.y == r.y }
}
a = Vec.new(1, 2)
b = Vec.new(3, 4)
c = (a + b) * 2 # Vec(8, 12)
Trait-method fallback. When a class does not define the explicit
__eq__ / __lt__ dunders, the comparison operators fall back to the
Eq and Comparable trait methods: == / != route through an
eq(other) the class writes, and < / <= / > / >= derive from
cmp(other) (the canonical Comparable method). This is what makes a
hand-written eq / cmp class usable with operators directly —
a == b and a < b work without writing __eq__ / __lt__.
Precedence is explicit dunder > trait method > default (structural
equality for ==; a type error for ordering an Object with neither).
Routing == through a written eq also keeps the operator consistent
with the key equality used by Object / Set lookups, which asks the
same eq. A derived eq is not written: it only says the instances
match by their fields, so == compares such a class by structure — the
default — as it does an enum variant (see Deriving methods).
The same precedence holds for every comparison that means "equal", not
the operator alone: the elements of two containers under ==, and the
searches contains and index_of (on an Array, a Tuple, an
iterator).
class Account {
new(id, note) { self.id = id; self.note = note }
eq(o) { self.id == o.id }
hash() { self.id }
}
let a = Account(7, 'as read from disk')
let b = Account(7, 'as typed in')
inspect(a == b) # => true
inspect([a] == [b]) # => true
inspect({owner: a} == {owner: b}) # => true
inspect([a].contains(b)) # => true
inspect([a].index_of(b)) # => 0
Subscripting. A class instance can define __index__(key) and
__setindex__(key, value) so obj[k] and obj[k] = v delegate to it —
a user collection wrapper then subscripts like a built-in. These fire
for keys the object doesn't hold as a direct property (so named-field
access obj["field"] still reads the field); slicing a user type
(obj[a..b]) is not routed through __index__. Both backends dispatch
identically.
class Grid {
new() {
self.d = [10, 20, 30]
}
__index__(i) {
self.d[i]
}
__setindex__(i, v) {
self.d[i] = v
}
}
let g = Grid.new()
g[1] = 99
inspect(g[1]) # => 99
Calling (__call__). A class instance can define __call__(*args)
so obj(args) invokes it — the twin of __index__. obj(x) is exactly
obj.__call__(x), with the instance bound as self. This gives the
model(x) idiom for layered/composable values (a model whose __call__
runs its sub-layers' __call__). Like the subscript hooks it fires only
on class instances, so a plain dict holding a __call__ key stays an
ordinary value. Both backends dispatch identically.
class Adder {
new(b) {
self.b = b
}
__call__(x) {
self.b + x
}
}
let add3 = Adder.new(3)
inspect(add3(10)) # => 13
inspect(add3.__call__(10)) # => 13
A class with __call__ structurally satisfies the Function type, so a
callable instance is a first-class function value: pass it to a
higher-order builtin (map / filter / reduce / sort_by) or bind it
to any Function-annotated parameter, and it is invoked through its
__call__.
class Scale {
new(k) {
self.k = k
}
__call__(x) {
x * self.k
}
}
inspect([1, 2, 3].map(Scale.new(10))) # => [10, 20, 30]
fn apply_twice(f: Function, x) {
f(f(x))
}
inspect(apply_twice(Scale.new(2), 5)) # => 20
__call__ accepts the full call form — positional, *args, and keyword
arguments — so obj(x: 1) binds x against __call__'s parameters like
any method call.
Custom string representation (__str__)
Defining a 0-arg __str__ method on an Object lets the value
supply its own display form for inspect / print / println, string
interpolation ("{x}"), and to_string(x). The method must
return a String; anything else is a type error. Objects
without __str__ use the default formatter ({key: value, ...}).
class Matrix {
new(r, c) { self.rows = r; self.cols = c }
__str__() { "Matrix {self.rows}x{self.cols}" }
}
m = Matrix.new(2, 3)
inspect(m) # Matrix 2x3
inspect("shape: {m}") # 'shape: Matrix 2x3'
to_string(m) # 'Matrix 2x3'
The dispatch is top-level only: when __str__ is involved inside a
nested structure (inspect([m])), the outer formatter's default
recursive walk still uses str() for each element, so the custom
form only appears for the operand passed directly to the display
hook. Deeper customization can be layered on once a concrete need
arises.
__str__ bodies should not recursively invoke inspect(self) or
interpolate "{self}" — there is no built-in recursion guard, so
that form loops until the call stack is exhausted. Produce the
final string via direct property access (self.x) instead.
Auto-reflection
For commutative operators (+, *, ==), when the LHS doesn't carry
the matching special method (e.g., it's a Long or Float) but the
RHS is an Object that does, the call reflects: lhs op rhs becomes
rhs.__op__(lhs).
The rhs-side method receives the scalar as its argument and is
expected to handle it (typically with a match on the argument type):
class Vec { ...
__mul__(r) {
match r {
n: Long => Vec.new(self.x * n, self.y * n), # scalar
_ => Vec.new(self.x * r.x, self.y * r.y) # elementwise
}
}
}
a = 2 * Vec.new(3, 4) # reflects → Vec.__mul__(2) → Vec(6, 8)
Non-commutative operators (-, /, %, **, @, <, <=) do
not reflect — rhs.__op__(lhs) would compute the wrong answer.
For those, the LHS must be an Object carrying the matching special method.
Auto-synthesized parameters()
Every class instance gains a 0-arg parameters() method that returns a
flat Array of every class-instance value reachable from its own
fields. The walker descends through Array elements and plain Object
dicts (no class: tag), and collects each class instance it encounters
as a leaf without recursing into it. Skipped during the walk: the
class: tag itself, and any field whose name starts with _ (treated
as private/cache state). Iteration is in insertion order on both
backends.
class Value { new(x) { self.x = x } }
class GPT {
new() {
self.layers = range(2).map(|_| {
W: range(4).map(|_|
range(4).map(|_| Value.new(0.0)).collect()
).collect()
}).collect()
self.wte = range(8).map(|_| Value.new(0.0)).collect()
}
}
let model = GPT.new()
model.parameters().size() # 40 — every Value collected as a leaf
A class instance is a leaf: the walker stops at it rather than
recursing into its fields. Use plain Object dicts (no class: tag)
for intermediate grouping that should be transparent to the
enumeration, and reserve class instances for the leaves you want to
treat as the actual parameters.
A user-defined parameters method (any property of that name) takes
precedence over the synthesized one — useful when the field walk is
not the intended enumeration. Both engines synthesize the method
through the same walker; output ordering is identical.
Tuples
A Tuple is an immutable, fixed-arity sequence — the hashable
counterpart of Array. Literal syntax uses parentheses with at least
one comma so it doesn't collide with a parenthesized expression:
let pair = (3, 4)
let single = (1,) # trailing comma marks a 1-tuple
let triple = (1, 'two', 3.14)
Element access is by index, same as Array:
pair[0] # 3
triple[-1] # 3.14
Equality is element-wise (and Array is value-equal too, so nesting
recurses):
(1, 2) == (1, 2) # true
([1], 2) == ([1], 2) # true (inner Arrays are value-equal)
Tuples are hashable when their elements are, so they double as Object and Set keys:
let grid = {(0, 0): 'origin', (1, 0): 'east'}
let visited = {(0, 0), (1, 1)} # Set of Tuples
Iteration with for x in t { ... } walks the elements in order. Both
backends implement these semantics identically.
Built-in methods:
| Method | Description |
|---|---|
size() | Element count (Long) |
empty() | Bool — is size() zero? |
presence() | The tuple unchanged if non-empty, else nil |
contains(x) | Bool — is x an element? |
to_array() | Fresh Array with the same elements |
iter() | Iterator yielding elements in index order |
Destructuring
A tuple binding let (a, b) = t unpacks the tuple into named slots.
Element count must match exactly; mismatches throw ValueError. Tuple
patterns also work in match arms and nest naturally:
let (x, y, z) = (10, 'hi', 3.14)
let (p, (q, r)) = (1, (2, 3)) # nested
let kind = match (1, 2) {
(0, 0) => 'origin',
(a, b) => 'other',
}
Parallel / swap assignment. Dropping let reassigns existing
variables instead of declaring new ones — the RHS is fully evaluated
before any binding, so swaps and rotates need no temporary:
mut a = 1
mut b = 2
(a, b) = (b, a) # swap → a == 2, b == 1
(x, y, z) = (y, z, x) # rotate
Each target behaves exactly as it would on its own: an existing binding
is reassigned (an immutable one raises ImmutableError), and a name
that is not visible is declared, just as bare x = v does. Array and
object patterns work too: [p, q, _] = xs, {g, h} = rec.
The pattern is matched against the value before any target is written,
so a shape mismatch (ValueError) writes nothing — [a, 5] = [9, 6]
leaves a alone rather than assigning 9 and then failing on the 5.
An ImmutableError is a different failure: it is raised as that target
is written, so targets to its left already hold their new values.
Index and property targets. A target may also be an index or property chain, so swapping elements needs no temporary either:
(p[i], p[j]) = (p[j], p[i]) # swap two elements
(m[0][1], m[1][0]) = (m[1][0], m[0][1])
(self.front, self.back) = (self.back, self.front)
(a, o.count) = f() # names and chains may mix
A chain target writes through the same rules as the single-target form
(p[i] = v, o.a = v) — including __setindex__ on a class instance,
and including the fact that an immutable binding holding a mutable
container still permits xs[0] = v.
The right-hand side is evaluated once, must be an Array or Tuple,
and must hold exactly one element per target; anything else raises
ValueError. Its elements are read before any target is written, so a
target may safely refer to the right-hand side itself
((p[1], p[0]) = p), and a mismatch writes nothing at all. Targets are
then written left to right, each evaluating its own receiver and index
at that point.
Targets are a flat list: nested patterns and ...rest belong to the
pattern forms above and do not mix with chains. A call is never a
target (f() = v is a SyntaxError).
Sets
A Set is an insertion-ordered collection of unique hashable values.
Literal syntax uses braces with at least two elements (or a trailing
comma) so it doesn't collide with the empty-Object literal {} or the
{key: value} Object shorthand:
let s = {1, 2, 3}
let one = {42,} # trailing comma forces a 1-element Set
let mixed = {1, 'two', 3.14, true}
Duplicate elements collapse on construction. Membership uses the same
key identity as Object keys, which distinguishes types rather than
numeric value — so {1, 1.0, true} keeps all three elements, exactly
as those three are three distinct Object keys.
Equality ignores order:
{1, 2, 3} == {3, 2, 1} # true
Built-in methods:
| Method | Description |
|---|---|
size() | Element count (Long) |
empty() | Bool — is size() zero? |
presence() | The set unchanged if non-empty, else nil |
contains(x) | Bool — is x a member? |
union(b) | New Set of all elements from this and b |
intersect(b) | New Set of elements present in both |
diff(b) | New Set of elements in this but not in b |
sym_diff(b) | New Set of elements in either but not both |
subset(b) | Bool — every element of self is in b |
superset(b) | Bool — every element of b is in self |
to_array() | Fresh Array with the members in insertion order |
iter() | Iterator yielding members in insertion order |
add(x) | Insert x; returns true if newly added |
remove(x) | Remove x; returns true if present |
The set-operation methods preserve the left operand's insertion order
for elements that survive. .add(x) and .remove(x) may shadow a
user-defined Object method of the same name (e.g. a Calculator.add(1)
class method), so both backends emit a runtime tag dispatch at the
call site: Set receivers go to the set primitive, Object receivers go
to the user property. Works on the VM and JIT (and AOT).
Set operations are method-only: union, intersect, diff,
sym_diff. There are no | / & / - / ^ operator forms —
the | close delimiter of lambda parameters made the operator
ambiguous to a stateless PEG parser, so all four operations route
through methods for consistency. Operators otherwise stay reserved for
math types (Long, Float, Tensor, user numeric classes via __add__
etc.); the one exception is +, which also concatenates strings (§8)
and Arrays (§9). A Set has no + — to_array() first, or union.
11. Functions and closures
Function literals
fn () { nil }
fn (x) { x + 1 }
fn (mut x) { x = x + 1; x }
fn (a: Long, b: Long) -> Long { a + b }
fn (name, greeting = 'hello') { "{greeting}, {name}" }
Functions are first-class values and the only way to define a reusable
piece of code. Bind a literal like any other value
(name = fn (...) { ... }), or use the declaration form
fn name(...) { ... }. Only the declaration form can be a
generator or take part in multimethod
dispatch, and only it gives fn.name a source-level
name.
Lambda sugar |x| ...
A lighter form is available for the common case of passing a tiny function to a higher-order call:
add = |x, y| x + y # expression body
sq = |x| x * x
noop = || 42 # zero params
scale = |x, y| { # block body
let factor = 10
x * y * factor
}
abs_v = |x| if x < 0 { -x } else { x } # if/while/for/match/try
# are expressions, so they
# work as the body
xs.map(|x| x * 2) # passes cleanly as a functor
The body is normally a single expression. For multiple statements,
intermediate let, or side effects, use a braced block:
clamp = |v, lo, hi| {
mut x = v
if x < lo { x = lo }
if x > hi { x = hi }
x
}
Both forms are semantically identical to fn (...) { ... }:
- Captures variables from the enclosing scope.
- Accepts the same
mut/ type-annotation / default-value parameter forms:|mut x = 10| x + 1works just likefn (mut x = 10) x + 1. - Self-reference via
fnworks the same as in a named function. - No return-type annotation slot (keep lambdas short; use
fnfor functions whose signature deserves the annotation).
The | delimiter is not ambiguous with pattern alternation (which
only appears inside match arms) or logical || (which never starts
an expression).
Parameters
-
muton a parameter makes that parameter binding mutable inside the body. Withoutmut, reassigning the parameter raisesimmutable variable. -
Optional type annotations are enforced on entry (§14).
-
A parameter may have a default value via
name = expr. When the caller omits the argument, the default is evaluated on each call in the function's definition environment, extended with the bindings the frame makes before its parameters — the receiverself, the recursion handlefn, and any earlier parameter — so bothfn (a, b = a + 1)andm(k = self.n)work. In a constructor the instance already exists but its field initializers have not run yet, soself's fields readnilthere. Default parameters must follow all required parameters. -
A parameter may be a destructuring pattern —
fn ({x, y}),fn ([a, b]),fn ((k, v)), and the lambda form|{a, b}|. The argument is matched against the pattern at entry and its names are bound in the body; a shape mismatch raisesValueError. Patterns nest (fn ({user: {name}})) and mix with normal params (fn (factor, {x, y})).fn dist({x, y}) { x * x + y * y } dist({x: 3, y: 4}) # → 25 -
A final parameter written
*nameis a positional catch-all: it collects every positional argument beyond the regular params into anArray(empty when there are none). It must be the last parameter.fn f(first, *rest) { [first, rest] } f(1, 2, 3) # → [1, [2, 3]] f(1) # → [1, []]A
*argsdeclaration also opts the function into variadic dispatch: it matches any call with at least the regular-param count, but a fixed-arity overload always wins a tie, and among variadic candidates the one with more regular params is the more specific.fn h(x: Long) { "exact" } fn h(*xs) { "variadic" } h(1) # → 'exact' (fixed-arity wins) h(1, 2) # → 'variadic'
Keyword arguments and ** splat
A call site may pass arguments by name (name: value) and/or expand
an Object as kwargs (**obj). Forms:
let f = fn (x, y = 10, z = 100) { x + y + z }
f(1, 2, 3) # positional
f(1, z: 5) # kwargs (y defaults)
f(z: 3, y: 2, x: 1) # any order
let opts = {y: 7, z: 8}
f(1, **opts) # splat an Object as kwargs
f(1, **opts, z: 100) # explicit kwarg overrides splat
f(1, **{y: 2}, **{y: 5}) # multiple splats: later wins
Catch-all **rest collects every kwarg that isn't claimed by a
declared parameter into an Object, bound to the rest parameter's
name. It must be the last parameter.
let route = fn (path, **opts) {
route_internal(path, opts.method, opts.headers, opts.body)
}
route('/users', method: 'GET') # opts = {method: 'GET'}
route('/users', method: 'POST', body: '…') # opts = {method, body}
Keyword-only parameters: a bare * in the parameter list marks the
boundary; parameters after it can only be passed by name.
let g = fn (x, *, y, z = 10) { x + y + z }
g(1, y: 2) # 13
g(1, y: 2, z: 3) # 6
g(1, 2) # TypeError: takes 1 positional argument but 2 given
g(1) # ArityError: missing required argument 'y'
Rules:
- Positional arguments must come before any kwarg or splat. Mixing
the other way (
f(x: 1, 2)) is aSyntaxError. - The same name may not appear twice as an explicit kwarg.
- A name cannot appear both positionally and as a kwarg.
- Unknown kwarg names are a
TypeErrorunless the callee declares a**restcatch-all, which absorbs them. - A
**splat operand must be anObjectwithStringkeys only. - Defaults are re-evaluated on every call, so a mutable default
(
fn f(xs = [])) is a fresh value each time rather than one accumulating across calls.
A keyword may skip a defaulted parameter: f(1, z: 3) against
(x, y = 10, z = 20) fills y from its default.
Standard library functions. A namespace function or bare global
function of the standard library (Math.clamp, JSON.stringify,
Canvas.rect, to_string) binds keyword arguments and ** splats by
the same rules, under the parameter names its reference entry prints —
a required parameter included:
Math.clamp(x: 5, lo: 0, hi: 3) # → 3
Canvas.rect(**{x: 1, y: 1, w: 6, h: 6, color: c})
Math.clamp(x: 5, lo: 0) # ArityError: missing required argument 'hi'
A function that collects *args (Math.max) has no names to bind and
rejects a keyword with a TypeError. The methods of the built-in value
types bind by position instead (see "Built-in methods bind
positionally").
Compile-time errors (e.g. positional argument follows keyword argument) are detected during compilation and bypass try/catch on
every engine.
Return
- The body is a block. The last expression of the block is the function's return value.
return exprreturns early;return(no expression) returnsnil.- If a return type is declared (
-> T), the returned value is checked againstTbefore control leaves the function, both for natural fallthrough and explicitreturn.
Recursion: fn
Within a function body, fn refers to the function value currently
being executed. Recursion without giving the function a name:
fib = fn (x) { if x < 2 { x } else { fn(x - 1) + fn(x - 2) } }
Methods: self
self is bound for the duration of a method call. Outside a method
call, self is not in scope; accessing it raises
undefined variable 'self'.
Closures
Function literals capture their lexical environment by reference.
- Variables read inside a nested function look up the scope chain until found.
x = vinside a nested function reassigns an outer binding ifxexists in an outer scope; otherwise creates a new local.- Captured mutable variables (e.g., via
mut xin an outer scope) are shared: changes are visible to every closure that captured them and to the outer scope itself — capture is by reference, not by value.
Example:
make_counter = fn () {
mut n = 0
fn () { n = n + 1; n }
}
c1 = make_counter()
c2 = make_counter()
inspect(c1()) # 1
inspect(c1()) # 2
inspect(c2()) # 1, independent
In the JIT, captured mutable variables are allocated in heap cells so that multiple closures can share the same slot. See §17.
Generators (yield)
A fn declaration whose body contains yield is a generator
function. Calling it does not run the body: it returns an Iterator
(§18.5), and the body advances one yield at a time as the consumer
pulls values out of it.
fn counter() {
yield 1
yield 2
yield 3
}
inspect(counter().collect()) # => [1, 2, 3]
yield expr is a statement, not an expression. It hands a value to
the consumer and evaluates to nothing itself, so let x = yield 1 is a
syntax error. Values travel out of a generator only — there is no
send() and no way to resume a suspended body with a value.
yield from expr delegates to any iterable — an Array, a String,
a lazy chain, or another generator — yielding each of its elements
before the delegating body resumes:
fn inner() {
yield 'x'
}
fn mixed() {
yield from [1, 2]
yield from inner()
yield 'done'
}
inspect(mixed().collect()) # => [1, 2, 'x', 'done']
Recursive traversals compose from it, which is the usual reason to reach for delegation:
fn walk(node) {
yield node.value
for kid in node.kids {
yield from walk(kid)
}
}
let leaf = {value: 3, kids: []}
inspect(walk({value: 1, kids: [{value: 2, kids: [leaf]}]}).collect()) # => [1, 2, 3]
The generator object. What the call returns is an ordinary
Iterator: it carries iter / has_next / next / dispose, so it
drives for-in, the lazy combinator set (§18.5), and any other
consumer of the protocol. Suspension makes unbounded sources practical:
fn nat() {
mut i = 0
while true {
yield i
i += 1
}
}
inspect(nat().map(|x| x * x).take(5).collect()) # => [0, 1, 4, 9, 16]
The body starts on the first has_next(), not at the call. A generator
is one-shot: once drained it stays exhausted, and iterating the same
object again produces nothing. Call the function again for a fresh run.
Inside the body, yield may appear at any depth in while, for,
if and { } bodies, and break / continue / return behave as they
do in a normal function — return ends the generation. A defer runs
when its scope is left, as in a normal function (§15): a loop body's
defer on every iteration, the body's own when the body finishes. A
for-in in the body closes its iterator the same way.
fn rows() {
for id in [1, 2] {
defer {
inspect("close {id}")
}
yield id
}
}
for r in rows() {
inspect(r)
}
# => |
# 1
# 'close 1'
# 2
# 'close 2'
A generator can also stop while suspended: the consumer leaves its
for-in by break, return or an exception, a terminal method
finishes early (§18.5), dispose() is called, or the last reference to
the generator goes away. The defers still pending run then, innermost
first — cleanup does not depend on the body reaching its end.
fn two() {
defer {
inspect('closed')
}
yield 1
yield 2
}
for v in two() {
inspect(v)
break
}
# => |
# 1
# 'closed'
Restrictions.
-
yieldmay not appear inside atry/catchor adeferblock. The parser rejects it withSyntaxError: yield cannot appear inside a try-catch or defer block.To guard a yielded value, put thetryin the expression (yield try { ... } catch e { ... }); to clean up, use adeferas above. -
Only
fn name(...) { ... }declarations are transformed into generators — at the top level or nested inside another function. Ayieldanywhere else — in a class method, in an object property's function, in afnexpression assigned to a variable, or at the top level of a file — is rejected at parse time:SyntaxError: yield can only appear inside a `fn name(...) { ... }` declaration body — a class method, an object property's function, or a fn expression cannot be a generator. Declare a named fn and call it instead.The check runs over what the transform pass leaves behind, so every backend rejects the same programs at the same position.
-
selfmay not be referenced in a generator's body. The body is lowered into methods of a synthesized state class, so a bareselfthere could only name that internal object — never a receiver (a generator is a named fn and cannot be a method). The parser rejects it:SyntaxError: self is not available inside a generator body (a function that uses yield) — bind it outside first (let me = self) and use that variable, or pass it as a parameter.The rule reaches a
fn/ lambda defined in the body, which would otherwise read that state object through the enclosing method'sself. One shape is exempt: a function that is an object property (yield {m: fn () { self.x }}), whoseselfis the dynamic receiver of the object it is called on (§10), never the state object. A property name, object key, or kwarg label spelledselfis not a reference either. To reach an enclosing receiver, capture it first:let me = selfoutside the generator, then usemeinside. Aneffect fnbody is the same shape and refuses the same way; ahandlebody is not — it is spliced where it was written, so an enclosing method'sselfis still that method's receiver (§16). -
A body local keeps plain-variable semantics even though the lowering stores it on the state object: a local holding a function is a value, not a method of that object, so
f == fstays true, calling it (f()) passes no receiver exactly as it would outside a generator, and passing it on leaves the receiver to whatever call follows (holder.f = f, thenholder.f()seesholder). The generator's own protocol methods are not locals and bind as usual (§10). Aneffect fnbody and ahandlebody promote their locals the same way, so this rule holds there too (§16) — it is only theselfrule above that treats the two differently. -
A generator body cannot
performa bare effect operation or declare aneffect fn; a self-containedhandle { ... }expression inside the body does work (§16).
Generators are compiled by a source-level transform shared by every lane, so an identical program yields identical values under the VM, the JIT, and an AOT binary.
12. Control flow
if
if cond { then_block }
if cond { then_block } else { else_block }
if c1 { b1 } else if c2 { b2 } else { b3 }
if init; cond { … }
cond must be truthy-convertible (Bool or Long). if is an
expression; its value is the taken branch's last expression, or nil
if no branch is taken (no else and the if was false).
Like while, if accepts an optional init clause — declarations
before the condition, split by ; — scoped to the whole if / else if / else chain (the same form as while):
# doctest: skip
if mut d = compute(n); d > threshold {
use(d) # d is in scope here …
} else {
fallback(d) # … and here, but not after the chain
}
Each binding must be a declaration (let / mut); a bare if x = 0; … is a SyntaxError. Multiple bindings use ,
(if mut a = f(), mut b = g(); …).
Statement modifiers (if / unless)
stmt if cond
stmt unless cond
A trailing if / unless (Ruby's statement modifiers) runs stmt
only when cond is (if) or is not (unless) truthy — unless is
exactly if !cond. It desugars to an ordinary if cond { stmt } /
if !cond { stmt } before either backend sees it, so it is nil
when the condition doesn't hold, same as a bodyless if:
mut i = 0
while i < 10 {
i = i + 1
break if i > 5
continue unless i % 2 == 0
println(i)
}
# => 2
# => 4
The modifier attaches to a whole statement, not an arbitrary
expression — f(1 if true) is a SyntaxError, exactly as in Ruby.
Its value only shows through when the modified statement is the last
one in a block or function body:
fn grade(v) {
return "small" if v < 10
return "big" unless v < 100
"medium"
}
inspect(grade(5)) # => 'small'
inspect(grade(500)) # => 'big'
unless is a reserved word only at the assignment-target position,
like the rest of the hard-reserved keywords — it stays a
valid parameter name, object key, and property name.
cond
cond { test => body, ..., _ => default }
The subjectless multi-way conditional (Elixir's cond, Kotlin's
argless when): arms are tried top to bottom, and the value of the
first arm whose test is truthy becomes the value of the whole
expression. _ is the always-match wildcard, conventionally last —
arms after a matched _ never run. If no arm matches, the value is
nil (like an unmatched match).
fn grade(n) {
cond {
n >= 90 => 'A',
n >= 80 => 'B',
_ => 'C',
}
}
inspect(grade(85)) # => 'B'
A test follows the same truthiness rule as if (Bool or Long,
zero falsy), and may be any expression — calls, &&/||,
comparisons. An arm body follows the same
expression-or-block rule as match arms: a bare
expression, or a brace block that yields its last statement's value
and forms its own scope.
Use cond in place of a match true { _ if … => ... } guard chain
when there is nothing to match on — it reads as a priority-ordered
list of conditions rather than a match against a dummy subject.
while
while cond { body }
while init; cond { body }
while is a statement; its value is nil. break and continue
work inside the loop body. The body is a fresh scope per iteration
(like for — see Scope).
An optional init clause — a comma-separated list of declarations
before the condition, split by ; — binds variables scoped to the
loop, so a counter no longer leaks into the enclosing scope:
while mut i = 0; i < len {
out.push(i)
i = i + 1
}
# `i` is not visible here
The init variables persist across iterations (a body i = i + 2
re-assigns) and are dropped when the loop exits by any path (normal,
break, or an exception). Multiple bindings use ,:
while mut i = 0, mut j = xs.size() - 1; i < j { i = i + 1; j = j - 1 }
Each binding must be a declaration (let or mut); a bare
while x = 0; … is a SyntaxError (it would reassign an outer x
rather than scope one to the loop). Destructuring binds too:
while mut (a, b) = pair; ….
for ... in
for var in iterable { body }
Iterates by calling iterable.iter() once, then driving the returned
iterator with has_next() / next(): while has_next() is true,
next() produces the element bound to var in a fresh scope per
iteration (see §18.5).
for x in [1, 2, 3] {
inspect(x)
}
for k in {b: 2, a: 1} {
inspect(k)
} # keys, ascending
for i in 0..10 {
inspect(i)
} # exclusive (0..9)
for i in 0..=10 {
inspect(i)
} # inclusive (0..10)
Range values a..b (exclusive) and a..=b (inclusive) iterate the
same lazy integer sequence as range. A bounded range (both Long
endpoints present) is iterable; an open-ended range used for slicing
(xs[2..]) has no iteration end and raises if iterated.
An endpoint may also be a Float. Such a range is a numeric interval
rather than a sequence: it can be built, displayed and compared, but
iterating it, slicing with it or passing it to grid raises
TypeError (expected Long, got Float) where that happens. An
endpoint that is not a number raises TypeError where the range is
built.
inspect(0.0..0.86) # => 0.0..0.86
inspect(..0.5) # => ..0.5
inspect(1..3 == 1.0..3.0) # => true
A range takes an optional by <step> clause to iterate by something
other than 1, including descending (step negative). The step is
always a Long:
for i in 0..10 by 2 {
inspect(i)
} # 0, 2, 4, 6, 8
for i in 10..0 by -2 {
inspect(i)
} # 10, 8, 6, 4, 2
step must not be 0 (raises ValueError when the range is
iterated). Slicing (xs[a..b by n]) ignores step — it only affects
iteration.
Building one from computed parts — Range(start, end, inclusive = false, step = 1). The literal states in syntax what it cannot state at runtime:
which end is open, and whether the range is inclusive. The constructor takes
both as values, and nil for an endpoint is the open form (a.., ..b,
..). A zero step raises ValueError, as the literal's does.
inspect(Range(1, 5)) # => 1..5
inspect(Range(1, 5, true)) # => 1..=5
inspect(Range(1, 9, false, 2)) # => 1..9 by 2
inspect(Range(2, nil)) # => 2..
inspect(Range(1, 5) == 1..5) # => true
A range is what a..b or Range(...) built, not whatever wears its fields:
an Object carrying the same keys is an ordinary Object, and slicing,
iterating or passing it to grid raises rather than treating it as a range.
Membership — r.contains(x) tests the interval: start <= x and
x < end (x <= end for ..=), where an open end bounds nothing. A
Long and a Float compare the way < compares them. An argument that
is not a number, and NaN, lie outside every range, so the call never
raises for its argument. A range with a by step other than 1 raises
ValueError.
inspect((0..10).contains(5.5)) # => true
inspect((0.0..0.86).contains(0.86)) # => false
inspect((..=0.5).contains(0.5)) # => true
inspect((0..10).contains('a')) # => false
Destructuring loop variable. The var may be a pattern, matched
against each element's shape (a mismatch raises ValueError).
Comma-separated targets without parens are sugar for a tuple pattern:
for k, v in xs means for (k, v) in xs. The bracket form [a, b] and
the tuple/bare form (a, b) / a, b are interchangeable — either matches
any indexed sequence (Array or Tuple) of the right length, so the
pattern's punctuation is style, not a type constraint. This holds for
assignment destructuring too: [a, b] = (1, 2) and (a, b) = [1, 2]
both bind.
# doctest: skip
for [a, b] in [[1, 2], [3, 4]] { inspect(a + b) } # array pattern
for (k, v) in [(1, 'a'), (2, 'b')] { inspect(k) } # tuple pattern
for k, v in {a: 1, b: 2} { inspect("{k}={v}") } # bare comma == (k, v)
for i, v in xs.enumerate() { ... } # (index, value) tuples
The iterator protocol (see §18.5) requires the target to be either an
Object (or subtype Array) with an iter method, or an object
already playing the iterator role with has_next / next methods.
Passing any other type raises type error.
Iterating a Set walks a snapshot of the members taken when the iterator
is created: mutations made by the loop body change the set but not the
walk. An Object snapshots only its keys — removed keys are skipped and
values are read live (see §18.5). An Array is read live, one index per
step.
for is a statement; its value is nil. Shadow rules apply to
var: if it would shadow a closure-captured name from an enclosing
function, the script is rejected (see §6).
JIT: for / break / continue compile under --jit for
direct iteration over Array, Object (yields (key, value) pairs in
insertion order), and String (UTF-8 scalar walk). Objects that carry their
own iter property (user-defined iterators, range,
String.code_points() / .graphemes(), iterator method chains) are
driven through the iterator protocol at runtime, with a native-loop
fast path preserved for the Array/String/keys cases.
break and continue
break # exit the innermost enclosing loop
continue # skip to the next iteration of the innermost loop
Valid only inside for or while. Using them outside a loop is a
SyntaxError, checked before the program runs. break / continue do
not carry a value (the loop's value remains nil).
Both are expressions (§3), so they need no statement position of
their own — a match arm (')' => break), a cond arm, or either side
of a ternary takes one directly. See Arm bodies for the
arm form, and Diverging expressions for what
this means in the middle of a larger expression.
Labels. A while or for may carry a label, and break / continue
may name one — the jump then leaves (or advances) that loop rather than
the innermost:
label: while cond { body }
label: for var in iterable { body }
break label # exit the loop named `label`
continue label # next iteration of the loop named `label`
let grid = [[1, 2], [3, 4], [5, 6]]
mut found = nil
search: for row in grid {
for cell in row {
if cell == 4 {
found = cell
break search
}
}
}
inspect(found) # => 4
The label must sit on the same line as its break / continue, so a bare
break followed by a line that starts with an identifier stays two
statements. A label reaches only the loops of the function it is written
in — a break outer inside a nested fn or a defer block is the same
SyntaxError a bare break there is. Naming a label no enclosing loop
carries, or reusing a label an enclosing loop already carries, is a
SyntaxError checked before the program runs.
A label is not a binding: the name stays available for a variable, and the same label may name a loop in a sibling nest.
Leaving a nest by label is the same exit as leaving one loop at a time —
the abandoned loops run their defer blocks (innermost first) and close
their iterators before the jump lands.
nobreak (loop-else)
while cond { body } nobreak { … }
for var in iterable { body } nobreak { … }
An optional nobreak { … } block after a while or for runs only
when the loop finishes normally — the condition became false, or the
iterator was exhausted — and is skipped when the loop exits via
break (and also by return or a throw, which unwind past it). These
nobreak names the condition it tests — it runs when no break
happened — so the block needs no comment to say when it fires. It
carries no value (the loop stays nil).
The canonical use is search — the block is the "not found" branch, with no flag variable:
fn find(xs, target) {
for x in xs {
if x == target {
return "found"
}
} nobreak {
return "not found" # only reached if the loop never broke
}
}
A while init clause is in scope in its nobreak block (the counter
survives to the post-loop step); a for loop variable is not (it is
per-iteration and already gone). A break / continue inside a
nobreak block belongs to an enclosing loop, since the block runs
after this loop has finished — including by label, so a labelled
break there leaves the nest the way it does anywhere else. A loop
abandoned by a labelled break from inside its body does not run
its own nobreak block: it was broken out of. nobreak is a contextual
keyword —
recognized only in this trailing position, so it stays usable as an
ordinary identifier elsewhere.
return
Valid only inside a function body. Exits the enclosing function with
the given value (or nil). return outside any function is a
SyntaxError (checked before the program runs) — a script ends at its
last statement, or exits early via Sys.exit.
Like break and continue, return is an expression, so
_ => return x is a legal match arm.
Diverging expressions
return, throw, break and continue share a property: none of them
produces a value. Reaching one transfers control elsewhere, so the
expression it sits inside never finishes evaluating. That is why all
four are expressions rather than statements — there is no value for the
surrounding expression to be given, and so no position where one would
be ill-typed:
# doctest: skip
let width = "v" + match n {
# the `+` never completes when n is 4
4 => break,
_ => "x",
}
Consequences worth knowing:
- Operands evaluated before the diverging one are discarded. In
f(a(), match n { 0 => break, _ => b() }),a()has already run when thebreakis taken; the call tofnever happens. deferblocks still fire on the way out, innermost first, exactly as they would for the statement spelling (§15).break/continuebind to the nearest enclosing loop, never to thematchorcondthey appear in — those are expressions, not control-flow scopes, so there is no "exit the match" reading to confuse them with.- The check that a
breakhas an enclosing loop, and areturnan enclosing function, runs before the program does and is aSyntaxErrorin expression position just as in statement position.
Being an expression does not make one bind tighter. return's operand
is a whole expression in either position, so return 1 + 2 returns 3
— in a match arm exactly as at the head of a statement.
debugger
When the program is run under --debug, encountering debugger
pauses execution and drops into a simple REPL debugger showing the
current source line. Without --debug, debugger is a no-op.
13. Pattern matching
match subject {
pattern1 (if guard1)? => body1,
pattern2 (if guard2)? => body2,
...
}
match is an expression. Arms are tried top-to-bottom; the first arm
whose pattern succeeds (and whose guard evaluates truthy) runs its
body and the result is the value of the match. If no arm matches,
the value is nil.
Like while and if, match accepts an optional init clause —
declarations before the subject, split by ; — scoped to the subject
and every arm (the same form as while):
match mut x = compute(n); x {
0 => "zero",
v if v > x - 1 => v, # init var `x` is in scope in guards …
_ => x # … and arm bodies, but not after the match
}
Each binding must be a declaration (let / mut); a bare
match x = 0; … is a SyntaxError. Multiple bindings use ,
(match mut a = f(), mut b = g(); a + b { … }). The init variables are
dropped when the match is left by any path (a matched arm, the no-match
fall-through, or an exception).
Patterns
| Form | Matches |
|---|---|
0, 'x', "x", nil, true | Literal equality (a string pattern is '...', a backtick string, or a "..." with no interpolation) |
-1, -2.5 | A negative numeric literal: equality, of the same type, like 1 |
lo..hi, lo..=hi, ..hi, lo.. | A Long or Float inside the interval, as Range#contains decides; the bounds are numeric literals |
name | Any value, binds it to name |
_ | Any value, no binding |
name: Type | Value whose type is Type; binds |
p1 | p2 | p3 | Any of the sub-patterns matches; alternatives cannot bind |
[p1, p2, ...] | Array of exactly the same length |
[p1, ...rest] | Array of ≥ n−1 elements; rest is a fresh Array of the remainder |
[a, ...m, z] | Rest can be in the middle; pre/post positions match fixed elements |
{k1, k2} | Object containing at least those keys; binds each ki to obj.ki (shorthand) |
{k1: p1, k2: p2} | Object whose k1 matches p1 and k2 matches p2. Nests freely ({user: {name}}). Mixes with shorthand. |
{} | Any Object (keys ignored) |
(p1, p2, ...) | Tuple of exactly the same arity; element-wise sub-patterns |
Ok(p1, ...) | Enum constructor: matches the variant Ok and destructures its positional payload against the sub-patterns. The qualified form Result.Ok(p) also requires the value's enum to be Result — how two enums that each declare an Ok are told apart. See "Sum types". |
Semantics
- A literal pattern must be a compile-time constant. An interpolating
"...{x}..."is a SyntaxError — match against the runtime value with a guard instead (s if s == x => ...). - A range pattern matches the numbers
(lo..hi).contains(x)accepts. ALongand aFloatcompare the way<compares them, so unlike a literal pattern it does not care which of the two the subject is:match 5.5 { 0..10 => … }matches,match 1.0 { 1 => … }does not. A subject that is not a number, or isNaN, falls through without raising. The bounds are numeric literals, optionally negated, and take noby. Leave a space in5.. =>:5..=>reads as..=.
kind = fn (x) {
match x {
-1 => 'minus one',
..0 => 'negative',
0..=9 => 'digit',
_ => 'other',
}
}
inspect(kind(-1)) # => 'minus one'
inspect(kind(-0.5)) # => 'negative'
inspect(kind(3.0)) # => 'digit'
inspect(kind(9.5)) # => 'other'
inspect(kind('a')) # => 'other'
- Patterns are tried left-to-right, and sub-patterns are evaluated depth-first.
- A binding
nameintroduced by the pattern is visible in the guard and the body. |(or) sub-patterns cannot bind: a name inside an alternative would exist only on the paths that took it, so any binding there is aSyntaxError(a | _,5 | a,Ok(x) | Err(x)). Write literals and_inside|, and one arm (or one pattern) per binding shape.Arraypatterns require an exact length unless a...restelement is present; objects do not require exact key sets — extra keys are ignored.- The rest array is a newly-allocated
Arraywith elements from the subject shallow-copied; mutating it does not affect the subject, but mutating a reference-typed element mutates the shared object. - The
matchsubject is evaluated exactly once.
Arm bodies
An arm body is normally a single expression. To run several statements, write a brace block — the arm evaluates to the block's last statement value:
let kind = match tok {
'+' => { let p = prec(tok); register(p); "op" },
_ => "atom",
}
A block arm is its own scope: pattern bindings and any let inside it are
visible only within the arm, and a defer fires when the arm's braces
close (LIFO, before the arm value is consumed). return / break /
continue inside a block arm behave as they would anywhere — and still run
the arm's pending defers on the way out.
An arm that does nothing but transfer control needs no block, since all four control-transfer forms are expressions (§12):
match tok {
')' => break,
'!' => throw "unexpected",
' ' => continue,
_ => tok,
}
The arm parser tries an expression first, so a brace that is a valid literal keeps its literal meaning, not a block:
match x {
0 => {}, # empty Object
1 => {a: v}, # Object literal
2 => {p, q, r}, # Set literal (two or more elements)
_ => { f(); g() } # block: runs f() then g(), yields g()'s value
}
One sharp edge falls out of object shorthand: a brace wrapping a single
bare identifier, { v }, is the object {v: v}, not a block yielding v.
Write _ => v (no braces) for that, or { v; } to force a block. Shorthand
reads a lone { break } / { continue } the same way, and since no variable
may be named break, that arm raises NameError — write the bare
_ => break, which is what you meant.
Exhaustiveness
No static exhaustiveness check is performed. If no arm matches, the
match yields nil. Add _ => nil or a typed fallback explicitly
if that bothers you.
14. Optional type annotations
Culebra is dynamically typed; type annotations are optional and enforce their invariant at three specific runtime points:
- On entry to a function, each parameter's value is checked against its annotation.
- On function return (either fallthrough or explicit
return), the result is checked against the declared return type. - On assignment with an annotated target, the RHS value is checked against the annotation.
Syntax
let x: Long = 10
let mut s: String = 'hi'
fn (a: Long, b: Long) -> Long { a + b }
Recognized type names
Nil Bool Long Float String Array Object Function Any
Any always matches. Unknown type names fail the check and raise
type error. Class names declared with class C { ... } are also
valid annotations and accept any instance of that class. An enum name
accepts any of its variants, a variant name accepts that variant of any
enum, and the qualified Enum.Variant accepts only that enum's — see
"Sum types".
Union types
An annotation may list alternatives separated by |. The runtime
check accepts the value when it matches any alternative.
fn show(x: Long | Float) -> String { to_string(x) }
show(1) # → '1'
show(2.5) # → '2.5'
show("hi") # !! type error
fn lookup(k: String) -> Long | Nil {
if k == "answer" { 42 } else { nil }
}
Union annotations are valid wherever a single type annotation is —
parameters, return types, and let / let mut declarations:
let id: Long | String = "u-42"
let mut count: Long | Nil = nil
count = 0 # OK: re-check is *not* run on reassignment
Whitespace around | is tolerated (Long|Float, Long | Float).
Class names compose with primitives (Square | Circle,
String | Nil). Single-alternative annotations remain unchanged in
behavior — Long and Long| are not equivalent (the latter is a
parse error).
Object inside a Union is a catch-all for class instances
(any value carrying a class: tag), not for every value: primitives
fall through it. So Long | Object accepts a Long and any class
instance, but a String still fails. Use Any if you want to accept
everything.
If any alternative is an undeclared class name, the runtime treats it the same as a class annotation that no instance can satisfy — the alt simply never matches, so it's effectively dead. Other alternatives in the Union still match normally.
Multimethod dispatch (§20) understands Union
parameter annotations by scoring each alternative and taking the
best match — fn area(s: Square | Circle) dispatches on either
exact class. A bare concrete type outranks any Union that contains
it: defining both fn pick(x: Long) and fn pick(x: Long | Float)
routes a Long arg to the concrete pick, while a Float still goes
through the Union version.
Optional types (T?)
A trailing ? on a type name is sugar for T | Nil — it makes nil an
accepted value while still enforcing the base type for non-nil values.
fn id(x: Long?) -> Long? { x }
id(5) # → 5
id(nil) # → nil
id("s") # !! rejected (not Long, not nil)
? works on any type name, including Generic outers (Array<Long>?)
and in let / parameter / return annotations (let a: String? = nil).
Introspection canonicalizes it to the Union form: fn.params[0].type
of x: Long? reads "Long | Nil". Pair with the null-safe operators
below (?., ?[], !!, ??, ??=) for ergonomic nil handling.
Function types (fn(T) -> U)
The bare Function type accepts any callable. To document a
higher-order parameter's shape — what it takes and returns — write a
function type: fn(T1, T2) -> R.
fn apply(f: fn(Long) -> Long, x: Long) -> Long { f(x) }
apply(|n| n * 2, 21) # → 42
fn make_adder(n: Long) -> fn(Long) -> Long { |x| x + n }
let add5 = make_adder(5)
add5(37) # → 42
The parameter list may be empty (fn() -> String), hold several types
(fn(Long, Long) -> Long), or nest (fn(fn(Long) -> Long) -> Long).
A callable class instance (one with a __call__ method) satisfies a
function type just like a closure does:
class Doubler { new() {} __call__(n) { n * 2 } }
apply(Doubler.new(), 21) # → 42
Like Generic element types, the parameter and return types are documentation in the MVP — the runtime checks only that the value is callable, not its arity or the types it actually accepts. A non-callable is rejected:
apply(99, 21) # !! type error: ... expects fn(Long) -> Long
The return is a single type, so a top-level | after -> belongs to
the surrounding Union: fn(A) -> B | C parses as (fn(A) -> B) | C (a
Union of a function returning B and C). An optional return uses
? on the return type — fn(A) -> B? is a function whose result is
B | Nil; the function itself is still required:
fn run(f: fn(Long) -> Long?) -> Long { f(0) ?? -1 }
run(|n| nil) # → -1
run(nil) # !! type error (a function is required)
Multimethod dispatch (§20) treats a function-type
parameter like Function: a closure or __call__ instance routes to
the fn(...) -> ... overload, while a concrete type (Long) routes to
its own overload.
Generic types
A type name may carry type parameters in angle brackets:
Array<Long>, Array<Array<Long>>, Array<Long | Float>. Nesting
is supported, and Generic args may themselves be Unions. The only
built-in container type today is Array; other Generic outer names
(Box<T>, Pair<K, V>, ...) come from user-declared classes
(see "Generic class declarations" below).
fn first(xs: Array<Long>) -> Long { xs[0] }
fn lookup(k: String) -> Array<Long> | Nil { ... }
Element-level runtime checks are no-ops: only the outer type is checked at the boundary — verifying the element type would mean walking the collection on every call. The args exist for documentation and for multimethod dispatch tie-breaks.
fn first(xs: Array<Long>) { xs[0] }
first([1, "two", 3]) # OK — element type is not enforced
Whitespace inside <> is tolerated and canonicalized:
Array < Long >, Array< Long > and Array<Long> all surface
as "Array<Long>" in fn.params[i].type.
Multimethod dispatch tie-break: a Generic param is more specific
than its bare outer-only form. With both
fn show(xs: Array) and fn show(xs: Array<Long>) defined, an
Array arg routes to the Generic version.
Where class declarations may appear
Classes are top-level constructs. class C { ... } is rejected
when it appears directly inside another class's body — declarations
must live at top level, or inside a fn / lambda / lexical-scope
({ ... }) block that re-opens a fresh scope. The restriction
keeps the mental model flat; namespacing is handled by import
and export, and inner helpers by free functions / UFCS.
class Bad {
m() {
class Inner { ... } # !! SyntaxError
}
}
# OK — fn body is a scope boundary
fn make_inner() {
class Inner { ... }
Inner.new()
}
Generic class declarations
A class can declare type parameters:
class Box<T> {
new (v: T) { self.v = v }
}
class Pair<K, V> {
new (k, v) { self.k = k; self.v = v }
}
An unbounded type parameter is documentation — it makes method
signatures readable but the runtime sees it as Any. A bounded
parameter (<T: Comparable>) is enforced; see "Bound constraints"
below. The class is bound under the outer name (Box, Pair), and
instances carry the outer class tag:
let b = Box.new(42)
type_of(b) # → 'Box'
Annotations referring to a Generic class use the same <...>
syntax as built-in Generic types and behave the same way (outer
match):
fn unbox(b: Box<Long>) -> Long { b.v }
Bound constraints
A type parameter may carry a bound — a trait the bound type must
conform to: <T: Comparable>. The bound applies on both free
functions and classes, and is enforced at the call boundary: a value
that does not conform to the bound is not a valid argument for that
parameter. Inside the body, the bound's default methods are available
on object arguments (a.gt(b) for a Comparable class).
class Money {
new(amount) { self.amount = amount }
cmp(other) { self.amount - other.amount } # conforms to Comparable
}
fn pick_max<T: Comparable>(a: T, b: T) {
if a.gt(b) { a } else { b } # gt is a Comparable default
}
pick_max(Money.new(10), Money.new(25)).amount # → 25
pick_max([1], [2]) # !! no matching method
Bounds are lowered to the bound trait at declaration time, so they reuse the ordinary trait-conformance machinery. Three consequences:
-
Dispatch specificity sits at the trait level: a concrete overload beats a bounded one, and a bounded one beats an unbounded
<T>catch-all.fn rank(x: Long) { "concrete" } fn rank<T: Comparable>(x: T) { "bounded" } fn rank<T>(x: T) { "unbounded" } rank(5) # → "concrete" (exact type) rank(1.5) # → "bounded" (Float conforms to Comparable) rank([1]) # → "unbounded" (Array conforms to neither) -
Lenient unification: a type parameter that appears in several positions (
a: T, b: T) checks the bound at each position independently — repeatedTis not forced to a single concrete type. A function taking(a: T, b: T)accepts(1, 2.0)(bothComparable). -
The bound enforces conformance, but the primitive method-surface limitation still applies:
a.gt(b)only dispatches whenais an object. For primitive arguments use the native</>operators in the body instead.
An unbounded <T> accepts any argument (it lowers to Any); it
exists to name a parameter and to lose to more specific overloads.
A composite bound <T: A + B> requires the argument to conform to
all parts (the all-of dual of a Union). Spacing is tolerated
(<T:A+B> works too).
fn both<T: Hashable + Stringer>(x: T) { x }
both(5) # OK — Long conforms to both
both([1, 2]) # !! Array is Stringer but not Hashable -> rejected
Known limitations
- Two overloads sharing an outer Generic name (
fn show(xs: Array<Long>)vsfn show(xs: Array<String>)) both score 4 in dispatch — calls with an Array arg are reported as ambiguous. Resolves naturally once element-level checks ship (Phase 2b). - Type parameters in introspection:
fn.params[i].typereports the declared form (T,Array<T>) for readability, but the runtime type check only sees the outer /Any. Don't rely on introspectedTas a real type for further dispatch decisions.
Planned (Phase 2b)
- Expression-position type args (
Box<Long>.new(42)) — the parser currently reserves<...>for type-annotation contexts only. - Opt-in element runtime check.
- Class declarations inside another class's body (currently a SyntaxError, see "Where class declarations may appear" above).
Nil and annotations (null-safety policy)
Culebra keeps nil universal but lets annotations opt into runtime
null safety, in three tiers:
- No annotation — anything goes, including
nil(the default dynamic behavior). T(non-optional) —nilis rejected. A function parameter typedx: Longwill not acceptnil(the call is reported as a dispatch/type error); alet x: Long = nilraises a type error.T?(optional) —nilis accepted, and the base type is still enforced for non-nil values (x: Long?takes a Long or nil, but not a String).
This is a runtime policy — there is no static null checker (by
design). The null-safe operators pair with it:
?. / ?[] navigate possibly-nil values, !! asserts non-nil
(raising NilError), and ?? / ??= supply or store fallbacks.
There is no late / lateinit keyword: in static languages that
feature exists to satisfy a static non-null checker for deferred
initialization, which culebra does not have. Use a plain (nullable)
field and assert with !! at the use site, or guard with ??. For
lazily-computed / memoized fields, the idiom is a nil-checked accessor
method — self._data ??= load() works directly (??= supports
obj.key targets, not just simple variables):
class Cache {
new() { self._data = nil }
data() {
self._data ??= load()
self._data
}
}
Where annotations are not checked
- Arithmetic / boolean operators check their own operand types; annotations on local intermediates do not add extra coverage.
- Object property values have no annotation slot — a property an
object grows outside a class declaration is unconstrained. A class's
field declared with a scalar type (
Float,Long,Bool, or a fixed-width spelling of one) is the exception: that annotation is checked on every write (§10). Every other field annotation,Stringincluded, is not checked. - Array element types are not tracked.
mut x: Ton a local, or amutparameter, does not re-check on later reassignment — the check happens once at the annotated declaration, or at the call.
Annotations are primarily a documentation and boundary check feature, not a type system.
What a primitive-typed parameter settles
A parameter annotated with one of the primitive type names — Long,
Float, Bool, String, Array, Object, Set, Tuple — is checked
on entry, so every read of that name in the body already satisfies it.
The compiler takes that as given: a built-in method call on such a
parameter resolves to the one receiver its type names instead of testing
which of the several types sharing that method name it holds, and
arithmetic on a Long or Float parameter drops its operand tests. The
answer is identical either way; the annotation only saves asking.
Two annotations settle nothing. A mut parameter can be reassigned, and
that is not re-checked (above). Function is satisfied structurally — a
class instance with a __call__ passes it — so it does not name one kind
of value. A class name settles the class, not the representation; §10
covers what that gives a reader.
fn width(s: String) {
s.size()
}
inspect(width('hello')) # => 5
# the entry check is what settles it, and it still refuses the rest. A
# typed parameter is an overload signature, so the refusal is a
# DispatchError.
# !! DispatchError
width(42)
# a `mut` parameter promises nothing about later reads: the reassignment
# is not re-checked, so the body sees whatever it was given
fn swapped(mut s: String) {
s = 42
type_of(s)
}
inspect(swapped('x')) # => 'Long'
What a class-typed name lets a reader assume
A name whose declared class is known — a parameter or a let whose
annotation names it, self inside that class's own members, or the next
step of a chain through a class-typed field — is checked where it is
bound, and that class's scalar-declared fields are checked on every write
(§10), so a read of one of those fields through that name knows what it
will find. The compiler uses that: p.x where p: Point and Point
declares x: Float compiles as the read of a Float, with no question
asked of the value. A field declared anything else is read the ordinary
way, since nothing checked what went into it. A mut binding takes no
such promise either: its reassignment is not re-checked, so there is
nothing to read the declaration off.
The entry check tests what built the value, not what it looks like: a
class's name lives on the meta every one of its instances reaches, and an
ordinary Object cannot acquire one by carrying a field. So an Object that
names a class is refused where it is passed, before any field is read. The
read keeps a check of its own for the receivers no entry check saw — a
method value moved onto a foreign object reads self.x the same way — and
it names the field and the type rather than answering with a value of the
wrong kind. For the same reason remove refuses a scalar-declared field: a
field that can vanish is no contract. A field declared anything else is
removable, since nothing was promised about it.
class Point {
x: Float
new(x) {
self.x = x
}
}
fn scaled(p: Point) {
p.x * 2.0
}
inspect(scaled(Point(1.5))) # => 3.0
# An Object that merely names the class is not one, so it never binds the
# parameter.
# !! DispatchError
scaled({class: 'Point', x: 'not a float'})
Sum types (enum)
An enum declares a sum type — a fixed set of variants, each
optionally carrying positional payload. Generic parameters are
supported.
enum Result<T, E> {
Ok(T),
Err(E),
}
enum Color { Red, Green, Blue } # nullary variants
enum Shape { Circle(Float), Rect(Float, Float), Origin }
Each variant lowers to a variant-as-class: the enum name is bound as a namespace, and constructing a value goes through it. Payload variants are constructors; nullary variants are singleton values.
let r = Result.Ok(5) # payload variant — call the constructor
let c = Color.Red # nullary variant — a singleton value
let s = Shape.Rect(3.0, 4.0)
Result.Ok # the constructor is a first-class value
A variant instance is tagged with both the variant name and the parent enum name, so type annotations / patterns match at either level:
fn handle(r) {
match r {
Ok(x) => x, # constructor pattern, binds payload
Err(e) => -1,
}
}
match v {
x: Ok => "the Ok variant", # variant type pattern
x: Result => "some Result", # enum type pattern (any variant)
_ => "other",
}
A bare variant name names the variant, not the enum: where two enums
each declare an Ok, Ok(p) and x: Ok take either. Qualify it to
pin the enum — Result.Ok(p) as a pattern, x: Result.Ok as a type —
and both halves are checked. Naming the enum alone (x: Result) takes
any of its variants.
Payload is positional and reachable as _0, _1, … (r._0), though
constructor patterns are the idiomatic accessor.
A variant prints the way it is written — in interpolation, println,
to_string() and inspect alike. type_of names the variant, and
class_of answers with the enum object itself, so class_of(v) == Shape
asks which enum a variant belongs to:
"{Color.Red}" # 'Color.Red'
inspect(Shape.Rect(3.0, 4.0)) # Shape.Rect(3.0, 4.0)
type_of(Shape.Origin) # 'Origin'
class_of(Shape.Origin) == Shape # true
A variant is Eq and Hashable by construction — its identity is the
enum, the variant, and each payload field — so it is an Object /
Set key and a hash(v) argument as-is, with no @derive. A payload
variant is a valid key exactly when
its payload is: Shape.Rect(2.0, 3.0) is, Result.Ok([1, 2]) raises
the Array's own unhashable type TypeError. Same-named variants of two
enums are two keys.
mut hits = {}
hits[Color.Red] = 1
hits[Shape.Rect(2.0, 3.0)] = 6
{Color.Red, Color.Green, Color.Red}.size() # → 2
Notes and limitations
- The canonical sum-type style is an untyped parameter plus a
match:fn area(s) { match s { Circle(r) => ..., ... } }. - Multimethod dispatch keys at any of the three levels:
fn f(x: Result.Ok)takes only that enum'sOk,fn f(x: Ok)any enum's, andfn f(x: Result)every variant of the enum. The most specific one declared wins, and the enum name outranks anOk | Errunion of the same variants. - No methods-in-enum block (use free functions + UFCS), no named
payload fields, no explicit discriminant values, and no static
exhaustiveness check (a non-matching
matchyieldsnil). - Nullary variants are matched with a type pattern (
o: Origin), not a parens-free constructor pattern.
Traits and protocols
A trait declares a named contract — the set of methods a value
must provide to be treated under that name in dispatch and Generic
Bound positions.
trait Greeter {
hello() -> String
}
Structural conformance: any class whose methods cover the trait's
required ones (name + arity) automatically satisfies the trait —
no impl Foo for Bar block needed: conformance is decided by the
methods a class actually has.
class Bob {
new(name) { self.name = name }
hello() { "hi, {self.name}" }
}
fn greet(x: Greeter) { IO.inspect(x.hello()) }
greet(Bob.new("Alice")) # → "hi, Alice"
A class missing required methods is rejected at the dispatch boundary (DispatchError).
Default methods
A trait method may carry a body. Conforming classes inherit it for free; defining the same name on the class overrides:
trait Counter {
current() -> Long
next() -> Long { self.current() + 1 } # default
}
class Zero {
new() {}
current() { 0 }
}
Zero.new().next() # → 1 (default)
class Five {
new() {}
current() { 5 }
next() { 99 } # override
}
Five.new().next() # → 99
A default is a method the conforming class inherits, so it ranks with
the class's own methods: only an own property or a declared method
outranks it, and it outranks everything the runtime supplies for an
object — the Object built-in methods (size, keys, ...), the
synthesized parameters(), and the duck-typed iterator method set.
A trait declares each method name once. Two same-name methods are a
SyntaxError: the contract and the default bodies are keyed by name,
so a trait has no overload set to merge them into (a class does — see
"Method overloading").
A trait declaration takes effect when it runs, like a class: a trait
inside a branch that is never taken registers nothing, and re-declaring
a name replaces the earlier contract and defaults outright.
Built-in traits
The runtime ships seven foundational traits as a preamble — they
are visible without import:
| Trait | Required | Defaults |
|---|---|---|
Stringer | to_s() -> String | — |
Eq | eq(other) -> Bool | neq |
Comparable | cmp(other) -> Long | lt, le, gt, ge |
StringLike | to_string_view() -> StringView | — |
Hashable | hash() -> Long | — |
Iterator | has_next() -> Bool, next() -> Any | — |
Iterable | iter() -> Iterator | — |
A class with only cmp automatically gets the six-way comparison
suite; a class with eq gets neq; a class with to_s is
displayable wherever a Stringer is expected. StringLike accepts
any byte-readable string value at API boundaries — String and
StringView both conform out of the box. Hashable is the
contract Object / Set keys check at insertion: a user class
becomes a valid key by defining hash() (returning Long) and
eq(other). An enum variant conforms to Eq and Hashable without
either — see "Sum types".
Iterator + Iterable formalize the for-in protocol. A class that
exposes iter() -> Iterator,
has_next() -> Bool, and next() -> Any is iterable and works
with every pipeline method (map / filter / take / collect /
...). has_next() is required to be idempotent on repeat calls —
runtime wrappers cache one lookahead so a has_next() peek doesn't
consume the next value. next() may yield any value including
nil (nil terminator designs lose this); pairing with has_next()
is the contract for end-of-stream detection.
Built-in primitives also conform via a hard-coded table — no class wrapper required:
| Primitive | Stringer | Eq | Comparable | StringLike | Hashable | Iterable |
|---|---|---|---|---|---|---|
| Nil / Bool | ✓ | ✓ (Bool) | ✓ (Bool only) | — | ✓ | — |
| Long / Float | ✓ | ✓ | ✓ | — | ✓ | — |
| String / StringView | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Tuple | ✓ | ✓ | — | — | ✓ | ✓ |
| Array / Set | ✓ | ✓ | — | — | — | ✓ |
| Object (bare literal) | ✓ | ✓ | — | — | — | ✓ |
| Tensor | ✓ | ✓ | — | — | — | — |
| Function | ✓ | — | — | — | — | — |
So fn show(x: Stringer) { to_string(x) } accepts show(42) and
show([1, 2, 3]) without requiring a class wrapper. Note that
trait method calls (x.to_s()) only resolve on class instances —
primitives have no method-dispatch surface, so the trait reaches
them via fn boundaries (fn show(x: Stringer)), not via the
.method() syntax. The same pattern applies to Hashable:
primitives go through the hash(v) global builtin (companion to
to_string); class instances also accept the direct x.hash()
call. For Hashable user classes used as Object / Set keys the
matching eq(other) method is required so the container's
equality check stays consistent with the hash.
Generic Bound
A type parameter may carry a single trait as its bound, declared inline:
fn min<T: Comparable>(a: T, b: T) -> T {
if a.lt(b) { a } else { b }
}
The bound is enforced at the call boundary (no compile-time
check; culebra's type checks are runtime) — see "Bound constraints"
under Generic types for the full semantics (lowering, dispatch
ordering, lenient unification). A composite bound <T: A + B> is
also supported — the argument must conform to all parts (the
all-of dual of a Union). where clauses are not planned (the inline
form already covers the cases).
Dispatch tie-break
A trait param scores between Object and concrete in multifn
dispatch: a method with x: Pri wins over a method with
x: Stringer when the arg is a Pri instance; x: Stringer is
chosen when only the trait path matches.
Trait inheritance
A trait may declare supertraits after a colon: trait Ord: Eq,
trait Both: Eq, Show. The supertrait's methods are flattened into
the subtrait, so a value conforming to Ord must supply both
Eq's and Ord's required methods. This works transitively
(Done: Ord pulls in Eq too) and a supertrait's default methods
become available on the subtrait.
trait Eq { eq(other) -> Bool }
trait Ord: Eq { cmp(other) -> Long } # conforming to Ord needs eq + cmp
fn sorted<T: Ord>(xs: Array<T>) { ... }
Supertraits must be declared before the trait that inherits them (declaration order resolves the flatten). Because conformance for built-in primitive types is name-based (a hard-coded table), trait inheritance affects structural (class) conformance only — a primitive does not retroactively satisfy a user trait via its supertrait.
Deriving methods (@derive)
The @derive(...) class decorator generates standard conformance
methods automatically, so a data class doesn't have to spell out
boilerplate eq / hash / to_s / cmp. Four names are derivable,
each mapping to one method:
| Derive | Generated method | Behavior |
|---|---|---|
Eq | eq(other) | true when other is the same class and every data field is equal as a key |
Hash | hash() | combines the class name and each data field's hash |
Show | to_s() | "ClassName(f1, f2, ...)" with each field's value repr |
Comparable | cmp(other) | lexicographic over data fields, in declaration order |
@derive(Eq, Hash, Show, Comparable)
class Point {
new(x, y) { self.x = x; self.y = y }
}
let a = Point.new(1, 2)
let b = Point.new(1, 2)
a.eq(b) # → true
a.hash() == b.hash() # → true
a.to_s() # → "Point(1, 2)"
a.cmp(Point.new(1, 3)) # → -1
Because @derive only supplies the trait's required method, the
trait defaults follow automatically: deriving Eq gives you neq,
and deriving Comparable gives you the full lt / le / gt / ge
suite. Deriving Eq + Hash makes the class usable as an Object /
Set key.
The generated methods are reflective — they walk the instance's own data fields at call time (skipping methods and the internal class tag), so they stay correct as fields change. Details:
- User definitions win. A derived method whose name the class
already declares is skipped, so you can derive most and hand-write
one (e.g.
@derive(Eq, Hash)plus a customto_s). @deriveis compiler-recognized, not a function —Eq/Hash/Show/Comparableare the only accepted names (anything else is a SyntaxError). It composes with regular decorators on the same class.- Classes only. An enum's variants are
EqandHashableby construction, so@deriveon anenumis a SyntaxError. - Nominal: derived
eqrequires the same class tag, so two classes with identical fields never compare equal. - A derived
eqstates no equality. It says the instances match by their fields, the way an enum variant's do. As keys they match strictly by type (1and1.0are two different keys), which is what letseqagree withhashand the class be a key at all. Under==they compare by structure — fields by==, which mixesLongandFloatby value — exactly as a class with noeqdoes, so derivingEqto become a key does not narrow==. The two differ only for such a mixed pair:P(1) == P(1.0)is true whileP(1)andP(1.0)stay two keys. Aneqor__eq__the class writes is stated, and==follows it as usual. cmpordering compares each field pair exactly as<does: equal fields (by==) move on to the next, and the first unequal pair decides. A pair<refuses to order — an Array, an Object, two different types — is the samecannot compareTypeError there, socmporders the numeric and string fields and no others.- A derived
cmpagrees with==. Both walk the fields by==, socmp(other)is0exactly when==holds. It is the key relationeqthat is stricter, for the reason above. eqandcmptakeother,hashandto_stake nothing; calling one without its argument is the ordinary missing-required ArityError.
Limitations (Phase 4 MVP)
impl Foo for Barblock is unsupported — conformance is purely structural. (Phase 4+: revisit if explicit conformance is wanted.)- Operators do not call trait defaults:
</<=/>/>=reach__lt__/__le__and thencmp, and==reaches__eq__and then a statedeq(Trait-method fallback). Comparable's defaultlt/le— and a class's own override of them — fire only when called asx.lt(y). - Same-name defaults across traits are non-deterministic: if
two registered traits both provide a
to_sdefault and an instance conforms to both, the dispatch picks the one the internal hash map iterates first. Avoid declaring same-named defaults across overlapping traits. - Self-recursive trait defaults are user responsibility:
trait X { foo() { self.foo() } }will stack-overflow if a conforming class doesn't overridefoo. No depth guard is installed yet. - JIT trait-default closure refcount: each declared default method is held at +1 in a per-Runtime table. Re-declaring a trait releases the displaced bodies, and the rest are released when the Runtime is destroyed, so long-running JIT hosts re-compiling many sessions do not accumulate them.
- multifn dispatch overhead: every call walks the full trait_registry × args matrix once per warmup, even for functions with no trait-typed parameters. Cache amortizes repeats, but ~3 × n_args probes per call show up in hot loops. Phase 4+ optimization can skip the warmup when no method references a trait.
- JIT trait-default fallback is not IC-cached: property
accesses that resolve through a trait default re-walk the
default table on each call. Hot loops on
instance.default()pay this overhead repeatedly. Phase 4+ adds an inline-cache shape for the trait-default slot. - AOT trait re-declaration: each method registers via per-
method runtime upserts, so re-declaring
trait Eq { eq(other) }after the preamble'sEq { eq, neq }leavesneqin place. REPL / hot-reload contexts should declare traits fully or restart. - Traits, like classes, may only be declared at top level.
15. Error handling
throw / try / catch
Culebra supports user-raised exceptions via throw and catching via
try/catch. The thrown value may be any Culebra value.
throw 'something bad'
throw {kind: 'io', msg: 'file not found', path: p}
try { risky() }
catch e { inspect("error: {e}") }
Semantics:
throw exprevaluatesexprand propagates it as a Culebra value until caught by the nearest enclosingtry. It crosses function boundaries and propagates through any number of frames.try { A } catch name { B }evaluatesAin a fresh scope; onthrowwithinA,Bis evaluated in another fresh scope withnamebound to the thrown value. The wholetry/catchis an expression yielding the value of whichever block ran last.- An uncaught
throwat the top level is reported withuncaught: ... at LINE:COL.— the position of thethrowitself — and the program exits with status 1. The exception is a caught runtime error thrown again (catch e { throw e }): it reports as that error,Kind: message at LINE:COL.at the position it was raised. throwis distinct fromreturn: a function's earlyreturnunwinds only that function; a userthrowtravels past enclosing functions.throwis an expression (§12), so it can be the whole of amatcharm (_ => throw "unexpected") with no block around it.
defer
A defer { BLOCK } statement registers BLOCK to run when the
enclosing lexical scope exits, in LIFO order. Defers run on
every exit path — normal completion, early return, or throw.
fn work() {
{
f = open_tmp()
defer { f.close() } # fires when this inner block exits
process(f)
} # ← f.close() runs here
more_work()
}
Block scope (not function scope) means:
deferin a loop body fires on every iteration — matches the programmer's intent of per-iteration cleanup.deferin a conditional block fires only if that block actually ran.- To tie cleanup to the function, place
deferat the top of the function body (the function body itself is a block). - Top-level
deferruns when the program exits — an imported module's top-leveldefertoo, after the entry module's.
A return inside a defer body exits only the defer closure, not the
enclosing function. throw inside a defer body aborts that defer and
propagates as a regular exception — after the defers still pending in
the same scope have run, in the same last-in-first-out order: a defer
always runs, whatever the ones registered after it did. If more than one
of them throws, the last throw is the one that propagates. The scope is
then left like any other scope a throw passes through — its bindings are
released, and the enclosing scopes run their own defers on the way out.
fn close_all() {
defer { inspect('file closed') }
defer { throw 'flush failed' }
inspect('body')
}
inspect(try { close_all() } catch e { e })
# => 'body'
# => 'file closed'
# => 'flush failed'
If the scope was already unwinding from another exception, the defer's
throw replaces it: the original is discarded, and the unwind goes on
from that scope outward with the new one, which an enclosing try can
catch. A try body's own defers run outside the reach of its catch,
on the normal exit and the throw exit alike — the next try out
receives what they throw:
fn close_fails() {
defer { throw 'close failed' }
throw 'read failed'
}
inspect(try { close_fails() } catch e { e }) # => 'close failed'
Interruption (Ctrl+C)
Ctrl+C raises a cooperative, catchable Interrupted rather than
killing the process outright:
try {
serve() # long-running loop
} catch e {
inspect("shutting down: {e.kind}") # → Interrupted
cleanup()
}
- The running computation stops at the next loop iteration or statement
boundary and throws
Interrupted. A tight loop (evenwhile true {}) is interruptible, as is a wait onIO.stdin().read()/IO.input(blocking on stdin), a blockingHttprequest (connect, response wait, or body transfer), or a blockingProc.run/Proc.all/Proc.race(the child is killed) — a single press breaks the wait, not just the second. - It unwinds like any exception, so
deferblocks run on the way out. - If you
catchit, execution resumes normally — the interrupt is one-shot, so a server / REPL can treat Ctrl+C as "cancel the current request" and keep going. - If it reaches the top uncaught, top-level defers run and the program
exits with status
130(128 + SIGINT), the conventional code. - A second Ctrl+C while the first is still pending (a wedged program that never reaches a safepoint) force-terminates with the default disposition.
In the REPL, Ctrl+C interrupts the running evaluation and returns to the
prompt instead of killing the session (the Interrupted is caught by the
read-eval loop); a second press during a wedged eval still force-quits.
To receive Ctrl+C as a value instead of a throw — the graceful-shutdown
pattern for a long-running service — register a channel with Signal.notify
(see the stdlib guide's Signal section).
The behaviour is identical under the VM, JIT, and AOT binaries.
Interrupted carries no source position (line/col are 0): the
interrupt is asynchronous, not tied to a particular expression.
Scope guard pattern
When cleanup must be registered from code that cannot place its own
defer (e.g., a callback that wants cleanup at the caller's scope),
a small helper object is enough — the language does not ship a
built-in for it:
make_guard = fn () {
mut fns = []
{
add: fn (f) { fns.push(f) },
run: fn () {
mut i = fns.size() - 1
while i >= 0 { fns[i](); i = i - 1 }
fns = []
}
}
}
process = fn (items) {
g = make_guard()
defer { g.run() }
mut f = nil
items.for_each(fn (item) {
if f == nil {
f = open('out')
g.add(fn () { f.close() })
}
write(f, item)
})
}
Runtime errors
Internal runtime errors (type mismatches, division by zero, missing
properties, ...) surface to user try/catch as structured Error
Objects with four properties:
| Property | Type | Description |
|---|---|---|
kind | String | Error category, e.g. 'TypeError'. See list below. |
message | String | Human-readable description (includes at L:C.). |
line | Long | 1-based source line of the offending AST node, 0 if unknown. |
col | Long | 1-based source column, 0 if unknown. |
User code can branch on e.kind:
try { let x = arr[100] }
catch e {
if e.kind == 'IndexError' { inspect("out of range at line {e.line}") }
else { throw e }
}
Standard kind values, with the exact trigger condition and
catchability for each. Every kind populates e.kind, e.message,
e.line, and e.col identically on the VM, the JIT, and AOT
builds (unless noted).
| Kind | Trigger | Catchable |
|---|---|---|
TypeError | Arithmetic / comparison on incompatible operand types; calling a non-callable; failing : T annotation; to_long/to_float on non-coercible value; __str__ returning non-String; * splat of non-Array / ** splat of non-Object; built-in arg-type check failure; mixed positional + keyword targeting the same parameter; duplicate keyword; more positionals than the * separator allows (takes N positional arguments but M given). | yes |
ZeroDivisionError | Integer /, %, ** with negative exponent collapsing to division; float / or % with RHS == 0. | yes |
NameError | Read of an undefined identifier; compound assignment (x += rhs) on undefined x; REPL global lookup miss. A name bound in no scope (and not a builtin) is caught before evaluation — see Compile-time errors; a name read before its own later declaration runs (use-before-def) stays a runtime error. | yes¹ |
ImmutableError | Assignment to a let (non-mut) binding; assignment to an immutable Object property or Dict entry; rebinding self inside a constructor body. | yes |
KeyError | Dict subscript on absent key; Object subscript on absent key. | yes |
IndexError | Array / String / Tensor index out of range; Tensor slice out of bounds; Tensor reduction axis out of range. | yes |
ValueError | Destructure pattern mismatch ([a, b] = ... shape mismatch); Tensor shape / dtype mismatch; [].min() or other empty-collection reductions; numeric conversion of malformed string; JSON parse failure; walking a value nested deeper than 1000 levels (see the value-nesting bound below). | yes |
AttributeError | Compound assignment (o.x += ...) on a missing property; reading or writing (=/??=) a member a builtin namespace doesn't have (namespaces are closed — a class or plain dict still reads/writes freely). | yes |
ArityError | Call missing a required argument — too few positional args to a function or a class constructor; more positional args than a built-in or namespace function accepts. Surplus positionals to a user function are not an error: they overflow into __ARGS__ (§19). | yes |
DispatchError | Multimethod call with no matching method or with ambiguous specificity tie (§20). | yes |
AssertionError | Matcher failure (assert_true / assert_eq / etc.) or user throw {kind: "AssertionError", ...}. Message names both operands for comparison matchers. | yes |
SyntaxError | Structural errors raised during AST lowering: **rest not last param, duplicate * separator, non-default param after default, compound let, break / continue outside loop. Surfaces at function decl evaluation, before that function runs. | yes |
ShadowError | Static shadow analyzer (§6) detected a binding that shadows a captured outer name. Fires before any user try block can observe it. | no (pre-eval analyzer) |
IOError | FS / File / stdlib file ops failing (missing path, permission, closed handle); Tensor.from_csv failure. | yes |
ProcessError | Proc.run spawn failure (e.g. the executable doesn't exist), or a non-zero exit / signal death under check: true. | yes |
SendError | A value that is not Sendable was passed across an isolate boundary (Isolate.spawn / tx.send) — a native handle, a Tensor, a closure capturing a mut, or a cyclic value. | yes |
ChannelError | tx.send on a channel whose receivers/senders have all gone (closed). | yes |
ParallelError | A Parallel.map / Parallel.each element threw; carries the failing element's index and cause (fail-fast). | yes |
DropContractError | drop / iter / has_next / next property bound to a function that takes arguments. | yes |
RecursionError | Function-call depth exceeded the fixed limit of 1000 frames. Every user-function entry counts one frame (fn, lambda, method, constructor — field initializers run inside the constructor's frame); built-in helpers and multimethod dispatch do not. The limit and the reported depth are identical on every backend, and reported at the call site. The count unwinds with throw, so a catch regains the full budget. | yes |
RuntimeError | Fallback when the engine catches an unconverted std::runtime_error from a not-yet-migrated throw site. e.line == 0 and e.col == 0 are possible in this case only. | yes |
Uncaught errors print as Kind: message at LINE:COL. and exit with
status 1 (an uncaught Ctrl+C prints interrupted and exits with 130),
from culebra and from a culebra build binary alike. User-thrown
values via throw expr print as
uncaught: {value} at LINE:COL., the position being the throw's own —
unless the value is the error a catch received for a runtime error,
which reports as that error at the position it was raised, even when it
crossed an isolate boundary first.
A value re-raised across an isolate boundary carries no position — the
throw that produced it ran on another thread — and prints without one.
A message that runs to several lines leads with the position instead
(Kind at LINE:COL: message), so it cannot be read as part of the last
line's value.
¹ The NameError for a name that is bound in no enclosing scope and is
not a builtin is the exception: it is caught before evaluation (see
Compile-time errors) and so is not catchable. Every other NameError —
notably use-before-def — is a runtime error and is catchable.
The value-nesting bound
A loop can build a value of unbounded depth (a = [a] repeated). Every
operation that walks a value's nesting — printing (inspect / print /
interpolation / to_string / join), == / !=, hash, membership
(contains / index_of), sending across an isolate boundary
(Isolate.spawn arguments, tx.send, Shared.new — where a chain of
closures nesting through captures counts the same way), and the
@derived methods — stops at 1000 levels and raises a catchable
ValueError (nesting too deep (limit 1000)) at the call site,
identically on every backend. A same-pointer comparison (a == a)
answers without walking. Building, indexing, and dropping a deeper value
stay safe at any depth — teardown is bounded internally, not by this
error. JSON / TOML apply the same limit to their own trees (see the
stdlib reference).
A Tensor's autograd graph is not value nesting and is not subject to
this bound: .backward() and dropping an unevaluated graph both stay
safe at any depth, with no ValueError cap. A computation graph has no
natural "too deep" — an RNN unrolled over a long sequence is a
legitimate graph, not malformed data — so both are internally bounded
without an artificial depth limit (see the Autograd section of the
stdlib reference).
Compile-time errors
Two checks run when the program is loaded, before any try block runs —
so user code cannot catch them. Both abort with the same Kind: message
format and run identically on every backend.
ShadowError— a binding that shadows a captured outer name; see "Shadow prohibition" in §6 for the rule.NameError(statically undefined) — a variable read whose name is bound in no lexical scope and is not a builtin. This is certain to fail at runtime, so — like an unknown name in a statically-checked language — it is reported before evaluation rather than waiting for the (possibly never-taken) path that reads it. The check is sound: it flags only names that resolve nowhere, so it never rejects a valid program.
The complementary cases stay at runtime, catchable by try/catch,
because they cannot be proven to fail statically: use-before-def (a
name is bound in scope but is read before its declaration executes — the
binding might run first on a later loop iteration), ImmutableError, and
missing / unknown kwargs.
Assertion API
There is no assert keyword or builtin. For tests, use the matcher
family (assert_true / assert_eq / etc., see §19 and docs/stdlib.md).
For production invariants, throw an Object:
# doctest: skip
if !cond {
throw {kind: "AssertionError", message: "..."}
}
Assertion control flow is written out rather than hidden behind a keyword that a build flag can switch off.
JIT support
The JIT lane supports throw / try / catch / defer with
semantics identical to the VM's:
throw/try/catchpropagate across function boundaries via LLVMinvoke/landingpadand the Itanium C++ personality.deferregisters a closure on a defer stack; a scope cleanup landingpad runs it on fall-through,return, and throw-unwind paths. This holds at every level — inside a lexical-scope block ({ defer { ... } ... }), directly in a function body, and at the top level — so cleanup fires even when an uncaught throw escapes the function or the program.
Internal runtime errors (TypeError, ZeroDivisionError, etc.) flow
through user try/catch as structured Error Objects on both backends
— see "Runtime errors" above.
16. Algebraic effects
An effect is an operation whose meaning is supplied by the dynamically
enclosing context rather than fixed at the call site. Code performs an
operation; a handle block installed higher on the call stack decides what
the operation does — and whether, and how many times, to resume the code
that performed it. One mechanism expresses generators, exceptions, cooperative
scheduling, and backtracking search without each needing dedicated syntax.
Effects lower, at parse time, to ordinary Culebra classes plus a small runtime (the same compile-time transform the generators use), so all three backends run identical code and behave identically.
Declaring an effect
effect fn introduces either an operation (no body) or an effectful
function (with a body):
effect fn log(msg) # operation: performed, handled elsewhere
effect fn greet(name) { # effectful fn: may perform operations
perform log("hi {name}")
name
}
An operation declaration only names the operation and its parameters; invoking
it directly is an error — it must be reached through perform.
An effectful function's body is lowered into a synthesized computation class,
the same shape a generator body takes, so self may not be referenced there
(self is not available inside an effect fn body) — pass what you mean as a
parameter, or bind it outside as let me = self. A handle { ... } body is
not restricted: it stays where it was written, so inside a method self is
still that method's receiver. Both promote their locals onto the synthesized
object, where a local holding a function stays a plain variable (§11).
Performing an operation
perform op(args) suspends the current computation and transfers to the
nearest enclosing handler for op:
effect fn ask()
let x = handle {
let n = perform ask()
n + 1
} with ask(resume) {
resume(10)
}
inspect(x) # => 11
Handling
handle { BODY } with op(params…, resume) { CLAUSE } runs BODY with a
handler installed for the duration. When BODY (or anything it calls)
performs op, the matching clause runs with the operation's arguments bound to
the leading parameters and the continuation bound to the last parameter
(resume above) — a first-class function that resumes the performing code with
the value passed to it. resume() with no argument resumes with nil.
A handler may resume, or not — declining to call resume discards the rest of
the performing computation:
effect fn fail()
let r = handle {
perform fail()
"unreachable"
} with fail(resume) {
"aborted"
}
inspect(r) # => 'aborted'
Multiple operations, and a return clause
One handle may carry several with clauses (one per operation) and an
optional with return(v) { … } that maps BODY's normal-completion value:
effect fn get()
effect fn put(v)
mut cell = 0
let out = handle {
let a = perform get()
perform put(a + 5)
perform get()
} with get(k) { k(cell) }
with put(v, k) { cell = v; k(nil) }
with return(v) { "final={v}" }
inspect(out) # => 'final=5'
The return clause applies only to normal completion. When a handler aborts
(never resumes), its own value is the result and return does not run.
Resuming more than once (multi-shot)
A continuation is multi-shot: a handler may call resume any number of times,
and each call independently re-runs the rest of the performing computation from
its suspension point. This expresses nondeterminism / backtracking:
effect fn choose(a, b)
let all = handle {
let x = perform choose(1, 2)
let y = perform choose(10, 20)
x + y
} with choose(a, b, k) {
[k(a), k(b)]
}
inspect(all) # => [[11, 21], [12, 22]]
Forks share referenced heap values (arrays, objects) — the fork is a shallow copy — while independent scalar state is copied per fork.
Dynamic scope and effectful calls
Handlers are dynamically scoped: a perform reaches the nearest handler on
the current call stack, not the lexically nearest. Calling an effectful
effect fn from a handled body delegates into it, so its performs reach the
same handlers:
effect fn ask2()
effect fn double() {
let n = perform ask2()
n * 2
}
inspect(handle { double() } with ask2(k) { k(21) }) # => 42
Plain functions and effects
Ordinary (non-effect) functions participate fully. A perform in a plain
fn dispatches straight off the dynamic handler stack, and a plain fn may call
an effect fn (and vice versa) — including through first-class uses like
.map(f):
effect fn ask()
fn greet() {
# a plain fn performing an effect
let name = perform ask()
"hi " + name
}
inspect(handle { greet() } with ask(k) { k("ana") }) # => 'hi ana'
What makes this work is a classification of each handler clause at parse time:
- tail-resumptive —
resume(v)called exactly once, as the clause's final statement. The clause runs like a plain function call at the perform point; the native call stack is the continuation. This is the common case (dependency injection, state, mocking, logging). No continuation is captured, so a long chain of tail performs uses constant stack, and adeferin the clause fires when the clause returns — before the performing code proceeds — under both plain-fn andeffect fndispatch. - abort — the clause never calls
resume. Its result becomes thehandle's result: the stack between the perform point and thehandleunwinds, runningdefers on the way, and the unwind is not observable bytry/catchin between (it is not an exception). - full-control — anything else (multi-shot, a non-tail
resume, storing or passingresume). The continuation must be captured, which only aneffect fnbody can support: reaching a full-control handler from a plain fn raisesEffectError. This is the one remaining case where theeffect fnmarker is required.
The classification is deliberately conservative: any use of resume other
than a lone tail call classifies the clause as full-control (which is always
safe — it only means direct dispatch from plain code refuses it).
Capturing an enclosing binding
A nested handle — written inside an effectful effect fn body or another
handle body — may read and write the enclosing computation's locals:
effect fn outer()
effect fn inner()
effect fn work() {
let base = perform outer()
handle {
base + perform inner()
} with inner(k) { k(100) }
}
inspect(handle { work() } with outer(k) { k(5) }) # => 105
Effects compose with generators in both directions — a handle expression
inside a generator body, and a named generator fn inside an effect body:
effect fn scale()
fn doubled() {
yield handle { perform scale() * 2 } with scale(k) { k(10) }
yield 7
}
inspect(doubled().collect()) # => [20, 7]
Semantics and limitations
- Performing an operation with no handler raises
EffectError. - Reaching a full-control handler clause (multi-shot / non-tail
resume) from aperformor effect-fn call in ordinary code raisesEffectError: the native frames in between cannot be captured as a continuation. Route such performs through aneffect fnrun under thehandle. - Inside an effect fn or
handlebody, aperformis supported only in statement position or an unconditionally-evaluated operand. Aperformin a short-circuit (&&/||/??) operand, a ternary arm, a method-chain receiver, or a control-flow condition is rejected at parse time (symmetrically on every backend). In a plain fn, aperformis an ordinary expression with no positional restrictions. - Errors inside an effect body report the line and column where the failing
code was written, as in a plain fn. An unhandled
perform'sEffectErrorcarries the perform's line ase.line. - Effects and generators compose: a named generator fn declared in an
effect fn/handlebody works (and may read the body's locals), a self-containedhandle { … }expression works inside a generator body (including in a yielded expression or a loop), and a bareperformin a generator body dispatches dynamically — against the handlers installed at the.next()call that runs it. The remaining boundaries, rejected at parse time (symmetrically): a bareyieldin an effect body (the body itself is not a generator — wrap the yield in a nested generator fn) and aneffect fndeclaration inside a generator body. - A
deferat an effect fn orhandlebody's statement level runs when the body is left by any path: normal completion, athrowunwinding through it, or an abort — whether a handler clause returns without resuming or an abort signal unwinds the driver from a plain-fnperformor a cross-handle abort. A tail or abort clause that exits bythrowwithout resuming also runs the suspended body's defers; only a full-control clause's throw leaves them pending — it may have keptresume, and a kept continuation runs its defers when a resumed fork completes. Adefernested in control flow, or aperforminside adefer, is rejected at parse time. - A named fn declared in an effect body must sit at the body's statement level (one buried in nested control flow is rejected at parse time); it may be a generator and may reference the body's locals.
- Handlers are isolate-local. A
handleinstalled on one thread is not visible to a spawned isolate; aperforminside the child reaches only handlers installed within that same isolate (an unhandled operation there raisesEffectError, which propagates to the parent onjoin). This follows the concurrency invariant that a script value — and a continuation is one — never crosses an isolate boundary.
17. Memory model
Reference counting
All heap-allocated types (Array, Object, Function, and in the
JIT internal Cells and Closures) are reference-counted.
- An object's refcount starts at 1 when created.
- Binding to a new variable, pushing into an array, or storing into an object property increments the refcount.
- Overwriting a variable, property, or array slot decrements the overwritten value's refcount.
- When the refcount hits zero, the object is freed and each of its held references is decremented.
Cycle collection
Pure reference counting cannot reclaim cycles (a.c = a). Both
backends register every refcounted heap object with an auxiliary
cycle collector that runs a mark-and-sweep periodically (an adaptive
threshold — see §25 for the exact schedule, which differs between
backends) and at program exit.
- Subject to collection:
Array,Object,Tuple,Set, and the environments captured by closures (plusClosure/Cellin the JIT). Both backends reclaim every container cycle shape, including one routed purely throughObjectproperty maps. Stringis not refcounted. String bytes come from a per-Runtimeslab, andStringis a traced-only value: it carries no refcount to hit zero, so the tracing sweep above is its sole reclaimer rather than a cycle-only backstop.
Cyclic data is retrieved and mutated normally; once no external root keeps the cycle alive, the collector frees it on the next cycle.
Display of cycles
Printing a cycle produces {...} / [...] at the re-entry point
rather than infinite recursion (see §8).
Auto-drop (RAII)
If an Object has a Function-typed property named drop that takes
no arguments, the runtime calls it automatically when the object's
properties map is about to be released. The contract is enforced at
assignment: binding drop to a non-function value, or to a function
with non-zero arity, raises a type error.
# doctest: skip
File.open = fn (path) {
h = _native_open(path)
{
read: fn () {
_native_read(h)
},
drop: fn () {
_native_close(h)
}, # called on scope exit
}
}
{
let f = File.open('data.txt')
process(f.read())
}
# f's drop ran here
Cascade: when a parent's drop returns, its properties map is
cleared, which decrements each child's reference count. Children
whose count reaches zero have their own drop invoked, and so on.
Parent-before-child order is guaranteed, and siblings under a single
parent release in property-declaration order (Array/Tuple/Set
elements in element order) — the same on every backend and platform.
Exceptions: an exception thrown from drop is logged to stderr
and swallowed so that the rest of the cleanup cascade proceeds.
Failed construction: a class instance exists from the moment
C.new(...) is entered — before its arguments bind and before its
field initializers run — so a constructor call that throws anywhere
(a missing or wrong-typed argument, a throwing default, a field
initializer, the body) still drops the half-built instance. A
dispatch miss on an overloaded new throws before any instance is
built, so nothing is dropped there.
Reentrant drop bodies: a drop body may mutate the object's own
fields, including releasing a reference that is itself part of a
cycle with the object being dropped (self.other = nil). This is
safe — it cannot cause drop to fire a second time or corrupt the
cascade; drop runs exactly once no matter what its own body
releases.
Replacement order: overwriting what a slot holds — a[i] = v,
o.x = v, o[k] = v, or reassigning a variable — stores the new
value first and releases the old one after. A drop that the release
runs therefore finds the new value in that slot, never the value being
dropped, and it may grow or rewrite the container freely: if it assigns
to the same slot, its own write is the one that stays.
mut box = [nil]
box[0] = {drop: fn () {
inspect(box[0])
}}
box[0] = 'next' # => 'next'
Cycles: a reference cycle does not block drop. A cycle whose
members are owned by a scope — created under it and unreachable from
outside when it exits — is dropped at that scope's exit, newest
member first, on both backends:
let log = []
make_thing = fn (id) {
{id: id, drop: fn () {
log.push(id)
}}
}
{
let a = make_thing('a')
let b = make_thing('b')
a.other = b
b.other = a
} # both drop at the block's exit, despite the cycle
inspect(log) # => ['b', 'a']
The owning scope is wherever the cycle last escaped to: a cycle that
rides out of a function as its return value drops at the caller's
scope exit once discarded. A resource captured by a sibling
closure, or cycled through its own closure slot, also drops at the
owning scope's exit. Two shapes miss the
deterministic window on every lane: a closure-only cycle that holds
the resource without the resource referencing it back (e.g. a
recursive local fn capturing it), and cyclic garbage discarded
directly at the top level (a scope that never exits). Those fire
at the GC backstop:
a collection finalizes every orphaned resource exactly once — before
reclaiming memory, while the structure is still intact — in the
spirit of Python's PEP 442. The background collections run on an
allocation threshold and find their roots conservatively, so when
one of them reaches a given orphan is not specified; GC.stat() runs
a collection that finds its roots from the reference counts, and it
finalizes every orphan unreachable at that call, on every backend.
Finalization order within one collection is unspecified. The
collection at program exit is the
exception: like top-level bindings (below), an orphan that survives to
exit is not finalized — its memory is reclaimed but drop does not
run, on every backend. When cleanup must happen at a known point,
break the cycle manually before the last reference is released (e.g.
a.other = nil), use an explicit .drop(), or defer (§15) at the
scope that owns the resource.
Very large cycles: the scope-exit cascade above has a size limit tuned to keep ordinary programs synchronous; a scope that drops an unusually large number of cycle-bearing objects at once (thousands, not the handful in typical code) may not resolve all of them at that scope's exit — the excess defers to the GC backstop instead, with the same unspecified-order, PEP-442-style finalization described above. Nothing is lost, only delayed; ordinary programs never approach this size.
Binding-scope caveat: drop fires reliably when the object is
held in a block-scoped binding whose drop function was produced
by a factory function (so the closure captures the factory's
call env, not the surrounding scope). The idiom mirrors tutorial §5
closures-as-objects and works out of the box:
make_thing = fn () {
{drop: fn () {
inspect('cleaned')
}} # captures make_thing's env
}
{
let t = make_thing() # block-scoped binding
} # drop fires here
Binding a drop-bearing object at the top level leaves it alive
until program exit without running drop, on every backend: every
backend suppresses drop where the top-level scope is released, and
an uncaught error is the same exit — the scopes the error passes
through release deterministically (each one's defers, then its own
bindings, innermost first), but the top level's own bindings still
leak un-dropped. For script-wide resources, prefer defer (§15) or
explicit cleanup, which run on both exits.
JIT: auto-drop fires under --jit with the same timing as the
VM — at scope exit, cycle members included, closure-held shapes too
(backstop-only shapes as above);
top-level orphans go to the GC backstop on both, and top-level
bindings still leak un-dropped at program exit. The well-known property contract
(a Function bound to drop/iter/has_next/next takes no arguments) is enforced at
assignment time on both backends, on every write surface: literal
properties, computed subscript keys (o[k] = v), native builders
(JSON.parse etc.), values rebuilt on the receiving side of a channel
transfer, class declarations (a class whose method template binds a
well-known name to a non-conforming function — wrong arity, or an
overload set — raises DropContractError when the declaration runs),
and trait default implementations, which reach every conforming
instance and are checked where the trait is declared. A body-less
trait method only states a requirement, so it binds nothing and is not
checked.
When the target slot is immutable, ImmutableError wins over the
contract check, and a failed check leaves the old value in place.
static members are a namespace, not part of the instance protocol:
a static named drop (any arity) is an ordinary function — it is not
contract-checked and is never auto-invoked, and C.drop() calls it like
any other static. The at-most-once guard of an explicit x.drop()
belongs to instances: on a class object, an enum or a namespace, drop
is a member name like any other, and one that is not there is the error
any missing member gives.
class Pool {
static drop() { 'pool emptied' }
drop() { inspect('connection closed') }
}
inspect(Pool.drop()) # => 'pool emptied'
let c = Pool()
c.drop() # => 'connection closed'
c.drop() # already dropped: nothing runs
The contract is about functions. A value that is not a Function is data
that happens to use the name, and binds like any other property — text a
program does not control decides its own keys:
let page = JSON.parse('{"items": [1, 2], "next": "/page/2", "drop": false}')
inspect(page.next) # => '/page/2'
inspect(page.keys()) # => ['items', 'next', 'drop']
Such an Object is not a protocol member of anything: nothing runs when
it goes away, and a for walks it by its keys like any other Object.
18. Built-in type methods
The methods below are part of the language: they are available on any
value of the corresponding type without any import, and cannot be
shadowed by user code (for String, which has no property store; for
Array/Object, a user-defined property of the same name wins and
the built-in is a fallback).
Global built-in functions (inspect, to_string, Math.*, IO.*,
matcher family assert_true / assert_eq / etc.) are specified
separately in docs/stdlib.md.
Well-known method names (protocols). Several method names are
recognized by the runtime and let plain Objects opt into language
features. They are checked by name on both backends:
| Method | Purpose | Defined in |
|---|---|---|
__add__, __sub__, __mul__, __div__, __mod__, __pow__, __matmul__, __neg__, __eq__, __lt__, __le__ | Operator overloading | §10 |
__str__ | Custom display form | §10 |
drop | RAII cleanup hook | §17 |
iter, next | Iterator protocol | §18.5 |
Conventions:
- Types follow §14.
Anydenotes any value. - Positional indices are zero-based. Negative indices where noted
count from the end (
-1is the last element). - Length and indexing for
Stringare in bytes (the Go model). Unicode-aware traversal is explicit:iter()/code_points()walk scalars andgraphemes()walks extended grapheme clusters.
18.1 String methods
All string methods accept String or StringView as receiver (and
where shown, as argument — StringLike). They return new values; the
receiver is never mutated.
| Signature | Description |
|---|---|
s.size() -> Long | Byte length. |
s.empty() -> Bool | Whether size() == 0. |
s.presence() -> String | StringView | Nil | s unchanged if non-empty, else nil — pairs with ??/?. for "use it if there's anything there" (x.presence() ?? default). |
s.upper() -> String | Uppercase, full Unicode case mapping (UAX #21). A mapping may change the length: 'ß'.upper() is 'SS'. |
s.lower() -> String | Lowercase, same mapping rules. |
s.capitalize() -> String | First letter titlecase, the rest lowercase (Python/Ruby capitalize). |
s.title() -> String | Every word's first letter titlecase, the rest lowercase. Word boundaries are UAX #29, so "o'neil".title() is "O'neil" — the apostrophe does not start a word. |
s.normalize(form: StringLike = "NFC") -> String | Unicode normalization: "NFC", "NFD", "NFKC", or "NFKD". Composed and decomposed spellings of the same text compare equal after normalize(); equality itself is by bytes (§8), so it does not normalize for you. Any other form is a ValueError. |
s.eq_ignore_case(other: StringLike) -> Bool | Default caseless matching (full case folding), so 'Straße'.eq_ignore_case('STRASSE') is true. Canonical equivalence is deliberately not folded in — normalize() both sides for that. |
s.reverse() -> String | Reversed by Extended Grapheme Cluster (UAX #29), so an emoji ZWJ sequence or a base+combining pair survives intact. A new String — unlike Array.reverse(), which mutates in place and returns nil. |
s.repeat(n: Long) -> String | n copies concatenated. n == 0 → ""; a negative n is a ValueError, as is a result too large to allocate. |
s.truncate(max: Long, ellipsis: StringLike = "...") -> String | s unchanged if it already fits in max bytes; otherwise cut so the result (content + ellipsis) is exactly max bytes. Byte-based like slice, so a multi-byte scalar can split at the cut point. max < 0, or too small to fit ellipsis, is a ValueError. |
s.trim() -> String | Remove leading/trailing Unicode whitespace (the White_Space property, so NBSP and U+3000 count, not just /\t/\n/\r). |
s.trim_start(chars: StringLike = "") -> String | Trim from the start. No arg → whitespace; chars → leading scalars in that set (no ranges). |
s.trim_end(chars: StringLike = "") -> String | Trim from the end. e.g. s.trim_end("\n"). |
s.tr(from: StringLike, to: StringLike) -> String | Per-scalar translation, character-list form — no a-z ranges or ^. Each scalar of s found in from becomes the scalar at the same position in to; a shorter to repeats its last scalar, an empty to deletes. s.tr("0123456789", "0123456789"). |
s.split(sep: StringLike, limit: Long = 0) -> Array<StringView> | Split on every occurrence of sep. Empty sep → [s]. Elements share a single source. limit caps how many pieces come back, the last one holding the remainder ('a.b.c'.split('.', 2) → ['a', 'b.c']); 0 is uncapped and a negative one is a ValueError. |
s.rsplit(sep: StringLike, limit: Long = 0) -> Array<StringView> | The same pieces, but limit is filled from the right: 'a.b.c'.rsplit('.', 2) → ['a.b', 'c']. The result stays in left-to-right order, and without a limit the two agree. |
s.split_whitespace() -> Array<StringView> | Split on runs of Unicode whitespace, with no empty piece at either end: ' a b '.split_whitespace() is ['a', 'b'], where split(' ') keeps the empties. |
s.split_once(sep: StringLike) -> Tuple | Nil | The two halves around the first sep, or nil when it does not occur — the shape a key=value parse wants, where split would break on a value containing sep. An empty sep never splits, so nil. |
s.rsplit_once(sep: StringLike) -> Tuple | Nil | The halves around the last sep, or nil. |
s.split_iter(sep: StringLike) -> Iterator<StringView> | Lazy variant of split. Short-circuits with .take(n) over huge inputs. |
s.lines() -> Array<StringView> | Split on \n, \r\n, or \r. The terminator is dropped and a trailing one yields no final empty element, so "a\nb\n".lines() is ["a", "b"] — the same line boundaries as File.lines(). Eager like split. |
s.replace(pat: String | Regex, repl: String | Function) -> String | Replace every occurrence; chains. pat a String → literal; pat a Regex (incl. re'…') → regex, so repl may be a $1 / $<name> template or a fn (Match) -> String. A stdlib helper reached by UFCS — s.replace(p, r) is replace(s, p, r). |
s.contains(sub: StringLike) -> Bool | Whether sub appears anywhere. Empty sub → true. |
s.count(sub: StringLike) -> Long | Non-overlapping occurrences of sub, so "aaaa".count("aa") is 2. Empty sub → 0, matching split("") yielding the receiver whole. Literal only — a compiled Regex counts its own matches. |
s.starts_with(prefix: StringLike) -> Bool | Whether s begins with prefix. |
s.ends_with(suffix: StringLike) -> Bool | Whether s ends with suffix. |
s.index_of(sub: StringLike, start: Long = 0) -> Long | Byte offset of the first sub at or after start, or -1 — the same "-1 means absent" convention Array.index_of uses. A negative start counts from the end, like slice's indices. |
s.last_index_of(sub: StringLike) -> Long | Byte offset of the last sub, or -1. |
s.strip_prefix(prefix: StringLike) -> String | s without that exact prefix, or s unchanged when it is not there. Unlike trim_start(chars), which trims a set of scalars, this matches the whole affix once. |
s.strip_suffix(suffix: StringLike) -> String | s without that exact suffix, or s unchanged. |
s.replace_first(pat: String | Regex, repl: String | Function) -> String | Like replace, but only the first occurrence. Same pat/repl rules. |
s.is_digit() -> Bool | Whether s is non-empty AND every scalar is a decimal digit (General_Category Nd, so fullwidth '123' counts). |
s.is_alpha() -> Bool | Non-empty and every scalar is Alphabetic. |
s.is_alnum() -> Bool | Non-empty and every scalar is Alphabetic or a number. |
s.is_space() -> Bool | Non-empty and every scalar is White_Space. |
s.is_ascii() -> Bool | Non-empty and every byte is below 0x80. Empty is false here too — one rule for the whole is_* family rather than a per-method exception. |
s.slice(start: Long, end: Long) -> StringView | Substring [start, end) borrowing from the receiver's bytes. Negative indices count from end; start clamped to [0, size()], end to [start, size()]. |
s.view() -> StringView | A view aliasing the receiver's bytes (no copy). |
s.to_string() -> String | Materialize an owning String (no-op for String; copies for StringView). |
s.iter() -> Iterator<StringView> | Lazy walk yielding one-scalar StringViews (UTF-8 scalar). What for c in s { ... } uses internally. Invalid bytes yield as one-byte substrings. |
s.code_points() -> Iterator<Long> | Lazy walk yielding Unicode scalar values as Long (U+0000–U+10FFFF). For numeric / range / classification work where the per-scalar allocation of iter is wasteful. Invalid bytes yield as 0–255. |
s.graphemes() -> Iterator<StringView> | Lazy walk yielding Extended Grapheme Clusters (UAX #29) — one user-perceived character per step (e.g. an emoji ZWJ sequence is a single element). |
s.words() -> Iterator<Tuple> | Lazy walk yielding UAX #29 word boundary segments as (text, is_word) pairs — every run between two boundaries is its own segment, including whitespace/punctuation runs; is_word is true when the segment holds at least one alphabetic or numeric scalar. Filter to real words: s.words().filter(|(_, w)| w).map(|(t, _)| t). |
s.sentences() -> Iterator<StringView> | Lazy walk yielding UAX #29 sentence boundary segments — one sentence per step, trailing whitespace stays attached to the sentence before it. |
s.bytes() -> Iterator<Long> | Lazy walk yielding the receiver's raw UTF-8 bytes as Long (0–255), one byte per step — no decoding, unlike code_points. For when the encoding itself is wanted (hashing, tokenizer vocabularies, wire formats). |
String.from_code_point(cp: Long) -> String | The inverse of code_points(): one Unicode scalar value in, a one-character String out. Raises ValueError for cp above U+10FFFF or in the surrogate range U+D800–U+DFFF — the same boundary the \u/\U literal escapes (§4.1) reject at parse time. |
String.from_code_points(cps: Array) -> String | The plural inverse of code_points(): an Array of Unicode scalar values in, a String out. Each element passes through the same gate as from_code_point, so String.from_code_points([cp]) == String.from_code_point(cp); a non-Long element is a TypeError, an out-of-range one a ValueError. |
String.from_bytes(bytes: Array) -> String | The inverse of bytes(): an Array of raw byte values (0–255) in, a String out. No UTF-8 validation, and no error on malformed input: culebra Strings tolerate invalid UTF-8 (same as iter()), so String.from_bytes(s.bytes().collect()) == s holds for every String, including ones with invalid sequences. A non-Long element is a TypeError, an out-of-range one (outside 0–255) a ValueError. |
# 'é' is 2 UTF-8 bytes, so 'café' is 5 bytes
inspect('café'.bytes().collect()) # => [99, 97, 102, 195, 169]
inspect(String.from_code_point(233)) # => 'é'
inspect(String.from_bytes([99, 97, 102, 195, 169])) # => 'café'
inspect(String.from_code_points([99, 97, 102, 233])) # => 'café'
StringView
StringView is a borrowed view over an owning String's bytes.
slice / split / view / iter / graphemes / sentences return
StringView instead of String to avoid per-call copies (words
yields a StringView too, inside each (text, is_word) tuple). The
view keeps its
source alive (shared ownership), so a view outlives the temp that
created it:
let v = 'hello world'.slice(6, 11)
inspect(v) # => 'world'
# Equality is by bytes, so a view compares equal to a String:
inspect(v == 'world') # => true
inspect(type_of(v)) # => 'StringView'
# to_string() materializes an owned String:
inspect(v.to_string()) # => 'world'
Use .to_string() when you need an owning String (storing in a
data structure, returning from a long-lived function, etc.). Most
APIs declared with StringLike accept either flavor directly.
Known limitation (cycle B): calling .contains() / .starts_with()
/ .ends_with() etc. on a StringView materializes a temporary cstr
copy per call. It is reclaimed by the tracing collector like any other
String (§17), not leaked, but hot-looping over a large sequence of
views still means avoidable per-call allocation — materialize the view
once with .to_string() in that case. (Object-key normalization
between String and StringView is not a limitation — see §18.3.)
inspect('hello'.size()) # => 5
inspect('HeLLo'.lower()) # => 'hello'
inspect(' hi '.trim()) # => 'hi'
inspect('a,b,c'.split(',')) # => ['a', 'b', 'c']
inspect('hello'.slice(1, 4)) # => 'ell'
inspect('hello'.slice(-3, -1)) # => 'll'
inspect('hello world'.truncate(8)) # => 'hello...'
inspect(''.presence() ?? 'default') # => 'default'
# Three views of the same string: bytes, scalars, clusters
inspect('café'.size()) # => 5
inspect('café'.code_points().count()) # => 4
inspect('café'.graphemes().count()) # => 4
# Emoji ZWJ sequence: 5 scalars, 1 grapheme
inspect('👨👩👧'.code_points().count()) # => 5
inspect('👨👩👧'.graphemes().count()) # => 1
# Numeric ops via code_points
upper = 'Hello World'.code_points().filter(fn (cp) {
cp >= 65 && cp <= 90
}).count() # 2 ('H', 'W')
# Word segments (UAX #29): whitespace/punctuation are their own
# segments too, so filter on `is_word` to keep only the real words
inspect('Hello, world!'.words().collect())
# => [('Hello', true), (',', false), (' ', false), ('world', true), ('!', false)]
inspect('Hello, world!'.words().filter(|(_, w)| w).map(|(t, _)| t).collect())
# => ['Hello', 'world']
# Sentence segments (UAX #29): trailing whitespace stays with the
# sentence before it
inspect('Hello. World.'.sentences().collect()) # => ['Hello. ', 'World.']
These iterators decode the source string on demand: each next()
only touches as much of the UTF-8 buffer as that step needs.
s.graphemes().take(3).collect() on a multi-megabyte s therefore
reads only enough bytes to resolve the first three clusters — words
and sentences share the same windowed-decode implementation, just
feeding it a different boundary rule. The per-iterator state (decode
offset, lookahead buffer) is independent, so multiple iterators
derived from the same String walk in parallel without interfering
with each other.
JIT: iter / code_points / graphemes / words / sentences
return iterator Objects that the JIT drives through the same protocol
path as user-defined
iterators. Semantics match the VM's; throughput is dominated
by the per-step closure dispatch. For maximum speed over Arrays and
direct String scalars, prefer for c in s { ... } (native loop).
18.2 Array methods
Methods marked mutating modify the receiver in place and return
nil (except pop and remove_at, which return the removed
element); others return a new Array and leave the receiver
unchanged.
A callback that mutates the receiver is allowed, and the walk follows
the live array: the size is re-read before every step, so the walk ends
where the array ends. Shrinking the receiver ends the walk early;
growing it keeps the walk going. This is the same rule for x in a
follows (§12), and it covers map, filter, for_each, reduce,
find, any, all, and flat_map.
mut a = [1, 2, 3, 4]
mut seen = []
a.for_each(fn (x) {
seen.push(x)
a.pop()
})
inspect(seen) # => [1, 2]
| Signature | Description |
|---|---|
a.size() -> Long | Number of elements. |
a.empty() -> Bool | Whether size() == 0. |
a.presence() -> Array | Nil | a unchanged (same Array, not a copy) if non-empty, else nil — pairs with ??/?. for "use it if there's anything there" (xs.presence()?.join(",") ?? default). |
a.push(x: Any) -> Nil (mutating) | Append x to the end. |
a.pop() -> Any (mutating) | Remove and return the last element. nil if empty. |
a.extend(other: Array) -> Nil (mutating) | Append every element of other. a.extend(a) appends the elements a had on entry. For Tuple / Set sources use spread (§9). |
a.insert(i: Long, x: Any) -> Nil (mutating) | Insert x at position i, shifting the rest right. Negative i counts from the end, like a[i]; i == a.size() is the append slot. Out of range raises IndexError. |
a.remove_at(i: Long) -> Any (mutating) | Remove the element at i and return it, shifting the rest left. Negative i counts from the end. Unlike insert, i == a.size() is out of range — raises IndexError, as does any i on an empty array. |
a.get(i: Long, fallback: Any) -> Any | The element at i (negative counts from the end, like a[i]), or fallback if out of range. Read-only, never throws. |
a.slice(start: Long, end: Long) -> Array | Shallow subarray [start, end). Same clamping as String.slice. |
a.join(sep: String) -> String | Concatenate elements via to_string (strings unquoted), separated by sep. |
a.contains(v: Any) -> Bool | Whether v == elem for some element — by =='s own rule, an element's __eq__ / eq included. |
a.index_of(v: Any) -> Long | Index of first equal element, else -1. |
a.reverse() -> Nil (mutating) | Reverse in place. |
a.map(f: Function) -> Array | New array of f(x) for each element. f must take one parameter. |
a.filter(f: Function) -> Array | New array of elements for which f(x) is truthy. f must take one parameter. |
a.for_each(f: Function) -> Nil | Call f(x) for each element for side effects. f must take one parameter. |
a.reduce(init: Any, f: Function) -> Any | Fold: start with init, apply acc = f(acc, x) for each element, return final acc. f must take two parameters. |
a.find(f: Function) -> Any | First element for which f(x) is truthy, else nil. f must take one parameter. |
a.any(f: Function) -> Bool | true if f(x) is truthy for any element, else false. f must take one parameter. |
a.all(f: Function) -> Bool | true if f(x) is truthy for every element (or if empty), else false. f must take one parameter. |
a.flat_map(f: Function) -> Array | Concatenate f(x) for each element; each f(x) must be an Array. f must take one parameter. |
a.sum() -> Long | Float | Sum of all elements, each Long or Float. Stays Long while every element is a Long, becomes Float once any element is one. Empty → 0. |
a.product() -> Long | Float | Product of all elements, promoting like sum. Empty → 1. |
a.min() -> Any | Smallest element, compared numerically across Long / Float. The element itself is returned, so its own type survives. Throws on empty. |
a.max() -> Any | Largest element. Same rule as min. |
a.min_by(f: Function) -> Any | Element whose key f(x) is smallest; f must take one parameter and return a Long or Float. Each key is computed once, and ties keep the earlier element. Throws on empty. |
a.max_by(f: Function) -> Any | Element whose key f(x) is largest. Same rules as min_by. |
a.to_set() -> Set | Fresh Set of the elements in first-seen order, duplicates dropped. Set literals aside, this is how a Set is built from a collection. Unhashable elements throw. |
a.to_object() -> Object | Fresh Object from (key, value) tuples — the inverse of Object.iter(), so a table can be built as an expression instead of a mut + loop. Keys keep first-seen order and a repeat overwrites in place (last value, first position). Entries are immutable, like group_by's; {...built} is the mutable copy. An element that is not a 2-tuple raises TypeError; an unhashable key throws like any other. |
a.group_by(f: Function) -> Object | Buckets elements into Arrays keyed by f(x), in first-seen key order; f must take one parameter and return a hashable key. |
a.partition(p: Function) -> Tuple | One-pass split into (matching, non_matching), order preserved in both halves. p must take one parameter. Destructures: let (yes, no) = xs.partition(p). |
a.unzip() -> Tuple | Split (a, b) pairs into (Array, Array) — the inverse of zip. Each element must be a 2-element Tuple or a {first, second} Object (either pair spelling is accepted); anything else raises TypeError. Destructures: let (xs, ys) = pairs.unzip(). |
a.sort(reverse: Bool = false) -> Nil (mutating) | Stable-sort in place in natural order — elements compare by the same rule as <, so an Object's __lt__ / cmp is honored (a Path array sorts) and incomparable elements throw (leaving the array as it was). Keyword-only reverse: true sorts descending (still stable). |
a.sorted(reverse: Bool = false) -> Array | Like sort but returns a new sorted Array, leaving the receiver unchanged — so it chains (xs.sorted().join(",")). reverse: true for stable descending. |
a.sort_by(key: Function, reverse: Bool = false) -> Nil (mutating) | Stable-sort in place using key(x) as the comparison key (ascending). key must take one parameter and return a comparable value (Long / String / Bool). Keyword-only reverse: true sorts descending (still stable). |
a.sorted_by(key: Function, reverse: Bool = false) -> Array | Like sort_by but returns a new sorted Array, leaving the receiver unchanged — so it chains (xs.sorted_by(f).join(",")). reverse: true for stable descending. |
mut a = [1, 2, 3]
a.push(4)
inspect(a.pop()) # => 4
a.extend([9, 8])
inspect(a) # => [1, 2, 3, 9, 8]
a.insert(0, 0)
inspect(a) # => [0, 1, 2, 3, 9, 8]
inspect(a.remove_at(-1)) # => 8
inspect([1, 2] + [3]) # => [1, 2, 3]
inspect([10, 20, 30, 40].slice(1, 3)) # => [20, 30]
inspect(['a', 'b', 'c'].join('-')) # => 'a-b-c'
inspect([1, 2, 3].contains(2)) # => true
inspect([10, 20, 30].index_of(99)) # => -1
inspect([].presence() ?? 'default') # => 'default'
inspect([1, 2, 3].map(fn (x) {
x * x
})) # => [1, 4, 9]
inspect([1, 2, 3, 4].filter(fn (x) {
x % 2 == 0
})) # => [2, 4]
inspect([1, 2, 3, 4].reduce(0, fn (acc, x) {
acc + x
})) # => 10
inspect([3, 1, 4, 1, 5].find(fn (x) {
x > 3
})) # => 4
inspect([1, 2, 3].any(fn (x) {
x > 2
})) # => true
inspect([1, 2, 3].all(fn (x) {
x > 0
})) # => true
inspect([1, 2, 3].flat_map(fn (x) {
[x, x * 10]
})) # => [1, 10, 2, 20, 3, 30]
mut words = ['banana', 'fig', 'apple']
words.sort_by(fn (s) {
s.size()
})
inspect(words) # => ['fig', 'apple', 'banana']
Callback arity. A higher-order method calls its callback with a fixed
number of arguments (one for map / filter / find / …, two for
reduce), and the callback must accept exactly that many. A function with
the wrong fixed arity is a TypeError — checked once before iterating, so it
fails even on an empty receiver:
[1, 2, 3].reduce(0, |x| x) # !! reduce expects a 2-parameter function
[1, 2, 3].map(fn (a, b) {
a
}) # !! map expects a 1-parameter function
A *args callback absorbs whatever it is given, so it works at any arity —
the per-call arguments arrive as the rest Array:
inspect([10, 20].map(fn (*xs) {
xs.size()
})) # => [1, 1]
inspect([1, 2, 3].reduce(0, fn (a, *xs) {
a + xs.size()
})) # => 3
This is why range / iota (variadic builtins) can be passed directly as
callbacks. The rule is identical under the VM, --jit, and AOT.
Callback parameter types. A type annotation on a callback parameter is
enforced on every invocation, exactly like a direct call — the first
wrong-typed element raises a TypeError:
[1, 'x'].map(fn (v: Long) {
v * 2
}) # !! parameter 'v' expects Long
18.3 Object methods
| Signature | Description |
|---|---|
o.size() -> Long | Number of own properties. |
o.empty() -> Bool | Whether size() == 0. |
o.presence() -> Object | Nil | o unchanged (same Object, not a copy) if non-empty, else nil — pairs with ??/?. for "use it if there's anything there". |
o.keys() -> Array | Array of keys in insertion order (matches display order — §8). |
o.values() -> Iterator | Lazy iterator of values in insertion order — the value-only view of o.iter() (which yields (key, value) pairs). Chains / collects like any iterator. |
o.has(key: String) -> Bool | Whether o has an own property named key, or (for a class instance) a method of that name. Ignores built-in method names. |
o.get(key, fallback) -> Any | The value for key, or fallback if absent. Read-only — never inserts. |
o.get_or_put(key, init) -> Any (mutating) | The value for key; on a miss, store init and return it (sharing storage, so o.get_or_put(k, [] ).push(x) grows the stored array). When init is a function it is called lazily — only a miss pays for it: `o.get_or_put(k, |
o.remove(key: String) -> Nil (mutating) | Delete the property named key if present. |
A String and a byte-equal StringView (e.g. s[0..2]) are the same key — they hit the same slot across every operation above and o[key] = v.
o = {b: 2, a: 1, c: 3}
# keys() and values() both walk in insertion order:
inspect(o.keys()) # => ['b', 'a', 'c']
inspect(o.values().collect()) # => [2, 1, 3]
inspect(o.has('a')) # => true
# An absent key yields the fallback and is not inserted:
inspect(o.get('z', 0)) # => 0
# Group items by their first letter (StringView keys):
mut groups = {}
for w in ['apple', 'avocado', 'banana'] {
groups.get_or_put(w[0..1], || []).push(w)
}
inspect(groups) # => {mut a: ['apple', 'avocado'], mut b: ['banana']}
mut p = {a: 1, b: 2}
p.remove('a')
inspect(p) # => {b: 2}
o.iter() yields (key, value) tuples and to_object() consumes them, so
a whole table is one expression — no mut accumulator to declare:
mut prices = {apple: 100, banana: 80}
# Remap the values:
inspect(prices.iter().map(|(k, v)| (k, v * 2)).to_object())
# => {mut apple: 200, mut banana: 160}
# Invert it:
inspect(prices.iter().map(|(k, v)| (v, k)).to_object())
# => {mut 100: 'apple', mut 80: 'banana'}
18.4 Special identifiers
| Identifier | Visibility | Meaning |
|---|---|---|
fn | Inside every function | The currently-executing function. |
self | Inside method calls | The method's receiver. |
18.5 Iterator protocol
for x in expr { ... } (§12) requires expr to participate in the
iterator protocol. The protocol uses three well-known method names
on Object (and its subtype Array), with a has_next gate in front
of each next:
| Method | Shape | Called on | Returns |
|---|---|---|---|
iter | fn () -> Object | an Iterable | an Iterator (may be self) |
has_next | fn () -> Bool | an Iterator | whether another element is available |
next | fn () -> Any | an Iterator | the next element (called only when has_next() was truthy) |
A built-in iterator called past its end yields nil rather than
raising, the same as a drained generator. Since nil is also a
perfectly good element, gate on has_next() to tell the two apart.
Contract, enforced at property assignment: binding iter,
has_next or next to a function that takes arguments raises
DropContractError: type error: '<name>' must be a Function taking no arguments. at the assignment site (mirrors the drop contract — §17).
A value that is not a Function is data and binds like any property; it
is never a protocol member, so an Object that merely has a next key
is not an iterator, and one whose iter key holds data is walked by its
keys. A lookalike field (has_next_at) was never affected.
Checked when the iterator is opened: the for-in head validates
before the first step, and a lazy chain validates when a terminal
(collect, count, ...) first drives it — building the chain pulls
nothing, so that is where the error appears:
| Situation | Error | Reported at |
|---|---|---|
iter() returned a non-Object | TypeError: type error: iter() did not return an Object | the iterable expression |
iterator missing has_next or next | TypeError: type error: iterator missing has_next()/next() | the iterable expression, or the terminal that drove the chain |
Optional dispose: if the iterator has a zero-arity dispose, it
is called on every exit path — normal drain, break, early return, or
an exception leaving the loop, whichever side raised it: the loop body,
has_next() / next(), or a generator body that throws while suspended.
Generators use it to run the defers registered inside a suspended body.
It runs after the iteration's own bindings are released, so the loop
variable's and the body's drops precede it on every exit alike. A
dispose that itself throws while an exception is already unwinding is
swallowed, so the enclosing catch still sees the original error — and
a loop variable whose pattern did not match the element is such an
exception, so the mismatch is the error that reaches the catch.
The same contract covers the terminal iterator methods: whoever
drives the protocol closes it. Every terminal — draining
(collect, count, ...), early-exiting (find, any, first, ...),
or one that throws mid-drain (from its callback or the producer) —
calls dispose exactly once on the way out, after the result is
computed and before it is handed back. On the throwing paths the
original error propagates and a throwing dispose is swallowed; on a
successful drain a throwing dispose propagates, like break.
A lazy chain is one iterator from the consumer's side, so closing it closes
its source: map / filter / take / zip / … forward dispose to the
iterator(s) they pull from (both, for the two-source combinators). So
for x in gen().map(f) { break } runs gen()'s defers just like
for x in gen() { break } does. A chain over a source with no dispose
carries none itself.
let countdown = fn (start) {
mut i = start
{
iter: fn () { self },
has_next: fn () { i > 0 },
next: fn () { let v = i; i = i - 1; v }
}
}
for v in countdown(3) { inspect(v) } # 3, 2, 1
Iterators are also Iterables: an iterator's iter should return
itself, so for x in some_iterator { ... } works without a separate
Iterable wrapper.
Built-in iterables:
| Type | iter() yields | Order |
|---|---|---|
Array | elements | index order (0..size-1) |
Object | (key, value) tuples | insertion order (matches o.keys()) |
String | one-scalar String per UTF-8 code point | byte order |
String iteration walks the UTF-8 buffer lazily, so iterating a
100 MB string with break after a few steps does not materialize the
rest. The yielded values are 1-scalar Strings, not integer code
points — use .map on the iterator to project into whatever shape
you need.
Iterating an Object yields (key, value) pairs, so the
natural loop is for k, v in obj and obj.iter() is the entries view
(it chains and collects like any iterator). The single-axis views are
obj.keys() (an Array) and obj.values() (a lazy iterator):
for k, v in {a: 1, b: 2} {
inspect("{k}={v}")
} # 'a=1' then 'b=2'
for k in {a: 1, b: 2}.keys() {
inspect(k)
} # 'a' then 'b'
for v in {a: 1, b: 2}.values() {
inspect(v)
} # 1 then 2
Object iter and mutation: iteration is over a snapshot of the keys
taken at loop start, so mutating the object in the loop body is safe —
there is no fail-fast guard. Keys added during the loop
are not visited; keys removed are skipped; each value is read live
at its step. This makes the setdefault pattern — adding derived entries
while iterating — work without copying. The for k, v in obj sugar uses
the same protocol.
mut o = {mut x: 1, mut y: 2}
for k, v in o.iter() {
o[k] = 99
} # update existing values
inspect(o.x) # => 99
mut books = {a: ('alpha', 1)}
for reading, n in books.values() {
# add aliases while iterating
if !books.has(reading) {
books[reading] = (reading, n)
}
}
inspect(books.has('alpha')) # => true
Iterator methods: any Object exposing the iterator interface —
next together with has_next, or iter — picks up the lazy iterator
method set below, which drives the receiver through has_next() /
next(). This means a user iter() result (a plain {has_next, next}
object) and a generator (§11 "Generators") chain the same combinators
as a built-in iterator, not just range/array iterators. Non-terminal methods return
a new Iterator; terminal methods consume the iterator and return a
concrete value.
fn nums() {
yield 1
yield 2
yield 3
yield 4
}
inspect(nums().filter(|x| x % 2 == 0).map(|x| x * 10).collect()) # => [20, 40]
| Non-terminal | Result | Notes |
|---|---|---|
it.map(f) | Iterator | yields f(x) for each upstream x |
it.filter(p) | Iterator | yields only x where p(x) is truthy |
it.take(n) | Iterator | first n elements, then done |
it.skip(n) | Iterator | discards first n elements on first next() |
it.take_while(p) | Iterator | yields until the first p(x) that is falsy |
it.skip_while(p) | Iterator | discards the leading run where p(x) is truthy, then yields the rest unconditionally |
it.step_by(n) | Iterator | the first element, then every n-th one after it; n must be at least 1 |
it.distinct() | Iterator | first occurrence of each element, later duplicates dropped (elements must be hashable) |
it.tap(f) | Iterator | runs f(x) for its side effect and passes x through unchanged — a probe for a lazy chain |
it.scan(init, f) | Iterator | running fold: yields acc = f(acc, x) at each step, starting from init. init itself is not yielded, so the output length matches the input's |
it.flatten() | Iterator | removes one level of nesting; each element must be iterable (same coercion as flat_map) |
it.chunk_by(f) | Iterator | groups adjacent elements sharing a key f(x) into Arrays. A key that reappears after a different one starts a new run — unlike group_by, which buckets by key regardless of position |
it.chunks(n) | Iterator | groups elements into Arrays of n (the last group may be shorter); n must be at least 1 |
it.windows(n) | Iterator | sliding window of the last n elements as an Array, advancing by one each step; n must be at least 1 |
it.flat_map(f) | Iterator | f(x) must return an iterable; results concatenated |
it.chain(other) | Iterator | yields it then other |
it.zip(other) | Iterator | yields {first, second} pairs; stops at the shorter side |
it.enumerate() | Iterator | yields (index, value) tuples with index starting at 0. Also a direct Array method (arr.enumerate()) returning a lazy iterator. |
| Terminal | Result | Notes |
|---|---|---|
it.collect() | Array | materialize into an Array |
it.join(sep) | String | concatenate elements with sep between them (each rendered as by to_string); like Array.join, so xs.map(...).join(",") needs no intermediate .collect() |
it.for_each(f) | Nil | invoke f(x) for side effects |
it.reduce(init, f) | Any | left fold: acc = f(acc, x) starting from init |
it.find(p) | Any | nil | first x where p(x) is truthy, else nil |
it.any(p) | Bool | true if any p(x) is truthy |
it.all(p) | Bool | true if every p(x) is truthy (empty → true) |
it.count() | Long | number of elements consumed |
it.first() | Any | nil | first element, else nil; pulls only one |
it.last() | Any | nil | last element, else nil |
it.nth(n) | Any | nil | element at 0-based n, else nil; pulls only n + 1. Negative n raises ValueError |
it.position(p) | Long | nil | 0-based index of the first x where p(x) is truthy, else nil (unlike Array.index_of, which answers -1) |
it.contains(v) | Bool | true if some element == v (as Array.contains) |
it.sum() | Long | Float | sum of all elements; Long while every element is a Long, Float once any element is one (empty → 0) |
it.product() | Long | Float | product of all elements, promoting like sum (empty → 1) |
it.min() | Any | smallest element, compared numerically; the element is returned, so its own type survives. Throws on empty |
it.max() | Any | largest element, same rule as min |
it.min_by(f) | Any | element with the smallest key f(x); ties keep the earlier one. Throws on empty |
it.max_by(f) | Any | element with the largest key f(x); ties keep the earlier one. Throws on empty |
it.to_set() | Set | members in first-seen order, duplicates dropped |
it.to_object() | Object | (key, value) tuples into an Object — the inverse of Object.iter(). Keys in first-seen order, a repeat overwriting in place; entries immutable. A non-2-tuple element raises TypeError |
it.group_by(f) | Object | buckets elements into Arrays keyed by f(x), in first-seen key order |
it.partition(p) | Tuple | (matching, non_matching) in one pass, order preserved in both halves |
it.unzip() | Tuple | (Array, Array) — the inverse of zip. Each element must be a 2-element Tuple or a {first, second} Object (either pair spelling is accepted); anything else raises TypeError |
Every terminal disposes the iterator it drove (see Optional
dispose above), including the early-exiting ones — so after
it.find(p) the source is closed, and a follow-up terminal on the same
it yields nothing ([] from collect), exactly as after a break:
fn g() {
yield 1
yield 2
yield 3
}
let it = g()
inspect(it.find(|x| x == 2)) # => 2
inspect(it.collect()) # => []
The find also ran g()'s defers; the empty second answer is the
disposed source reporting done.
Eager vs lazy: Array has its own eager map / filter /
for_each / reduce / find / any / all / flat_map (§18.2)
which all return a new Array. Calling them on an Array dispatches
to the eager versions; call .iter() first to opt into the lazy
chain. The eager form is the default because a chain that materialises
each step is easier to reason about; laziness is what you ask for when
the sequence is large or unbounded.
# doctest: skip
# Eager: allocates two intermediate Arrays
arr.map(f).filter(g)
# Lazy: single pass, no intermediate Arrays
arr.iter().map(f).filter(g).collect()
# Lazy with early termination: only touches 4 elements
range(1000000).filter(f).map(g).take(4).collect()
User-defined example:
countdown = fn (start) {
mut i = start
{
iter: fn () {
self
}, # Iterator is its own Iterable
has_next: fn () {
i > 0
},
next: fn () {
v = i
i -= 1
v
},
}
}
for x in countdown(3) {
inspect(x)
} # 3, 2, 1
JIT: everything in this section — for-in driving the protocol,
user-defined iterators, lazy iterator method chains, range,
String.code_points() / .graphemes(), and Array-eager method
chaining on [...].map(...).filter(...) — runs under --jit with
VM-equivalent semantics. The Array/String/keys fast paths
stay native (one load per element); iterator-protocol driving pays a
per-step closure dispatch, which is inherent to a dynamic-language
iterator chain. Use iota + Array.map / .filter / .reduce
when you want eager materialization and maximum throughput; use
range + lazy methods when you want constant-memory streaming.
19. Core built-in functions
The functions below are part of the language proper: they are bound
into every execution environment as global names and cannot be
replaced. The first group (to_long / to_float / to_string,
type_of) is tied to language semantics — source-position errors,
type introspection, and the display convention. The second group
(range, iota, grid) provides the canonical integer-sequence
factories; both backends recognise range/iota for fusion /
specialisation, and they are the standard form used in for-in
loops throughout the language. repeat is a related but separate
eager Array constructor — n copies of one value rather than an
integer sequence. The matcher family (assert_true /
assert_eq / assert_throws / assert_close / etc.) is a third
group of globals — see
docs/stdlib.md for the full reference. The broader
standard library (namespaced under Math, IO, Sys) is also
documented in stdlib.md. Output primitives inspect, print, and
println are CLI-installed globals (§22).
All of these globals are first-class values: bind one to a variable or hand it to a higher-order function and it behaves like any closure, on both backends.
inspect([1, 2, 3].map(type_of)) # => ['Long', 'Long', 'Long']
inspect([1, 2, 3].map(range).map(|r| r.collect())) # => [[0], [0, 1], [0, 1, 2]]
let f = range
inspect(f(0, 10, step: 2).collect()) # => [0, 2, 4, 6, 8]
A direct call is still the fast path; the closure form is used only
when the name appears in value position. range / iota accept their
1-2 positional bounds and range's keyword-only step through that
closure too (range(0, 10, **{step: 2}) works), with the same
ArityError / unknown-keyword diagnostics as the direct call.
to_long(v: Any, *, base: Long = 10) -> Long
Convert v to Long:
Long→ itself.Float→ truncated toward zero.to_long(3.7) == 3,to_long(-3.7) == -3.String→ parsed as a signed integer inbase; leading/trailing whitespace is allowed, anything else fails.- Other types raise
type error.
base is keyword-only and ranges over 2–36. For base 16 / 8 / 2 the
matching 0x / 0o / 0b prefix is accepted, so a literal copied out
of source parses as itself. Naming a base for a non-String v is an
error rather than a silently ignored argument — keyword-only also keeps
to_long a one-parameter function, so map(to_long) still binds.
Throws: type error at L:C. on an unparseable string, a
non-numeric / non-string argument, or a base given for one;
ValueError for a base outside 2–36.
inspect(to_long('42')) # => 42
inspect(to_long('-7')) # => -7
inspect(to_long(3.9)) # => 3
inspect(to_long('ff', base: 16)) # => 255
inspect(to_long('0b1010', base: 2)) # => 10
to_float(v: Any) -> Float
Convert v to Float:
Float→ itself.Long→ promoted toFloat(exact for absolute values up to 2⁵³; larger magnitudes may lose precision).String→ parsed as a decimal or exponent-form float; leading / trailing whitespace is allowed.- Other types raise
type error.
inspect(to_float(3)) # => 3.0
inspect(to_float('1.5')) # => 1.5
inspect(to_float('1e-5')) # => 1e-05
to_string(v: Any) -> String
Convert v to its display form (same formatting as interpolation
inserts — strings come through unquoted). See §8 for the display
convention. Float uses the shortest round-trip decimal and always
carries either a decimal point or an exponent, so the type is
visually distinguishable from Long.
inspect(to_string(42)) # => '42'
inspect(to_string(1.0)) # => '1.0'
inspect(to_string(1e-5)) # => '1e-05'
inspect(to_string([1, 2])) # => '[1, 2]'
inspect(to_string('hi')) # => 'hi'
class_of(v: Any) -> Object?
The class object that built v, or nil. type_of answers what a value
is called; this answers with the thing itself — which is where a class's
statics live, and where a decorator that took the class and returned it put
whatever it added. It is the only way to reach either from an instance,
since neither is copied onto one.
let mark = fn (cls) {
cls.marked = true
cls
}
@mark
class Point {
new(x, y) {
self.x = x
self.y = y
}
}
let p = Point(4, 5)
inspect(class_of(p) == Point) # => true
inspect(class_of(p).marked) # => true
inspect(class_of(p).name) # => 'Point'
inspect(class_of(p).name == type_of(p)) # => true
A class object answers class_of with itself, so the read is idempotent.
An enum variant answers with its enum — the object its variants are read
from — and type_of names the variant, so the two together say which
variant of which enum a value is:
enum Mirroring { Horizontal, Vertical }
enum Shape { Circle(Float), Rect(Float, Float) }
inspect(class_of(Mirroring.Vertical) == Mirroring) # => true
inspect(class_of(Shape.Circle(1.0)) == Shape) # => true
inspect(type_of(Shape.Circle(1.0))) # => 'Circle'
A variant names its enum without keeping it alive, so once nothing else
holds the enum — one declared inside a function that has returned, and
handed out none of its constructors (a constructor read off the enum
holds it) — its variants answer nil. So does a variant rebuilt elsewhere, received from
a Channel or read back from a SharedBuffer, as a class instance
received from a Channel does. Either way it still prints and compares
as its variant.
Everything with no class object of its own answers nil, which is more
than it sounds: a plain Object, a primitive, and a Range. So
class_of(v).name is not a second spelling of type_of(v): only
class-sugar values have both.
inspect(class_of(42)) # => nil
inspect(class_of({a: 1})) # => nil
inspect(class_of(1..3)) # => nil
type_of(v: Any) -> String
Return the name of v's type, in the vocabulary type annotations speak:
one of 'Nil', 'Bool', 'Long', 'Float', 'String',
'StringView', 'Array', 'Object', 'Function', 'Tensor',
'Tuple', 'Set', 'Range' — or, for a value some class built, that
class's own name (an enum variant answers with the variant's). A class
object itself is a 'Class'. So the set is open, and type_of(v) == 'T'
answers the same question v: T asks of a parameter, except that an
annotation also accepts a subtype (Object accepts any instance) where
this names the exact one.
inspect(type_of(42)) # => 'Long'
inspect(type_of(1.5)) # => 'Float'
inspect(type_of('hi')) # => 'String'
inspect(type_of([1, 2])) # => 'Array'
inspect(type_of((1, 2))) # => 'Tuple'
inspect(type_of({1, 2})) # => 'Set'
inspect(type_of({a: 1})) # => 'Object'
inspect(type_of(1..3)) # => 'Range'
range(n: Long, *, step: Long = 1) -> Iterator / range(start: Long, end: Long, *, step: Long = 1) -> Iterator
Lazy integer-sequence factory: returns an Iterator (§18.5) that
yields integers one at a time. Use with for-in or iterator method
chains to iterate in constant additional memory regardless of
the range size.
range(n)yields0, 1, ..., n-1. Ifn <= 0, the iterator completes immediately.range(start, end)yieldsstart, start+1, ..., end-1. Ifstart >= end, completes immediately.step:is keyword-only. Positive values count up (exclusive end); negative values count down (exclusive end).step: 0raisesValueError.
for i in range(5) {
inspect(i)
} # 0, 1, 2, 3, 4
for i in range(2, 6) {
inspect(i)
} # 2, 3, 4, 5
for i in range(0, 10, step: 2) {
inspect(i)
} # 0, 2, 4, 6, 8
for i in range(5, 0, step: -1) {
inspect(i)
} # 5, 4, 3, 2, 1
# Constant memory even for huge bounds
for i in range(1000000000) {
if i > 3 {
break
}
inspect(i)
}
JIT: range returns a JIT-native iterator Object, and the
range(N).<HOF>(...) method-chain pattern is fused into a direct
counter loop. See §18.5.
iota(n: Long) -> Array / iota(start: Long, end: Long) -> Array
Eager counterpart to range: materialise an Array of the same
sequence. Named after APL / C++ std::iota / Scheme SRFI-1. Prefer
range for for-in loops; use iota when you actually need the
full Array (e.g. to index into it, or to pass to a function that
expects an Array).
iota(n)returns[0, 1, ..., n-1]. Ifn <= 0, an empty array.iota(start, end)returns[start, start+1, ..., end-1]. Ifstart >= end, an empty array.
inspect(iota(3)) # => [0, 1, 2]
inspect(iota(2, 5)) # => [2, 3, 4]
inspect(iota(5, 2)) # => []
repeat(n: Long, value: Any) -> Array
Materialise an Array of n copies of value. n < 0 raises
ValueError, and so does an n too large to allocate (repeat() result is too large); n == 0 returns []. Every copy is the same
value — if it's a mutable reference (Array/Object), all n
slots alias one instance, the same sharing [value] * n gives in
Python or Array(n).fill(value) in JavaScript.
inspect(repeat(3, 0)) # => [0, 0, 0]
inspect(repeat(0, "x")) # => []
let shared = repeat(2, [])
shared[0].push(1)
# Both slots are the same Array
inspect(shared) # => [[1], [1]]
grid(x_range: Range, y_range: Range) -> Iterator
Lazy cartesian-product factory over two bounded integer ranges:
returns an Iterator (§18.5) that yields (x, y) Tuples, x varying
fastest — the same order the nested loop for y in y_range { for x in x_range { ... } } walks, in one line and constant additional
memory. Both arguments must be bounded Range values (a..b /
a..=b, by step); an open-ended range or step: 0 raises the
same errors for-in over a Range does. If either range is empty,
the whole product is empty.
for (x, y) in grid(0..3, 0..2) {
inspect((x, y))
} # (0,0), (1,0), (2,0), (0,1), (1,1), (2,1)
inspect(grid(0..2, 0..2).collect())
# => [(0, 0), (1, 0), (0, 1), (1, 1)]
# An empty x or y range makes the whole product empty.
inspect(grid(0..0, 0..3).collect()) # => []
__ARGS__ (variadic catch-all binding)
Inside any function body, the implicit local __ARGS__ is bound to
an Array of positional arguments that overflowed the declared
parameters. Use it when a fn takes a variable number of trailing
values without declaring an explicit **rest (which catches
keyword args, not positional).
let logger = fn (level) {
inspect("[{level}] " + __ARGS__.join(' '))
}
logger('info', 'building', 'fizzbuzz') # → '[info] building fizzbuzz'
Because that overflow always has somewhere to land, passing a user
function more positional args than it declares is not an error —
unlike a built-in, which raises ArityError (§15). A bare * in the
parameter list is the way to cap positionals on a user function:
overflow past it is a TypeError rather than reaching __ARGS__.
__ARGS__ does not receive keyword arguments — those go through
the explicit param list or **rest. Prefer the explicit *args
parameter (see Parameters) when you want a named overflow binding
and variadic dispatch; __ARGS__ is the implicit form that always
accompanies it.
Function introspection
Every Function value exposes three read-only properties for
introspection:
| Property | Type | Description |
|---|---|---|
fn.name | String | Source-level declaration name (fn name(...)) or "" for anonymous fns |
fn.return_type | String | Return-type annotation (fn f() -> X) or "" if unannotated |
fn.params | Array<Object> | Per-parameter metadata; each entry has name / mut / type / has_default / kw_only / kwargs_rest |
fn greet(name: String, *, prefix = "hi") {
"{prefix}, {name}"
}
inspect(greet.name) # => 'greet'
inspect(greet.return_type) # => ''
let ps = greet.params
inspect(ps.size()) # => 2
inspect(ps[0].name) # => 'name'
inspect(ps[0].type) # => 'String'
inspect(ps[1].kw_only) # => true
inspect(ps[1].has_default) # => true
fn.params returns a fresh Array each access; mutating it has no
effect on the function. Multifn dispatchers expose the first
registered method's signature (overloads after the first do not
appear).
20. Multimethods
Multiple top-level fn name(params) body declarations sharing a name
form a multimethod when their parameters have differing type
annotations (§14). At a call site, the most specifically matching
method is selected based on the runtime types of the arguments.
fn area(s: Long) { s * s }
fn area(s: Float) { s * s }
fn area(s: String) { s.size() }
area(5) # → 25 (Long)
area(5.0) # → 25.0 (Float)
area("hello") # → 5 (String)
Anonymous function expressions let f = fn(...) {...} are unaffected.
Multimethods only apply to top-level fn name(...) declarations.
The overloads of a name belong to the scope that declares them. An if
or cond arm is not a scope of its own (§6), so a fn name written in
one joins the overloads of the scope around it, from the moment the arm
runs:
fn describe(x: Long) { 'a number' }
if true {
fn describe(x: String) { 'a string' }
}
inspect(describe(1)) # => 'a number'
inspect(describe('hi')) # => 'a string'
Default parameters and arity. A method with default parameters matches any call whose positional-argument count is between its required count (params without a default) and its total param count; the unsupplied tail is filled from the defaults. Among equally type-specific matches, the one that fills fewer parameters by default wins (a more exact arity is more specific).
A keyword argument may also cover a required parameter that the
positional arguments didn't fill — keywords contribute to which
methods are applicable. Selection itself still scores on the
positional arguments only, so two overloads that differ solely by a
keyword (or keyword-supplied type) are ambiguous, not silently
disambiguated. A required parameter supplied by neither a positional
argument nor a keyword leaves the call unmatched (DispatchError).
fn at(a, b = 10) { a + b }
at(1) # → 11 (b defaulted)
at(1, 2) # → 3
at(1, b: 2) # → 3 (b by keyword)
at(a: 1, b: 2) # → 3 (a — required — covered by keyword)
at(b: 9) # !! DispatchError — required `a` not supplied
Syntax
MULTIFN_DECL <- 'fn' IDENTIFIER PARAMETERS RETURN_TYPE? BLOCK
When fn name(...) appears in statement position, this rule takes
precedence. Anonymous fn(...) {...} continues to be an expression
(closure), as before.
Specificity rules
For each argument i, the parameter annotation and the runtime type of the argument yield a specificity score:
| Annotation | Argument | Score |
|---|---|---|
Any (no annotation) | anything | 0 |
Object | class instance (e.g. Square, Circle) | 1 |
Exact match (Long, Float, ..., concrete class name) | same type | 2 |
| otherwise | — | no match |
Three more tiers sit between the Object row and an exact match,
from least to most specific: a trait the argument conforms to, a union
(Ok | Err), and the enum name of a variant argument (Result). Two
sit above an exact match: a qualified variant (Result.Ok beats a bare
Ok), then a generic that carries type arguments (Array<Long> beats
a bare Array). T? scores at the union tier whatever T would score
above it, so a bare T always wins over T?.
The per-argument scores form a tuple. A method is selected when its score tuple is at least as large as the other in every position and strictly greater in at least one. If two tuples compare equal, an ambiguous dispatch error is raised. If two tuples are incomparable (each wins on some position), the earlier-declared method takes precedence.
fn label(x: Long, y) { "long-any" }
fn label(x, y: Long) { "any-long" }
fn label(x: Long, y: Long) { "long-long" } # most specific
fn label(x, y) { "any-any" }
label(1, 2) # → "long-long"
label(1, "x") # → "long-any"
label("x", 1) # → "any-long"
label("x", "y") # → "any-any"
Class instance dispatch
Instances created by class declarations (§10) dispatch on their
class name:
class Square { new(side) { self.side = side } }
class Circle { new(r) { self.r = r } }
fn shape_area(s: Square) { s.side * s.side }
fn shape_area(c: Circle) { 3.14 * c.r * c.r }
shape_area(Square.new(4)) # → 16
shape_area(Circle.new(2)) # → 12.56
This relies on the class: String property that class sugar attaches
to each instance (§10). The Object annotation is looser and matches
any class instance.
Same-name, same-signature redeclaration
A subsequent declaration whose parameter type sequence matches an existing entry exactly overwrites that entry's body. The semantics are intended for REPL iteration.
fn greet(name: String) { "hi, {name}" }
fn greet(name: String) { "hello, {name}" } # overwrites
greet("alice") # → "hello, alice"
Overwriting is confined to one execution of the declaring scope. A
declaration inside a function or a loop body builds a fresh overload set
every time it runs, so each body keeps the captures of the activation that
declared it — an earlier fn value is never re-pointed at a later run's
bodies. The same holds for the class overload sets below.
fn make(tag) {
fn m(a) { "{tag}-one" }
fn m(a, b) { "{tag}-two" }
m
}
let a = make("A")
let b = make("B")
[a(1), b(1)] # → ["A-one", "B-one"]
Coexistence with existing features
- Ordinary local bindings
let f = fn(...) {...}continue to work unchanged. - Method dispatch on
obj.method()is unaffected. Methods are still defined inside aclassbody and called throughself.. - A method whose parameters are all
Anyserves as a catch-all.
Constraints
- Top-level / free functions only. Nested declarations inside a block, and class methods, are not subject to this mechanism.
- Errors. With no matching method the runtime raises
no matching method; with a tie in specificity it raisesambiguous dispatch. Both halt the program immediately rather than surfacing as catchable runtime exceptions (§15, §24).
Keyword arguments and multimethods
Dispatch picks on the positional argument types only (Julia-style
kwsorter). Keyword arguments and ** splats flow through dispatch
into the picked method's signature, where they bind via the regular
kwargs rules (§11).
fn paint(s: String, *, color = "red") { "{color} {s}" }
fn paint(n: Long, *, color = "blue") { "{color} {n}" }
paint("circle") # → "red circle" (String)
paint(7) # → "blue 7" (Long)
paint("box", color: "green") # → "green box"
paint(7, **{color: "gold"}) # → "gold 7"
Each method's own kw-only defaults, **rest catch-all, and parameter
names are independent — only the positional signature participates in
dispatch.
Method multidispatch (instance methods)
A class may declare several instance methods with the same name but
different positional-param-type signatures. They merge into one
dispatcher; obj.method(args) then picks the overload on the runtime
types of the explicit arguments. The receiver is fixed by the property
lookup — only the arguments are scored — and self is bound into the
picked overload:
class Calc {
new() {}
go(x: Long) {
"long: {x}"
}
go(x: String) {
"string: {x}"
}
go(x: Long, y: Long) {
"sum: {x + y}"
}
}
let c = Calc.new()
c.go(1) # → "long: 1"
c.go("a") # → "string: a"
c.go(2, 3) # → "sum: 5"
The same scoring, default-param, *args, kwarg, and **rest rules as
free-function multimethods apply (a keyword argument flows into the
picked overload). A call with no matching overload raises a catchable
DispatchError, exactly like a free-function multimethod.
A class declaring a name once keeps a plain method (no dispatcher,
no overhead). The following stay compile-time errors: two methods with
an identical signature, a field and a method sharing a name, and a
duplicate field. Constructors (new) and operator/dunder methods
(__add__, __eq__, __call__, …) overload too (see below).
Method multidispatch (static methods)
Static methods overload the same way. Same-named static methods with
different positional-param-type signatures merge into one dispatcher on
the class object; Cls.method(args) picks on the argument types. No
self participates — a static call binds none:
class Vec {
static make(x: Long) {
"long: {x}"
}
static make(x: String) {
"str: {x}"
}
static make(x: Long, y: Long) {
"pair: {x}, {y}"
}
}
Vec.make(5) # → "long: 5"
Vec.make("hi") # → "str: hi"
Vec.make(1, 2) # → "pair: 1, 2"
Static and instance overload sets are independent: a static and an instance method may share a name, each with its own overloads. Two static methods with an identical signature, and a static field clashing with a static method name, stay compile-time errors.
Constructor multidispatch (new)
A class may declare several new bodies with different
positional-param-type signatures. C(args) (or C.new(args)) dispatches
on the runtime types of the arguments, exactly like method overloads:
class Point {
new(x: Long, y: Long) {
self.tag = "xy"
self.x = x
self.y = y
}
new(s: String) {
self.tag = "str"
self.x = s.size()
self.y = 0
}
new() {
self.tag = "empty"
self.x = -1
self.y = -1
}
}
Point(3, 4).tag # → "xy"
Point("hi").tag # → "str"
Point().tag # → "empty"
The overload is picked before any instance is allocated, so a call
matching no overload raises a catchable DispatchError with no side
effects — no instance is built and no field initializer (§10) runs. When
an overload is picked, its declared field initializers run once, then its
body, with self immutable throughout (a self = … reassignment raises
ImmutableError in every overload, as for a single new).
Default parameters, keyword arguments, and *args follow the same rules
as method overloads. A class declaring new once keeps a plain
constructor (no dispatcher, no overhead); two new bodies with an
identical signature stay a compile-time error.
Dunder / operator multidispatch
Operator methods (__add__, __eq__, __lt__, __index__, __call__,
…, §10) are ordinary instance methods reached through the operator-lookup
path, so several with distinct operand-type signatures merge into one
dispatcher; the operator then picks the overload on the operand's runtime
type:
class Vec {
new(x: Long, y: Long) {
self.x = x
self.y = y
}
__add__(o: Vec) {
Vec(self.x + o.x, self.y + o.y)
} # elementwise
__add__(n: Long) {
Vec(self.x + n, self.y + n)
} # scalar
}
let v = Vec(1, 2)
v + Vec(10, 20) # → Vec(11, 22) (Vec overload)
v + 5 # → Vec(6, 7) (Long overload)
Commutative auto-reflection is unaffected: 5 + v still reflects to
v.__add__(5) and picks the Long overload. An operand matching no
overload raises a catchable DispatchError (kind, message, and position
agree across backends) — the same rule as a method with a typed parameter
that the operand doesn't satisfy. A class declaring a dunder once keeps
a plain method; two dunder bodies with an identical signature stay a
compile-time error.
21. Decorators
A @expr line before a fn or class declaration wraps the
declared value through expr before binding it to the original name.
Stacked decorators apply bottom-up — the one closest to the
declaration runs first, the topmost runs last:
# doctest: skip
@a
@b
fn foo() { ... }
is sugar for foo = a(b(<original fn value>)).
Decorator expression
Each @ is followed by any expression of the form name,
name.attr, or name(args) — i.e. a CALL chain. The expression
is evaluated once at declaration time in the enclosing scope;
its result must be callable (a function or a class with a single
positional parameter). Multi-arg factories — @combine(a, b) — work
by returning a one-argument decorator from the factory call.
Bindings
The decorator receives the unbound declared value and returns the value that ends up in the variable:
let tag = fn (f) {
fn () {
"[{f()}]"
}
}
@tag
fn greet() {
"hello"
}
greet() # "[hello]"
For a class, the decorator gets the class object and returns the object the variable will hold:
let mark = fn (cls) {
cls.marked = true
cls
}
@mark
class Point {
new(x, y) {
self.x = x
self.y = y
}
}
Point.marked # true
Interaction with multimethods
A decorated fn name(...) does not participate in multimethod
dispatch — the decorator's return value binds directly to name.
Combine the patterns by writing the multimethod first and then a
separate decorated wrapper if you need both.
Built-in decorator: @value
@value is not a callable — it marks a class as a value type. The
reason to reach for it is cost. A small data class used in a hot loop pays
an allocation for every intermediate result it produces; a @value class
whose fields are all scalars pays none, because the loop compiles to
machine slots and builds no instance at all:
@value
class Vec2 {
x: Float
y: Float
new(x: Long | Float, y: Long | Float) {
self.x = to_float(x)
self.y = to_float(y)
}
__add__(o) {
Vec2.new(self.x + o.x, self.y + o.y)
}
__mul__(s) {
Vec2.new(self.x * s, self.y * s)
}
}
let DT = 1.0 / 60.0
let G = Vec2.new(0, -9.8)
let mut p = Vec2.new(0, 100)
let mut v = Vec2.new(12, 0)
for _ in 0..600 {
v += G * DT
p += v * DT
}
inspect(Math.round(p.x)) # => 120
Delete the @value line and the loop prints the same number, but each
iteration now builds and frees four intermediate instances — G * DT,
v + …, v * DT, p + … — and reaches __add__ and __mul__ through
the ordinary runtime dispatch. With it, p and v are four machine slots
the loop never leaves. Where those allocations do not matter, an ordinary
class is the right choice: what @value gives up to earn the loop is what
the rest of this section describes.
What it gives up is identity: two instances holding the same fields are
the same value, and nothing a program can do tells them apart. It is §5's
value/reference split, applied to a class — an instance becomes what a
Long already is, its contents and nothing else, and a value the program
cannot tell apart from a copy is one the compiler may keep in registers,
take apart, and rebuild. An ordinary instance has an identity, so two names
for it are two handles on one mutable thing:
class Mutable {
x: Float
new(x: Float) { self.x = x }
}
let p = Mutable.new(1.0)
let q = p
q.x = 5.0
inspect(p.x) # => 5.0
The same program over a @value class does not get that far. The write is
refused, which is what leaves two names for one value nothing to disagree
about:
@value
class Frozen {
x: Float
new(x: Float) { self.x = x }
}
let p = Frozen.new(1.0)
let q = p
inspect(p == Frozen.new(1.0)) # => true
q.x = 5.0 # !! immutable property 'x'
The contract has five clauses, each closing one way a program could observe an instance's identity:
| Clause | Checked |
|---|---|
Every instance field is declared, and its type is Long, Float, Bool, or another @value class | At the declaration (SyntaxError) |
| No member writes a field the class does not declare — two instances of one value class cannot differ in shape | At the declaration (SyntaxError), and while new runs (ImmutableError) |
No drop method or field — a destructor observes which instance died | At the declaration (SyntaxError) |
Not also @packable — a packed instance aliases shared bytes | At the declaration (SyntaxError) |
Frozen once new returns: no field written, added, or removed | At the write (ImmutableError) |
Inside new the instance is an ordinary object, so a constructor assigns
its declared fields the usual way. What it cannot do is add one:
# doctest: skip
@value
class Bad {
x: Float
new(k: String) {
self.x = 1.0
self.z = 2.0 # SyntaxError at the declaration: `self.z` writes a field
# the class does not declare
self[k] = 2.0 # ImmutableError while `new` runs — the declaration cannot
# see a computed key, so the instance refuses it instead
}
}
The freeze closes the rest of the shape, not just the write above:
# doctest: skip
v.z = 9 # ImmutableError: cannot add property 'z' to a @value instance
v.remove('x') # ImmutableError: cannot remove property 'x' from a @value instance
A @value class gets eq and hash derived from its fields — the same
pair @derive(Eq, Hash) generates — so it is a key matched by its fields
and == compares it by structure. Its fields are declared scalars, each
of one type, so the two never part the way they can for a derived class
holding 1 beside 1.0: Set membership and Object keys agree with
==. hash is the half an ordinary class does not have: without it an
instance is unhashable, so it cannot be a Set member or an Object key at
all. A class that writes its own __eq__ or eq keeps it and opts out of
both, which is what keeps the pair consistent.
Nothing else about the class changes: methods, operators, getters,
statics, match type patterns, keys() and display all behave as they do
for any class (§10). The field types are deliberately narrow — String,
Array, Object and closures carry a body or an identity of their own —
and a field type naming another @value class must name one declared
earlier.
The unboxed form is decided per binding and is all-or-nothing: one use that
is not a field read, a same-class operand or a reassignment — passing v
to an untyped function, storing it in an array — puts that binding on the
ordinary path for its whole scope.
Built-in decorator: @packable
@packable is not a callable — it marks a class as a fixed-layout
struct. Its fields carry a scalar type annotation with an optional
default (x: Float32 = 0.0), and the decorator fixes their byte layout
(C-ABI alignment). Packable classes back SharedBuffer, the zero-copy
buffer shared across isolates — see
stdlib §12 SharedBuffer.
Why @ doesn't collide with the matmul operator
@ is also the binary matrix-multiplication operator (PEP 465).
The two uses never overlap: decorator @ only matches at statement
prefix position, while matmul @ is parsed inside an expression
between two operands and never crosses a newline. So
}\n@deco fn ... is unambiguously two statements (a decorator on a
fresh declaration), not a matmul continuation of the preceding
expression.
22. Command-line interface
culebra [flags] [script.cul | -] [arg ...]
culebra <command> [command-flags]
Everything after the script path is captured verbatim and exposed to the
script as Sys.argv — no -- needed; see
docs/stdlib.md. A standalone -- before the script is an
optional escape hatch: it stops flag parsing, so the next argument is taken
as the script even if it begins with a dash.
- reads the script from stdin instead of a file, e.g.
curl ... | culebra -. Sys.script is nil there, same as in the REPL; a
file actually named - is still reachable by spelling the path, e.g. ./-.
Flags
| Flag | Effect |
|---|---|
--shell | Start the REPL (the default when no script is given). |
--ast | Print the parsed AST instead of running it. |
--debug | Print debug diagnostics while running. |
--jit | Use the LLVM ORC JIT instead of the bytecode VM. |
--jit-faststart | Like --jit, but skips both the IR and the machine-code optimizers, cutting JIT warmup time 3–5x. The steady-state cost tracks how much of the hot work the optimizer can reach: ~0% when it is in the C++/BLAS runtime, ~30% on call-heavy script code, several-fold on tight scalar arithmetic. Implies -O0. Output matches --jit. |
--emit-llvm | With --jit, print the generated IR and exit. |
-O0..-O3 | With --jit, select the LLVM optimization level. Default -O2. |
-h, --help | Print the option / command summary and exit. |
--version | Print the version, the commit a development build came from, and the available backends, then exit. A build of a release tag with a clean checkout prints the bare version. |
Subcommands
The binary carries the toolchain as well. The development subcommands
are documented in tooling.md and the packaging ones in
deployment.md.
| Command | Effect | Reference |
|---|---|---|
test [paths...] | Run tests and doctests (--filter, --doc, --reporter, --bail, --list). | tooling.md §1 |
lint [paths...] | Report static errors and warnings without running (--fix removes unused imports). | tooling.md §2 |
fmt [paths...] | Reformat source to the canonical style (-i in place, -l list, --check). | tooling.md §3 |
dap | Speak the Debug Adapter Protocol over stdin/stdout; an editor launches it. | tooling.md §4 |
build <in.cul> -o <out> | Compile ahead-of-time into a standalone executable (§24 covers the module graph it bundles; culebra build --help lists the cross-compile flags). | deployment.md §1 |
wrap | Build an extended culebra binary that exposes your own C++ classes as builtins. | deployment.md §3 |
toolchain [status|install|uninstall] | Report, install or remove what build needs to link on this host. | deployment.md §1 |
If no script is provided, the REPL is launched automatically. It
always runs on the VM's executor — a prompt line is never a hot loop,
so --jit applies to scripts only and passing it here just prints a
note. Session state is preserved across inputs: let, mut, and
fn declarations from one input are visible to subsequent inputs,
including from inside nested closures.
Each accepted input's result is echoed, except when it is nil —
println(...), a fn declaration and an if without an else branch
all evaluate to nil, and echoing it would double every line of output.
Output written by the program itself (println, print) is unaffected.
The REPL persists input history across sessions. The path is
$CULEBRA_HISTFILE if set, otherwise $XDG_STATE_HOME/culebra/history
when defined, otherwise ~/.culebra_history. History is rewritten
after each accepted line so a crash mid-session doesn't lose it.
CLI-installed globals
The CLI binary adds three globals to the script environment before running user code:
| Global | Aliased to |
|---|---|
inspect | IO.inspect |
print | IO.print |
println | IO.println |
These are convenience shortcuts for the most common output calls;
they point to the same function values that live under IO, so
inspect(x) and IO.inspect(x) are fully equivalent. Both engines bind
them unconditionally, so a program embedded in a C++ host sees the same
three names a script does.
23. Known limitations
- No big integers or bignums;
Longoverflow wraps. - Single-quoted
'...'strings are raw (no escapes, no interpolation, no embedded apostrophes). Double-quoted"..."strings recognize a fixed escape set (\n \r \t \\ \" \{); the{{/}}form for literal braces is not supported. Stringis byte-indexed (size/slicecount bytes); Unicode work goes throughcode_points()/graphemes()iterators.Array/Object/Tuple/Setequality is structural (by value); onlyFunction/Tensorcompare by reference identity.- Type annotations are enforced only at function boundaries and annotated assignments; they do not make the language static.
- Pattern matching has no exhaustiveness check.
- Dot-form property names are identifiers only (
obj.foo). Non-String hashable keys (Long,Float,Bool,Nil,Tuple, enum variants) reach the Object via the subscript path (obj[k]) and live in a sidecar map. RuntimeStringkeys viaobj[k]unify with the shape —obj['x']andobj.xreach the same slot. See §10 "Subscript assignment". - Only
fn name(...)declarations can be generators;yieldanywhere else — a class method, an object property's function, afnexpression, or the top level of a file — is a parse-timeSyntaxError(§11 "Generators"). A method that needs to yield delegates to a named fn declared beside it.
24. Modules
A program may span multiple files. One file imports another by
name; the imported file may export selected bindings to expose
them to importers. The system is static and minimal — paths are
literal strings, imports and exports appear only at the top level,
and the runtime never resolves a name across files at call time.
Module syntax
A module's source is an ordinary culebra file. It exposes bindings
through one or more export statements:
# math_utils.cul
let add = fn (a, b) { a + b }
let sub = fn (a, b) { a - b }
let helper = fn () { ... } # internal, not exported
class Pair {
new (x, y) { self.x = x; self.y = y }
}
export { add, sub, Pair }
A consumer of the module imports it through a single name:
# main.cul
import math from './math_utils.cul'
math.add(2, 3) # 5
let p = math.Pair.new(1, 2)
math.helper # nil — not on the export Object
Statements
| Form | Meaning |
|---|---|
import NAME from 'path' | Loads the module at path and binds its export Object under NAME in the current file. |
export { N1, N2, ... } | Adds each listed local name to the exporting module's export Object. |
The path is a STRING literal (single-quoted, non-interpolated). It is resolved relative to the importing module's directory. Absolute paths are accepted unchanged.
Top-level only
Both import and export may only appear as top-level statements
of a module. They are syntactically rejected inside functions,
if branches, blocks, and so on:
let f = fn () {
import lib from './lib.cul' # SyntaxError
}
This rule guarantees that the dependency graph and the set of exported names are determined at parse time, which is required for both bundled AOT builds and tree-shaking analysis.
Evaluation rules
- Dependency resolution. The driver loads the entry module,
walks its
importstatements to find dependencies, recurses, and records absolute paths in load order. The resulting graph is topologically sorted (every module appears after its dependencies). - Per-module scope. Each dependency is evaluated in its own fresh scope. Top-level bindings inside the module are visible only within that file plus its export Object — they do not leak into the importer.
- Caching. The same absolute path is evaluated at most once per program; subsequent imports retrieve the cached export Object.
- Export Object. After a dependency's body finishes, the
runtime collects every name listed in its
EXPORT_STMTs into a single Object (immutable properties). Multipleexportstatements are merged in source order — a file may issueexportmore than once. - Entry module. The entry module evaluates last and shares the caller-facing scope, so its top-level bindings remain visible to whoever invoked the program.
Exports
- A module without any
exportstatement produces an empty Object.import x from "..."still succeeds; missing members read back asnilper the standard Object property rule (§10). Useful for side-effect-only modules. - Multiple
export { ... }statements in a single file are merged. This lets a file declare a few helpers, export them, declare more, export more. - Listing the same name twice — whether in one
EXPORT_STMTor spread across multiple — is aSyntaxError(parse-time). The name must be defined as a local binding in the module before the export references it.
Errors
| Trigger | kind |
|---|---|
import from a file that doesn't exist | IOError |
| Imported file fails to parse | SyntaxError |
| Circular import (A imports B which imports A) | ImportError |
import or export outside top level | SyntaxError |
Duplicate name in one export statement | SyntaxError |
export { foo } but foo is not defined in the module | NameError |
(Missing-member lookup on an imported namespace is not an
error — mod.unknown returns nil, matching the rule for
plain Objects in §10.)
All errors carry kind / message / line / col and are
surfaced identically by both backends. IOError, SyntaxError,
and NameError join the existing catalog in §15; ImportError
is added for the circular-dependency case (catchable).
Backend equivalence
Both engines produce identical observable behavior for any
module program, and both wire dependencies the same way: every
module's body compiles into one scope-isolated program. After each
dependency's body, the compiled code calls
culebra_runtime_module_register with the absolute path and the
export Object; IMPORT_STMT compiles to culebra_runtime_module_get
with the same path and binds the result.
The module table is a thread_local keyed by absolute path, so
the same module compiled into different JIT sessions on the
same thread observes the same caching rules without sharing
state between threads.
25. Appendix: VM ↔ JIT divergence
This document is normative. The VM's executor and the LLVM lowering
consume the same bytecode from the same compiler and are required to
produce identical observable behavior for every program — same
return values, same side-effect ordering, same error kind / message
/ location. Internal representation is free to differ as long as the
externally observable behavior matches; any deviation in observable
behavior is a bug. A released binary is the independent second
implementation the gates compare against.
Identical observable behavior
The following are guaranteed identical across VM, JIT, and AOT builds:
- Numerics (§7). Long overflow wrap, Float IEEE-754 (NaN, inf,
signed zero), division/modulo by zero raising
ZeroDivisionError, comparison and truthiness rules. - Argument evaluation order. Left-to-right in source order for
positional, keyword, and
**-splat arguments, with||/&&/??short-circuiting at the first decisive operand. The same rule applies to array and object literals. - Method dispatch and UFCS (§10), operator special methods (§10),
__str__display (§10), and the iterator protocol (§18.5). throw/try/catch/defer(§15). Including function-level and top-leveldeferfiring on the throw-unwind path. The JIT lowers throws to LLVMinvoke/landingpadwith the Itanium personality; observable propagation is identical.- Auto-drop (§17). Fires on every refcount-to-zero transition, whether triggered by scope exit, property overwrite, or array rebind. Cascade order — parent before child — is the same. Scope-owned cycle members drop at their scope's exit on both backends; closure-held and top-level cycles fire at the GC backstop (collection-timed, order unspecified — see §17).
- Error reporting. Every
kindlisted in §15 fires under the same trigger condition on both backends;e.message,e.line, ande.colare populated identically. Uncaught errors print asKind: message at L:C.. - Class sugar (§10),
staticmethods, immutableself, and the auto-synthesizedparameters()reflection. - Module-scope evaluation. Statement order, top-level closures, forward-reference resolution, and decorator application all follow the same rules.
Permitted internal optimizations
These change how the program executes but not what it observes:
- Object property storage. Both engines share a process-interned hidden-class (Shape) layout with vector-backed slots; the JIT adds per-callsite inline caches. Iteration order over an Object's keys is insertion order on both.
- HOF inlining. Under
--jit,array.map / filter / for_each / reduceandrange(N).<HOF>(...)/iter.map(λ).collect()compile their lambda bodies directly into the iteration loop. Side-effect order is preserved. - Class method storage. Methods live on a shared per-class meta
object reached via prototype delegation.
obj.mreturns a bound function on either engine. - Cycle-collector scheduling. Both engines share one full,
non-generational mark-sweep on an adaptive threshold (re-armed
from the surviving live set and live bytes), so the collect
timing differs between runs. What a program observes does not:
cycle members still fire
dropin the pre-sweep pass (§17). - Forward-reference pre-allocation. Both engines scan each
function body and pre-allocate capture cells, so closures compiled
before their
letdeclaration still see the post-declaration value.
Technical differences without behavioral impact
These do not affect any program-visible behavior, but operators embedding Culebra should be aware:
- Cycle collector internals. Both engines share one conservative
stack-scanning mark-sweep. It reclaims every cycle shape —
including one routed purely through
Objectproperty maps. - Thread safety. Most runtime state — the garbage collector,
the defer stack, the interrupt flag — lives in
a
thread_localRuntime, so concurrent isolates on separate threads do not share it. The two process-global intern tables, theShapeRegistryand the trait registry, are mutex-guarded. Execution within a singleRuntimeis single-threaded: an embedder that drives oneRuntimefrom multiple threads must serialize the calls itself.
Top-level drop note (§17)
Top-level bindings to drop-bearing objects live until program exit
without running drop, on either engine (see §17 for the worked
example): the engines suppress drop at the top-level scope release.
Use defer or a factory function for script-wide resources
regardless of which lane you run on.
When in doubt, this document is authoritative — a deviation in observable behavior on any lane is treated as a bug, and the two lanes plus the released-binary differential are what detect it.