(rontolisp) docs

Compiled read/load Limitations

The reader compiled into a JVM class or WASM module parses the same syntax as the interpreter's reader: integers, big integers (JVM), floats, strings, symbols, nil/t, lists (dotted pairs included), 'quote, #'function, ratios (1/3), radix integers (#x10/#o17/#b101), character literals (#\a, #\Space, ...), vectors (#(1 2)), rank-n arrays (#2A((1 2) (3 4))), bit vectors (#*101, read as a general vector like the frontend), packed float arrays (#f(...)/#d(...)), structure literals (#S(NAME :SLOT value ...)), pathname literals (#P"dir/file", read as the pathname VALUE carrying that namestring), ; line comments and nesting #| ... |# block comments. A token no # dispatch claims reads as a symbol (#foo is the symbol #FOO), exactly as it does in source.

Three # forms are permanent exceptions, because they need an evaluator or the feature set at read time: #. read-time evaluation, #+/#- feature conditionals, and #n=/#n# reader labels. The compiled reader SIGNALS a catchable error on them instead of misreading (the interpreter's runtime read still resolves #. and #+/#- like the frontend, and reads labels). read and a read-from-string call passing more than the string are the exception for #+/#-: they resolve a guard themselves on every backend, against the live *features*; only the one-argument read-from-string signals on one.

A #S(...) datum resolves against the structure types the compiled program defines; a slot the datum omits takes its nil initform, a simple constant initform (numbers, strings, characters, symbols, quoted lists, nested literals) is re-read from its baked printed form, and an initform outside that set signals rather than silently substituting a wrong value.

In every backend read consumes exactly the characters of one S-expression and leaves the stream positioned after them: whitespace and comments before the datum are skipped, a datum may span lines, a second datum on the same line survives, and EOF signals end-of-file unless eof-error-p is explicitly nil (it defaults to t), in which case the eof-value is answered. The interpreter's runtime reader is also stricter about malformed tokens -- dot-only tokens (..), #:a:b, #< and #n* over/under-fill signal reader-error there -- while the compiled reader reads those shapes as symbols. It is written in rontolisp over read-char / unread-char / read-from-string, so the syntax it accepts is exactly read-from-string's on that backend -- including the exceptions listed above and the backquote note below.

The WASM reader has a hand-written parser and is narrower in its NUMBERS and its error MESSAGES:

  • Integers and ratios parse exactly at any magnitude. Integer and radix tokens promote through the same boxed-integer tiers the frontend uses (a value past the 31-bit fixnum range becomes a boxed integer, past the signed 64-bit range a limb-based big integer), and so do both sides of a ratio token (3000000000/7 reads as that ratio). Digits followed by a final . and nothing else (5., -5., +5., 123456789012345678901234567890.) are a decimal integer, as in Common Lisp.
  • Floats parse with the frontend's exponent markers. A decimal token (optional leading - or +, digits, at most one . that is not the last character unless an exponent follows, at least one digit, e.g. 1.0, -2.5, .5) parses to an f64-backed float, with or without a Common Lisp exponent suffix (one marker e/s/f/d/l in either case, an optional sign, and at least one exponent digit -- 1e3, 2E3, 1d-8, 1.5f3, .5e2, 1.e5); a token with two dots, a marker with no exponent digits, or any other non-digit (e.g. 1.2.3, foo.bar, 1e, 1e+) stays a symbol (a sign stays part of such a symbol: +5x is the symbol +5X). The value is the double nearest the decimal the token spells -- the one the frontend's Double.parseDouble reads, subnormals and ties included -- and a magnitude past the double range reads as Infinity or 0.0.
  • Error messages are static. A reader error signals (catchably under handler-case), but the message is a fixed text without the offending name interpolated -- the JVM and the interpreter carry the frontend's exact messages.
  • Symbol interning is runtime-backed. Symbols that appear in the compiled program resolve to the same offset the compiled eval uses; symbols seen only at runtime (e.g. a lambda parameter inside a loaded file) are interned in a runtime table so repeated occurrences stay consistent.
  • load requires a preopened directory. It opens the file via WASI path_open, resolving the path against the preopened directories: a relative path against the first one (fd 3), an absolute path against the preopened directory whose name is its longest prefix. Either way, run with --dir -- and with a --dir that COVERS an absolute path, since --dir . preopens a directory named ., which covers none. A path with a .. component resolves against the SHORTEST covering preopened directory instead, since a preopened directory refuses a .. that leaves it; a relative one climbs above the first directory only when the host names that directory after its absolute path (--dir "$PWD::$PWD" --dir /, which is what a native executable's runner does).

Backquote templates and a source's own feature announcement (see Data Types) are frontend read-time constructs the runtime reader of compiled output does not resolve: a file read at runtime via read/read-from-string/a computed load must not use them. (The *features* VARIABLE is an ordinary special and works everywhere; what the runtime reader does not do is let a push in the text it is reading affect its own #+.) (#| ... |# block comments ARE skipped; #+/#- behave as above.)

require/provide are compile-time directives on the compiled backends, so they are not understood by the runtime load of compiled output: a file read at runtime via a computed or nested load must not contain them (only a literal, top-level require/provide works, consumed at compile time) — the same limitation as a runtime-loaded file's package directives (see Packages).

read-from-string reuses the same runtime reader, so on the compiled backends it parses the same syntax as read, and (read-from-string (prin1-to-string x)) round-trips for every printable-readable value kind on all backends. Its second value, the stop index, is answered on every backend; *read-suppress* is honored by the interpreter's reader only, since the runtime reader of compiled output has no suppressed mode. parse-integer is independent of the reader. Both work on every backend with their keyword and optional arguments, as first-class values too (#'parse-integer, #'read-from-string); in a call, a keyword name must be a literal.