(rontolisp) docs

Compile to JVM Bytecode

Give rontolisp an output path ending in .class with -o, and it compiles the source straight to JVM bytecode instead of interpreting it -- no ASM or other library, the bytecode is emitted by hand. The output extension is what selects the backend (.class or .jar for JVM, .wasm for WASM).

echo '(print (+ 1 2))' > hello.lisp
rontolisp hello.lisp -o Hello.class
java Hello

The generated class is named after the output file, so the name you pass to java is the file's stem: -o Hello.class produces a class Hello you run with java Hello. A directory in the path becomes the class's Java package: -o com/example/Kernels.class produces com.example.Kernels, which you run with java -cp . com.example.Kernels from the directory the path started in (the missing directories are created). A directory that could not BE a package is just a directory instead: an absolute path (whose leading / would open the name with an empty package, which no JVM loads) and a ./ or ../ segment leave only the file's stem, so -o /tmp/out/Hello.class is again the class Hello that java -cp /tmp/out Hello runs. --class-name is how an absolute path names a package anyway. The program's top-level forms become the class's entry point and run in order when you launch it.

-o out.jar writes a jar rather than a bare class: the class, the runtime classes that have to travel with it, a manifest, and -- with --maven-coordinates -- the Maven metadata that lets a consumer install it with no flags at all. The jar is executable, so it needs nothing else to run:

rontolisp hello.lisp -o hello.jar
java -jar hello.jar

A jar path names no class, so the class inside takes its name from the file's stem in CamelCase (hello.jar -> Hello, my-app-1.0.0.jar -> MyApp100) and the manifest's Main-Class points at it. That name only matters if you call the class directly: --class-name sets it, and it is REQUIRED for a --no-main library jar, whose class is the artifact's Java API rather than an entry point. It works for .class output too, where it replaces the name the path would give.

One class file's constant pool holds at most 65534 entries, and a program that splices several large libraries can need more (a mito program does). Such a program comes out as its class plus Hello$Part1.class, Hello$Part2.class, ... written beside it in its package directory, and a jar carries them too. Run the class as before; the part files only have to stay next to it.

A program that uses java:, the geom: kernels, --simd, --blas, --gpu, objc: or ffi: gets the bridge each one needs the same way: Hello$JavaBridge.class, Hello$SimdBridge.class, ... beside the class (for --gpu, objc: and ffi: with a renamed copy of the binding library, Hello$Gpu*.class and so on), and inside a jar. Nothing is defined at run time, so such a jar also builds into a GraalVM native image with native-image -jar. A --blas, --gpu or objc: jar carries the native-image metadata its foreign calls need and builds as it is; reflective java: calls and the ffi: foreign calls need the metadata the tracing agent records from one java -jar run (Java interop). A java:reify or java:proxy object, and a function passed where an interface is expected, is an instance of a class generated at compile time -- Hello$Reify0.class, Hello$Proxy0.class, ... and their common Hello$Implementation.class, written beside the class the same way -- and needs no such metadata.

A class can also be a library Java code calls directly: rontolisp:jvm-export declares a typed, Java-callable static method for a defun, and --no-main drops the main entry point entirely. See Export a JVM library.

Example (hello.lisp):

3

Recursion Depth

The compiled main runs the program on a thread of its own with a 16 MiB stack, the size the interpreter runs on, so -Xss no longer decides how deep a program can recurse. The system property rontolisp.stack sets another size, in MiB (0 leaves it to the JVM, that is to -Xss):

java -Drontolisp.stack=64 Hello
java -Drontolisp.stack=64 -jar hello.jar

A program that uses objc: or appkit: stays on the launcher's first thread, which AppKit requires, and so does a class whose top level runs when it is initialized (one with a rontolisp:jvm-export).

A function that calls itself in tail position -- a defun by name or through #'name, a labels function, Clojure's recur -- jumps back to its own start instead of calling, so a loop written as tail recursion runs at any depth, whatever its lambda list. Tail position reaches through if, let, cond and the other built-in macros, the body of an inline ((lambda ...) ...), and the value of a return/return-from that leaves the function, from a loop body too. Functions that call each other in tail position -- defuns, or the functions of one labels form (a Clojure letfn) -- jump to each other the same way: the method a call enters holds the code of the functions the cycle runs through. A cycle whose code would pass the 8,000 bytes of bytecode HotSpot compiles in one method keeps a frame per call instead. A tail call through a function value -- from a defun, a lambda, an flet or labels function, an apply -- runs in constant stack too: a closure calling itself through the variable that holds it, closures in a table calling each other, a continuation handed to a function that calls it. Such a call stays an ordinary call, which the JIT can inline, until 64 of them have run in a row since the nearest call that is not one -- an adapter closure or a composed function never gets that deep -- and past that the chain continues through the class's trampoline. It goes through a copy of the dispatch method for its argument count that only such calls use, so a JIT inlining a chain of them sees only the functions they reach; where that dispatch method is split across several methods (a program with many functions of that argument count), they use the shared one, so the class never carries its largest dispatch methods twice. Every thread counts its own. A recursion that is not a tail call keeps the frames of such a chain on each of its levels. A call inside a special binding is not in tail position, so it keeps a frame.

Optimize (Dead-Code Elimination)

Compilation drops every method unreachable from main, along with any static field only they referenced, and compacts the constant pool accordingly. You get that without asking:

echo '(defun fact (n) (if (<= n 1) 1 (* n (fact (- n 1)))))
(print (fact 10))' > fact.lisp
rontolisp fact.lisp -o Fact.class
java Fact

For a small program like fact the class is ~6.5 KB. Pass --optimize=off and the class instead embeds the entire runtime (printer, numeric, reader and eval helper methods, plus a first-class wrapper for every built-in) regardless of what the program actually uses, which for the same fact is ~190 KB. The elimination is behavior-preserving: reachability follows the actual invoke instructions in the bytecode, so anything a first-class function value, funcall, or an embedded eval/load can dispatch to is kept, and the java: interop bridge's reflective entry point survives as an explicit root. Every rontolisp:jvm-export typed method is an explicit root too — its caller is Java code the bytecode cannot show — which is what lets a compiled library keep the default size. The same levels also tree-shake the WASM output.

The dispatch methods funcall goes through list only the functions your program can actually obtain as a value — #'name, a quoted 'name designator, a lambda, or (while the program holds a symbol builder such as intern or find-symbol) a string or keyword constant spelling the name — so everything else becomes ordinary dead code the shaker removes. A lambda counts only while code that survives the shake makes it: a closure made only inside a function nothing calls goes with that function. A #'name makes the function a value, not its name a designator: a symbol finds a function at run time only when the program spells that name as a constant, or a package walk (do-symbols, apropos-list, ...) hands the symbol back, so a function taken as a value only inside a function nothing calls goes with it as well. That listing switches off, and every function stays reachable, only when the program can name a function out of data this compile never sees: any use of eval, read, read-from-string, a runtime load or a ~/name/ format directive — including one inside a library you loaded — as does --dynamic. Compile with -Drontolisp.debug.dispatchgate=true to have the compiler name the operator responsible.

One carve-out follows from that: a designator assembled at run time out of computed pieces — (funcall (intern (concatenate 'string "gre" suffix))) — is no constant the compiler can read, so the call signals the ordinary "undefined function" error, whether or not the program also takes that function as a #'name value. --dynamic is the way back. --optimize=off is not: the listing is not part of what the level switches, so declining the optimizer does not bring such a name back.

--optimize takes an optional level, shared with the WASM backend. --optimize and --optimize=default both spell what an absent flag already selects — everything above — for a build script that wants it written down. --optimize=off declines it, and emits what a build before the flag was on by default emitted. --optimize=size asks for the smallest output a backend can give. On this backend it declines the three emissions that spend bytes on speed. One is the typed numeric loop: a dotimes whose body reads and writes packed single/double-float arrays through fixnum index math, let temporaries, + - * /, (length a) of such an array, the unary math functions and if/when/unless tests compiles by default to a primitive long/double loop over raw float[]/double[] accesses, behind a check at the loop's entry that falls back to the ordinary emission whenever the variables are not what the typing assumed -- the same values either way, several times faster, and a larger class because the body is emitted more than once. Another is integer expression-tree fusion: a nested + - * mod rem logand logior logxor lognot ash tree compiles by default into a shared method that runs the whole tree as raw long arithmetic and boxes only the result, with the generic per-operation chain kept alongside as the fallback for anything that is not a machine-word integer at run time -- again the same values, with output and errors in the interpreter's order, and a class that carries each tree twice. The third is the copy of the dispatch method that tail calls through a function value go through (above), one per argument count those calls use. --optimize=size keeps only the ordinary emissions; a program with none of these shapes compiles to the same class at both levels, and the same program's JVM bytecode is about a third the size of its WASM to begin with. So one build script can pass --optimize=size for every target.

--optimize=off exists for two jobs, and neither of them is making a program work: comparing an artifact against one built before a compiler change, and bisecting a suspected shaker bug by asking whether the unshaken class behaves differently. A program whose functions are reached only through a name the compiler cannot read needs --dynamic, as above.

Independently of the level, compilation always tree-shakes the libraries it splices in: the bundled Lisp-source ones (linalg:, vec:, JSON, URL, equalp/string<) and every system loaded with asdf:load-system / ql:quickload. A function, variable or constant your program never mentions -- by name anywhere in the source, including quoted symbols and string literals -- is not compiled in. Your own code is never pruned, and neither is anything a load/require splices in: only a library that came from a system is subject to it.

Classes, generic functions, methods, conditions and structures are pruned by the same rule: a class nothing references leaves together with its methods, and a method on a generic your program does call is still dropped when no reachable code can create an instance of the class it specializes on. Methods on the standard protocol names (initialize-instance, print-object, close, ...) follow their class alone, since those calls are implicit.

The one consequence: a library function whose name is only assembled at runtime from computed strings and called through eval/apply signals the usual "undefined function" error. Compile with --no-prune (or --dynamic) to keep every library definition in that case.

The generated .class file targets Java 17 (class version 61), so running it requires a Java 17 or newer JRE. Beyond java.lang and java.io, the emitted runtime helpers reference java.math (BigInteger/BigDecimal/MathContext, for the overflow-promoting integer and exact ratio arithmetic) and java.util (ArrayList/Arrays, and HashMap for hash tables); a program that calls rontolisp:fetch additionally references java.net/java.net.http, and rontolisp:await / rontolisp:futurep represent futures as java.util.concurrent futures -- all of which are part of Java 17, so none of these raise the requirement. The one exception is a program that uses the java: interop package. Its java: calls are resolved at compile time wherever the program text allows, against a JDK's class files, and become direct calls; the class is then stamped for that Java release (the compiling JDK's own by default), so it needs a JRE of that release. A call left to run time goes through a reflection bridge (compiled with the project's own Java release) the compiler writes beside the class, which needs a JRE at least as new as the one rontolisp was built with.

--java-release N and the program's class path (--java-classpath, --java-dep) choose which class files, --warn-java-reflection reports the calls left to run time, and --java-static makes each of them a compile error, so the class carries no reflection and GraalVM native-image builds it with no reachability metadata (the guide's Resolving calls before they run). A program jar carries its class path beside it and names it in its manifest, a war in WEB-INF/lib (the guide's Java libraries).

Skip the JIT Warm-Up with an AOT Cache

A compiled program starts as bytecode, so its first few dozen milliseconds run in the JVM's interpreter and first-tier compiler while the JIT works out which methods are hot. For a long-running program that cost disappears into the run. For a short one it can be most of it: bench-report's mandelbrot finishes in about 95 ms on its first in-process run and about 22 ms on its third, and every new process starts over at 95.

JDK 25 can persist what that first run learned. Compile to a jar, do one training run, build a cache from it, and pass the cache to every run after:

rontolisp mandelbrot.lisp -o app.jar
java -XX:AOTMode=record -XX:AOTConfiguration=app.aotconf -jar app.jar
java -XX:AOTMode=create -XX:AOTConfiguration=app.aotconf -XX:AOTCache=app.aot -cp app.jar
java -XX:AOTCache=app.aot -jar app.jar

On a 64-core Linux box with GraalVM 25 that takes mandelbrot's first run from 95 ms to 51 and matmul's from 94 to 57 (medians). The gain is the JIT not re-deriving a profile it has already been handed, so it shows up where warm-up was a large share of a short run and nowhere else: of the ten bench-report programs, only those two move more than 10%.

Four things to know before relying on it.

It needs -o app.jar. The cache cannot be built from a classpath that contains a directory, so a bare -o Prog.class run under java -cp . Prog cannot be trained. The jar is the same compiled class plus a manifest, so this costs nothing else.

The training run has to do the real work. The cache stores a profile, not compiled code, and a profile only exists once the training run has itself warmed up. Training mandelbrot on a quarter-size grid -- 32 ms of work -- buys the full run nothing. Train on a representative workload, not a smoke test.

The cache belongs to one jar file. It records the jar's path and timestamp, so rebuilding the program invalidates it. Nothing breaks when that happens: the JVM prints a few [error][aot] lines on stderr, ignores the cache and runs the program correctly at the usual speed. Rebuild the cache whenever you rebuild the jar, or drop the flag. A cache is around 11 MB.

On GraalVM, use the two-step flow above rather than the one-command -XX:AOTCacheOutput. That shortcut assembles the cache in a child JVM which loses GraalVM's jdk.internal.vm.ci module, and the cache it writes is rejected at load -- with the error only on stderr, so the program still runs, just with no gain at all.

The same works on rontolisp itself, where it is worth roughly 3x on start-up (java -jar rontolisp.jar on a program that computes nothing: 476 ms to 144 ms; a cache trained on a compile invocation halves -o out.class instead, 978 ms to 480 ms). If you have a GraalVM to hand, the native binary is the better answer to the same problem -- it does those in 12 ms and 109 ms with no cache file and no training step.