Compile to JVM Bytecode
Give rontolisp an output path ending in .class with -o, and it compiles the
source straight to JVM bytecode instead of interpreting it -- no ASM or other
library, the bytecode is emitted by hand. The output extension is what selects the
backend (.class or .jar for JVM, .wasm for WASM).
echo '(print (+ 1 2))' > hello.lisp
rontolisp hello.lisp -o Hello.class
java Hello
The generated class is named after the output file, so the name you pass to
java is the file's stem: -o Hello.class produces a class Hello you run with
java Hello. A directory in the path becomes the class's Java package:
-o com/example/Kernels.class produces com.example.Kernels, which you run
with java -cp . com.example.Kernels from the directory the path started in
(the missing directories are created). A directory that could not BE a package is
just a directory instead: an absolute path (whose leading / would open the name
with an empty package, which no JVM loads) and a ./ or ../ segment leave only
the file's stem, so -o /tmp/out/Hello.class is again the class Hello that
java -cp /tmp/out Hello runs. --class-name is how an absolute path names a
package anyway. The program's top-level forms
become the class's entry point and run in order when you launch it.
-o out.jar writes a jar rather than a bare class: the class, the runtime
classes that have to travel with it, a manifest, and -- with
--maven-coordinates -- the Maven metadata that lets a consumer install it with
no flags at all. The jar is executable, so it needs nothing else to run:
rontolisp hello.lisp -o hello.jar
java -jar hello.jar
A jar path names no class, so the class inside takes its name from the file's
stem in CamelCase (hello.jar -> Hello, my-app-1.0.0.jar -> MyApp100) and
the manifest's Main-Class points at it. That name only matters if you call the
class directly: --class-name sets it, and it is REQUIRED for a --no-main
library jar, whose class is the artifact's Java API rather than an entry point.
It works for .class output too, where it replaces the name the path would give.
One class file's constant pool holds at most 65534 entries, and a program that splices
several large libraries can need more (a mito program does). Such a program comes out as
its class plus Hello$Part1.class, Hello$Part2.class, ... written beside it in its
package directory, and a jar carries them too. Run the class as before; the part files
only have to stay next to it.
A program that uses java:, the geom: kernels,
--simd, --blas, --gpu, objc: or ffi: gets the bridge each one needs the same
way: Hello$JavaBridge.class, Hello$SimdBridge.class, ... beside the class (for
--gpu, objc: and ffi: with a renamed copy of the binding library,
Hello$Gpu*.class and so on), and inside a jar. Nothing is defined at run time, so
such a jar also builds into a GraalVM native image with native-image -jar. A --blas,
--gpu or objc: jar carries the native-image metadata its foreign calls need and builds as
it is; reflective java: calls and the ffi: foreign calls need the metadata the tracing
agent records from one java -jar run (Java interop).
A java:reify or java:proxy object, and a function passed where an interface is
expected, is an instance of a class generated at compile time -- Hello$Reify0.class,
Hello$Proxy0.class, ... and their common Hello$Implementation.class, written beside
the class the same way -- and needs no such metadata.
A class can also be a library Java code calls directly:
rontolisp:jvm-export
declares a typed, Java-callable static method for a defun, and --no-main
drops the main entry point entirely. See
Export a JVM library.
Example (hello.lisp):
3
Recursion Depth
The compiled main runs the program on a thread of its own with a 16 MiB stack, the
size the interpreter runs on, so -Xss no longer decides how deep a program can recurse.
The system property rontolisp.stack sets another size, in MiB (0 leaves it to the
JVM, that is to -Xss):
java -Drontolisp.stack=64 Hello
java -Drontolisp.stack=64 -jar hello.jar
A program that uses objc: or appkit: stays on the launcher's first thread, which AppKit
requires, and so does a class whose top level runs when it is initialized (one with a
rontolisp:jvm-export).
A function that calls itself in tail position -- a defun by name or through #'name, a
labels function, Clojure's recur -- jumps back to its own start instead of calling, so a
loop written as tail recursion runs at any depth, whatever its lambda list. Tail position
reaches through if, let, cond and the other built-in macros, the body of an inline
((lambda ...) ...), and the value of a return/return-from that leaves the function,
from a loop body too. Functions that
call each other in tail position -- defuns, or the functions of one labels form (a
Clojure letfn) -- jump to each other the same way: the method a call enters holds the code
of the functions the cycle runs through. A cycle whose code would pass the 8,000 bytes of
bytecode HotSpot compiles in one method keeps a frame per call instead. A tail call through a
function value -- from a defun, a lambda, an flet or labels function, an apply --
runs in constant stack too: a closure calling itself through the variable that holds it,
closures in a table calling each other, a continuation handed to a function that calls it.
Such a call stays an ordinary call, which the JIT can inline, until 64 of them have run in a
row since the nearest call that is not one -- an adapter closure or a composed function never
gets that deep -- and past that the chain continues through the class's trampoline. It goes
through a copy of the dispatch method for its argument count that only such calls use, so a
JIT inlining a chain of them sees only the functions they reach; where that dispatch method is
split across several methods (a program with many functions of that argument count), they use
the shared one, so the class never carries its largest dispatch methods twice. Every thread
counts its own. A recursion that is not a tail call keeps the frames of such a chain on each of
its levels. A call inside a special binding is not in tail position, so it keeps a frame.
Optimize (Dead-Code Elimination)
Compilation drops every method unreachable from main, along with any static field
only they referenced, and compacts the constant pool accordingly. You get that
without asking:
echo '(defun fact (n) (if (<= n 1) 1 (* n (fact (- n 1)))))
(print (fact 10))' > fact.lisp
rontolisp fact.lisp -o Fact.class
java Fact
For a small program like fact the class is ~6.5 KB. Pass --optimize=off and the
class instead embeds the entire runtime (printer, numeric, reader and eval
helper methods, plus a first-class wrapper for every built-in) regardless of what the
program actually uses, which for the same fact is ~190 KB. The elimination is
behavior-preserving: reachability follows the actual invoke instructions
in the bytecode, so anything a first-class function value, funcall, or an embedded
eval/load can dispatch to is kept, and the java: interop bridge's reflective
entry point survives as an explicit root. Every
rontolisp:jvm-export typed
method is an explicit root too — its caller is Java code the bytecode cannot
show — which is what lets a compiled
library keep the default size. The same levels also
tree-shake the WASM output.
The dispatch methods funcall goes through list only the functions your program
can actually obtain as a value — #'name, a quoted 'name designator, a
lambda, or (while the program holds a symbol builder such as intern or
find-symbol) a string or keyword constant spelling the name — so everything
else becomes ordinary dead code the shaker removes. A lambda counts only while
code that survives the shake makes it: a closure made only inside a function
nothing calls goes with that function. A #'name makes the function a value,
not its name a designator: a symbol finds a function at run time only when the
program spells that name as a constant, or a package walk (do-symbols,
apropos-list, ...) hands the symbol back, so a function taken as a value only
inside a function nothing calls goes with it as well. That listing switches off,
and every function stays reachable, only when the program can name a function
out of data this compile never sees: any use of eval, read,
read-from-string, a runtime load or a ~/name/
format directive — including one inside a
library you loaded — as does --dynamic. Compile with
-Drontolisp.debug.dispatchgate=true to have the compiler name the operator
responsible.
One carve-out follows from that: a designator assembled at run time out of
computed pieces — (funcall (intern (concatenate 'string "gre" suffix))) —
is no constant the compiler can read, so the call signals the ordinary
"undefined function" error, whether or not the program also takes that function
as a #'name value. --dynamic is the way back. --optimize=off is
not: the listing is not part of what the level switches, so declining the
optimizer does not bring such a name back.
--optimize takes an optional level, shared with the WASM backend.
--optimize and --optimize=default both spell what an absent flag already
selects — everything above — for a build script that wants it written down.
--optimize=off declines it, and emits what a build before the flag was on by
default emitted. --optimize=size asks for the smallest output a backend can
give. On this backend it declines the three emissions that spend bytes on speed.
One is the typed numeric loop: a dotimes whose body reads and writes packed
single/double-float arrays through fixnum index math, let temporaries,
+ - * /, (length a) of such an array, the unary math functions and
if/when/unless tests compiles by
default to a primitive long/double loop over raw float[]/double[]
accesses, behind a check at the loop's entry that falls back to the ordinary
emission whenever the variables are not what the typing assumed -- the same
values either way, several times faster, and a larger class because the body is
emitted more than once. Another is integer expression-tree fusion: a nested
+ - * mod rem logand logior logxor lognot ash tree compiles by default into a
shared method that runs the whole tree as raw long arithmetic and boxes only
the result, with the generic per-operation chain kept alongside as the fallback
for anything that is not a machine-word integer at run time -- again the same
values, with output and errors in the interpreter's order, and a class that
carries each tree twice. The third is the copy of the
dispatch method that tail calls through a function value go through (above), one
per argument count those calls use. --optimize=size keeps only the ordinary
emissions; a program with none of these shapes compiles to the same class at both
levels, and the same program's JVM bytecode is about a third the size of
its WASM to begin with. So one build script can pass --optimize=size for every
target.
--optimize=off exists for two jobs, and neither of them is making a program
work: comparing an artifact against one built before a compiler change, and
bisecting a suspected shaker bug by asking whether the unshaken class behaves
differently. A program whose functions are reached only through a name the
compiler cannot read needs --dynamic, as above.
Independently of the level, compilation always tree-shakes the libraries it
splices in: the bundled Lisp-source ones (linalg:, vec:, JSON, URL,
equalp/string<) and every system loaded with
asdf:load-system / ql:quickload. A function,
variable or constant your program never mentions -- by name anywhere in the
source, including quoted symbols and string literals -- is not compiled in. Your
own code is never pruned, and neither is anything a load/require splices in:
only a library that came from a system is subject to it.
Classes, generic functions, methods, conditions and structures are pruned by
the same rule: a class nothing references leaves together with its methods,
and a method on a generic your program does call is still dropped when no
reachable code can create an instance of the class it specializes on. Methods
on the standard protocol names (initialize-instance, print-object,
close, ...) follow their class alone, since those calls are implicit.
The one consequence: a library function whose name is only assembled at runtime
from computed strings and called through eval/apply signals the usual
"undefined function" error. Compile with --no-prune (or --dynamic) to keep
every library definition in that case.
The generated .class file targets Java 17 (class version 61), so running it
requires a Java 17 or newer JRE. Beyond java.lang and java.io, the emitted
runtime helpers reference java.math (BigInteger/BigDecimal/MathContext,
for the overflow-promoting integer and exact ratio arithmetic) and java.util
(ArrayList/Arrays, and HashMap for hash tables); a program that calls
rontolisp:fetch additionally references java.net/java.net.http, and
rontolisp:await / rontolisp:futurep represent futures as
java.util.concurrent futures -- all of which are part of Java 17, so none of
these raise the requirement. The one exception is a program that uses the
java: interop package. Its java: calls are
resolved at compile time wherever the program text allows, against a JDK's
class files, and become direct calls; the class is then stamped for that Java
release (the compiling JDK's own by default), so it needs a JRE of that
release. A call left to run time goes through a reflection bridge (compiled
with the project's own Java release) the compiler writes beside the class,
which needs a JRE at least as new as the one rontolisp was built with.
--java-release N and the program's class path (--java-classpath, --java-dep)
choose which class files, --warn-java-reflection reports the calls left to run time, and
--java-static makes each of them a compile error, so the class carries no reflection and
GraalVM native-image builds it with no reachability metadata (the guide's Resolving
calls before they run). A
program jar carries its class path beside it and names it in its manifest, a war in
WEB-INF/lib (the guide's Java libraries).
Skip the JIT Warm-Up with an AOT Cache
A compiled program starts as bytecode, so its first few dozen milliseconds run in
the JVM's interpreter and first-tier compiler while the JIT works out which
methods are hot. For a long-running program that cost disappears into the run.
For a short one it can be most of it: bench-report's mandelbrot finishes in
about 95 ms on its first in-process run and about 22 ms on its third, and every
new process starts over at 95.
JDK 25 can persist what that first run learned. Compile to a jar, do one training run, build a cache from it, and pass the cache to every run after:
rontolisp mandelbrot.lisp -o app.jar
java -XX:AOTMode=record -XX:AOTConfiguration=app.aotconf -jar app.jar
java -XX:AOTMode=create -XX:AOTConfiguration=app.aotconf -XX:AOTCache=app.aot -cp app.jar
java -XX:AOTCache=app.aot -jar app.jar
On a 64-core Linux box with GraalVM 25 that takes mandelbrot's first run from
95 ms to 51 and matmul's from 94 to 57 (medians). The gain is the JIT not
re-deriving a profile it has already been handed, so it shows up where warm-up
was a large share of a short run and nowhere else: of the ten bench-report
programs, only those two move more than 10%.
Four things to know before relying on it.
It needs -o app.jar. The cache cannot be built from a classpath that
contains a directory, so a bare -o Prog.class run under java -cp . Prog
cannot be trained. The jar is the same compiled class plus a manifest, so this
costs nothing else.
The training run has to do the real work. The cache stores a profile, not
compiled code, and a profile only exists once the training run has itself warmed
up. Training mandelbrot on a quarter-size grid -- 32 ms of work -- buys the
full run nothing. Train on a representative workload, not a smoke test.
The cache belongs to one jar file. It records the jar's path and timestamp,
so rebuilding the program invalidates it. Nothing breaks when that happens: the
JVM prints a few [error][aot] lines on stderr, ignores the cache and runs the
program correctly at the usual speed. Rebuild the cache whenever you rebuild the
jar, or drop the flag. A cache is around 11 MB.
On GraalVM, use the two-step flow above rather than the one-command
-XX:AOTCacheOutput. That shortcut assembles the cache in a child JVM which
loses GraalVM's jdk.internal.vm.ci module, and the cache it writes is rejected
at load -- with the error only on stderr, so the program still runs, just with no
gain at all.
The same works on rontolisp itself, where it is worth roughly 3x on start-up
(java -jar rontolisp.jar on a program that computes nothing: 476 ms to 144 ms;
a cache trained on a compile invocation halves -o out.class instead, 978 ms to
480 ms). If you have a GraalVM to hand, the native binary
is the better answer to the same problem -- it does those in 12 ms and 109 ms
with no cache file and no training step.