BuildNot yet confirmed elsewhere1 publisher2 min readPublished
jitLLM claims 90% of llama.cpp speed by compiling Java bytecode to CUDA kernels
TornadoVM's University of Manchester team says its jitLLM engine runs LLM inference on NVIDIA GPUs from Java at about 90% of llama.cpp's speed. The figure comes from the project's own RTX 5090 recordings, so a JVM team still needs a run on its own GPUs before it retires a Python sidecar.
The Engineer · Build desk
What happened
- TornadoVM, the open-source JDK plugin underneath jitLLM, compiles annotated Java bytecode into OpenCL, PTX/CUDA or SPIR-V kernels.
- jitLLM's transformer loop is written in Java and loads GGUF models including Llama 3, Qwen 3, Phi-3, IBM Granite and Devstral 2, with no call out to llama.cpp or Ollama.
- LangChain4j lists jitLLM as an official inference engine, so its AI Services, tool-calling agents and retrieval pipelines can run on a local GPU.
- Red Hat is named as the collaboration partner on jitLLM alongside the Manchester TornadoVM group.
Why it matters
- constraint Until someone reproduces the 90% figure on mid-range cards and Q4 models, it cannot support a capacity plan or a case for retiring an existing Python inference service.
- decision Since LangChain4j callers code against an interface, a team can trial jitLLM and switch back to its current backend without rewriting call sites.
- exposure Dropping the sidecar takes PyTorch and a second container out of the patch cycle, and puts a JDK plugin that generates GPU code at runtime into the application's own security review.
On NVIDIA hardware with the newer toolchains, TornadoVM's JIT emits CUDA C and cuTile code at runtime. It can also bind existing CUDA libraries directly from Java [8]. It handles device transfers itself and reuses Java memory where it can [9]. According to the dev.to write-up, that removes most of the boilerplate a team would otherwise write by hand against the CUDA API [9].
We think the strongest design choice here is also the plainest one. The backend is picked at runtime, so the same Java code can target CUDA, OpenCL or SPIR-V, and Apple Silicon is in scope [10]. TornadoVM has been in development at Manchester for years [7]. jitLLM started as GPULlama3.java, a GPU-accelerated fork of the pure-Java Llama3.java, and it keeps a CPU path for machines without a GPU [11].
The speed claim spread after the project reached r/LocalLLaMA on October 8 [5]. The number comes from the project itself. It was shown on an RTX 5090 in the project's own recordings, and no independent benchmark has confirmed it [2]. The baseline is hard to beat. llama.cpp is hand-tuned C++ and CUDA with years of kernel optimization behind it [15], so a 10% gap [18] from compiled Java would be a strong result. Before the figure means anything for a given shop, it has to hold on smaller cards, on the quantization formats services actually run, and across the models jitLLM lists. The author of the dev.to analysis, who has not run jitLLM on their own hardware [14], wrote: "Treat it as promising until someone replicates it on a plain RTX 4060 with a Q4 model, which is what most of us actually run." [3]
That replication is cheap to set up. jitLLM reads standard GGUF files, the same format llama.cpp uses [11], so both engines can load one model file on one card. Wiring it into a service takes one dependency and one builder call: add `langchain4j-jitllm`, then construct a `JitLLMChatModel` with `onGPU(true)` [17]. Quarkus users get the model as a CDI bean through `quarkus-langchain4j-jitllm` [13].
In the dev.to write-up's account, the sidecar jitLLM is aimed at brings PyTorch and a CUDA toolkit with it, plus its own Dockerfile, its own set of CVEs and a REST call between two halves of one application [16]. jitLLM drops the Python process and ships no C++ code [1]. That describes the repository. On the host, the kernels still run through NVIDIA's CUDA stack, and jitLLM can call existing CUDA libraries from Java [8]. The write-up does not say which NVIDIA toolchain components a deployment still needs when kernels are generated as CUDA C at runtime [8].
What to watch
- An independent jitLLM versus llama.cpp run on the same GGUF file, on an RTX 4060-class card with a Q4 model.
- Project documentation listing which NVIDIA toolchain components a jitLLM deployment needs for runtime CUDA C and cuTile generation.
- Published numbers for jitLLM's OpenCL or SPIR-V backends, which would show whether the speed claim holds off NVIDIA hardware.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives55
- Confidence35
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
jitLLM compiles Java bytecode down to CUDA and cuTile and runs LLM inference on NVIDIA GPUs at a claimed 90% of llama.cpp performance, with no C++ code, no Python process and no sidecar.
ReportedSupportedSource: project claim as reported by dev.to2 sources— create a free account to open themView cited source - [2]
The roughly 90%-of-llama.cpp figure was demonstrated on an RTX 5090 in the project's own recordings; it is the project's claim, not an independent benchmark.
ReportedSupportedSource: dev.to analysis2 sources— create a free account to open themView cited source - [3]
"Treat it as promising until someone replicates it on a plain RTX 4060 with a Q4 model, which is what most of us actually run."
ReportedSupportedSource: author of the dev.to analysis2 sources— create a free account to open themView cited source - [4]
jitLLM was built by the TornadoVM team at the University of Manchester, with Red Hat as a collaboration partner.
- [6]
jitLLM runs Llama 3, Mistral, Qwen 2.5, Qwen 3, Phi-3, IBM Granite and Devstral 2 models in GGUF format; it does not shell out to llama.cpp or call Ollama over HTTP; the transformer inference loop is written in Java and the heavy math is JIT-compiled to GPU kernels at runtime.
- [7]
TornadoVM is an open-source plugin to the JDK that compiles annotated Java bytecode to OpenCL, PTX/CUDA or SPIR-V, and has been under active development for years at the University of Manchester.
- [8]
On NVIDIA with the newer toolchains, TornadoVM generates CUDA C and cuTile code at runtime; it can also bind existing CUDA libraries directly from Java code.
- [9]
TornadoVM reuses Java memory where possible and handles device transfers, which removes most of the boilerplate otherwise written by hand against the CUDA API.
- [10]
Backend choice in TornadoVM is a runtime property: the same code can target CUDA, OpenCL or SPIR-V, and Apple Silicon and other OpenCL-capable devices are in scope.
- [11]
jitLLM started as GPULlama3.java, a GPU-accelerated fork of the pure-Java Llama3.java project; it reads standard GGUF files, the same format llama.cpp uses, and falls back to a CPU path when no GPU is present.
- [12]
jitLLM is an official LangChain4j inference engine; with it, LangChain4j AI Services, tool-calling agents and retrieval pipelines run on the user's own GPU.
- [13]
A Quarkus extension, quarkus-langchain4j-jitllm, exposes the model as a CDI bean.
- [14]
The dev.to author has not run jitLLM on their own hardware; the piece draws on the project's pages, the TornadoVM blog and discussion around the announcement.
- [15]
llama.cpp is a heavily hand-tuned C++/CUDA codebase with years of kernel optimization.
- [16]
The dev.to write-up describes the typical Python inference sidecar as PyTorch, a CUDA toolkit, a second Dockerfile, a second set of CVEs, and a REST hop between the two halves of the application.
- [17]
To use it from LangChain4j, add langchain4j-jitllm and construct a JitLLMChatModel with onGPU(true); because of the ChatLanguageModel abstraction, the rest of the code does not know or care where inference happens.
- [18]
jitLLM's claimed performance gap to llama.cpp is 10%.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toThis Framework Compiles Java Straight to CUDA: Can It Really Match llama.cpp?
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- GPU programming on the JVMFollow
- Local LLM InferenceFollow