Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

jitLLM claims 90% of llama.cpp speed by compiling Java bytecode to CUDA kernels

TornadoVM's University of Manchester team says its jitLLM engine runs LLM inference on NVIDIA GPUs from Java at about 90% of llama.cpp's speed. The figure comes from the project's own RTX 5090 recordings, so a JVM team still needs a run on its own GPUs before it retires a Python sidecar.

The Engineer · Build desk

How we use AISend a correction

What happened

  • TornadoVM, the open-source JDK plugin underneath jitLLM, compiles annotated Java bytecode into OpenCL, PTX/CUDA or SPIR-V kernels.
  • jitLLM's transformer loop is written in Java and loads GGUF models including Llama 3, Qwen 3, Phi-3, IBM Granite and Devstral 2, with no call out to llama.cpp or Ollama.
  • LangChain4j lists jitLLM as an official inference engine, so its AI Services, tool-calling agents and retrieval pipelines can run on a local GPU.
  • Red Hat is named as the collaboration partner on jitLLM alongside the Manchester TornadoVM group.

Why it matters

  • constraint Until someone reproduces the 90% figure on mid-range cards and Q4 models, it cannot support a capacity plan or a case for retiring an existing Python inference service.
  • decision Since LangChain4j callers code against an interface, a team can trial jitLLM and switch back to its current backend without rewriting call sites.
  • exposure Dropping the sidecar takes PyTorch and a second container out of the patch cycle, and puts a JDK plugin that generates GPU code at runtime into the application's own security review.

On NVIDIA hardware with the newer toolchains, TornadoVM's JIT emits CUDA C and cuTile code at runtime. It can also bind existing CUDA libraries directly from Java [8]. It handles device transfers itself and reuses Java memory where it can [9]. According to the dev.to write-up, that removes most of the boilerplate a team would otherwise write by hand against the CUDA API [9].

We think the strongest design choice here is also the plainest one. The backend is picked at runtime, so the same Java code can target CUDA, OpenCL or SPIR-V, and Apple Silicon is in scope [10]. TornadoVM has been in development at Manchester for years [7]. jitLLM started as GPULlama3.java, a GPU-accelerated fork of the pure-Java Llama3.java, and it keeps a CPU path for machines without a GPU [11].

The speed claim spread after the project reached r/LocalLLaMA on October 8 [5]. The number comes from the project itself. It was shown on an RTX 5090 in the project's own recordings, and no independent benchmark has confirmed it [2]. The baseline is hard to beat. llama.cpp is hand-tuned C++ and CUDA with years of kernel optimization behind it [15], so a 10% gap [18] from compiled Java would be a strong result. Before the figure means anything for a given shop, it has to hold on smaller cards, on the quantization formats services actually run, and across the models jitLLM lists. The author of the dev.to analysis, who has not run jitLLM on their own hardware [14], wrote: "Treat it as promising until someone replicates it on a plain RTX 4060 with a Q4 model, which is what most of us actually run." [3]

That replication is cheap to set up. jitLLM reads standard GGUF files, the same format llama.cpp uses [11], so both engines can load one model file on one card. Wiring it into a service takes one dependency and one builder call: add `langchain4j-jitllm`, then construct a `JitLLMChatModel` with `onGPU(true)` [17]. Quarkus users get the model as a CDI bean through `quarkus-langchain4j-jitllm` [13].

In the dev.to write-up's account, the sidecar jitLLM is aimed at brings PyTorch and a CUDA toolkit with it, plus its own Dockerfile, its own set of CVEs and a REST call between two halves of one application [16]. jitLLM drops the Python process and ships no C++ code [1]. That describes the repository. On the host, the kernels still run through NVIDIA's CUDA stack, and jitLLM can call existing CUDA libraries from Java [8]. The write-up does not say which NVIDIA toolchain components a deployment still needs when kernels are generated as CUDA C at runtime [8].

What to watch

  • An independent jitLLM versus llama.cpp run on the same GGUF file, on an RTX 4060-class card with a Q4 model.
  • Project documentation listing which NVIDIA toolchain components a jitLLM deployment needs for runtime CUDA C and cuTile generation.
  • Published numbers for jitLLM's OpenCL or SPIR-V backends, which would show whether the speed claim holds off NVIDIA hardware.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence30
Adoption
Insufficient
Hype gap+35
Incentives55
Confidence35
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    jitLLM compiles Java bytecode down to CUDA and cuTile and runs LLM inference on NVIDIA GPUs at a claimed 90% of llama.cpp performance, with no C++ code, no Python process and no sidecar.

    ReportedSupportedSource: project claim as reported by dev.to2 sources— create a free account to open themView cited source
  2. [2]

    The roughly 90%-of-llama.cpp figure was demonstrated on an RTX 5090 in the project's own recordings; it is the project's claim, not an independent benchmark.

  3. [3]

    "Treat it as promising until someone replicates it on a plain RTX 4060 with a Q4 model, which is what most of us actually run."

    ReportedSupportedSource: author of the dev.to analysis2 sources— create a free account to open themView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 10, 2026

    This Framework Compiles Java Straight to CUDA: Can It Really Match llama.cpp?

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories