jitLLM compiles Java bytecode straight to CUDA for LLM inference
The TornadoVM team at the University of Manchester, with Red Hat, unveiled jitLLM, a Java inference engine that compiles bytecode to CUDA and cuTile without C++ or Python. It claims about 90% of llama.cpp performance on an RTX 5090, with no independent benchmarks yet.
- jitLLM runs Llama 3, Mistral, Qwen 2.5/3, Phi-3, Granite and Devstral 2 in GGUF format
- Claimed ~90% of llama.cpp performance on an Nvidia RTX 5090
- Ships an official LangChain4j engine and a Quarkus extension
- Java code compiles to CUDA, OpenCL or SPIR-V via TornadoVM
Read next
Software