Our #Java #AI inference engine just got another major update with a huge performance leap and native #CUDA integration!
GPULlama3 v1.0.0 is out - now #CUDA-native ⚡ ✅ TornadoVM CUDA backend w/ tensor-core (MMA) batch prefill ✅ FP16 & Q8_0 · Llama, Qwen, Devstral & more models ✅ An OpenAI-compatible server (llama-tornado --server) Pure #Java. No JNI. github.com/beehive-lab/GPULlama3.java #opensource #AI #LLM #GPU