Emilio Melis / aatricks
  • build 2025.09
  • stack · 8 components
  • published 2025-09-18

llmedge

stack // Kotlin / JNI / C++ / Android NDK / OpenCL / Vulkan / ONNX / GGUF

architecture flow

01

A Kotlin layer owns the app-facing API, coroutine entrypoints, and lifecycle-safe sessions.

02

ModelRepository handles downloading models, validating them, resuming interrupted fetches, and where they get cached.

03

RuntimePool and RuntimeCoordinator reuse native engines, check what a device can do, and skip backends that have already failed on it.

04

JNI bridges connect the Kotlin side to llama.cpp, stable-diffusion.cpp, whisper.cpp, bark.cpp, and the ONNX embedding utilities.

engineering notes

llmedge is an Android library for running AI models on the phone. One Kotlin API over four native engines (llama.cpp, stable-diffusion.cpp, whisper.cpp, bark.cpp) plus ONNX for embeddings. It is published on Maven Central as io.github.aatricks:llmedge and powers on-device chapter summaries in Emaki.

On a real phone

one Kotlin API over llama.cpp / sd.cpp / whisper.cpp. stock Galaxy S22: 35 tok/s text (qwen3-0.6B, Q4_K_M), ~40s for a 128×128 SD1.5 image (20 steps, dpmpp2m). no server, no cloud call.

What it does

  • Text. GGUF models through llama.cpp (an ik_llama.cpp build, so BitNet b1.58 IQ2_BN runs). Batched blocking and streaming generation, native KV-cache reuse across turns, separate prompt and generation thread counts, a ChatSession that replays transcripts for reasoning models, reasoning on/off controls.
  • Tool calling. edge.text.toolAgent(...) lets the model call app-defined tools through a JSON envelope. Read-only tools run automatically. Action tools need an explicit policy. A bash tool exists for JVM hosts.
  • Speech. whisper.cpp with timestamps, language detection, streaming transcription, SRT output. bark.cpp for text-to-speech with ARM optimisations.
  • Image. stable-diffusion.cpp with EasyCache and LoRA, FLUX.2 Klein 4B (distilled DiT), per-step progress streaming, ESRGAN upscaling (Remacri).
  • Video. Wan 2.1, 4 to 64 frames, components loaded sequentially so it fits.
  • Vision. LLaVA-style models and a SmolVLM2-256M preset. OCR through ML Kit.
  • RAG. PDF indexing, ONNX embeddings, vector search, Q&A, all on the phone.
  • Model files. Hugging Face download with resume, progress, private repos, validation, old-mirror redirects. Big files can go through Android’s DownloadManager so they stay off the Dalvik heap. Context window is read from the model and capped to what the heap allows (2K to 8K).
  • On-device conversion. safetensors to GGUF on the phone (Llama arch, GPT2-BPE tokenizer), with optional Q8_0 / Q4_K_M / IQ2_BN quantisation. So a model that only exists as safetensors can still run.

How it’s built

  • LLMEdge is a thin facade that lazy-creates text, speech, image, vision, rag clients on first access. Each client is also constructible on its own.
  • ModelRepository owns download, validation, and caching. Inference clients never fetch files themselves.
  • RuntimePool and RuntimeCoordinator cache native runtimes across calls and pick the backend. Order is OpenCL, then Vulkan, then CPU. A backend that fails on a device gets recorded in a verdict store and is not retried on that device.
  • RuntimePoolProfile lets each domain say how its pool is keyed, sized, and loaded without duplicating the pool code.
  • Native loading is explicit and overridable, so the Kotlin layer runs in JVM tests without an emulator. 161 test files, including Linux end-to-end runs of the native engines.

Distribution & docs

Published on Maven Central as io.github.aatricks:llmedge. Emaki uses it in production for local chapter summaries. The llmedge-examples repo demonstrates each modality in standalone sample apps, and the documentation site covers setup, architecture, and runtime configuration.