LLM AI Server with llama.cpp
7.3 MB
File Size
- Security
Everyone
Android 7.0+
Android OS
About LLM AI Server with llama.cpp
LLM AI server providing advanced local generation controls with API/WebUI.
Unlimited AI, zero cloud. Your phone is the server.
Run powerful LLMs locally, privately, and freely — with a lightweight UI anyone can use.
One Android device becomes your personal AI backend, accessible from PCs, tablets, and other devices on the same network.
---
1. Overview
LLM AI Server with llama.cpp is a fully local LLM server for Android, enabling private, offline text and multimodal generation.
It supports Hugging Face model search, GGUF downloads, local model loading, and optional MTP speculative decoding for faster inference.
Supported model families include Gemma‑4 / Gemma‑4 Vision, Qwen / Qwen Vision, Mistral, LLaMA, Phi, Bonsai, and others.
A built‑in API server is compatible with the Ollama, OpenAI, and Anthropic (Claude) Messages APIs — providing chat, completion, generation, embedding, tokenize/token‑count, and model‑listing endpoints, so existing OpenAI, Ollama, and Claude clients and SDKs can point straight at your phone.
The bundled WebUI runs on the same port and supports model switching, parameter editing, log viewing, and PWA installation.
MCP and Function Calling enable tool‑augmented workflows, and structured output is supported via GBNF and JSON Schema.
Profile backup & restore allow saving and transferring configuration sets.
---
2. Intended Users and Supported Devices
Designed for users who want a private, fully local LLM environment:
- Offline‑first users
- Developers integrating a local backend
- Advanced users tuning sampling and performance
- Researchers testing inference behavior
- Privacy‑focused users
Performance can be tuned via context size, threads, batch size, GPU offload, KV cache quantization, and optional MTP decoding.
---
3. Key Features
- Hugging Face GGUF search & download
- Local GGUF loading
- Fully offline LLM server
- Gemma‑4 Vision & Qwen Vision projector detection
- MTP speculative decoding
- Detailed parameter control (Mirostat, DRY, XTC, Min‑p, Typical‑p, penalties, dynamic temperature)
- Expanded KV cache quantization options
- Improved nPredict handling
- Integrated Ollama / OpenAI / Anthropic (Claude) Messages‑compatible API server
- Claude‑compatible /v1/messages (+ token counting) — drop‑in for Anthropic SDK clients
- Embedding API
- Tokenizer API
- Automatic prompt template selection
- Streaming / non‑streaming output
- Timestamped logs
- Enhanced WebUI with multi‑language support
- PWA support
- MCP and Function Calling
- GPU load stability improvements
- Profile backup & restore support
Available endpoints — Ollama: /api/chat, /api/generate, /api/embed(dings), /api/tokenize, /api/tags · OpenAI: /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models, /v1/responses/input_tokens · Anthropic (Claude): /v1/messages, /v1/messages/count_tokens.
Full API docs, examples & an interactive tester:
https://micklab.web.app/llama-tester-en.html
---
4. Getting Started
You can begin using the app immediately in either of the following ways:
A) Start API/WebUI → Open in browser
1. Tap Start API/WebUI on the main screen.
2. Tap Open in browser to launch the bundled WebUI.
3. You can chat and run the model directly from the WebUI.
- Note: On first use, if the model is not downloaded yet, the WebUI will download it when you send the first message.
Large models may take time to download.
B) Direct Run
1. Open the Direct Run section on the main screen.
2. Enter a prompt and tap Send to run the model directly inside the app.
- If the model is not downloaded yet, the app will show a notice before downloading.
C) or Use Settings to prepare a model first
1. Open Settings.
2. Search Hugging Face for a GGUF model or enter a direct URL or import a local GGUF file.
3. Adjust parameters (context size, temperature, penalties, MTP, GPU offload, etc.).
4. Tap Save Config, then SAVE & CLOSE to load the model.
5. Optionally create a profile backup for later restoration.
What's new in the latest 1.590
- API server keeps running during screen-off / low-power state, with auto-restart and preserved state
- Fixed shutdown-related issues; more stable large-model loading (mmap), clearer crash messages, new profiles: default / Chat / Vision
LLM AI Server with llama.cpp APK Information
Old Versions of LLM AI Server with llama.cpp
Super Fast and Safe Downloading via APKPure App
One-click to install XAPK/APK files on Android!






