About Local LLM
On Device LLM
local LLM runs open-weight AI models directly on your phone. There is no server, no account, and no internet connection needed once a model is downloaded. Your conversations never leave the device.
WHY ON-DEVICE
Every prompt you type is answered by a model running on your own hardware. Nothing is uploaded, nothing is logged on a server, and nothing is used to train anyone's model. Put the phone in aeroplane mode and it still works.
CHOOSE YOUR MODEL
Seven models are available in-app, from a 369 MB SmolLM2 that runs comfortably on modest hardware up to Llama 3.2 3B for stronger answers:
• SmolLM2 360M — fastest, lowest RAM
• Qwen2.5 0.5B — very fast, solid quality
• TinyLlama 1.1B — lightweight classic
• Llama 3.2 1B — balanced default
• Qwen2.5 1.5B — stronger reasoning
• Gemma 2 2B — high quality, slower
• Llama 3.2 3B — best quality, needs headroom
Each shows its exact download size before you commit. You can also paste a direct link to any GGUF model you like, and switch models at any time without leaving your chat.
CHAT THAT REMEMBERS
Conversations are saved in a private database on your device. Titles are written by the model itself, so your history is easy to scan. Reopen any past chat, or clear them all in one tap. Markdown renders properly, with syntax-highlighted code blocks and one-tap copy.
A LOCAL API FOR YOUR OTHER DEVICES
Turn on the optional API service and your phone serves the model over an OpenAI-compatible HTTP endpoint on your local network. Point any OpenAI client at it — curl, the Python SDK, your own scripts — and get completions from the model in your pocket, over USB or Wi-Fi. Streaming is supported. It stays off until you enable it, and runs with a visible notification while active.
BUILT FOR PRIVACY
• No account, no sign-in, no ads
• Chats stored only on your device
• Message text is never transmitted anywhere
• Anonymous usage analytics only, never message content
A NOTE ON AI OUTPUT
Responses come from open-weight models running locally. They can be wrong, outdated, or unsuitable, and are not reviewed by anyone. Please don't rely on them for medical, legal, or financial decisions.
REQUIREMENTS
Android 7.0 or later, 64-bit device. A one-time model download of 369 MB or more, and roughly 1-3 GB of free RAM depending on the model you pick. Larger models are noticeably slower on older phones.
What's new in the latest 1.0
• Chat with AI models that run entirely on your phone. No cloud, no account, and nothing you type ever leaves the device.
• Choose from seven models, from a 369 MB SmolLM2 up to Llama 3.2 3B, or paste a link to your own GGUF file.
• Conversations are saved on-device, with titles written by the model itself.
• Optional local API, OpenAI-compatible, so your laptop can use the model over USB or Wi-Fi.
Local LLM APK Information
Old Versions of Local LLM
Super Fast and Safe Downloading via APKPure App
One-click to install XAPK/APK files on Android!





