Content rating
Everyone
10+
Downloads
Content rating
Everyone
Learn more
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image

About this app

local LLM runs open-weight AI models directly on your phone. There is no server, no account, and no internet connection needed once a model is downloaded. Your conversations never leave the device.

WHY ON-DEVICE

Every prompt you type is answered by a model running on your own hardware. Nothing is uploaded, nothing is logged on a server, and nothing is used to train anyone's model. Put the phone in aeroplane mode and it still works.

CHOOSE YOUR MODEL

Seven models are available in-app, from a 369 MB SmolLM2 that runs comfortably on modest hardware up to Llama 3.2 3B for stronger answers:

• SmolLM2 360M — fastest, lowest RAM
• Qwen2.5 0.5B — very fast, solid quality
• TinyLlama 1.1B — lightweight classic
• Llama 3.2 1B — balanced default
• Qwen2.5 1.5B — stronger reasoning
• Gemma 2 2B — high quality, slower
• Llama 3.2 3B — best quality, needs headroom

Each shows its exact download size before you commit. You can also paste a direct link to any GGUF model you like, and switch models at any time without leaving your chat.

CHAT THAT REMEMBERS

Conversations are saved in a private database on your device. Titles are written by the model itself, so your history is easy to scan. Reopen any past chat, or clear them all in one tap. Markdown renders properly, with syntax-highlighted code blocks and one-tap copy.

A LOCAL API FOR YOUR OTHER DEVICES

Turn on the optional API service and your phone serves the model over an OpenAI-compatible HTTP endpoint on your local network. Point any OpenAI client at it — curl, the Python SDK, your own scripts — and get completions from the model in your pocket, over USB or Wi-Fi. Streaming is supported. It stays off until you enable it, and runs with a visible notification while active.

BUILT FOR PRIVACY

• No account, no sign-in, no ads
• Chats stored only on your device
• Message text is never transmitted anywhere
• Anonymous usage analytics only, never message content

A NOTE ON AI OUTPUT

Responses come from open-weight models running locally. They can be wrong, outdated, or unsuitable, and are not reviewed by anyone. Please don't rely on them for medical, legal, or financial decisions.

REQUIREMENTS

Android 7.0 or later, 64-bit device. A one-time model download of 369 MB or more, and roughly 1-3 GB of free RAM depending on the model you pick. Larger models are noticeably slower on older phones.
On Device LLM
Updated on
Sep 11, 2026

Data safety

Safety starts with understanding how developers collect and share your data. Data privacy and security practices may vary based on your use, region, and age. The developer provided this information and may update it over time.
  • No data shared with third parties
    Learn more about how developers declare sharing
  • No data collected
    Learn more about how developers declare collection

What’s new

First release of local LLM.

• Chat with AI models that run entirely on your phone. No cloud, no account, and nothing you type ever leaves the device.
• Choose from seven models, from a 369 MB SmolLM2 up to Llama 3.2 3B, or paste a link to your own GGUF file.
• Conversations are saved on-device, with titles written by the model itself.
• Optional local API, OpenAI-compatible, so your laptop can use the model over USB or Wi-Fi.