Llama.cpp Edge Gallery AI

Contains ads
24 reviews
Content rating
Everyone
1K+
Downloads
Content rating
Everyone
Learn more
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image

About this app

Run capable AI models directly on your Android phone with Llama.cpp Offline AIβ€”the ultimate private, local LLM solution built for on-device performance.

Download a compatible GGUF model, load it once, and start chatting locally with CPU-first inference. Your prompts and generated responses stay strictly inside the app's local chat flow instead of being sent to a cloud AI service. Enjoy total data privacy, zero subscription fees, and complete control over your on-device AI experience.

⚑ Why Choose Llama.cpp Offline AI?

β€’ πŸ”’ Private AI Chat: Run your offline chatbot completely air-gapped. No account required, no sign-ups, no data harvesting, and no cloud processing.
β€’ ⚑ Fast Local LLM Performance: Powered by llama.cpp to deliver optimized, efficient on-device AI directly on mobile hardware.
β€’ πŸ“‚ GGUF Management: Effortlessly download, store, and organize GGUF model files directly within the app.
β€’ πŸ’¬ Real-Time Token Streaming: Watch your offline AI stream responses line-by-line as the LLM generates answers in real time.
β€’ βš™οΈ CPU LLM Thread Controls: Fine-tune thread settings to balance model speed, battery usage, and device thermal states.
β€’ πŸ“Š Benchmark Screen: Measure performance, compare token generation speeds (tok/s), and optimize settings across different GGUF quantizations.
β€’ πŸ’‘ Smart Chat Features: Enjoy quick-start prompts for empty chats, collapsible thinking sections, and bold response highlights.
β€’ πŸ’Ύ RAM Guidance: Clear indicators help you choose the right model size and quantization level for your specific phone RAM capacity.
β€’ ⭐ Favorites & Organization: Bookmark your favorite GGUF models, monitor active downloads, and switch between models seamlessly.

πŸ“± Built for Real Android Hardware

Llama.cpp Offline AI is engineered to bring practical, lightweight Android AI to modern smartphones. Running large language models on mobile devices requires smart memory management, which is why this app focuses on quantized GGUF models tailored for phones with varied RAM capacities.

CPU inference serves as the stable default across Android devices, preventing overheating while delivering predictable output speed. Actual generation speeds depend on your phone processor, available RAM, thermal state, model parameter size (such as 1B, 3B, or 7B models), and quantization type (e.g., Q4_K_M, Q5_K_M).

πŸš€ Everyday Applications for Local AI

Take your private AI assistant everywhere you go. Because Llama.cpp Offline AI requires no internet connection once your model is loaded, you can use it on airplanes, in remote areas, or anywhere you need instant offline intelligence:

β€’ ✍️ Creative Writing: Draft emails, refine prose, outline essays, and brainstorm creative stories entirely offline.
β€’ πŸ“ Summarization & Study Tools: Paste long text to create concise summaries, flashcards, and study guides.
β€’ 🌐 Off-Grid Translation: Translate text across multiple languages without relying on web-based translation APIs.
β€’ πŸ’» Code Assistance: Analyze code snippets, debug logic errors, and write simple functions on the go using specialized coding GGUF models.
β€’ 🧠 Private Brainstorming: Ask sensitive questions or explore new concepts with zero fear of data logging.
β€’ πŸ§ͺ On-Device AI Experiments: Test different quantized LLM architectures, compare llama.cpp optimizations, and explore open-weight model capabilities.

βš–οΈ Model Responsibility & Licensing

All language models accessed through or loaded into the app are provided by third-party authors or hosting platforms (such as Hugging Face). Users are responsible for reviewing and adhering to each model's license and usage terms prior to downloading. Llama.cpp Offline AI is an independent client interface and does not claim ownership of third-party models.

πŸ“© Support & Feedback

Email: contact@shannsingh.com
Private local AI chat & on-device LLM. Run offline GGUF models with llama.cpp.
Updated on
Sep 10, 2026

Data safety

Safety starts with understanding how developers collect and share your data. Data privacy and security practices may vary based on your use, region, and age. The developer provided this information and may update it over time.
  • This app may share these data types with third parties
    App info and performance and Device or other IDs
  • No data collected
    Learn more about how developers declare collection
  • Data isn’t encrypted
  • Data can’t be deleted

Ratings and reviews

23 reviews
Chris Drake
August 21, 2026
there's no way to let you specify where you stored the download, and "scan" finds nothing
Did you find this helpful?
FreeRouter Team
August 21, 2026
we've got a dedicated spot for that in the shared preferences! We'll definitely work on improving and updating the app for you. Thank you
Jason
August 12, 2026
please allow to save chat so that can continue chat when re enter app
Did you find this helpful?
FreeRouter Team
August 23, 2026
New feature alert πŸš€ Thanks for your feedback! Available tomorrow. If you love it, maybe a 5-star update? ⭐⭐⭐⭐⭐
Abdel Basit
August 9, 2026
easy to use
Did you find this helpful?
FreeRouter Team
August 13, 2026
Thank you for your support

What’s new

Offline Tool Calling
Rich Chat Rendering
Storage Model Scanner
Performance & Stability