Content rating
Everyone
100+
Downloads
Content rating
Everyone
Learn more
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image

About this app

Turn your Android phone into a powerful, private AI server — no cloud, no subscription, no data leaving your device.

LLM Hoster lets you download and run large language models (LLMs) directly on your phone using the llama.cpp inference engine, and exposes a fully compatible API over your local network. Use it with any client that supports the baseAPI/v1API — Open WebUI, Continue.dev, Chatbox, LM Studio, or your own apps.

━━━━━━━━━━━━━━━━━━━━

🔒 COMPLETELY PRIVATE
Your data never leaves your device. No internet required for inference. No API keys, no cloud costs, no telemetry. What happens on your phone stays on your phone.

📥 ONE-TAP MODEL DOWNLOADS
Browse and download GGUF models from Hugging Face with a single tap. Just paste a repo name like "author/model_name-GGUF" — the app automatically finds and downloads the best quantized version for mobile.

🌐 /v1-COMPATIBLE API
Expose a real REST API on your LAN:
• GET /v1/models
• POST /v1/chat/completions
• POST /v1/completions
Connect from your laptop, tablet, or any device on the same network.

💬 BUILT-IN CHAT
Chat with your loaded model directly inside the app. No external client needed for quick conversations.

⚡ PERFORMANCE TUNING
Fine-tune inference for your device:
• CPU core allocation (2–8 threads)
• KV cache quantization (q4_0, q8_0)
• Context window size (2K–16K tokens)
• Batch size control
• GPU layer offloading
• Flash Attention support

🔧 RUNS IN THE BACKGROUND
A foreground service keeps the server alive even when you switch apps. Stop it anytime from the notification shade.

━━━━━━━━━━━━━━━━━━━━

SUPPORTED MODELS
Works with any GGUF-format model. Recommended starting points:
• Llama 3.2 (1B / 3B)
• Phi-3 Mini
• Gemma 2 (2B)
• Qwen 2.5 (1.5B / 3B)
• Mistral 7B (Q4 quantized)

Smaller quantized models (Q4_K_M, Q4_0) run best on phones. Models are stored locally and persist across sessions.

━━━━━━━━━━━━━━━━━━━━

USE CASES
• Private AI assistant with zero cloud dependency
• Development & testing of LLM-powered apps
• Offline AI in areas with limited connectivity
• Learning and experimenting with open-source models
• Serving AI to multiple devices on your home network

━━━━━━━━━━━━━━━━━━━━

No ads. No tracking. No accounts. Just AI on your phone.

Download LLM Hoster and start hosting your own private AI server today.
Run AI models locally LLM's on your phone & serve over LAN.
Updated on
Jun 2, 2026

Data safety

Safety starts with understanding how developers collect and share your data. Data privacy and security practices may vary based on your use, region, and age. The developer provided this information and may update it over time.
  • No data shared with third parties
    Learn more about how developers declare sharing
  • No data collected
    Learn more about how developers declare collection

What’s new

• Advanced Settings: New controls for GPU offloading, thread batches, and speculative decoding.
• Stability improvements and general UI cleanup.
Content rating
Everyone
Learn more

App support

About the developer
Sachin Ramesh Jha
joblessaideveloper@gmail.com
India