LLM Server — On-Device AI API

Content rating
Everyone
1+
Downloads
Content rating
Everyone
Learn more
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image
Screenshot image

About this app

★ Turn Your Phone Into an AI Server ★

LLM Server lets you run large language models directly on your Android device — and expose them as a local API endpoint that any app, script, or AI agent can connect to.

No cloud. No subscriptions. Your phone becomes the AI backend.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━

🧠 ON-DEVICE INFERENCE
• Run open-weight LLMs (Llama, Gemma, Phi, Qwen, Mistral and more)
• Optimized for mobile hardware with GGUF quantization
• Full privacy — your data never leaves your device
• Works completely offline once models are downloaded

🔌 BUILT-IN API SERVER
• OpenAI-compatible REST API served from your phone
• Connect any tool that speaks the OpenAI API format
• Stream responses in real-time via Server-Sent Events
• Serve multiple clients on your local network

🤖 PERFECT FOR AI AGENTS
• Let your AI assistants call your phone as an LLM backend
• Hermes Agent, LangChain, AutoGPT — anything with HTTP
• Tailscale/ZeroTier support for secure remote access
• No API keys needed — you own the model

📱 HOW IT WORKS
1. Download a model from the built-in model browser
2. Tap "Start Server" to launch the API endpoint
3. Point any client to http://your-phone-ip:8080/v1/chat/completions
4. That's it — your phone is now an AI server

🔧 TECHNICAL DETAILS
• llama.cpp backend for fast CPU/GPU inference
• Supports GGUF models (Q4_K_M, Q5_K_M, Q8_0 and more)
• OpenAI-compatible /v1/chat/completions endpoint
• Configurable context length, temperature, and sampling
• Background service keeps serving even when app is minimized
• Battery-optimized inference scheduling

━━━━━━━━━━━━━━━━━━━━━━━━━━━━

💡 USE CASES
• Personal AI assistant that runs 100% on your device
• Development & testing — prototype AI features without cloud costs
• AI agent backend — let autonomous agents use your phone as their brain
• Privacy-first AI — sensitive data never touches the cloud
• Offline AI — works on planes, in tunnels, anywhere

🔒 PRIVACY & SECURITY
• All inference happens on-device
• No data collection, no telemetry, no cloud dependency
• API server only accessible on your local network by default
• Open-source model ecosystem — no vendor lock-in

━━━━━━━━━━━━━━━━━━━━━━━━━━━━

LLM Server bridges the gap between powerful open-source AI models and the devices you carry every day. Stop paying for cloud API calls — run your own AI, on your own terms.

Download now and turn your phone into an AI powerhouse.
Run AI models on your phone and serve them as an API endpoint to any app
Updated on
Aug 31, 2026

Data safety

Safety starts with understanding how developers collect and share your data. Data privacy and security practices may vary based on your use, region, and age. The developer provided this information and may update it over time.
  • No data shared with third parties
    Learn more about how developers declare sharing
  • No data collected
    Learn more about how developers declare collection

What’s new

v1.7.0 — Improved first-launch experience with automatic starter model setup. New users get a working AI model within a minute of opening the app.
Content rating
Everyone
Learn more

App support

Phone number
+12135339019
About the developer
chan hing tat
c_matrix50@hotmail.com
38 Russell St 銅鑼灣 香港島 Hong Kong

More by magnus2026