★ Turn Your Phone Into an AI Server ★
LLM Server lets you run large language models directly on your Android device — and expose them as a local API endpoint that any app, script, or AI agent can connect to.
No cloud. No subscriptions. Your phone becomes the AI backend.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🧠 ON-DEVICE INFERENCE
• Run open-weight LLMs (Llama, Gemma, Phi, Qwen, Mistral and more)
• Optimized for mobile hardware with GGUF quantization
• Full privacy — your data never leaves your device
• Works completely offline once models are downloaded
🔌 BUILT-IN API SERVER
• OpenAI-compatible REST API served from your phone
• Connect any tool that speaks the OpenAI API format
• Stream responses in real-time via Server-Sent Events
• Serve multiple clients on your local network
🤖 PERFECT FOR AI AGENTS
• Let your AI assistants call your phone as an LLM backend
• Hermes Agent, LangChain, AutoGPT — anything with HTTP
• Tailscale/ZeroTier support for secure remote access
• No API keys needed — you own the model
📱 HOW IT WORKS
1. Download a model from the built-in model browser
2. Tap "Start Server" to launch the API endpoint
3. Point any client to http://your-phone-ip:8080/v1/chat/completions
4. That's it — your phone is now an AI server
🔧 TECHNICAL DETAILS
• llama.cpp backend for fast CPU/GPU inference
• Supports GGUF models (Q4_K_M, Q5_K_M, Q8_0 and more)
• OpenAI-compatible /v1/chat/completions endpoint
• Configurable context length, temperature, and sampling
• Background service keeps serving even when app is minimized
• Battery-optimized inference scheduling
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
💡 USE CASES
• Personal AI assistant that runs 100% on your device
• Development & testing — prototype AI features without cloud costs
• AI agent backend — let autonomous agents use your phone as their brain
• Privacy-first AI — sensitive data never touches the cloud
• Offline AI — works on planes, in tunnels, anywhere
🔒 PRIVACY & SECURITY
• All inference happens on-device
• No data collection, no telemetry, no cloud dependency
• API server only accessible on your local network by default
• Open-source model ecosystem — no vendor lock-in
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
LLM Server bridges the gap between powerful open-source AI models and the devices you carry every day. Stop paying for cloud API calls — run your own AI, on your own terms.
Download now and turn your phone into an AI powerhouse.
Run AI models on your phone and serve them as an API endpoint to any app