š¤ Fluent AI ā Private Offline LLM + Claude, GPT-4 & Gemini
Run AI entirely on your device ā no cloud, no account, no data sent anywhere. Then switch to Claude, GPT-4 or Gemini when you need more power. One app. Every AI. Always private.
⨠WHAT'S NEW IN v1.3
š„ MEDICAL AI (MedGemma)
⢠Google's MedGemma 4B ā clinical Q&A and biomedical text, 100% on-device
⢠Requires accepting Google's Health AI Developer Foundation Terms
⢠Not a substitute for professional medical advice
š¤ AGENTIC MODE
⢠On-device AI agent with 12 built-in skills
⢠Runs tasks autonomously: calendar events, web research, document digest, trip planning
⢠Agent Task Inspector ā see every reasoning step in real time
⢠3 free agent runs/day ā no subscription needed to start
⢠Scheduled tasks available with Premium
ā” LITERT MTP ā UP TO 2Ć FASTER
⢠Gemma 4n E2B/E4B with Multi-Token Prediction on Android GPU
⢠Speculative decoding ā more tokens per step, same quality
⢠Tok/s display measures decode-phase speed only for accurate results
šļø ON-DEVICE VISION (Android)
⢠Attach photos using Gemma 4n ā processed entirely on-device
⢠No image uploaded to any server, ever
š PRIVACY FIRST
⢠Conversations stay on your device
⢠Optional local models = zero cloud data
⢠API keys encrypted with AES ā never stored in plain text
⢠No mandatory account required
š§ LOCAL AI MODELS
⢠GGUF / llama.cpp: Gemma 3/4, Qwen 3.5, Phi-4, Llama, DeepSeek R1, Nemotron, MedGemma
⢠LiteRT (Android GPU/NPU): Gemma 4n E2B/E4B ā vision + MTP speculative decoding
⢠Apple MLX: Native Metal on Apple Silicon and iOS 18+ (A17 Pro+)
⢠Q5_K GPU acceleration on Qualcomm Adreno (alongside Q4_0)
⢠Device-aware model recommendations based on your RAM and chipset
⢠Browse, download, and manage models in-app ā no sideloading needed
⢠Import custom GGUF from HuggingFace URL or device storage
āļø CLOUD AI
⢠Claude (Anthropic), GPT-4 (OpenAI), Gemini (Google)
⢠OpenRouter ā 200+ models via a single API key
⢠Streaming, vision, and tool calling across all providers
š ONLINE SERVERS
⢠Ollama Cloud and self-hosted Ollama
⢠LM Studio, vLLM, LocalAI, and any OpenAI-compatible /v1 API
⢠Multiple server profiles with per-profile encrypted auth headers
š¤ VOICE MODE
⢠5 conversation modes: Normal, Interview, Learning, Storytelling, Translation
⢠Animated waveform, voice commands (speed, repeat, stop)
⢠Quick-capture mic button directly in the chat input bar
š KNOWLEDGE BASES (RAG)
⢠Import PDFs, TXT, and Markdown ā AI references your docs when answering
⢠Semantic search for relevant context, topic and project organisation
š§ POWER FEATURES
⢠Tool calling: Calculator, DateTime, Weather, Web Search, mem0 Memory
⢠MCP servers: GitHub, Slack, Notion, Supabase, and 20+ presets
⢠Code execution: Python, Bash, Node.js from code blocks (desktop + mobile JS)
⢠Model benchmarking: tok/s, TTFT, MMLU-50 quality score, shareable PNG cards
⢠Slash commands: /agent, /clear, /export, /voice, /template and more
⢠Per-chat thinking toggle for Qwen3, DeepSeek R1, Nemotron reasoning models
⢠URL context injection ā paste a link, AI reads the page for context
⢠Polish Before Send ā AI rewrites your draft before you hit send
⢠Continue button ā resumes responses cut off at the token limit
š CHAT ORGANISATION
⢠Folders, tags, and cross-chat full-text search across every message
⢠HuggingFace model browser with bookmarks and memory fitness badges
⢠Conversation branching and message reactions
š PREMIUM (OPTIONAL)
⢠Ad-free experience
⢠Scheduled agent tasks (recurring or one-time)
⢠Priority feature access and advanced analytics
š± PERFECT FOR
ā Privacy-focused users ā local models, zero cloud data
ā Android power users ā LiteRT GPU/NPU with MTP acceleration
ā Developers ā benchmark GGUF, LiteRT, and MLX side-by-side
ā Healthcare researchers ā MedGemma on-device, no upload needed
ā Students ā knowledge bases for study documents and materials
ā Professionals ā agentic tasks, document Q&A, and tool calling
AI chat assistant with offline models - private, customizable, multimodal