br-AI-n is a secure, private, and offline-first AI assistant for your Android device. It allows you to download and run state-of-the-art open-source Large Language Models (LLMs) directly on your phone's hardware with zero cloud dependencies and complete data privacy.
Because br-AI-n executes inference locally, your chats, voice recordings, and image attachments never leave your phone. And whenever you need up-to-the-minute real-world information, optional live Web Search grounds your answers with real-time sources and cited footnotes.
š Key Features
š Offline-First & 100% Private by Default
No cloud APIs, no remote AI servers, and zero data tracking. Your model weights, chat sessions, and prompts remain strictly on your device.
š Optional Live Web Search & Citations
Break free from knowledge cutoff limits! Enable live web search to ground answers with real-time news and verified web sources. Explore cited claims with interactive, superscript footnotes ([1]) and direct source links.
š Customizable & Creative Personas
Tailor your AIās personality! Choose from built-in personas (Code Expert, Creative Writer, Researcher, Joker, and more) or create your own custom personas. Customize custom prompts, names, and emoji icons with real-time system prompt updates.
š§ Deep Reasoning & Scratchpad
Watch the model think in real time with transparent reasoning traces and multi-hop search synthesis inside an expandable thinking drawer.
šļø Voice Transcription & Real-Time Speech
Speak naturally with local voice transcription, and listen to low-latency streaming text-to-speech (TTS) that reads responses aloud sentence-by-sentence as they generate.
šļø Vision & Document OCR
Attach images, receipts, diagrams, or documents to transcribe, inspect, and discuss with multimodal visual models.
ā” Real-Time Hardware & Inference Stats
Monitor real-time token generation speeds (tokens/sec), RAM utilization, context length limits, and storage headroom.
š Saved Local Chat History
All conversations are organized locally in a persistent SQLite database. Resume past discussions, search topics, or clear context with ease.
š¤ Broad Open-Source Model Support
Download optimized GGUF quantizations of leading open-weight architectures, including Llama 3.2, Qwen 2.5, Gemma 2, and Phi series directly from Hugging Face.
š”ļø No Accounts, No Subscriptions
No logins, no tracking cookies, and no recurring API fees. Full on-device AI power right in your pocket.
āļø Technology
Powered by a high-performance, mobile-optimized C++ inference engine (llama.cpp), br-AI-n leverages your phone's CPU and GPU architecture to deliver state-of-the-art quantized local intelligence.
Note: Generation speed, loading times, and memory requirements depend on your device's processor and total RAM. Check the built-in Hardware Profile in the Model Manager to pick the optimal model size for your device.
Private on-device AI chat with optional web search, custom personas & voice.