Vanta: Local AI & Self-Hosted app icon

Vanta: Local AI & Self-Hosted

EZEL BAYRAKTAR

★ 5.03 ratingsDeveloper Tools$5.99

iPhone Screenshots

App screenshotApp screenshotApp screenshotApp screenshotApp screenshotApp screenshot

Description

Bring your own AI to iPhone and iPad. Connect to Ollama, vLLM or LM Studio, download supported models to run on-device, or use cloud providers with your own API keys. Ask questions about PDFs, compare models and switch models without switching apps. Buy Vanta once. No Vanta account. No Vanta subscription. Your purchase covers the app; cloud APIs and optional third-party services may charge separately. CONNECT TO YOUR SELF-HOSTED MODELS Use Ollama, vLLM, LM Studio or any compatible chat API endpoint on your computer, homelab or remote server. Add its address and any required API key, choose a model, and chat from your phone or tablet. Your server does the processing, so you can use models that do not fit on your phone. Vanta gives you a mobile interface for chatting, working with files and using tools. RUN MODELS DIRECTLY ON YOUR IPHONE OR IPAD No server required. Browse Hugging Face inside Vanta and download supported GGUF and MLX models, or import your own GGUF file. Run models from families including Llama, Qwen, Gemma, Phi and Mistral, plus DeepSeek distills. Once downloaded, on-device models work offline with no API key or per-token fees. Model compatibility and performance depend on your device and available memory. You can also chat with the existing Apple Intelligence on-device model when available, without a separate model download in Vanta. CONNECT CLOUD APIS AND CUSTOM ENDPOINTS Add your own API keys for your cloud providers. You are not limited to a fixed provider list: Vanta supports custom compatible chat API endpoints. Switch providers and models mid-conversation. CHAT WITH PDFS, DOCUMENTS AND PHOTOS Ask questions about a PDF, summarize documents, compare files, or analyze photos and screenshots with a compatible vision model. Use on-device or self-hosted models, or choose a cloud provider. CONNECT TOOLS AND SEARCH THE WEB Add Model Context Protocol (MCP) servers for web search, documentation, live data and your own custom tools. Choose which servers to connect and use their tools with compatible models. Deep Search researches a topic across web sources and brings back answers with citations inside your conversation. PUT MULTIPLE MODELS IN ONE CONVERSATION Assign roles, compare answers and let models challenge each other. Build a panel from on-device, self-hosted and cloud models instead of copying questions between separate apps. MEMORY AND VOICE Let Vanta remember useful details across conversations. Keep memory on-device, configure how it works, or turn it off. Speak instead of typing with Whisper-based, on-device speech recognition. Transcription works offline. Optional ElevenLabs text-to-speech adds spoken replies. CONTROL YOUR PROMPTS AND CONTEXT Generate system prompts from plain English, then edit, save and reuse them. Adjust temperature, top-p, penalties and context limits where supported. Inspect message, tool and system token usage, and automatically summarize long conversations to manage context. Markdown, syntax-highlighted code, themes and appearance controls are included. KNOW WHERE YOUR DATA GOES For offline chats with on-device models, your prompts and responses stay on your device. Server and cloud chats go to the endpoints you choose, with no Vanta AI server in between. External tools and online services receive the data needed to handle your requests. GETTING STARTED Choose an on-device model, connect a reachable server, or add a provider API key. You only need one of these to start. Apple Intelligence requires a supported device and an available system model. Features such as vision and tools depend on your selected model and provider.