BastionChat
Chat with Documents Privately
USD 11.99 · Designed for iPad
Open-source AI Agents in your pocket — on a boat, on a plane, or on Mars. Works offline. Index docs, run models, agentic search. No cloud. Completely yours.
Your data belongs to you. Your AI should too.
Bastion Chat puts a powerful, private AI assistant in your pocket — no subscriptions, no usage limits, no servers watching over your shoulder. Everything runs locally on your device, so your conversations stay completely confidential. Even offline.
——— KEY FEATURES ———
YOUR PRIVATE AI, YOUR RULES
Experience top-tier reasoning and "thinking" models without compromising an ounce of privacy. All processing happens on-device. Your conversations are never uploaded, never analyzed, never sold. Just you and your AI.
CURATED MODEL CATALOG
Browse and download a hand-picked selection of the best open-source models — from compact 1B models that run on any iPhone, to powerful 27B models on iPad and Mac. Thinking models, vision models, tool-use specialists — all pre-configured and ready to run.
VISION & MULTIMODAL AI
Ask questions about images, analyze screenshots, describe photos, or extract text from documents. Bastion Chat supports multimodal models that see and understand the world alongside you.
AGENTIC AI — YOUR AI THAT ACTS
Go beyond chat. Enable Agent Mode to let Bastion Chat reason through complex tasks, plan multi-step solutions, and execute actions autonomously using built-in tools. Your private AI that doesn't just talk — it does.
FULL RAG & DOCUMENT INDEXATION
A complete knowledge pipeline in the palm of your hand. Import PDFs, web pages, Wikipedia articles, or plain text. Bastion Chat indexes your personal knowledge base using full retrieval-augmented generation so answers are always grounded in your content.
AI-POWERED SUMMARIES IN SECONDS
Overwhelmed by lengthy articles or dense reports? Bastion Chat distills complex content into clear, actionable summaries. Grasp the key points instantly and save hours of reading.
SMART WEB SEARCH, ON DEMAND
Need real-time information? Enable the optional web search tool to ground your answers in live results — without sacrificing your privacy or your workflow.
BRING YOUR OWN MODEL
Import any GGUF model directly from Hugging Face or your local files. Capabilities like advanced reasoning, tool use, and vision are detected and configured automatically on import.
SCALES WITH YOUR DEVICE
On iPhone, run fast and efficient models for everyday tasks. On iPad and Mac, unleash frontier-class models like Qwen 27B for deep research, long-form writing, and complex reasoning — all still 100% local.
ONE PURCHASE. YOURS FOREVER.
No usage caps. No API costs. Pay once and run unlimited queries for as long as you own your device.
Ratings & Reviews
- This app has not received enough ratings or reviews to display an overview.
This is the biggest update to Bastion Chat yet.
VISION & IMAGE UNDERSTANDING
Ask questions about photos and images. Multimodal models can now see, describe, and analyze images directly in your conversations. Image history is preserved across sessions.
AGENTIC MODE
Bastion Chat can now act, not just chat. Enable Agent Mode for multi-step reasoning with autonomous tool use — your AI plans, executes, and delivers results.
WEB SEARCH
Optionally connect to the web for real-time information. Ground answers in live results without leaving the app or compromising your privacy.
WIKIPEDIA IMPORT
Pull any Wikipedia article directly into your knowledge base and query it instantly with full RAG.
BUILT-IN KNOWLEDGE PACKS
Curated, system-managed knowledge packs (first one available: Pet Care) come pre-loaded and searchable — no setup required.
SMARTER MODEL BROWSER
Redesigned with collapsible sections, family grouping, inline quantization picker, and quality tiers so you always know which model fits your device best.
FAVORITES
Star your most-used models for instant access at the top of the model list.
PER-MODEL SAMPLING PROFILES
Fine-tune temperature, context length, and inference settings independently for each model.
SMARTER CUSTOM MODEL IMPORT
Auto-detects capabilities (reasoning, tool calls, vision) on import. No manual configuration needed.
TURBOQUANT KV CACHE
New TurboQuant memory modes (3.5-bit and 2-bit) let you run larger models with dramatically less RAM — without sacrificing response quality.
FLASH ATTENTION
Flash Attention is now enabled by default on all Apple Silicon devices, significantly reducing memory usage during long conversations and improving inference speed.
FULL METAL GPU OFFLOAD
All model layers now offload to the Metal GPU by default, with device-aware thread and micro-batch tuning for the fastest possible inference on your iPhone, iPad, or Mac.
GEMMA 4 THINKING SUPPORT
Native support for Gemma 4's thinking block format, including proper stop sequence handling and streamed thinking output.
Bug fixes: MCP crash fix, Metal SET_ROWS crash on hybrid GDN models (Qwen 3.6 27B), atomic interrupt flag, download integrity checks, and Qwen context length handling.
The developer, Frederic Ayala Heinrich, indicated that the app’s privacy practices may include handling of data as described below. For more information, see the developer’s privacy policy .
Data Not Collected
The developer does not collect any data from this app.
Accessibility
The developer has not yet indicated which accessibility features this app supports. Learn More
Information
- Provider
- Frederic Ayala Heinrich
- Size
- 82.5 MB
- Category
- Productivity
- Compatibility
Requires iOS 18.0 or later.
- iPhone
Requires iOS 18.0 or later. - iPad
Requires iPadOS 18.0 or later. - Mac
Requires macOS 15.0 or later and a Mac with Apple M1 chip or later. - Apple Vision
Requires visionOS 2.0 or later.
- iPhone
- Languages
English and 7 more
- English, French, German, Italian, Japanese, Korean, Portuguese, Spanish
- Age Rating
4+
- 4+
- Copyright
- © Frederic Ayala Heinrich (c) 2026

