FlowDown is a fast, minimal AI chat app that makes working with large language models smooth, focused, and fully under your control.
Why FlowDown
Lightweight and efficient
Focus on your conversation, not the app. FlowDown launches quickly, stays out of your way, and keeps long chats responsive, even on busy machines.
Rich formatting with Markdown
Get clear, readable answers. FlowDown renders Markdown beautifully, so code blocks, bullet lists, tables, and headings in AI responses are easy to scan and use.
Works with the tools you already have
Connect to any OpenAI-compatible API, from major cloud providers to services running in your own home lab. If it speaks the OpenAI API, FlowDown can talk to it.
Talk with text, images, and audio
Show, don't just tell. Paste screenshots, drop in images, or record your voice to ask questions. Use it to:
- Transcribe meetings or lectures
- Summarize voice notes
- Ask about what's in a photo
Conversations that feel instant
Blazing-fast text rendering makes responses stream in fluidly, so you can start reading and refining your prompt without waiting for the model to finish.
Automatic chat titles
Stop hunting for the right thread. FlowDown names your conversations automatically based on their content, so you can quickly find that "API bug fix" or "Trip planning" chat later.
Shortcut integrations
Use FlowDown in your workflows. With Shortcuts support, you can:
- Send selected text from another app to your model
- Generate drafts, summaries, or replies with one tap
- Build custom automations that call your favorite model
Privacy by design
Your data is yours. FlowDown does not collect your chats. For maximum confidentiality, you can run models locally on your own machine with no external server involved.
Open source and transparent
Every line of code is available on GitHub. Inspect it, audit it, or adapt it to your own setup. You can see exactly how your AI client works.
The source code is available at: https://github.com/Lakr233/FlowDown
Local models
For offline, private use, FlowDown can run models locally on Apple Silicon devices with sufficient memory. Performance varies by device and model variant.
Supported model families include:
- Llama, Mistral, Mixtral, and Mistral 3
- Phi, Phi-3, and PhiMoE
- Gemma, Gemma 2, Gemma 3, Gemma 3n, and Gemma 4
- Qwen 2, Qwen 3, Qwen 3 MoE, Qwen 3 Next, Qwen 3.5, and Qwen 3.5 MoE
- DeepSeek V2 and V3
- GLM-4 and GLM-4 MoE
- Granite and Granite MoE Hybrid
- LFM2 and LFM2 MoE
- OLMo 2, OLMo 3, and OLMoE
- MiMo and MiMo V2 Flash
- MiniCPM (v1/v2/v4; v3 uses a different architecture)
- Cohere, InternLM 2, OpenELM, and StarCoder 2
- Falcon H1, BitNet, SmolLM 3, ERNIE 4.5, and EXAONE 4
- Baichuan M1, Bailing MoE, AfMoE, and MiniMax
- GPT-OSS, Jamba, Mamba 2, Apertus, and Hunyuan
- AceReason, Nanbeige, NanoChat, Nemotron H, Nemotron Labs Diffusion, and Lille 130M
This gives you:
- Full control over your data
- No dependency on external servers
- The ability to experiment with different open models
Please be aware that running these models may require significant device resources.
Remote models
We support a wide range of OpenAI-compatible APIs. For consistent, reliable use, we recommend:
- Bringing your own API, or
- Hosting your own inference server on local or home-lab hardware
FlowDown supports all OpenAI-compatible endpoints, making it ideal for connecting to tools like Ollama and LM Studio running on your own machine for maximum privacy.
Some setups require technical knowledge. If you are new to self-hosted AI or custom endpoints, please read our documentation carefully; the app can be challenging without it.
Important note
AI-generated content can be inaccurate or misleading. Always verify critical information and use FlowDown responsibly; you are responsible for how you use the output.
© 2026 FlowDown Team. All rights reserved.
more