AI · 2024 · Live
AI Chat
Ultra-secure, privacy-first AI client with local LLM execution capabilities.
- Role
- Creator & Full-Stack Engineer
- Duration
- 6 weeks
- Team size
- Solo
Project overview
What this product is
AI Chat is a privacy-first desktop client for AI conversations. It talks to cloud models when you allow it and to local models through Ollama when you don't — with conversation memory, a prompt library and zero telemetry by default.
I built it in six weeks as a tool I wanted to exist: one fast, native-feeling window for every model I use, where the data genuinely stays mine.
The problem
Why it existed
AI conversations are work products now — code reviews, drafting, thinking out loud — but they live inside browser tabs bound to cloud accounts, with opaque retention policies and no way to work offline.
Who experienced it
Developers and privacy-conscious professionals who use AI daily but don't want their work product logged in someone else's cloud by default.
Why current solutions weren't enough
Cloud chat apps are convenient but opaque; local-model frontends existed but were hobby-grade — clunky, ugly, and missing the small things (history, prompts, sync between machines) that make a tool daily-driver material.
The idea
Vision and constraints
Original vision
One client, two modes: cloud when you want reach, local when you want privacy — with the mode visible at all times, never silently downgraded.
The product contract was explicit: no telemetry, no account required, conversations stored locally in SQLite and exportable at any time. Privacy should be the default state, not a pricing tier.
Development process
01
Research
Two weeks of dogfooding my own AI usage — where the cloud-only workflow leaked data or broke offline.
02
Architecture
Chose Tauri over Electron for footprint; designed the provider trait so a new backend is one file.
03
Development
Streaming chat first, then providers, then memory and prompts — every week ended with a usable build.
04
Testing
Tested against flaky networks, huge histories and mid-stream model swaps; local models tested on modest hardware.
05
Launch
Released as an open download with signed installers and a public changelog.
My role
Exactly what I worked on
Solo project — I owned everything from product decisions to the packaging and auto-update pipeline.
Desktop App
Tauri shell with a React front end — small binary, native performance, tiny memory footprint.
Model Layer
A provider abstraction that treats Ollama, OpenAI and compatible endpoints as interchangeable sources with streaming everywhere.
Local AI
Ollama integration with model download status, context-window awareness and offline detection.
Data & Memory
SQLite storage, conversation threading, and a per-conversation memory summary that travels with the thread.
UX Polish
Keyboard-first navigation, markdown + code rendering, and a prompt library with variables.
Tech stack
The tools that shipped it
Key features
What makes it useful
Local & Cloud Models
Ollama for private, offline work; cloud APIs when a task needs reach. The active mode is always visible.
- Streaming responses from every provider
- Model switcher with context-window hints
- Offline mode that keeps working
Conversation Memory
Threads carry a compact running summary, so long conversations stay coherent without resend-everything token costs.
- Per-thread memory summaries
- Pinned facts the model always sees
- Full-text search across all history
Prompt Library
A personal set of reusable prompts with variables — turn repeated instructions into one-keystroke templates.
- Variables with fill-in prompts
- Import/export as plain files
- Per-model prompt variants
Design & UX
How it feels to use
The interface borrows from terminals and document editors: calm typography, code that renders like code, and no chat-app clutter.
Mode visibility
A persistent local/cloud indicator — privacy states should never be ambient or assumed.
Code-first rendering
Code blocks get syntax highlighting and one-tap copy; responses feel like engineering artifacts, not chat bubbles.
Zero chrome
One sidebar, one thread, one input. The app stays out of the conversation's way.
Technical challenges
And how they were solved
Provider fragmentation
Every provider streams differently. The provider abstraction normalises chunks into one event shape, so the UI never knows which backend is talking — adding a provider became a one-file change.
Long-conversation token costs
Resending full history doesn't scale on local models with small context windows. The memory-summary design keeps threads coherent within tight context budgets.
Offline reliability
Local model servers crash, ports conflict, downloads fail mid-way. Every failure state got a human-readable recovery path instead of a spinner of doom.
What I learned
Honest takeaways
The smallest project here taught the sharpest lessons about constraints and respect for user data.
- Local-first isn't harder — it's differently hard. Failure modes replace scaling problems.
- A visible privacy state builds more trust than a privacy policy.
- Provider abstractions are worth building on day one, not day forty.
- Six weeks of focused solo work can produce a real product if the scope is honest.
Results
Outcomes where available
AI Chat is a self-distributed tool without marketing metrics. The figures below are honest counts from the build and release process.
6 wks
Idea to release
Solo, part-time
<20MB
Installer size
Tauri shell vs ~100MB+ Electron equivalent
3
Model backends
Ollama, OpenAI, any OpenAI-compatible endpoint
0
Telemetry events
Nothing phones home, by design
Final takeaway
Building AI Chat taught me…
AI Chat taught me that privacy is a product feature you can design visibly — and that a tight six-week scope, held honestly, beats a loose six-month one.
Gallery
Conversation — streaming markdown and code
Model switcher — local and cloud, mode visible
Prompt library — variables and variants
Desktop — the zero-chrome window
Keep exploring



