The desktop GUI is gone — llama-server + browser UI
The biggest upgrade yet: CodAI now runs the official llama-server engine with a browser chat UI in a single portable Windows folder, and a controller that auto-restarts a crashed engine.
- Replaced the desktop GUI with the official llama.cpp llama-server engine — chat streams token by token locally.
- New browser chat UI with markdown rendering, code copy buttons, a stop button, and conversation memory.
- New controller (Codai.exe) with health monitoring, auto-restart, hardware-aware RAM/CPU tuning, rotating logs, and crash reports.
- Single portable folder: unzip, add a model, double-click run.bat.
- Ships with Codai.exe, engine\llama-server.exe, run.bat, and kill.bat.
- Windows 10/11 (64-bit), 4 GB RAM minimum — 2 GB works with the smallest models.
Highlights of CodAI v3.0.0
The biggest upgrade yet: the desktop GUI is gone. CodAI now runs the official llama-server engine with a browser chat UI in a single process, with a controller that auto-restarts a crashed engine.
New architecture
- Real llama.cpp engine: llama-server runs locally — chat streams token by token.
- Browser UI: markdown rendering, code copy buttons, stop button, conversation memory.
- Controller: health monitoring, auto-restart, hardware-aware tuning (RAM/CPU tiers), rotating logs and crash reports.
- Single portable folder: unzip, add a model, double-click run.bat.
In the box
Codai.exe(packaged controller) andengine\llama-server.exe(official llama.cpp build).run.batstarts everything, waits for readiness, and opens your browser.kill.batforce-cleans Codai processes and locks.
First run
Unzip anywhere writable (not C:\Program Files), drop one model into models\ — the default is
gemma-3-1b-it-Q4_K_M.gguf (0.81 GB) — then double-click run.bat and wait for [READY].
System requirements
Windows 10/11 (64-bit) · 4 GB RAM minimum (2 GB works with the smallest models) · ~1 GB free for the app + 0.5–5 GB for one model.