Changelog

Every release, in the open. New features, performance numbers, and fixes as they land.

v3.0.0· latest

The desktop GUI is gone — llama-server + browser UI

The biggest upgrade yet: CodAI now runs the official llama-server engine with a browser chat UI in a single portable Windows folder, and a controller that auto-restarts a crashed engine.

  • Replaced the desktop GUI with the official llama.cpp llama-server engine — chat streams token by token locally.
  • New browser chat UI with markdown rendering, code copy buttons, a stop button, and conversation memory.
  • New controller (Codai.exe) with health monitoring, auto-restart, hardware-aware RAM/CPU tuning, rotating logs, and crash reports.
  • Single portable folder: unzip, add a model, double-click run.bat.
  • Ships with Codai.exe, engine\llama-server.exe, run.bat, and kill.bat.
  • Windows 10/11 (64-bit), 4 GB RAM minimum — 2 GB works with the smallest models.

Highlights of CodAI v3.0.0

The biggest upgrade yet: the desktop GUI is gone. CodAI now runs the official llama-server engine with a browser chat UI in a single process, with a controller that auto-restarts a crashed engine.

New architecture

  • Real llama.cpp engine: llama-server runs locally — chat streams token by token.
  • Browser UI: markdown rendering, code copy buttons, stop button, conversation memory.
  • Controller: health monitoring, auto-restart, hardware-aware tuning (RAM/CPU tiers), rotating logs and crash reports.
  • Single portable folder: unzip, add a model, double-click run.bat.

In the box

  • Codai.exe (packaged controller) and engine\llama-server.exe (official llama.cpp build).
  • run.bat starts everything, waits for readiness, and opens your browser.
  • kill.bat force-cleans Codai processes and locks.

First run

Unzip anywhere writable (not C:\Program Files), drop one model into models\ — the default is gemma-3-1b-it-Q4_K_M.gguf (0.81 GB) — then double-click run.bat and wait for [READY].

System requirements

Windows 10/11 (64-bit) · 4 GB RAM minimum (2 GB works with the smallest models) · ~1 GB free for the app + 0.5–5 GB for one model.

v2.1.0

Graphite Design System, Benchmark Profiles & Theme Parity

A ground-up visual overhaul to a flat graphite-and-amber system, pre-benchmarked GGUF model profiles, blog hero and filter improvements, and full light/dark contrast parity.

  • Rebuilt the entire interface on a flat graphite design system: 1px rules instead of floating cards, a single amber accent, whisper shadows, and no decorative effects.
  • Added pre-configured, benchmarked profiles for Qwen2.5-Coder-1.5B (~45 tok/s), Llama-3.2-1B (~55 tok/s), and DeepSeek-R1-Distill-1.5B (~38 tok/s).
  • Achieved full light/dark contrast parity across every surface, with tabular-nums metrics throughout.
  • Redesigned the blog index and post heroes with topic filters.
  • Zero telemetry and strict air-gap compliance verified across all supported operating systems.

Highlights of CodAI v2.1.0

CodAI v2.1.0 is a comprehensive overhaul of both the site and the local execution stack.

1. Graphite design system

Every surface was rebuilt on a flat, rule-based system: warm graphite paper in light and dark, one phosphor-amber accent used sparingly, Space Grotesk display headings, and spec tables with tabular numbers. No glass, no glows, no gradient text.

2. Local model runtime benchmarks

Pre-configured profiles for compact coding models, each tied to a measured RAM budget:

  • Qwen2.5-Coder-1.5B: ~45 tokens/sec on standard consumer laptops (~1.2 GB RAM).
  • Llama-3.2-1B-Instruct: ~55 tokens/sec with a sub-1 GB footprint.
  • DeepSeek-R1-Distill-1.5B: step-by-step logical reasoning at ~38 tokens/sec.
  • Qwen2.5-0.5B: ~75 tokens/sec for 2 GB-RAM machines.

3. Zero-telemetry, verified

The release was audited against the air-gap guarantees: no network calls during inference, no analytics in the app, no accounts anywhere in the flow.

v1.2.0

Interface Refresh, FAQ Rich Results & Changelog System

A bolder visual identity, FAQPage JSON-LD structured data across the site, and a dedicated markdown changelog system.

  • Redesigned core interface elements with a bolder, high-contrast visual identity.
  • Integrated FAQPage JSON-LD schema across the homepage and FAQ section for search rich results.
  • Introduced a dedicated changelog system with markdown release support.

What’s New in v1.2.0

1. Interface refresh

A distinct, high-contrast visual identity with tactile interactive states across the app and the site.

2. FAQ structured data

Every FAQ item now generates Google-compliant FAQPage JSON-LD schema, improving search visibility and rich-snippet eligibility.

3. Changelog system

Product updates, release milestones, and technical improvements are now tracked on /changelog — this page.

v1.1.0

Offline Model Engine

The foundational offline inference runtime: quantized GGUF model support for Qwen2.5 and Llama 3.2, zero-telemetry local execution, and the dark/light theme system.

  • Added a quantized GGUF model runner supporting Qwen2.5-0.5B and Llama-3.2-1B.
  • Implemented the zero-telemetry local inference runtime — no network calls during operation.
  • Added the dark/light theme system with persistent preference memory.

What’s New in v1.1.0

Version v1.1.0 established the foundational offline execution runtime for CodAI.

Highlights

  • Local inference engine: run 0.5B–1.5B parameter models locally on CPU/GPU with no cloud API dependencies.
  • Privacy guarantees: zero network telemetry — all user data stays on local hardware.

Want to see what’s coming next?