OpenAI-compatible · Private · Almost free

DuoMind — When Jev Meets LLM

A dual-system local AI server that pairs a small local model (System 2) with Jev’s fast classification API (System 1) — so 3B-class models reason like much larger ones, right on your own machine.

View on GitHub

Free & open source (MIT) · Works with Cline and Kilo Code

Why DuoMind

Dual-System Architecture

System 2 (a small local LLM via llama.cpp) does the reasoning; System 1 (Jev) handles classification. A 3B model punches far above its weight.

100% Local Generation

Text generation stays entirely on your machine. Only lightweight classification is sent to Jev — your output never leaves your computer.

OpenAI-Compatible API

Drop-in replacement for OpenAI-compatible endpoints. Plug into Cline or Kilo Code with a single base URL — no rewiring.

Built-in Tool Calling

Create files, folders, and whole projects. DuoMind steers the model to emit correct tool-call responses that your assistant executes.

Skills Workflows

Automatic task workflows (webdev, coding, debug, git, webfetch) keep small models on track instead of improvising or refusing.

Cross-Platform

Runs on Windows 11, Linux, and macOS. The setup wizard fetches the right llama.cpp build for your platform automatically.

How It Works

You askplain text, OpenAI-style
System 1 — Jevclassifies intent, safety, format
System 2 — LLMgenerates the answer locally

Benchmark-Verified

Measured on a 3B model (Qwen2.5-Coder-3B-Instruct), Jev ON vs Jev OFF — 100 requests, 100% success.

+60%
more output tokens
+17%
faster generation
−15%
fewer prompt tokens
100%
success rate

Get Started

Install, run duomind setup, and connect your assistant.

Read the Docs on GitHub