A dual-system local AI server that pairs a small local model (System 2) with Jev’s fast classification API (System 1) — so 3B-class models reason like much larger ones, right on your own machine.
View on GitHubSystem 2 (a small local LLM via llama.cpp) does the reasoning; System 1 (Jev) handles classification. A 3B model punches far above its weight.
Text generation stays entirely on your machine. Only lightweight classification is sent to Jev — your output never leaves your computer.
Drop-in replacement for OpenAI-compatible endpoints. Plug into Cline or Kilo Code with a single base URL — no rewiring.
Create files, folders, and whole projects. DuoMind steers the model to emit correct tool-call responses that your assistant executes.
Automatic task workflows (webdev, coding, debug, git, webfetch) keep small models on track instead of improvising or refusing.
Runs on Windows 11, Linux, and macOS. The setup wizard fetches the right llama.cpp build for your platform automatically.
Measured on a 3B model (Qwen2.5-Coder-3B-Instruct), Jev ON vs Jev OFF — 100 requests, 100% success.
Install, run duomind setup, and connect your assistant.