Skip to content
localcode

Architecture

localcode is a small Python launcher. It starts three things, all on 127.0.0.1:

Launcher → Supervisor + llama-server → Runtime + plugin
↓
Tools (read / edit / bash / grep / websearch / MCP)
↓
http://127.0.0.1:PORT/v1
llama-server, llama.cpp fork + TurboQuant KV compression
↓
GGUF weights on disk
  1. Launcher - localcode picks free ports, starts the supervisor, writes a per-session config in the run directory, and runs the interface binary in your project. It writes nothing into the project. See Configuration.
  2. Supervisor - a Python process that owns llama-server. The model server listens on 127.0.0.1:8123 or the next free port. The supervisor serves the model picker and status API on a control port from 8323. It downloads a quant when you pick one, reloads the server on the same port when you switch models, and shuts the server down when you exit.
  3. Runtime + plugin - the interface, localcode-ui, is a fork of opencode branded localcode. Every cloud feature is removed. It talks only to http://127.0.0.1:PORT/v1. The config loads localcode’s discipline plugin from the package.
  4. Tools - file reading and editing, glob and grep, shell commands, todo list, web search and fetch, language servers you install through /lsp, and any MCP servers you configure.

One model server runs per user. Opening localcode in another terminal starts a separate interface session attached to that server. Close either window without ending the other session; use /models to switch the shared model.

Small local models finish a task when the loop makes them. The plugin adds localcode’s completion discipline to the runtime:

  • Orient from a snapshot - the first message of a session carries a snapshot of the project’s layout, and later messages report files that changed outside the session, so the model does not spend steps listing directories.
  • Plan first - for work of three or more steps the model lays out the steps in the todo list before editing, and marks each one done as its last edit lands.
  • Keep going - a turn does not end while todo items are still open.
  • Prove it - after edits, the plugin runs the project’s own typecheck and tests and feeds failures back to the model.
  • Audit stubs - placeholder code and TODO bodies are flagged before the task counts as done.
  • No foreground servers - a command that would block the session, such as a dev server, is refused with a hint to run it in the background.
  • Stop cleanly - when the model only repeats itself across rounds, the plugin asks it to wrap up; if that does not help it refuses further tool calls with the reason, so the model reports what it has and the turn ends. A task that keeps covering new ground is never stopped.

localcode is designed to enable agentic coding with local models on consumer hardware. The prompts, the loop, and the model server are tuned for small quantised models rather than a frontier model:

  • Long context on 16 GB - the llama.cpp fork compresses the KV cache with TurboQuant (about 3.8x smaller than f16), so long contexts fit on small machines. See Unified Memory.
  • Fast multi-turn - the server keeps the prompt prefix between turns, so the next turn does not re-read it.
  • Compact tool definitions - the runtime shortens built-in tool descriptions and large MCP schemas before sending them to the model. This leaves more room for the task and project context, especially on a 16 GB Mac.
  • Hidden reasoning is off by default - the server starts with --reasoning off. Small models finish faster and follow tool calls better without it.
  • Vision on demand - the image projector for a vision-capable model is a separate download through /vision. Nothing is fetched until you ask.

The launcher, supervisor, model server, and interface all bind to 127.0.0.1. There is no telemetry, no update check, and no account. The network is used only for downloads you request in the picker, /vision, /lsp, or voice, and for the web tools and MCP servers when the model calls them. See Network Boundary.