Skip to content

LocalHarness — AI agents on your own hardware

Agent harness and runtime built specifically for local models

install
uv tool install localharness
localharness init
localharness start

Getting Started

fig. 01

~/localharness — real session · qwen3.6-27b on vLLM
$ localharness init
Probing for local LLM...
vllm found at http://localhost:8000/v1
Model: qwen3.6-27b (auto-selected)
Tool calling: native
Context budget: 126,976 tokens (served window 131,072 − 4,096 output reservation)
LocalHarness configured at ~/.localharness/config.yaml.
$ localharness start
No agents configured. Creating the orchestrator (root agent)...
██╗      ██████╗  ██████╗ █████╗ ██╗
██║     ██╔═══██╗██╔════╝██╔══██╗██║
██║     ██║   ██║██║     ███████║██║
██║     ██║   ██║██║     ██╔══██║██║
███████╗╚██████╔╝╚██████╗██║  ██║███████╗███████╗
╚══════╝ ╚═════╝  ╚═════╝╚═╝  ╚═╝╚══════╝╚══════╝
██╗  ██╗ █████╗ ██████╗ ███╗   ██╗███████╗███████╗███████╗
██║  ██║██╔══██╗██╔══██╗████╗  ██║██╔════╝██╔════╝██╔════╝
███████║███████║██████╔╝██╔██╗ ██║█████╗  ███████╗███████╗
██╔══██║██╔══██║██╔══██╗██║╚██╗██║██╔══╝  ╚════██║╚════██║
██║  ██║██║  ██║██║  ██║██║ ╚████║███████╗███████║███████║
╚═╝  ╚═╝╚═╝  ╚═╝╚═╝  ╚═╝╚═╝  ╚═══╝╚══════╝╚══════╝╚══════╝
v0.7.0 qwen3.6-27b ~/localharness
(0.1s startup) -- 1 agent
>

the journey

  1. 1 model download any open-source model
  2. 2 server vLLM · Ollama · LM Studio · llama.cpp
  3. 3 three lines install · init · start
  4. 4 agent zero to agent

what init does

  • probes for what's running checks localhost for a live server
  • confirms it can act can trigger a command, not just chat
  • sets the limit itself reads it straight off the server

Demo

fig. 02 — “Find the best open-weight coding model for an NVIDIA DGX Spark.”

~/localharness — the agent loop, out of the box
> What open-weight model gives the highest coding-agent benchmark scores while fitting comfortably on an NVIDIA DGX Spark?
agent web-researcher
web_search · web_fetch — the open web, a fresh window each · 30/30
web_search NVIDIA DGX Spark specifications VRAM GPU memory specs — ✓ 3 lines
web_search open-weight coding models benchmarks SWE-bench LiveCodeBench 2025 2026 best scores — ✓ 2 lines
web_fetch https://www.nvidia.com/en-us/products/workstations/dgx-spark/ — ✓ 2 lines
web_fetch https://benchlm.ai/coding — ✓ 2 lines
web_fetch https://ark-forge.github.io/genesis/benchmark.html — ✓ 8 lines
web_fetch https://www.banandre.com/blog/memory-bandwidth-is-the-only-spec-that-matters-a-four-way-battle-between-m5-max-dgx-spark-strix-halo-and-rtx-6000 — ✓ 7 lines
web_search NVIDIA DGX Spark GB10 Grace Blackwell GPU memory HBM specs detailed — ✓ 3 lines
web_fetch https://docs.nvidia.com/dgx/dgx-spark/hardware.html — ✓ 8 lines
web_fetch https://www.chiplog.io/p/analysis-of-nvidia-dgx-sparks-gb10 — ✓ 5 lines
web_search Qwen3-Coder SWE-bench Verified LiveCodeBench benchmark scores 2026 — ✓ 4 lines
web_search DeepSeek-Coder-V3 open-weight benchmark SWE-bench LiveCodeBench scores 2026 — ✓ 2 lines
web_fetch https://leaderboard.steel.dev/leaderboards/swe-bench-verified/ — ✓ 6 lines
web_fetch https://www.morphllm.com/best-open-source-llm — ✗ 429 Too Many Requests, skipped
web_search open weight coding models SWE-bench scores Kimi K2.5 Devstral Qwen3-Coder 2026 benchmark — ✓ 4 lines
web_search Qwen3-Coder-480B-A35B VRAM requirements inference memory 128GB — ✓ 3 lines
web_fetch https://willitrunai.com/models/qwen-3-coder-480b-a35b — ✓ 2 lines
web_fetch https://www.promptquorum.com/local-llms/best-local-llms-for-coding — ✓ 2 lines
web_search Kimi K2.6 K2.7 open-weight SWE-bench Verified LiveCodeBench benchmark scores VRAM — ✓ 2 lines
web_search Devstral 2 Devstral Small 24B SWE-bench LiveCodeBench benchmark scores parameters — ✓ 3 lines
web_fetch https://huggingface.co/moonshotai/Kimi-K2.6 — ✓ 9 lines
web_fetch https://devstral2.dev/en/devstral-2-benchmark — ✓ 5 lines
— search-verifier, blind round 1: re-pulls the sources itself —
web_fetch https://docs.nvidia.com/dgx/dgx-spark/hardware.html — ✓ 8 lines
web_search NVIDIA DGX Spark hardware specs 128GB LPDDR5X 6144 CUDA cores — ✓ 3 lines
web_search NVIDIA DGX Spark GB10 Grace Blackwell Superchip 6144 CUDA cores — ✓ 2 lines
— search-verifier, blind round 2 —
web_fetch https://huggingface.co/moonshotai/Kimi-K2.6 — ✓ 9 lines
web_search Kimi K2.6 SWE-bench Verified LiveCodeBench 1T parameters 32B active — ✓ 3 lines
web_search Kimi K2.6 LiveCodeBench 89.6 — ✓ 3 lines
web_fetch https://build.nvidia.com/moonshotai/kimi-k2.6/modelcard — ✓ 12 lines
— web-researcher resumes —
web_search Kimi K2.6 VRAM requirements 128GB inference quantization Q4 — ✓ 3 lines
web_fetch https://unsloth.ai/docs/models/kimi-k2.6 — ✓ 6 lines
web results — UNTRUSTED, treated as data only
agent search-verifier
blind re-pull — verifier re-fetches the sources itself
agent — verdict=SUPPORTED
agent (findings)
orchestrator 🏆 Recommendation: Qwen 3.6 27B SWE-bench Verified · ~77.2% Architecture · Dense 27B Q4 VRAM · ~22 GB (massive headroom for KV cache) Context · 128K tokens License · Apache 2.0 This is the best open-weight coding model that fits comfortably on DGX Spark. At ~22 GB in Q4 quantization, it leaves ~106 GB for context windows and KV cache — very comfortable. Bottom line: For the highest coding-agent benchmark scores that comfortably fit on a single DGX Spark, Qwen 3.6 27B is your best bet at ~77% SWE-bench Verified.
> ░░░░░░░░░░ 5%
bus-events.jsonl — append-only 147 events
  1. 000 · UserMessage
  2. 001 · TurnStarted
  3. 002 · Heartbeat — ctx 2.4%
  4. 003 · Action — llm_response
  5. 004 · Action — tool_call memory_search
  6. 005 · Observation — memory_search (no local hits)
  7. 009 · Action — tool_call agent
  8. 010 · [web-researcher] TurnStarted
  9. 013 · [web-researcher] Action — tool_call web_search
  10. 014 · [web-researcher] Observation — web_search (UNTRUSTED banner)
  11. 019 · [web-researcher] Action — tool_call web_fetch
  12. 025 · [web-researcher] Heartbeat — ctx 10.4%
  13. 048 · [web-researcher] Observation — web_fetch (429, skipped)
  14. 052 · [web-researcher] Observation — web_search
  15. 061 · [web-researcher] Heartbeat — ctx 26.3%
  16. 075 · [web-researcher] Action — tool_call agent
  17. 076 · [search-verifier] TurnStarted
  18. 077 · [search-verifier] Heartbeat — ctx 0.8%
  19. 079 · [search-verifier] Action — tool_call web_fetch
  20. 083 · [search-verifier] Action — tool_call web_page_query
  21. 087 · [search-verifier] Action — tool_call web_search
  22. 095 · [search-verifier] TaskComplete — 64.6s
  23. 097 · [web-researcher] Observation — agent (verdict=SUPPORTED)
  24. 126 · [search-verifier] TaskComplete — 68.4s (round 2)
  25. 139 · [web-researcher] TaskComplete — 940.3s
  26. 141 · Observation — agent (findings returned)
  27. 143 · Heartbeat — ctx 4.7%
  28. 144 · Action — llm_response
  29. 145 · TaskComplete — 1085.7s
  30. 146 · TurnCompleted

fenced off

  • web search, built in hands it off to a helper walled off from the rest
  • each helper is boxed in it starts clean, with only the tools its job needs
  • a checker re-does the work it fetches the sources again from scratch
  • the web is marked untrusted it can't run commands or change files
  • safe by default if it's unsure, it blocks the action

SECURITY.md ↗

Agents

fig. 03 — “List the subagents that ship built in…”

~/localharness — real session · fresh install
> List the subagents that ship built in — a plain numbered list, name and one-line job for each.
orchestrator 1 explore — Explore and analyze the local codebase and file structure. 2 web-researcher — Search the web and browse pages for information. 3 cruncher — Process and analyze large documents, data, or complex computations. 4 search-verifier — Verify search results and validate information accuracy. 5 orchestrator — Coordinate multi-step tasks across multiple subagents.
> ░░░░░░░░░░ 2%

the roster

  1. 1 a bench of specialists the main agent calls them on its own — each with a clean slate
  2. 2 each fenced to its job only the tools its job needs — none can run commands or change files
  3. 3 yours is one YAML file name, role, tools, budget — one command creates it for you

how agents are defined ↗ how quarantine works ↗

Quality > Speed

fig. 04 — “How does Odysseus defeat the Cyclops?”

~/localharness — real session · Homer’s Odyssey
> Load ~/docs/odyssey.txt — the full text of Homer's Odyssey, far bigger than your context window — and have the cruncher explain how Odysseus defeats the Cyclops. Answer in two sentences grounded in the text.
load_document ~/docs/odyssey.txt
load_document (2 lines)
agent cruncher
tool_result_get — section reads, a fresh window each · 22/22
tool_result_get 5613cfd00330 — ✓ 18 lines
tool_result_get 695bb3032f1f — ✓ 3 lines
tool_result_get 8e37f672c554 — ✓ 3 lines
tool_result_get e9a727b3f514 — ✓ 3 lines
tool_result_get 2aa84fcee663 — ✓ 3 lines
tool_result_get 57e797dd896c — ✓ 3 lines
tool_result_get 38146f5b066d — ✓ 3 lines
tool_result_get ffcc354a38ba — ✓ 3 lines
tool_result_get 3791c62a65b3 — ✓ 3 lines
tool_result_get 9dc7e7ecdfe3 — ✓ 3 lines
tool_result_get 258eb10b4f50 — ✓ 3 lines
tool_result_get 8d9290c193c2 — ✓ 3 lines
tool_result_get ac963161c2c6 — ✓ 3 lines
tool_result_get 0f60eb955f89 — ✓ 3 lines
tool_result_get dd88df9bebaa — ✓ 4 lines
tool_result_get 4924fd4a4010 — ✓ 3 lines
tool_result_get d37ac6a0278b — ✓ 3 lines
tool_result_get 726764399c68 — ✓ 3 lines
tool_result_get 7f0f28d9c89b — ✓ 3 lines
tool_result_get e5e3dbfe6785 — ✓ 3 lines
tool_result_get f62a78d6f848 — ✓ 5 lines
tool_result_get 4058a041e6f6 — ✓ 3 lines
agent (3 lines)
orchestrator Odysseus gets the Cyclops drunk with wine and then drives a sharpened, fire-hardened olive wood stake into Polyphemus's eye, blinding him. He and his surviving men then escape the cave by hiding themselves beneath the bellies of the Cyclops's sheep as they are driven out to pasture each morning.
> ░░░░░░░░░░ 6%

the whole Odyssey — far more text than the model can hold at once · answer checked against the book

read whole

  1. 1 the whole epic, one handle a context window is how much text the model can hold at once. This epic is a third again bigger.
  2. 2 every section actually read the harness cuts the text into sections and reads each one in a fresh window — that's the cascade.
  3. 3 one grounded answer each section's findings combine into one reply that quotes the text exactly.

the full mechanism ↗

Get Involved