Harbor on Vercel Sandbox: parallel Terminal-Bench without your laptop
Harbor is the open-source eval harness behind Terminal-Bench. As of 17 September 2026, harbor run --env vercel places each trial in its own Firecracker microVM on Vercel Sandbox. That is the production change: you can parallelize agent benchmarks beyond your laptop without collapsing network policy and secrets into the same guest as the agent under test.
If you already measure coding agents with a fixed bench, this update changes where the trials run — not whether you still need a written acceptance gate. Pair it with a durable evaluation loop, not a one-off screenshot of pass rates.
This is an excerpt. Read the full post at otf-kit.dev/blog/harbor-evals-vercel-sandbox — full-stack kits your AI coding agent can actually ship to production. Browse the kits →
