Skip to main content

Command Palette

Search for a command to run...

Harbor on Vercel Sandbox: parallel Terminal-Bench without your laptop

Updated
1 min readView as Markdown
Harbor on Vercel Sandbox: parallel Terminal-Bench without your laptop

Harbor is the open-source eval harness behind Terminal-Bench. As of 17 September 2026, harbor run --env vercel places each trial in its own Firecracker microVM on Vercel Sandbox. That is the production change: you can parallelize agent benchmarks beyond your laptop without collapsing network policy and secrets into the same guest as the agent under test.

If you already measure coding agents with a fixed bench, this update changes where the trials run — not whether you still need a written acceptance gate. Pair it with a durable evaluation loop, not a one-off screenshot of pass rates.


This is an excerpt. Read the full post at otf-kit.dev/blog/harbor-evals-vercel-sandbox — full-stack kits your AI coding agent can actually ship to production. Browse the kits →

More from this blog

O

OTF — kits your AI coding agent can ship to production

609 posts

Engineering notes on shipping production apps with AI coding tools — Claude Code, Cursor, Codex, Lovable, Bolt — and the stack underneath: React Native, Expo, Next.js, Supabase. Honest takes on what works, what breaks, and the full-stack kits that get you to production faster. By OTF.