Vercel Sandbox Adds Support for Harbor Benchmarking Framework
Vercel Sandbox now supports running Harbor evaluations, including Terminal-Bench, by utilizing isolated Firecracker microVMs. This integration requires Harbor version 0.22.0 or later and enables parallelized benchmarking across multiple AI models via AI Gateway.
Verified State Diff
Impact & Verification Analysis
AI engineers and developers using Harbor for benchmarking LLMs and agentic workflows.
This feature removes the hardware bottleneck for running large-scale AI benchmarks, allowing developers to scale evaluations significantly while maintaining security through isolated microVMs and firewall-level credential injection.
Full Fact Overview
Vercel has integrated support for the Harbor benchmarking framework within Vercel Sandbox. Users can now execute Harbor commands using the --env vercel flag, which triggers the execution of trials within isolated Firecracker microVMs. This architecture allows for massive parallelization of benchmarks like Terminal-Bench, SWE-bench, tau3-bench, and OSWorld. The integration includes firewall-level network policy enforcement and secure credential injection for outbound requests. Additionally, users can leverage AI Gateway to switch between hundreds of models by modifying the --model parameter, such as using vercel_ai_gateway/openai/gpt-5.6-luna.