VeBench

Agent Infra · ByteDance · Active development

VeBench is an agent evaluation framework built on Volcengine infrastructure. It brings task preparation, cloud execution, scoring, failure recovery and result analysis into one workflow, so developers can spend more time improving agents and less time keeping benchmarks running.

I lead its evaluation engineering work in the Volcengine Ark · Doubao LLM team. My architectural work connects task contracts, versioned environments, cloud execution, independent verification and trace analysis into an evaluation system that can be operated and improved as a whole.

VeBench in action

An introduction to the evaluation workflow, from starting an experiment to cloud execution and result analysis.

VeBench overview · 1:41 · In Chinese · Open video ↗

Removing infrastructure noise from evaluation

Running an agent benchmark involves much more than calling a model. Tasks need the right datasets, dependencies, images and runtimes. Local machines limit parallelism, while slow downloads and unreliable external services can introduce failures unrelated to the agent’s capabilities.

VeBench moves each case into an isolated cloud sandbox and prepares its assets ahead of execution. Images, benchmark data, runtime bundles and evaluation artifacts are managed centrally, making it easier to repeat an experiment under consistent conditions.

Architecture & executionOpen full view ↗
Explore the task preparation, scheduling, sandbox and scoring boundaries. The lower control loop adjusts admission for new trials.

Architecture decisions and their trade-offs

A task contract that survives runtime changes

A Harbor Task defines instructions, environment requirements and tests. The manifest binds that task version to prepared assets; the agent bundle is delivered separately. This makes a harness update possible without rebuilding every task image, while keeping the experiment’s task and environment identity explicit. It shifts setup work into a deliberate asset-publication and prewarming stage.

A control plane with replaceable execution infrastructure

The CLI and scheduler own task resolution, trial admission, concurrency and retries. The environment adapter connects these decisions to sandbox provisioning and runtime operations. The boundary keeps benchmark logic out of infrastructure-specific code and lets orchestration remain accessible to both people and agents.

Isolation at the execution, model-access and scoring boundaries

The agent runtime and model gateway have distinct responsibilities. The runtime executes the task; the gateway mediates model access and records traces. In separate-scoring mode, only the task’s declared outputs are handed to an independent verifier. These boundaries clarify what each component may observe or change. Separate scoring adds an environment and an artifact handoff, in exchange for a clearer verification boundary.

Capacity control that protects experiments already in flight

The adaptive control loop observes usage and rate-limit feedback, estimates available headroom, and adjusts permits for new trials. Keeping running trials intact avoids turning a capacity adjustment into an uncontrolled change to the experiment itself.

Evidence that connects a score to its cause

Reward, task artifacts, harness logs and model-call traces are collected as related outputs of a trial. This makes it possible to investigate whether a failure came from the model, the harness, the verifier or the infrastructure, and use that evidence in regression checks and release decisions. Keeping raw traces alongside normalized results preserves the detail needed for that investigation.

What happens during one trial

VeBench separates sandbox provisioning from runtime control. veFaaS creates the sandbox; the E2B integration handles connection, command execution and file transfer. A trial then moves through a defined lifecycle, from a fixed task version to a recorded result and resource cleanup.

One trial, end to endOpen full view ↗
The full lifecycle: resolve a task, provision and connect, run the agent, hand off artifacts, score independently, then collect results and release resources.

Execution produces more than a final reward. Logs, declared artifacts and model-call traces are collected alongside the score, so a failure can be investigated through the evidence behind it.

Built for repeatable experiments

CLI, TUI and Evaluation Job interfaces give users consistent control over starting runs, observing progress, stopping tasks, recovering failures and collecting results. A command-oriented workflow also makes evaluation easier to incorporate into agent-driven iteration and ablation studies.

The task adapters cover public benchmarks including SWE-bench Verified, Terminal-Bench, WideSearch and SkillsBench. Multiple harness adapters allow model and harness changes to be compared through the same execution workflow.

Concurrency control accounts for model capacity and rate-limit feedback. The aim is to use available resources effectively while keeping running trials intact.

My engineering contribution

  • Evaluation framework and environment delivery: integrate benchmark tasks, manage versioned assets, and automate image preparation, prewarming and validation.
  • Execution and capacity control: connect task orchestration with cloud sandboxes, independent runtimes, retry handling and resource management.
  • Observability: build a Rust model gateway and unified trace pipeline that connect model requests, tool calls and harness logs.
  • Experiment analysis: distinguish model failures, harness defects, scoring issues and infrastructure faults, and use that evidence for regression checks and release decisions.

Alongside VeBench, I contribute to Agent Harness development and optimization. The evaluation system provides a way to assess changes to context handling, tools and execution behavior against repeatable tasks.

VeBench is under active development and currently used through a managed evaluation workflow.