VeBench / Trial lifecycle

How one task completes an evaluation

VeBench orchestrates · veFaaS provisions · E2B connects, executes and transfers

01Load the task

Harbor Task
instruction.md
What the agent should accomplish
task.toml / environment
The execution environment
tests
How the result is verified

Lock the task version → create one Trial.

A task is a content contract, not a container image.

02Create & connect

TOS mapping → prewarmed image

Pinned task identity → exact image

Prepared image + mounts + gateway
VeBench Environment
veFaaS provisioning
Created sandbox

E2B connect(sandbox_id)

Provision with veFaaS first, then connect through E2B.

03Run the agent

Agent Sandbox · veFaaS

Agent bundle

Read-only runtime mount

Agent runtime

Task environment

localhost
Gateway

Model access

Read task → invoke model → execute tools

E2B uploads inputs and starts the agent process.

04Hand off artifacts

Collect the outputs declared by the task and hand them to the verifier.

Agent outputs

Answer files + task artifacts

download
Declared artifacts

Only the task-required outputs

upload
Scoring inputs

Ready for verification

An explicit artifact contract connects execution and scoring.

05Score independently

Independent scoring sandbox

Artifacts

From step 04

Task tests

From the pinned task

Run the verifier

Execute task tests → reward and diagnostics

The grader uses a separate, prewarmed environment.

06Collect & clean up

Download the result

Reward · logs · declared outputs

Collect gateway traces

Associate model calls with the trial

Save and clean up

Persist evidence and release sandbox resources

Results → reports / comparisons / trace analysis

One trial = agent execution + independent scoring + result collection
TaskInstructions, environment and testsE2BConnection, commands and filesveFaaSSandbox provisioning and cleanup

Drag to pan · Pinch or Ctrl / ⌘ + scroll to zoom · Arrow keys to move