How one task completes an evaluation
VeBench orchestrates · veFaaS provisions · E2B connects, executes and transfers
01Load the task
- instruction.md
- What the agent should accomplish
- task.toml / environment
- The execution environment
- tests
- How the result is verified
Lock the task version → create one Trial.
A task is a content contract, not a container image.
02Create & connect
Pinned task identity → exact image
E2B connect(sandbox_id)
Provision with veFaaS first, then connect through E2B.
03Run the agent
Agent Sandbox · veFaaS
Read-only runtime mount
Task environment
Model access
Read task → invoke model → execute tools
E2B uploads inputs and starts the agent process.
04Hand off artifacts
Collect the outputs declared by the task and hand them to the verifier.
Answer files + task artifacts
Only the task-required outputs
Ready for verification
An explicit artifact contract connects execution and scoring.
05Score independently
Independent scoring sandbox
From step 04
From the pinned task
Execute task tests → reward and diagnostics
The grader uses a separate, prewarmed environment.
06Collect & clean up
Reward · logs · declared outputs
Associate model calls with the trial
Persist evidence and release sandbox resources
Results → reports / comparisons / trace analysis