VeBench evaluation architecture
Versioned tasks · controlled scheduling · isolated execution
01Tasks & images
Original tasks
Reuse directly
- Instructions
- instruction.md
- Environment
- task.toml / environment
- Verification
- tests
One task contract; existing tasks stay reusable.
Build or synchronize
Validate the exact image
Task identity + digest → image + task configuration
Task = definition. Image = environment.
02Jobs & scheduling
Select case · agent · model
Resolve task → lock versions
Trial queue · concurrency · retry
Execution configuration
Resolve the exact image and task version
E2B: connect · execute commands · transfer files
Scheduling stays separate from sandbox execution.
03veFaaS execution
Prewarmed image + resources + mounts + gateway
One trial · one Agent Sandbox
Read-only mount from TOS
Upgrade the harness without rebuilding task images.
Task environment + agent runtime
Read task → call model → run tools
Provider access + model requests
Collect raw execution traces
Shared loopback · separate filesystems
Ark / other providers
04Scoring & results
Run the task tests
Separate scoring sandbox
Reward · logs · outputs · traces
Reports · leaderboard · trace analysis
Scoring images are prepared and prewarmed too.
Adaptive concurrency · optional control loop
Model capacity, usage and rate-limit feedback
Estimate headroom and workload demand
Adjust admission for new trials
Running trials continue; control changes apply to new admissions.