BENCHMARK ATLAS

One robot.
Many ways to fail.

A first-party field guide to choosing tests that expose capability, brittleness, and operating cost.

01

Task execution

Success, completion, sequence length, and intervention count.

LIBERO · CALVIN · Bridge
02

Real operations

Units per hour, mean time between failures, and recovery burden.

PhAIL-style field runs
03

Embodied brain

Intent, perception, planning, affordance, and failure analysis.

RoboBench dimensions
04

Real-time behavior

Control-loop deadlines and behavior under representative system load.

OpenNav-style missions
05

Generalization

Fresh objects, poses, scenes, instructions, and embodiments.

OXE · DROID · fresh sets
06

Deployment economics

Human attention, energy, hardware utilization, and cost per useful action.

Operator baselines
CITATION-READY RESOURCE

The Robot Benchmark
Selection Checklist

Before publishing a score, disclose the hardware, control rate, sample count, uncertainty, intervention policy, object novelty, data overlap, failure taxonomy, raw-run availability, and human baseline.

01Hardware + sensors02Task manifest03Episode count04Confidence interval05Intervention rule06Fresh test set07Raw failure runs08Human baseline

Suggested citation: Embodied Arena, “Robot Benchmark Selection Checklist,” public beta, 2026.