Task execution
Success, completion, sequence length, and intervention count.
LIBERO · CALVIN · BridgeA first-party field guide to choosing tests that expose capability, brittleness, and operating cost.
Success, completion, sequence length, and intervention count.
LIBERO · CALVIN · BridgeUnits per hour, mean time between failures, and recovery burden.
PhAIL-style field runsIntent, perception, planning, affordance, and failure analysis.
RoboBench dimensionsControl-loop deadlines and behavior under representative system load.
OpenNav-style missionsFresh objects, poses, scenes, instructions, and embodiments.
OXE · DROID · fresh setsHuman attention, energy, hardware utilization, and cost per useful action.
Operator baselinesBefore publishing a score, disclose the hardware, control rate, sample count, uncertainty, intervention policy, object novelty, data overlap, failure taxonomy, raw-run availability, and human baseline.
Suggested citation: Embodied Arena, “Robot Benchmark Selection Checklist,” public beta, 2026.