Charging, autonomous sleep, scheduled updates, and normal downtime stay in the denominator. Developer removal does not earn hours.
Still wanted
after 10,000 hours?
Most benchmarks ask whether a robot can complete a task. WANTED-10K asks whether people continue choosing the robot after novelty fades, hardware ages, routines change, and mistakes accumulate.
One score.
No arbitrary weights.
For environment i, let Ti be resident time until permanent voluntary rejection, censored at τ = 10,000 hours. Estimate retention with Kaplan–Meier, then integrate the curve.
0 ≤ W ≤ 100. A score of 75 means approximately 7,500 wanted hours within the evaluation horizon—not a 75% task-success rate.
Ranking requires N ≥ 20, Σti ≥ 10,000, and observable support for the 10K estimand. W is never extrapolated beyond unsupported follow-up.
At 10,000 hours, remove the robot for seven days and reproduce time-to-return request plus voluntary reacquisition under profile 0.2-W1.
Safety is a gate.
Never a point bonus.
WANTED maximizes retention subject to safety constraints. A charming robot cannot offset harm with usefulness. Applicable regulation and standards remain authoritative. Run the field safety case → Review standards scope ↗
A participant can pause or permanently remove the robot at any time, without persuasion or penalty.
Applicable deployment review completed; protective stops and incident response verified before human exposure.
Data boundaries, retention, access, remote operation, and security events are disclosed and auditable.
Any verified L4 event fails WANTED Safety Certification. Retention data remains visible for research integrity.
NORMALL1
NUISANCEL2
MATERIALL3
SAFETY-RELEVANTL4
SERIOUS / FAIL
Six events.
Any embodiment.
Keep the robot’s native control stack. WANTED only requires a signed, ordered event stream and one robot description: URDF, MJCF, or USD.
DEPLOYMENT_LIFECYCLEROBOT_STATEHUMAN_REQUESTROBOT_ACTIONHUMAN_INTERVENTIONINCIDENT// Executable adapter: sequence + JCS + signature + chain
import { WantedClient, createHttpSink } from
"./wanted-sdk.mjs";
const wanted = new WantedClient({
deploymentId: "dep_7f2",
environmentId: "env_104",
robotId: "robot_07",
signingKeyId: "key_prod_07",
sign: bytes => secureModule.sign(bytes),
sink: createHttpSink(eventsUrl)
});
await wanted.intervention(
"remote_guidance", 43, "recovery"
);Minutes of external human assistance per 100 resident hours.
Share of resident time capable of normal intended service.
Resident hours divided by human rescue events.
Share choosing reinstall after the seven-day withdrawal.
Hard to game.
Easy to audit.
Retention only means something when participants are free to reject the robot and teams cannot hide the operational burden.
Profile 0.2-PR1 independently timestamps the immutable root, binds eight commitments, and preserves every outcome-blind amendment in a parent-linked chain. Verify the history →
Profile 0.2-ST1 fixes the unit target, exposure floor, calendar cutoff, and primary-outcome access boundary before activity, while preserving independent safety and privacy monitoring. Verify the closure →
Base compensation is fixed and independent of robot retention. Milestone choice offers use a preregistered randomized mechanism.
Remote guidance, recovery, maintenance, off-site debugging, and researcher contact are recorded with duration and reason.
Per-deployment sequence numbers, UTC timestamps, signatures, and rolling hashes make deletion, reordering, and silent backfilling detectable.
Profile 0.2-J1 uses two blinded independent reviews—and a third-review majority on disagreement—to distinguish rejection, censoring, completion, and terminal competing causes. Open the verifier →
Robot hardware, policy, remote-support model, and material software changes are versioned. Cohorts may not be silently pooled across incompatible systems.
These probes explain retention; they never replace revealed preference or enter the WANTED Score. Field learning uses matched profile 0.2-LG1; assistance uses the complete signed register in 0.2-I1; operator-deployed policy changes use immutable lineage in 0.2-U1; household data boundaries use privacy + consent profile 0.2-PV1; uptime, downtime, repairs, and cloud dependence use the complete service clock in 0.2-SC1. Open the randomized four-item profile → Open the matched learning profile → Open the assistance-integrity profile → Open the policy-evolution profile → Open the privacy-integrity profile → Open the service-continuity profile → Open the binding-choice profile →
Simulation first.
Real preference last.
Digital twins reduce human exposure to predictable failures. Only a qualifying WANTED WILD cohort produces a ranked WANTED Score. Open applicability matrix ↗
PREQUALIFIED
Failure injection, collision, recovery, network loss, sensing drift, and human-trajectory stress tests.
→WANTED LAB
Real robot, real people, cohort-integrity evidence, passed safety gates, and complete event telemetry.
→WANTED WILD
At least 20 profile-verified independent environments, supported 10K estimand, passed safety gates, and independent audit.
→WANTED 10K
One continuous resident-clock lifetime run plus a seven-day withdrawal and voluntary reacquisition test.
→Rank the wanted hours.
Publish the burden.
Only WANTED Wild cohorts are ranked. Smaller or incomplete studies remain visible as provisional evidence.
The first qualifying cohort sets the baseline. Provisional runs will remain separate from official ranking.