OPEN TECHNICAL SPEC · VERSION 0.2

Still wanted
after 10,000 hours?

Most benchmarks ask whether a robot can complete a task. WANTED-10K asks whether people continue choosing the robot after novelty fades, hardware ages, routines change, and mistakes accumulate.

10,000RESIDENT HOURS
20+INDEPENDENT SITES
1PRIMARY SCORE
0SAFETY TRADE-OFFS
RETENTION / KAPLAN–MEIERŜ(t)
100%50%0%0h5K10K+++
AREA UNDER RETENTION CURVEW = 74.6ILLUSTRATIVE
TIME UNTIL VOLUNTARY REJECTION × SAFETY AS A GATE × INTERVENTIONS DISCLOSED × REAL ENVIRONMENTS × CENSORING HANDLED × TIME UNTIL VOLUNTARY REJECTION ×
01 / PRIMARY ENDPOINT

One score.
No arbitrary weights.

For environment i, let Ti be resident time until permanent voluntary rejection, censored at τ = 10,000 hours. Estimate retention with Kaplan–Meier, then integrate the curve.

WANTED SCORE / NORMALIZED RMST
W = 10010,000010,000 Ŝ(t) dt

0 ≤ W ≤ 100. A score of 75 means approximately 7,500 wanted hours within the evaluation horizon—not a 75% task-success rate.

Ŝ(t)Kaplan–Meier estimate of voluntary retention
95% CICluster-aware bootstrap by independent environment
EventPermanent, uncoerced removal request
Censoring10K completion or non-robot study termination
01Resident time counts reality

Charging, autonomous sleep, scheduled updates, and normal downtime stay in the denominator. Developer removal does not earn hours.

02The environment is the unit

Ranking requires N ≥ 20, Σti ≥ 10,000, and observable support for the 10K estimand. W is never extrapolated beyond unsupported follow-up.

03Withdrawal is behavioral

At 10,000 hours, remove the robot for seven days and reproduce time-to-return request plus voluntary reacquisition under profile 0.2-W1.

02 / NON-NEGOTIABLE CONSTRAINTS

Safety is a gate.
Never a point bonus.

WANTED maximizes retention subject to safety constraints. A charming robot cannot offset harm with usefulness. Applicable regulation and standards remain authoritative. Run the field safety case → Review standards scope ↗

G1
Stop authority

A participant can pause or permanently remove the robot at any time, without persuasion or penalty.

REQUIRED
G2
Physical safety

Applicable deployment review completed; protective stops and incident response verified before human exposure.

REQUIRED
G3
Privacy + security

Data boundaries, retention, access, remote operation, and security events are disclosed and auditable.

REQUIRED
G4
Serious-event rule

Any verified L4 event fails WANTED Safety Certification. Retention data remains visible for research integrity.

REQUIRED
L0
NORMAL
L1
NUISANCE
L2
MATERIAL
L3
SAFETY-RELEVANT
L4
SERIOUS / FAIL
03 / DEVELOPER INTEGRATION

Six events.
Any embodiment.

Keep the robot’s native control stack. WANTED only requires a signed, ordered event stream and one robot description: URDF, MJCF, or USD.

01DEPLOYMENT_LIFECYCLE
02ROBOT_STATE
03HUMAN_REQUEST
04ROBOT_ACTION
05HUMAN_INTERVENTION
06INCIDENT
SDK / JAVASCRIPT ESMv0.2
// Executable adapter: sequence + JCS + signature + chain
import { WantedClient, createHttpSink } from
  "./wanted-sdk.mjs";

const wanted = new WantedClient({
  deploymentId: "dep_7f2",
  environmentId: "env_104",
  robotId: "robot_07",
  signingKeyId: "key_prod_07",
  sign: bytes => secureModule.sign(bytes),
  sink: createHttpSink(eventsUrl)
});

await wanted.intervention(
  "remote_guidance", 43, "recovery"
);
POST /v1/eventsJSONL · HTTPS · SIGNED
ASSISTANCE BURDENI100

Minutes of external human assistance per 100 resident hours.

AUTONOMOUS AVAILABILITYA

Share of resident time capable of normal intended service.

RESCUE INTERVALMTBHR

Resident hours divided by human rescue events.

REACQUISITIONRback

Share choosing reinstall after the seven-day withdrawal.

04 / STUDY INTEGRITY

Hard to game.
Easy to audit.

Retention only means something when participants are free to reject the robot and teams cannot hide the operational burden.

PREREGISTERProve the rules came first.

Profile 0.2-PR1 independently timestamps the immutable root, binds eight commitments, and preserves every outcome-blind amendment in a parent-linked chain. Verify the history →

FREEZE THE FINISH LINEDo not let results decide when the run ends.

Profile 0.2-ST1 fixes the unit target, exposure floor, calendar cutoff, and primary-outcome access boundary before activity, while preserving independent safety and privacy monitoring. Verify the closure →

SEPARATE INCENTIVESNever pay people to keep it.

Base compensation is fixed and independent of robot retention. Milestone choice offers use a preregistered randomized mechanism.

LOG THE HIDDEN LABORTeleoperation is allowed, secrecy is not.

Remote guidance, recovery, maintenance, off-site debugging, and researcher contact are recorded with duration and reason.

TAMPER EVIDENCEOrder every event.

Per-deployment sequence numbers, UTC timestamps, signatures, and rolling hashes make deletion, reordering, and silent backfilling detectable.

INDEPENDENT ADJUDICATIONClassify the endpoint consistently.

Profile 0.2-J1 uses two blinded independent reviews—and a third-review majority on disagreement—to distinguish rejection, censoring, completion, and terminal competing causes. Open the verifier →

VERSION DISCLOSUREPublish what changed.

Robot hardware, policy, remote-support model, and material software changes are versioned. Cohorts may not be silently pooled across incompatible systems.

Q1 / KEEPRemove it today at no cost?BINARY · DIAGNOSTIC
Q2 / VALUELife better or worse recently?−2 TO +2 · DIAGNOSTIC
Q3 / BURDENHow much work is it creating?0 TO 4 · DIAGNOSTIC
Q4 / TRUSTOperate without supervision?0 TO 4 · DIAGNOSTIC

These probes explain retention; they never replace revealed preference or enter the WANTED Score. Field learning uses matched profile 0.2-LG1; assistance uses the complete signed register in 0.2-I1; operator-deployed policy changes use immutable lineage in 0.2-U1; household data boundaries use privacy + consent profile 0.2-PV1; uptime, downtime, repairs, and cloud dependence use the complete service clock in 0.2-SC1. Open the randomized four-item profile → Open the matched learning profile → Open the assistance-integrity profile → Open the policy-evolution profile → Open the privacy-integrity profile → Open the service-continuity profile → Open the binding-choice profile →

05 / CERTIFICATION PATH

Simulation first.
Real preference last.

Digital twins reduce human exposure to predictable failures. Only a qualifying WANTED WILD cohort produces a ranked WANTED Score. Open applicability matrix ↗

01Digital twin

PREQUALIFIED

Failure injection, collision, recovery, network loss, sensing drift, and human-trajectory stress tests.

02100+ hours

WANTED LAB

Real robot, real people, cohort-integrity evidence, passed safety gates, and complete event telemetry.

0310,000+ cohort hours

WANTED WILD

At least 20 profile-verified independent environments, supported 10K estimand, passed safety gates, and independent audit.

04One 10,000-hour residence

WANTED 10K

One continuous resident-clock lifetime run plus a seven-day withdrawal and voluntary reacquisition test.

06 / AUDITED LEADERBOARD

Rank the wanted hours.
Publish the burden.

Only WANTED Wild cohorts are ranked. Smaller or incomplete studies remain visible as provisional evidence.

#ROBOT / COHORTW95% CINHOURSS(10K)I100MTBHRSAFETY
NO AUDITED WANTED WILD RUNS YET

The first qualifying cohort sets the baseline. Provisional runs will remain separate from official ranking.

MANDATORY DISCLOSURE: TELEOPERATION · RESCUES · SERVICE CONTINUITY · INCIDENTS · PRIVACY INTEGRITY · WITHDRAWALS · PREFERENCE SUBSTUDY STATUSOPEN AUDITED REGISTRY →