Every rate travels with exposure, trials, events, labels, and missingness.
Explain the score.
Never dilute it.
W says how long people voluntarily retain the robot. This profile shows what life with it cost: hidden assistance, owner labor, failures, recovery, interruptions, learning, privacy response, and whether proactive behavior was actually welcome.
One retention score.
Many honest explanations.
A diagnostic is only useful when teams cannot improve it by hiding downtime, dropping unresolved labels, or redefining the denominator after seeing results.
No rescue or failure produces a right-censored lower bound: at least the observed exposure.
Stop and privacy response report N, P50, P95, P99, and maximum—not only an average.
Diagnostics explain retention. They are never combined into a weighted index or used to offset safety.
Operational burden.
Behavioral adaptation.
Each formula is computed per environment first. Official cohort uncertainty resamples independent environments, not individual robot actions.
All remote guidance, teleoperation, rescue, and research assistance.
Configuration, teaching, maintenance, cleanup, and rescue imposed on the user.
With zero rescues, report “≥ observed resident hours,” never infinity.
Unresolved actions stay in the denominator and label coverage is published.
Windows and task-matching rules are frozen before hour one.
Publish both components; suppress the ratio if familiar performance is zero.
Paste counters.
Export evidence.
The synthetic example demonstrates the aggregate input format. Calculations remain in the browser and contain no participant-level records.
Context travels
with every metric.
Task mix, embodiment, population, environment, support model, and robot version affect every diagnostic. WANTED requires disclosure instead of pretending those differences disappear.
Only voluntary retention determines rank.
Safety cannot be traded for usefulness.
Always contextual and non-ranking.