How we score
Scores are navigational shorthand: every number is anchored to observable behavior with cited evidence, and every change is logged on the specimen page.
The dimensions (0–10)
- Wobble
- Expressive, delightful, unnecessary physical motion.
- Creature Factor
- How readily humans begin treating it as a creature rather than a device.
- Kid Gravity
- Likelihood children return to it after the novelty period.
- Dad Gravity
- Likelihood the adult purchaser continues "testing."
- Household Fit
- Noise, setup, durability, maintenance, charging.
- Longevity
- Hardware and software survival probability — read with the support-risk flag.
- Magic
- Whether interaction occasionally produces surprise or attachment.
- Developer Platform
- For open robots and kits: API surface, SDK quality, docs, simulation support, community vitality. Closed companions read "closed platform," not a punitive number.
The Gift Test
Relevant robots additionally get the household assessment: who it's actually for, age bands (3–4 · 5–7 · 8–12 · 13+ · adult · collector), setup time, whether it works immediately, whether a parent must operate it, whether the opening moment lands, whether it's still interesting at one month, and every recurring cost. Manufacturer age recommendations are reported separately — they answer a liability question, not ours.
Evidence rules
Scores appear only after the editor confirms them against the record: owner-community findings need multiple independent sources; anything from our own household testing is labeled as such; unverifiable dimensions stay unscored rather than guessed. Every score change gets a dated changelog entry on the specimen page.