A humanoid robot can match a reference pose frame by frame and still look unmistakably wrong. Its feet may skate across the floor, a hand may arrive late, or its balance may unravel over several steps. A new robotics benchmark argues that these visible failures expose a measurement problem: the numbers commonly used to judge motion tracking do not always capture the qualities people notice first.
HumanTracker, submitted August 13 and accepted to ECCV 2026, pairs a much larger motion test set with HumanScore, a learned metric designed to predict expert preferences between robot-motion videos. The work comes from researchers affiliated with Nankai University, Tsinghua University, Galbot, Shanghai Jiao Tong University, Peking University, and the Shanghai Qi Zhi Institute.
A broader test for robot movement
The benchmark contains about 153 hours of optical motion capture from 24 professional performers, divided into roughly 25,000 clips. Instead of treating movement as one undifferentiated pool, the authors organize it into four families: daily activity, highly dynamic movement, interaction, and ground-level motion.
That division matters because each family stresses a different control problem. Walking and turning reveal drift. Jumps and fast footwork test impact handling and rapid support changes. Interaction motions demand coordinated hands, arms, and body movement. Kneeling, sitting, rolling, and recovering create multiple contacts and a low center of mass.
The paper compares that structure with earlier resources. AMASS unified more than 40 hours of existing motion-capture data into a common representation. PHUMA later assembled a 73-hour corpus focused on physically reliable humanoid locomotion, explicitly filtering artifacts such as floating, penetration, and foot skating. HumanTracker’s distinct contribution is to turn a larger, categorized collection into a diagnostic test set and connect it to preference judgments.
Why pose error misses the awkward moments
Conventional metrics average differences between a simulated robot and a reference motion across joints and time. They remain useful for locating pose error, but averaging can hide an event that dominates perception. One mistimed touchdown or progressive slide may occupy only part of a sequence while making the entire movement look unstable.
HumanScore was trained from comparisons produced by four humanoid trackers. Six doctoral researchers specializing in humanoid robotics annotated 6,000 original trajectory pairs; mirrored versions expanded the development set to 12,000 records. Annotators judged balance, contact, tracking stability, jitter, foot sliding, locomotion consistency, and whole-body naturalness.
The model evaluates five-second windows using simulator state, control, contact, and motion features. On held-out comparisons, the paper reports 90.83 percent alignment with human preferences. The strongest conventional metric in the reported comparison reached 84.05 percent. Longer temporal context improved alignment, supporting the idea that sliding, drift, and recovery are events rather than isolated poses.
TENS analysis: the scorecard shapes the robot
The central lesson is not simply that a learned score beats a hand-written metric. It is that robot progress is partly defined by what a benchmark makes visible. A single average rewards broad pose fidelity; category-level results can expose a controller that succeeds during upright daily motion but breaks during ground transitions.
The benchmark changes the evaluation question from “How close was every joint?” to “Did the complete movement remain convincing under its hardest contact regime?” That shift matters for teleoperation and imitation because a robot that reproduces a gesture while losing stable support has not reproduced the useful behavior.
HumanScore should be treated as an additional instrument, not a universal replacement. Success rate still detects catastrophic failure, joint error identifies local mismatch, and contact diagnostics can locate a specific defect. The learned score contributes a trajectory-level judgment that combines qualities people perceive together.
The comparison with AMASS and PHUMA also separates three jobs that are often conflated: collecting diverse human motion, producing physically plausible robot references, and evaluating the resulting controller. More data can improve coverage without guaranteeing physical realism; cleaner data can improve training without proving that a test metric reflects perceived quality.
The limits keep the claim grounded
The authors are explicit about boundaries. HumanScore learned from four trackers on one 29-degree-of-freedom humanoid in the MuJoCo simulator. Its inputs include privileged simulator state and contact quantities that may not be available on hardware. Each preference pair received one primary judgment, and ground-level clips are much less numerous than daily or interaction clips.
Those limits make the next tests clear: new robot bodies, unseen controller families, real-world rollouts, rarer contact regimes, and repeated independent judgments. The paper also warns against directly optimizing HumanScore without safeguards, because a controller could learn to exploit weaknesses in the learned metric.
That caution is the most useful framing for the field. Human-aligned evaluation can reveal what pose averages miss, but a learned judge can develop blind spots of its own. Better humanoid benchmarks will need both perception and diagnosis: a score that notices when motion looks wrong, plus measurements that explain why.
Sources: HumanTracker research paper; PHUMA research paper; AMASS research paper.
Featured image: NASA’s Valkyrie R5 humanoid robot. Photo: NASA, Bill Stafford, James Blair, and Regan Geeseman via Wikimedia Commons. Public domain, United States government work. Top-aligned crop from 4,000 × 3,000 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative or substantive alteration.


