Motion-to-Text · HumanML3D

Motion captioning under one population, one reference policy, and one semantic candidate protocol.

OFFICIAL TEST HUMANML3D-263 20 FPS MOTION → TEXT
Loading snapshot Protocol humanml3d_tm2t_v2
Evaluation clips
-
Full motions and tagged subclips
Baselines
-
GT is not ranked
References
-
Fixed captions per clip
Candidate group
-
Semantic R-Precision

Method comparison

Higher is better

Complete results

Best and second-best styling excludes GT and incomplete runs.

MethodStatusBLEU-1 ↑BLEU-4 ↑ROUGE-L ↑CIDEr ↑BERT raw ↑BERT rescaled ↑R@1 ↑R@2 ↑R@3 ↑Matching ↓
Population
3,969 + 431

Full motions plus tagged subclips under the released TM2T loader semantics.

Frame policy
[40, 200) → 196

TM2T source acceptance with evaluator-side truncation.

Language metrics
Two explicit tracks

TM2T token references reproduce its paper; raw captions expose reference-style sensitivity.

Semantic metrics
32 × 1 pass

Generated text queries against motion candidates with the official HumanML3D matching network.

All-case comparison

Search, play, and compare every one of the 4,400 evaluated clips. Motion assets load only for the selected case.

Quick caption audit

Four fixed records for a compact first look; use the all-case comparison above for the complete population.