Evaluation clips
-
Full motions and tagged subclips
Motion captioning under one population, one reference policy, and one semantic candidate protocol.
Higher is better
Best and second-best styling excludes GT and incomplete runs.
| Method | Status | BLEU-1 ↑ | BLEU-4 ↑ | ROUGE-L ↑ | CIDEr ↑ | BERT raw ↑ | BERT rescaled ↑ | R@1 ↑ | R@2 ↑ | R@3 ↑ | Matching ↓ |
|---|
Full motions plus tagged subclips under the released TM2T loader semantics.
TM2T source acceptance with evaluator-side truncation.
TM2T token references reproduce its paper; raw captions expose reference-style sensitivity.
Generated text queries against motion candidates with the official HumanML3D matching network.
Search, play, and compare every one of the 4,400 evaluated clips. Motion assets load only for the selected case.
Four fixed records for a compact first look; use the all-case comparison above for the complete population.