Hand-action recognition on any video with Hiera-Hand (ChildPlay-Hand, ECCVW 2024). It tracks people, localizes their hands, and labels every hand in every frame as grasp, hold, operate, or release.
⚠️ This is a standalone demo based on the original research work; for training, evaluation, and the dataset, see idiap/childplay_hand.
./setup_env.sh # needs uv; Python 3.11 venv in .venv
source .venv/bin/activateFetch the checkpoints from Zenodo into checkpoints/:
python src/download_checkpoints.py # manipulation (~0.4 GB download)
python src/download_checkpoints.py --task all # + object (~0.8 GB download)python src/demo.py input.mp4 --output result.mp4Works best on clips where people are fully visible, without camera cuts.
- Track people and their pose: YOLO26m-pose + BoT-SORT, in one pass.
- Find hands: each hand box sits just past the wrist, along the elbow→wrist direction.
- Recognize: for every hand and frame, a ~1 s window (32 frames, 16 sampled) is cropped around the hand at 224×224 and fed to Hiera-Base (51M params, 205 MB; MAE-pretrained on Kinetics-400, fine-tuned on ChildPlay-Hand), which outputs background / grasp / hold / operate / release.
- Display: smoothed hand boxes and hand actions.
Note: The paper used HRNet-W32 for pose; this demo uses YOLO26m-pose for speed, so predictions may differ slightly from the reported results.
- Code: GPL-3.0, based on idiap/childplay_hand (© Idiap Research Institute).
- Checkpoints: CC BY-NC 4.0 (non-commercial), from Zenodo.
- Pose model: Ultralytics YOLO26, AGPL-3.0.
- Hiera architecture (src/hiera/): Apache-2.0, © Meta.
@inproceedings{Farkhondeh_ECCVW_2024,
author = {Farkhondeh*, Arya and Tafasca*, Samy and Odobez, Jean-Marc},
title = {ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild},
booktitle = {Proceedings of the European Conference on Computer Vision (ECCV) Workshops},
year = {2024},
note = {* Equal contribution}
}