HARBench Leaderboard

A comprehensive benchmark for evaluating HAR foundation models.

Assesses domain robustness, position robustness, few-shot learning, and zero-shot transfer using 32 datasets (2.4M+ hours).

IEEE PerCom 2026 · [Paper] [Cite]

Datasets

32 public datasets (14 unlabeled for pretraining, 18 labeled for evaluation) totaling over 2.4 million hours of sensor data.

Daily Activities (10)

ForthTrace, HARTH, IMWSHA, PAAL, PAMAP2, RealWorld, SelfBack, UCAE-HAR, USC-HAD, WARD

Exercise (4)

DSADS, MEX, MHEALTH, RealDisp

Industry (4)

LARa, OpenPack, Exoskeletons, VTT-ConIoT

Evaluation Protocol

5 Evaluation Dimensions

  • Overall: Average F1 across all 18 datasets
  • Domain: Balanced performance across Daily/Exercise/Industry
  • Position: Consistency across 8 sensor positions
  • Few-Shot: Learning with 1-50% labeled data
  • Zero-Shot: Transfer to unseen datasets

Standardized Protocol

  • Sampling Rate: 30Hz
  • Window Size: 5 seconds (150 samples)
  • Validation: 4-fold person-independent CV
  • Metric: Macro F1 Score