HEAR-large
Pretraining Effectiveness
RQ: How does pretraining improve early optimization and final downstream performance relative to supervised training from scratch, and how does this benefit vary with the amount of labeled downstream data?
Pretraining Benefit Analysis across EEG FMs
We compare full fine-tuning (Full-FT) with architecture-matched scratch supervised training (Scratch-ST).
HEAR-base
ST-EEGFormer-small
LEAD
CSBrain
CBraMod
NeuroRVQ
CodeBrain
EEGMamba
REVE-base
Final-Epoch Gains across Label Budgets
Final-epoch gains increase with the amount of labeled downstream data, suggesting that pretraining enhances the models' ability to exploit additional labeled data.
Labeled downstream data
Pretraining Benefit Matrices
| Dataset | CBraMod | CodeBrain | CSBrain | EEGMamba | HEAR-base | HEAR-large | LEAD | NeuroRVQ | REVE-base | ST-EEGFormer-small |
|---|---|---|---|---|---|---|---|---|---|---|
| HGD | — | — | ||||||||
| Meng2019 | ||||||||||
| OpenBMI | — | — | — | |||||||
| PhysioNet-MI | — | |||||||||
| TUSL | — | — | ||||||||
| ERP-BCI | — | — | — | |||||||
| Kaggle-INRIA | — | — | ||||||||
| SEED | ||||||||||
| EEGMat | — | — |
Benchmark Score is Balanced Accuracy for binary tasks and Macro F1 for multiclass tasks.


