Pretraining Effectiveness

RQ: How does pretraining improve early optimization and final downstream performance relative to supervised training from scratch, and how does this benefit vary with the amount of labeled downstream data?

Pretraining Benefit Analysis across EEG FMs

We compare full fine-tuning (Full-FT) with architecture-matched scratch supervised training (Scratch-ST).

Full-FTScratch-ST
HEAR-large
20406080151015EpochAverage performance (%)
HEAR-base
20406080151014EpochAverage performance (%)
ST-EEGFormer-small
2040608015101520EpochAverage performance (%)
LEAD
2040608015101518EpochAverage performance (%)
CSBrain
20406080151013EpochAverage performance (%)
CBraMod
20406080151015EpochAverage performance (%)
NeuroRVQ
204060801510EpochAverage performance (%)
CodeBrain
20406080151015EpochAverage performance (%)
EEGMamba
20406080151013EpochAverage performance (%)
REVE-base
20406080151013EpochAverage performance (%)

Final-Epoch Gains across Label Budgets

Final-epoch gains increase with the amount of labeled downstream data, suggesting that pretraining enhances the models' ability to exploit additional labeled data.

-10-50+5+10+15+202%5%10%25%50%100%Labeled downstream dataFinal-epoch gain (%)10/10 positive

Labeled downstream data

Pretraining Benefit Matrices

DatasetCBraModCodeBrainCSBrainEEGMambaHEAR-baseHEAR-largeLEADNeuroRVQREVE-baseST-EEGFormer-small
HGD——
Meng2019
OpenBMI———
PhysioNet-MI—
TUSL——
ERP-BCI———
Kaggle-INRIA——
SEED
EEGMat——

Benchmark Score is Balanced Accuracy for binary tasks and Macro F1 for multiclass tasks.