Within-subject
Full-FTLinear Probe
RQ: Under a fixed model architecture, how does pretraining data scale affect downstream task performance?
Downstream performance under within-subject and cross-subject settings for Full-FT and Linear Probe across different pretraining data scales.
Relative parameter displacement and model output representational change induced by Full-FT. Layer-wise representational change across layers and pretraining data scales.
30k hours · 0.037 displacement · 0.901 change
30k hours · Layer 22 · 0.583