Before reaching for an EEG foundation model for an SSVEP speller, compare it with CCA, which needs no training at all.
BETA, 40 targets, 70 new people, eight posterior electrodes (chance 2.5%; descriptive 95% intervals in brackets): standard CCA reached 63.1% balanced accuracy (57.2–69.0%). Frozen LaBraM and CBraMod with ridge heads reached 10.8% (9.5–12.2%) and 33.7% (29.8–37.6%); EEGNet trained from scratch, 55.8% (50.1–61.3%). 16 further foundation-model checkpoints in the same frozen recipe reached at most 54.8% (49.4–60.0%), and none lies above CCA with eight electrodes or with four (57.6%). These are frozen probes with one fixed recipe, not tuned ceilings for any model.
Before deciding how much of an EEG foundation model to fine-tune, check how strong its frozen readout already is.
LaBraM Base on mental arithmetic vs rest (EEGMAT, 36 new people, no labels from the test person, three seeds; descriptive 95% intervals in brackets): updating the last block reached 65.7% balanced accuracy (61.6–69.7%) and rank-4 LoRA 64.1% (60.2–67.8%), against 56.6% (54.2–59.0%) for a short head-only fit. A frozen encoder with a ridge readout, published separately on the same people and folds, reached 64.6% (60.5–68.6%): a different head, not a paired comparison. LoRA minus last block was −1.6 pp (−4.4 to +1.0), so no ordering. And LoRA, with 38,802 trainable parameters against 482,882, took about 128 s to train against 65 s (15 fits each, one Apple-silicon Mac, this implementation).
Can a Jev-style decision interface work on EEG when the questions are asked in words? In our tests, only for questions the model was trained on.
A small head over frozen EEG features (encode a window once, answer several typed questions about it) was asked each question by a question number, a label template or a description. On SSVEP (BETA, 70 people), asking a seen question with a label template instead of its number cost −5.96 pp of balanced accuracy on a plain spectrum (interval −6.84 to −5.08 pp), the 2 pp margin not met. On flicker frequencies the head was never trained on, it reached 28.7%, where training-free CCA reached 80.9% (chance 12.5%).
Motivated by Jev-style typed decision interfaces; independent evaluation, not affiliated with TypeSafe, AutoTrust, ggml-org or the authors of any Jev paper.
On BETA (70 people, 40-target SSVEP, cue-paced lab trials), one calibrated threshold — accept when the top-class probability is at least 0.8 — accepted very different shares of trials: standard CCA 28.4% (95% interval 22.9–34.2%), EEGNet 16.6% (12.5–21.0%), frozen CBraMod 1.5% (0.8–2.3%), with nothing accepted at all for 38 of the 70 people under CBraMod. A learned reject option did not beat the model's own calibrated confidence, and a certified error level did not reliably hold on new people. These compare methods; they are not deployment error rates.