Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data
This paper discusses a new method for improving how we understand and interpret brain activity related to perception, especially in individuals with visual impairments.
Content & Liability Disclaimer
This article and its accompanying video are automated summaries derived from the original research paper by Unknown authors. The original research was conducted solely by the paper's authors; PDFdigest did not conduct any of the research and makes no claims of ownership over the underlying scientific work.
The video narration is generated by artificial intelligence and references the paper's authors for attribution. The video is not narrated by any of the paper's authors. This content may contain inaccuracies, omissions, or misinterpretations of the original research. First-person language (e.g., "we found", "our results") reflects the original authors' voice, not PDFdigest's. Always read the original paper for accurate, verified information before making any decisions based on this content.
This content is provided "as is" without any warranties, express or implied. Simulated systems OÜ, its officers, directors, employees, and agents shall not be liable for any direct, indirect, incidental, special, consequential, or punitive damages arising from your use of, reliance on, or access to this content, including but not limited to errors, omissions, or misinterpretations of the original research. This disclaimer applies to the fullest extent permitted by applicable law.
- 1 Training Objective We combine two loss terms to learn representations that generalize between sessions while preserving neural dynamics.
- 2 In clinical neuroprosthetics, such data are scarce: participant availability is limited, experimental sessions are short, and subjective reports are inherently difficult to measure.
- 3 Here, we demonstrate that self-supervised learning in spontaneous V1 activity can improve perception decoding.
- 4 Yet this is precisely where supervised decoders fail: with limited labeled data, they cannot distinguish the subtle neural states that determine perception versus non-perception at threshold.
Introduction
Decoding subjective perception from neural activity requires labeled training data, pairing neural recordings with perceptual reports. Meanwhile, tens to hundreds of hours of spontaneous neural activity accumulate during rest periods and between experimental blocks.
This creates a fundamental imbalance that limits neuroengineers from training decoding models: abundant unlabeled recordings, limited labeled examples.
This data bottleneck is most severe at perceptual threshold.
Limitations and Open Questions Our main limitation is that these findings are from a single subject.
This creates a fundamental imbalance that limits neuroengineers from training decoding models: abundant unlabeled recordings, limited labeled examples.
Research Question
Training Objective We combine two loss terms to learn representations that generalize between sessions while preserving neural dynamics.
Methodology
On a general psychometric task, the pretrained features achieved 84.1% accuracy. More critically, when evaluated on perceptual report trials at the 50% threshold, where identical stimuli produced variable outcomes, pretraining improved decoding accuracy achieving 64.0%, demonstrating that spontaneous dynamics contain structure relevant to perception.
Study Design
Three datasets gave us complementary views of the same neural population: spontaneous activity to learn dynamics, psychometric curves to map perception, and threshold trials where identical stimuli produced variable perception.
We then trained L2-regularized logistic regression to predict perception in the psychometric and threshold level datasets.
Also, our current 128ms analysis window might be too brief to capture these slower changes, which is an important area for future work.
How PDFdigest Helps You Understand Research
Instant Paper Analysis
Get structured summaries and key findings from dense PDFs in seconds.
Visual Explanations
Turn complex methods, figures, and results into clearer visual breakdowns.
AI-Powered Q&A
Ask focused questions and get answers grounded in the paper.
Results & Findings
In clinical neuroprosthetics, such data are scarce: participant availability is limited, experimental sessions are short, and subjective reports are inherently difficult to measure. Here, we demonstrate that self-supervised learning in spontaneous V1 activity can improve perception decoding.
- In clinical neuroprosthetics, such data are scarce: participant availability is limited, experimental sessions are short, and subjective reports are inherently difficult to measure.
- Here, we demonstrate that self-supervised learning in spontaneous V1 activity can improve perception decoding.
- Yet this is precisely where supervised decoders fail: with limited labeled data, they cannot distinguish the subtle neural states that determine perception versus non-perception at threshold.
- We hypothesize that spontaneous cortical activity explores the same low-dimensional manifolds that govern stimulusevoked responses.
- Masked autoencoders (MAEs) work by reconstructing randomly masked input, achieving strong results in multiple domains , including recent work on neural data .
In clinical neuroprosthetics, such data are scarce: participant availability is limited, experimental sessions are short, and subjective reports are inherently difficult to measure.
Yet this is precisely where supervised decoders fail: with limited labeled data, they cannot distinguish the subtle neural states that determine perception versus non-perception at threshold.
Practical Applications
Some trials simply could not be classified from the neural state alone, as nearly identical states could lead to different perceptual outcomes. We tested if abundant, unlabeled spontaneous neural activity could overcome the data bottleneck in perception decoding.
Data Preprocessing
The preprocessing of data involved extracting multi-unit activity and applying different normalization strategies based on the presence of stimulation artifacts. Continuous z-score normalization was used for resting state data, while robust z-score normalization was applied for stimulation data.
Self-Supervised Pretraining
This section details the architecture and training of the masked autoencoder, including modifications made to the standard setup. It describes the training objective, which combines reconstruction loss and cross-session loss to learn generalizable representations of neural dynamics.
Figures Explained
The paper’s visual material highlights the workflow and the main system components.
- Figure 1 :: Figure 1: Self-supervised pretraining architecture. Input signals are randomly masked. The encoder processes only visible patches. The decoder receives encoder outputs, learned mask tokens, and a session embedding.
- Figure 2 :: Figure 2: t-SNE projection of spontaneous activity features. (Left) Colored by mean neural activity. (Right) Colored by recording day.
- Figure 3 :: Figure 3: Learned representations capture spatial and perceptual structure. t-SNE projection of MAE features for psychometric trials. (Left) Colored by electrode and stimulation amplitude. Spatially adjacent electrodes cluster together, probably reflecting V1’s retinotopic organization. (Right) Colored by perception outcome. Perceived (orange) and missed (dark blue) trials show partial separation, indicating learned features reflect perceptual states despite no explicit supervision during pretraining.
- Figure 4 :: Figure4: Per-session t-SNE of stimulation responses showed clean separation (orange=perceived, blue=missed). Mixing only occurred in the 20-80% psychometric range where perception itself was uncertain.
Frequently Asked Questions
But perception at threshold is inconsistent for identical stimulation sometimes produces a phosphene, sometimes it does not. Training Objective We combine two loss terms to learn representations that generalize between sessions while preserving neural dynamics.
More critically, when evaluated on perceptual report trials at the 50% threshold, where identical stimuli produced variable outcomes, pretraining improved decoding accuracy achieving 64.0%, demonstrating that spontaneous dynamics contain structure relevant to perception. We then trained L2-regularized logistic regression to predict perception.
In clinical neuroprosthetics, such data are scarce: participant availability is limited, experimental sessions are short, and subjective reports are inherently difficult to measure. Here, we demonstrate that self-supervised learning in spontaneous V1 activity can improve perception decoding.
Some trials simply could not be classified from the neural state alone, as nearly identical states could lead to different perceptual outcomes. We tested if abundant, unlabeled spontaneous neural activity could overcome the data bottleneck in perception decoding.
In clinical neuroprosthetics, such data are scarce: participant availability is limited, experimental sessions are short, and subjective reports are inherently difficult to measure. Yet this is precisely where supervised decoders fail: with limited labeled data, they cannot distinguish the subtle neural states.
This paper discusses a new method for improving how we understand and interpret brain activity related to perception, especially in individuals with visual impairments.