The Mind Reader Has Seen a Lot of Pictures

AI reconstructs images from brain scans, but learned visual patterns can fill the gaps. Impressive results do not yet prove that machines can read our minds.

On October 1, MIT Technology Review reported on an AI system that reconstructs images people are looking at by analysing their brain scans. Researchers trained it using photographs and the brain activity recorded while volunteers viewed them. The results suggest progress in decoding visual perception. The accompanying discussion reaches considerably further: could such technology eventually reveal our dreams, memories, and private thoughts?

Before accepting that invitation, we should ask where the reconstructed images actually come from. The system combines information from brain scans with patterns learned from extensive image training. How much of the result reflects what the person saw, and how much reflects what the AI expects to see?

One revealing failure captures the problem. A photograph of a dog in a bathtub became a goat in a bathtub. The bath survived. The species did not. That is an intriguing result in visual decoding—and a useful warning against mistaking a plausible picture for a faithful record of someone’s experience.

Brain-IT, developed by Michal Irani’s team at the Weizmann Institute, deserves serious attention. Its paper was accepted at ICLR 2026, and the researchers have released code and model checkpoints. Dismissing the work as unscientific would be lazy. The sharper criticism concerns the distance between a controlled reconstruction experiment and the suggestion that someone has opened a window into private thought. Project repository

One result is particularly useful: the authors report that an hour of fMRI calibration data can yield performance comparable to earlier methods trained with forty hours. Reducing that burden could make laboratory research more practical. Study

Start with the central distinction: information recovered from a brain signal and information supplied by a model’s previous training. A generative system can turn incomplete clues into a plausible scene. The plausibility comes partly from knowing what scenes usually contain. If the clues suggest an animal in a bathroom, convincing fur does not establish that those particular hairs were encoded in the measurement. Photographic detail can conceal evidential poverty. Earlier reconstruction research

This makes the training-data question essential. There are two different concerns: whether test images leaked into training, and whether familiar categories make reconstruction look more general than it is. Brain-IT reserves images for testing. That addresses ordinary decoder training leakage; it does not, by itself, audit every pretrained component’s exposure. I found no demonstrated test-image contamination. Claiming that the system merely retrieves memorised photographs would therefore exceed the evidence. Study, methods

The broader concern already has experimental support. In Spurious reconstruction from brain activity, Shirakawa and colleagues found that some earlier text-guided methods failed to generalise beyond familiar image distributions. Their analysis showed how apparent reconstruction could combine category prediction with generated embellishment. Brain-IT uses a different architecture, including structural information, so that paper does not refute it. It does explain why attractive examples deserve harder questions. Shirakawa et al.

Synthetic training data deserve similar care. The article reports that roughly 70 percent of the training data involves images originally lacking measured brain scans. An encoder predicts responses; those predictions help train the decoder. This is legitimate augmentation, but simulated observations cannot become independent biological evidence through repetition. Agreement between two models can reflect shared assumptions. The decisive test remains performance against withheld, real measurements. Technology Review

Then consider the scoreboard. Several impressive percentages measure discrimination between two candidates, not the percentage of an image faithfully recovered. In a harder 1,000-candidate CLIP retrieval test, Brain-IT scores 39.3 percent for one subject. That substantially exceeds chance and the reported competitors. It also leaves considerable uncertainty. The authors themselves explain the limitations of easier metrics. A responsible account should give those qualifications the prominence normally reserved for the prettiest reconstruction. Study, Appendix B

The paper’s own ablation gives the caution teeth: its low-level branch scores better on pixel correlation and structural similarity than the complete pipeline. Adding generative sophistication improves other measures while reducing those particular measures of fidelity. Study, Table T4

The main benchmark averages four participants. Thousands of pictures cannot turn four people into a representative population. The paper also includes qualitative results on unusual synthetic stimuli, which deserves credit, but selected examples do not establish broad reliability. For applications involving individuals, we need performance across people, sessions, conditions, and failures, with uncertainty attached. A compelling gallery establishes that something can work; it does not establish how reliably it will work for you. Study

The most consequential leap is from seeing to imagining. The article moves toward dreams, traumatic memories, and eventual decoding through EEG devices. These are research ambitions, not capabilities established by this experiment. A model trained against externally displayed photographs has a convenient answer key. Dreams do not arrive with reference images. Testing the accuracy of a claimed dream reconstruction is therefore a separate scientific problem, not merely the next software upgrade. Technology Review

Funding belongs in this discussion, with the same demand for evidence. The acknowledged ERC project, MindReading, received approximately €2.5 million under an agreement signed in 2024. That chronology rules out portraying this particular grant as the reward for yesterday’s headline. It does not remove incentives to attract attention. ERC project record

My concern is how the story rewards escalation. Better reconstruction becomes mind reading; possible applications become liberation for patients; speculative misuse becomes a looming threat. Promise and fear both make excellent publicity. Neither establishes accuracy. We can scrutinise that narrative without pretending to read the researchers’ motives, an achievement even their own system has not demonstrated.

A stronger demonstration would use newly created images, unusual combinations of objects, audited training overlap, and independent replication. It should compare correct brain signals with shuffled signals and category-only baselines, report every failure, and show how outputs vary across repeated generations. Most importantly, it should identify which details remain supported when the generator’s assumptions change.

Until then, the sensible verdict is qualified respect for the engineering and sustained scepticism toward the advertised meaning. The scientific question is how much visual information survives decoding. The publicity question is how close we are to reading minds. Confusing them gives a goat an excellent bath and the audience an unjustified sense of certainty.

No comments yet