Peanuts—mostly accused of being nuts, when in reality their pedigree is populated by legumes. [1] When we first encounter a contradictory fact, a feeling of cognitive dissonance strikes us. Not long after fact-checking the statement, we accept the bit of information and have learned some new trivia.
Somehow, learning that bit of trivia must have unleashed a change in brain structure. The brain’s reaction to new experiences is called neuroplasticity, and it pushes the brain to adapt its structure via rewiring or the reinforcement of neurological connections. [2] But neuroplasticity does not adapt the brain uniformly. The adaptation depends on the type of teaching signal to which different learning rules are applied. Moldakarimov and Sejnowski summarized three learning rules in their paper: unsupervised, reinforcement, and supervised learning. [3] Supervised learning is what enabled the previous learning of the peanut trivia, as a state of knowledge was compared against a ground truth and updated accordingly.
This essay argues that learning from experience can be induced in not just one, but several ways, and that the difference in signal type determines what each rule can learn and how the physical structure of the brain adjusts.
Unsupervised Learning
Unsupervised learning takes a broad view of the entirety of experiences made and uses statistical patterns to uncover structure. It is independent of any teaching signal; all learning emerges from structure in the input. [4]
Every skill acquired without any feedback can be considered to be learned through unsupervised learning. Because most skills provide some form of natural feedback, identifying a task that is learned entirely without feedback is nontrivial. Vision, though, has almost no feedback when learning. [5]
Meet frankensteinin’ ferrets. In their paper, Melchner et al. asked whether regions of the brain are hardwired for specific tasks. To answer, they relocated the right retinal axon of neonatal ferrets to their auditory thalamus. The rewired and control group ferrets were raised identically to adulthood. Once adulthood was reached, the researchers trained the ferrets to fetch rewards from one of two tubes depending on whether they saw a light or heard a noise. The training was conducted solely on the unaltered hemisphere of the brain. After training concluded, the ferrets were exposed to the same experiment, simulating the opposite hemisphere. When the rewired eye was exposed to light, the ferrets chose the correct tube. Had their brain interpreted the signal as sound, they would have entered the incorrect tube. This showed that the auditory thalamus was repurposed by the brain to process visual information. Further, the researchers structurally ablated the lateral geniculate nucleus and auditory cortex to cement their hypothesis. [6] The ferrets were not subjected to specific training; their brain adjusted its structure by detecting the patterns of the input, without any teaching signal. The auditory thalamus of the rewired ferrets had developed pinwheels, which are primarily found in the visual cortex. [7] This shows how the brain updates structures to handle novel signals.
Reinforcement Learning
Where unsupervised learning lacks a teaching signal, reinforcement learning can be understood as a carrot-and-stick policy. The learner attempts to master a task through trial-and-error and is punished or rewarded depending on their actions. [8]
A motor skill is always learned by reinforcement learning, as it is a process of trial-and-error. Learning how to throw a basketball and scoring or not turns into automatic reward or punishment for the learner.
Vassiliadis et al. trained participants to learn a hand-pressure-based motor tracking skill to explore how patterns of oscillatory activity within the human striatum influence learning. They exploited a novel non-invasive deep brain stimulation technique called transcranial temporal interference stimulation (tTIS). Participants held a pressure-based controller that, when pressed, would move a cursor. The objective was to match the cursor’s position to a target on a screen.
Each participant went through multiple blocks of the motor learning task, cycling through combinations of two main variables: a learning condition and a stimulus condition. The learning condition included two modes, with or without external reward. The stimulus condition applied the tTIS and consisted of 80 Hz high-gamma stimulation, intended to inhibit reward-based learning, 20 Hz as a control, and a sham condition.
When exposed to the 80 Hz high-gamma stimulation, participants’ ability to utilize reinforcement feedback to improve their tracking accuracy was entirely abolished. They performed no better than when practicing the task completely devoid of external rewards. Conversely, the 20 Hz and sham conditions left the reinforcement learning process intact. [9] Had the striatum merely been a passive relay, the 80 Hz inhibition would have impaired all motor execution or general learning. Instead, it specifically severed the neural communication required to translate the teaching signal into improved performance. This indicates that reinforcement learning is fundamentally mediated by frequency-specific oscillatory activity within the striatum. The participants were subjected to the exact same physical trial-and-error, but their brains could not update their internal models without the correct neural rhythm to process the reward. Just as the ferrets’ auditory thalamus structurally adapted to the statistical patterns of visual inputs without a teacher, the human brain dynamically updates its motor pathways by synchronizing specific neural frequencies in response to explicit feedback. Disrupt this highly specific high-gamma bridge, and the teaching signal is rendered useless; the learner experiences the task, but the brain cannot translate the “carrot” into acquired knowledge.
Supervised Learning
While unsupervised learning detects structure without a teaching signal, and reinforcement learning relies on a carrot-and-stick trial and error, supervised learning relies on an explicit teaching signal to compare an outcome against a definitive ground truth. [3]
The human brain increases its gray matter volume and density through intense fact learning. [10] Kwok et al. retested this statement. Their experiment involved participants learning newly invented names for shades of green and blue. The learners were shown shades of color and explicitly corrected until they could categorize the new visual boundaries. After two hours of this supervised training, MRI scans revealed a structural increase in gray matter volume specifically within the left visual cortex, a region dedicated to color vision. Just as the ferrets’ auditory thalamus structurally repurposed itself based on raw visual input, and the striatum synchronized its rhythms to process rewards, the human visual cortex physically expanded its gray matter density in direct response to an explicit, supervised teaching signal. The brain grew to host the new ground truth. [11]
Conclusion
Together, these case studies demonstrate that neuroplasticity is not a universal mechanism. The brain requires distinct learning rules to process different types of inputs. Unsupervised learning relies entirely on statistical patterns when no teaching signal is available, physically repurposing neural regions by developing structures like visual pinwheels in the auditory thalamus. Reinforcement learning uses trial-and-error feedback to master actions without a ground truth, relying on the synchronization of specific, high-gamma oscillatory rhythms within the striatum to bridge the gap between action and reward. Supervised learning drives structural expansion by comparing internal models against explicit facts, measurably increasing gray matter volume and density to host the new ground truth. Each rule occupies a specific niche, ensuring that whether through raw observation, reward, or explicit correction, the brain can physically adapt to encode new knowledge.
These rules depend entirely on the nature and type of the teaching signal. From the surprising pedigree of a peanut to our motor skills, neuroplasticity ensures that, by applying the appropriate learning rule, every encounter is stamped directly into the neural networks of our minds.
References
[1] Y. Zhao et al., “Nuclear phylotranscriptomics and phylogenomics support numerous polyploidization events and hypotheses for the evolution of rhizobial nitrogen-fixing symbiosis in Fabaceae,” Mol. Plant, vol. 14, no. 5, pp. 748–773, May 2021, doi: 10.1016/j.molp.2021.02.006.
[2] P. Gazerani, “The neuroplastic brain: current breakthroughs and emerging frontiers,” Brain Res., vol. 1858, p. 149643, Jul. 2025, doi: 10.1016/j.brainres.2025.149643.
[3] S. Moldakarimov and T. J. Sejnowski, “Neural Computation Theories of Learning,” in Learning and Memory: A Comprehensive Reference, Elsevier, 2017, pp. 579–589, doi: 10.1016/B978-0-12-809324-5.21026-2.
[4] P. Dayan, “Unsupervised Learning,” in The MIT Encyclopedia of the Cognitive Sciences, R. A. Wilson and F. C. Keil, Eds., Cambridge, MA: The MIT Press, 2001, pp. 1–7.
[5] G. Matteucci, E. Piasini, and D. Zoccolan, “Unsupervised learning of mid-level visual representations,” Curr. Opin. Neurobiol., vol. 84, p. 102834, 2024, doi: 10.1016/j.conb.2023.102834.
[6] L. Von Melchner, S. L. Pallas, and M. Sur, “Visual behaviour mediated by retinal projections directed to the auditory pathway,” Nature, vol. 404, no. 6780, pp. 871–876, Apr. 2000, doi: 10.1038/35009102.
[7] J. Sharma, A. Angelucci, and M. Sur, “Induction of visual orientation modules in auditory cortex,” Nature, vol. 404, no. 6780, pp. 841–847, Apr. 2000, doi: 10.1038/35009043.
[8] L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement Learning: A Survey,” 1996, doi: 10.48550/ARXIV.CS/9605103.
[9] P. Vassiliadis et al., “Non-invasive stimulation of the human striatum disrupts reinforcement learning of motor skills,” Nat. Hum. Behav., vol. 8, no. 8, pp. 1581–1598, May 2024, doi: 10.1038/s41562-024-01901-z.
[10] E. A. Maguire, K. Woollett, and H. J. Spiers, “London taxi drivers and bus drivers: A structural MRI and neuropsychological analysis,” Hippocampus, vol. 16, no. 12, pp. 1091–1101, Dec. 2006, doi: 10.1002/hipo.20233.
[11] V. Kwok et al., “Learning new color names produces rapid increase in gray matter in the intact adult human cortex,” Proc. Natl. Acad. Sci., vol. 108, no. 16, pp. 6686–6688, Apr. 2011, doi: 10.1073/pnas.1103217108.