Learning Like a Neural Network: Evidence for Vectorized Credit Assignment in the Brain

Reading time: 11 Minutes

One of the central ideas behind modern artificial intelligence is surprisingly simple: different parts of a network should not all learn the same lesson.

Artificial intelligence systems are built from networks of artificial neurons, loosely inspired by the brain. While the information flows forward through the network, it is transformed at each step until it produces an output such as a sentence, a prediction, or an image label. But intelligence is not just about producing outputs; it is about improving them.

When a chatbot gets an answer wrong or an image-recognition system mistakes a dog for a cat, the system must somehow trace the error back through its many layers and determine which internal components were responsible. To improve performance, the system needs to identify not just that a mistake occurred, but also how much each neuron and connection contributed to it, so that appropriate adjustments can be made during learning (Figure 1a). This process, known as vectorized credit assignment, is central to powerful learning algorithms such as backpropagation.

For decades, neuroscientists have wondered how the brain solves the problem of determining which specific neurons and synapses are responsible for a successful or unsuccessful outcome that also confronts artificial intelligence systems. Much of the evidence suggests that the brain relies on signals that are broadcast widely rather than targeted to individual neurons. Chemicals such as dopamine can carry information about whether an outcome was better or worse than expected, or whether something is important or rewarding. These signals, however, do not tell each neuron its contribution to the result. Yet, learning is most effective when the correct neurons and synapses are strengthened or weakened. If all neurons receive the same signal, the brain cannot precisely identify which ones caused the success or error.

The key question is whether all neurons involved in a behaviour receive essentially the same feedback, or whether the brain somehow delivers more specific information to different neurons based on their roles.

A recent study from the Harnett lab at MIT set out to address this question. The researchers examined how feedback signals are distributed across populations of cortical neurons. Different parts of these neurons receive different kinds of information. Sensory inputs from the outside world arrive mainly near the cell body, while feedback from other brain regions, including information about predictions, expectations, or behavioural outcomes, arrives on long branching extensions called apical dendrites (Figure 1b).

 If learning is driven primarily by global signals, neurons should receive largely similar feedback. If the brain supports more precise forms of credit assignment, however, different neurons should receive distinct feedback signals that reflect their different contributions to neural computations.

Their results point toward the latter possibility. Rather than appearing as a uniform broadcast, feedback in the cerebral cortex seemed to vary substantially between different neurons. This suggests that the brain may provide more individualized learning signals than previously appreciated. So, how can neural circuits generate and coordinate so many distinct feedback signals simultaneously? And how could a single neuron determine what feedback is relevant to its own activity? 

Before addressing these questions experimentally, the researchers first established what evidence would be required for a signal to qualify as a neuron-specific teaching signal. A candidate learning signal should convey information that cannot be inferred solely from the neuron’s own activity. It should carry information about success or failure, but in a form that differs across neurons according to their roles in the computation. Most importantly, it should be causally involved in learning, meaning that disrupting the signal should impair the brain’s ability to improve performance. 

These requirements define a vectorized teaching signal. In such a signal, feedback can, in principle, assign credit to individual neurons rather than broadcasting the same message across an entire population.

Testing for a learning signal presents a major challenge. To determine whether a neuron receives feedback related to its contribution to a task, the first step is to know what the contribution actually was. In natural behaviours, this is difficult because success or failure emerges from the activity of vast neural populations interacting with sensory inputs, movements, and internal states. The influence of any single neuron is therefore hard to isolate.

To address this problem, the team at MIT designed an experiment that gave them precise control over the relationship between neural activity and behaviour. They combined a brain–computer interface with two-photon imaging in mice, allowing them to simultaneously manipulate a learning task and record activity from individual neurons.

In this experiment, the mice were trained to control a pattern on a screen using only their brain activity. When the pattern moved toward a target position, the animal received a reward. Unlike most learning experiments, success did not depend on running, pressing a lever, or making any other physical movement. Instead, it depended directly on the activity of a small group of neurons whose signals were continuously monitored by the researchers.

As the activity of these neurons directly determined whether the animal succeeded or failed, the researchers could measure how much each neuron contributed to the outcome. In effect, they created a learning task in which the “credit” assigned to individual neurons could be tracked much more precisely than in natural behaviour.

The researchers first used this setup to estimate each neuron’s contribution to performance. To understand how learning changes the brain, the researchers tracked the same neurons over many days as mice practised the task. They found that learning did not affect all neurons equally. Some neurons that supported the correct response remained active (P+), while others that promoted the wrong response gradually became less active (P-). This selective tuning made the brain’s activity pattern more efficient over time (Figure 2). 

The results suggest that learning works by strengthening useful neural signals and suppressing less helpful ones, creating a sparser and more energy-efficient network. This is similar to how machine-learning systems improve performance by reinforcing the components that contribute to success while reducing the influence of those that do not.

This experiment provided a neuron-by-neuron map of which cells were helping and which were hindering success. With this information, the researchers could ask whether dendrites carry vectorized teaching signals.

If neurons receive vectorized teaching signals, then feedback cannot simply act on the neuron’s overall output. It must interact with the parts of the neuron where learning-related signals are computed. In other words, where in the neuron is this information represented? 

The cell body and axon mainly reflect a neuron’s final output: whether it fires and what signal it sends onward, whereas dendrites are the branching structures that receive inputs from thousands of other neurons.  Dendrites thus provide a promising place to look for feedback related to learning.

To test this idea, the team recorded activity simultaneously from neuronal cell bodies and dendrites. They then estimated how much of the dendritic activity could be explained simply by the neuron’s output and subtracted it away. If any signal remained, it would suggest that dendrites were carrying information beyond the neuron’s final response. Such signals could represent the kind of neuron-specific feedback that theories of learning predict is necessary to solve the credit assignment problem.

They found that the neuron’s output and the residual dendritic signal often occurred at the same time, but their amplitudes varied independently. Sometimes the dendritic activity was much stronger than expected (amplification), and sometimes weaker (attenuation), even when the cell body response was similar (Figure 3).

To test whether these dendritic differences could be predicted from the state of the surrounding network, they looked at the activity of nearby neurons just before each event. They then used a machine-learning method to try to predict whether the dendrite would be amplified or attenuated. 

They found that in many neurons, the activity of nearby neurons could reliably predict whether a dendritic signal would be amplified or suppressed, better than chance. The same network activity also predicted the magnitude of that amplification or suppression. This means the network doesn’t just influence whether dendrites behave differently, but also how strongly they differ.

In other words, the dendrites appeared to contain information that was not simply a reflection of the neuron’s overall output. 

To understand where these additional signals were coming from, the team tested whether the signals could be selectively suppressed. They found that the dendritic signals were strongly reduced under anaesthesia and when a specific class of inhibitory neurons targeting apical dendrites was activated (Figure 4). Importantly, these manipulations affected dendritic signals without simply shutting down the neuron’s output. 

Together, these results suggest that these signals depend on computations occurring within dendrites, rather than merely reflecting activity arriving from elsewhere in the neuron. This is important because it implies that dendrites are not just passive cables carrying information to the cell body. Instead, they may actively compute signals within the neuron itself. 

With this established, the next question was what information these dendritic signals carry during behaviour. The researchers examined dendritic activity during behaviour by simultaneously recording signals from the apical dendrites and the cell bodies of the same neurons. 

They then related these recordings to behaviour by continuously tracking both neural activity and the performance of the brain–computer interface task in real time. Each attempt was labelled as a success or failure depending on whether the cursor reached the target and a reward was delivered.

By comparing dendritic activity during successful and unsuccessful attempts, the researchers could directly test whether these signals varied with behavioural outcome. Their analysis revealed that dendritic activity reliably distinguished successful from unsuccessful trials in real time, increasing or decreasing depending on whether the animal received a reward.

To assess whether these signals were actually used for learning, the researchers again selectively disrupted dendritic processing while animals performed the task. When this was done, the outcome-related signal largely disappeared, and the animals learned much more slowly under the same conditions (Figure 5).

These findings suggest that dendrites carry outcome-related information that can influence future learning 

If this is true, then these signals should not remain static; that is, they should change as the animal learns and its behaviour improves. To explore this, the researchers followed how dendritic activity evolved over the course of training.

They already knew which neurons tended to help the animal succeed and which tended to hinder performance. This allowed them to ask a more precise question: as learning progressed, did dendritic signals shift differently in these two groups?

What they observed was a clear and systematic pattern. As the animal improved and made fewer mistakes, dendritic signals grew stronger in neurons that supported successful performance (Figure 6). When performance deteriorated and errors became more frequent, the pattern reversed, with stronger signals appearing in neurons associated with poorer performance. Moreover, these changes could not be explained by overall neuronal activity and appeared to be specific to dendrites.

The feedback pattern is difficult to reconcile with a single global teaching signal, which would broadcast the same message to all neurons regardless of their role. Instead, the feedback appeared to vary across neurons in a way that depended on whether their activity was helping or hindering behaviour. In effect, neurons contributing positively and negatively to performance seemed to receive different learning-related information.

Interestingly, these signals did not encode the magnitude of error. Instead, they appeared to reflect direction: whether a neuron’s contribution was moving behaviour toward or away from success. This is important because learning does not necessarily require a detailed account of every mistake. What matters is identifying which changes improve performance. In this sense, the dendritic signals look less like error reports and more like instructions guiding improvement.

Taken together, these findings suggest that dendritic signals carry structured information about behavioural outcomes, and that this information is not static but changes systematically as learning unfolds. Rather than simply reflecting whether an action succeeded or failed, these signals appear to track how individual neurons influence performance, and how those influences evolve over time.

If this interpretation is correct, the brain may not rely on a single, global teaching signal after all. Learning could instead be far more personal: each neuron receives its own instruction, shaped by its role in the circuit and expressed through the language of dendrites.

Leave a Reply