Abstract
Abstract
Neurons in the brain are often many synapses away from motoneurons, yet if a movement results in error, each distant neuron needs a teacher that considers its specific contribution to production of that movement. This credit assignment problem is solved in machine learning via gradient descent of a loss function, where the loss defines the subjective cost incurred by error. Does the brain use gradient descent to teach individual neurons? We trained marmosets to make saccades to visual targets and then varied the loss by assigning reward value to each target. The climbing fibers, which are the teachers of Purkinje cells (P-cells) in the cerebellum, used a multiplicative encoding to scale the spatial properties of the error vector with its reward properties, incorporating reward prediction errors. Using spike-triggered suppression of P-cells, we quantified the potent vector that mapped each P-cell's output to eye movements and discovered that the climbing fibers were not merely transmitting errors. Rather, they were providing a signal that was, on average, proportional to the dot product of the reward dependent error vector upon the P-cell's potent vector. Thus, the climbing fibers solved the credit assignment problem by providing the gradient of a loss function with respect to the output of their individual P-cells.