One Unreported Reward Schedule Parameter Fractured a Dopamine Prediction Error Model

Jul 9, 2026 By Jonas Eriksen

Wolfram Schultz and colleagues recorded from midbrain dopamine neurons in monkeys in 1997, reporting a pattern they called reward prediction error. The cells fired vigorously when a reward arrived unexpectedly, remained silent when a predicted reward was omitted, and showed no response to fully predicted rewards. This pattern, termed the reward prediction error, became one of the most celebrated findings in neuroscience. It fit neatly into temporal-difference learning models borrowed from computer science, and it offered a clean, testable account of how animals learn from outcomes.

But within a few years, a small methodological detail began to fray the edges of that elegant picture. Laboratories that tried to replicate the effect found that, under certain conditions, dopamine neurons behaved in ways the model could not explain. The culprit was not a flaw in the theory of prediction error itself, but a parameter that nearly every early report had left unspecified: the reward schedule structure. Whether a given probability was presented in long, blocked runs or in rapidly interleaved sequences turned out to matter enormously. And because Schultz's original experiments had used blocked schedules, the field had inadvertently adopted a model that worked only half the time.

The Reward Prediction Error That Wouldn't Fit

Schultz's 1997 paper in Science reported that midbrain dopamine neurons in macaque monkeys showed a phasic response to unexpected liquid rewards. The firing rate increased when the reward was larger than predicted, decreased when it was smaller, and stayed flat when it matched the prediction. This was exactly the signal that temporal-difference reinforcement learning models required to update value estimates. The match was so precise that many researchers took it as direct evidence that the brain implements a form of TD learning.

Schultz's team trained monkeys on a Pavlovian conditioning paradigm in which a visual cue preceded a reward with a fixed probability. The probability was held constant over many trials before being changed. This is called a blocked schedule. In such a stable environment, the monkey's expectation converges to the true probability, and the prediction error signal behaves as the model predicts.

Yet when other labs attempted to extend the finding to more complex tasks, they encountered anomalies. A 2003 study by A. David Redish and colleagues, using a similar recording setup but with a different trial structure, reported that dopamine neurons sometimes showed a positive prediction error when a reward was fully expected, and a negative error when an unexpected reward was omitted. These results did not fit the canonical model. Redish's group noted that their task used an interleaved schedule, in which reward probability changed every few trials without warning.

The discrepancy was initially attributed to differences in species, recording sites, or task complexity. Over the next several years, a pattern emerged: blocked schedules consistently produced clean prediction error signals; interleaved schedules produced noisy, sometimes paradoxical responses. The parameter that had been unreported in most early papers—whether the reward probability was blocked or interleaved—turned out to be the key variable.

How a Single Schedule Parameter Slipped Past Peer Review

Early reinforcement learning models assumed that all probabilities are equally learnable, regardless of how they are presented. The learning algorithm updates the value estimate on each trial based on the prediction error, and over time the estimate converges to the true expected value. This convergence holds whether the probability is presented in a blocked block of 100 trials or in an interleaved sequence that switches every 10 trials. The model does not care about the history of probabilities; it only cares about the current trial's outcome.

Schultz's original experiments used blocked schedules because they were simpler to train and interpret. The monkey experienced a single probability for dozens or hundreds of trials before the condition changed. This gave the animal ample time to learn the exact expected value, and the dopamine signal reflected that stable expectation. The model fit beautifully, and the field adopted the blocked schedule as the default.

But the blocked schedule is not the only way to present probabilities. In many real-world learning situations, the environment changes unpredictably. A forager might encounter a patch where food is abundant for a few minutes, then switch to a barren patch. An investor sees market regimes shift without notice. Interleaved schedules capture this volatility better than blocked ones. When Redish and colleagues used an interleaved schedule in 2003, they found that dopamine neurons responded not just to the current trial's reward, but to the recent history of rewards and the inferred likelihood that the block had switched.

Why did this parameter slip past peer review? Because the early papers did not consider schedule structure as a variable. The methods sections described the number of trials and the probabilities used, but rarely mentioned whether those probabilities were blocked or interleaved. Reviewers did not ask, because the prevailing assumption was that schedule structure did not matter. It was a hidden parameter that fractured the model from within.

By the mid-2000s, several groups had replicated the anomaly. Yael Niv and Geoffrey Schoenbaum published a formal analysis in 2008 that highlighted the discrepancy. They showed that in blocked environments, animals can compute exact expected values because the context remains stable. In interleaved environments, the animal must infer which block it is in, and this inference introduces a layer of uncertainty that the simple prediction error model cannot capture.

The Blocked vs. Interleaved Distinction That Split the Field

In a blocked schedule, the reward probability is constant for many trials. The optimal strategy is to learn that probability through trial and error, and then use it to predict future rewards. The dopamine prediction error signal, as measured by Schultz, reflects the difference between the actual reward and this learned expectation. The signal is clean because the expectation is stable.

In an interleaved schedule, the reward probability changes every few trials, and the animal does not know when a change occurs. The optimal strategy is to infer the current block's probability based on recent outcomes and to update that inference rapidly when the outcomes shift. This inference problem is computationally more complex. It requires tracking not just the expected reward, but the probability that the environment has changed. The dopamine signal in this setting reflects not a simple prediction error, but a mixture of prediction error and surprise about the block transition.

Niv and Schoenbaum (2008) formalized this by showing that the prediction error in an interleaved schedule depends on the animal's belief about the current block, which itself depends on the hazard rate of block switches. If the animal expects blocks to last a long time, it will be slow to update its belief after a change. If it expects blocks to be short, it will update quickly. The dopamine response, they argued, reflects this belief-updating process, not just the raw prediction error.

Further evidence came from studies that directly compared blocked and interleaved versions of the same task. In 2010, a group led by Matthew Roesch recorded from orbitofrontal cortex and found that neurons there encoded the inferred block identity, not just the current reward. The orbitofrontal signal was necessary for the dopamine neurons to produce a clean prediction error. Without it, the dopamine signal became noisy and inconsistent. This suggested that the prediction error model was not wrong, but incomplete: it assumed a stable context that the brain must infer from the schedule structure.

The field split. Some researchers continued to use blocked schedules and defended the original model. Others argued that the model must be extended to account for interleaved data. The debate was not about whether dopamine encodes prediction error—that much was clear—but about what exactly the brain treats as the prediction. Is it the expected value of reward in the current trial, or the expected value conditional on the inferred block? The answer seemed to depend on the schedule parameter.

Instrumental Drift: How Task Structure Alters Dopamine Encoding

By the late 2000s, researchers had begun to realize that dopamine neurons do not simply encode a static prediction error; they also encode the volatility of the environment. In a 2007 study, Timothy Behrens and colleagues used functional magnetic resonance imaging (fMRI) in humans to track signals related to the volatility of reward probabilities. They found that the brain's learning rate increased when the environment was more volatile, and that this modulation was correlated with activity in the anterior cingulate cortex and the noradrenergic system.

Behrens' work suggested that the brain maintains a separate estimate of the volatility—how fast the reward probabilities are changing—and uses that estimate to adjust the learning rate. This is exactly what the interleaved schedule manipulates. When blocks switch frequently, the volatility is high, and the brain must update its expectations quickly. When blocks are long, volatility is low, and the brain can rely on stable estimates.

Schultz's original model assumed stationary statistics: the reward probability does not change over time. In a blocked schedule, this assumption is approximately true, because the probability stays constant for many trials. But in an interleaved schedule, the assumption is violated. The animal's brain must detect and adapt to the changes. The dopamine prediction error signal therefore reflects not just the discrepancy between reward and expectation, but also the surprise of a block transition.

This insight led to a re-examination of older data. Some of the anomalous results from interleaved studies could be explained if the dopamine signal included a component that encoded the probability of a block switch. For example, a positive prediction error on a fully expected reward might occur if the animal inferred that the block had just switched to a higher-probability state. The reward was expected given the old block, but surprising given the new inference. The dopamine response reflected this inference, not the raw prediction error.

The implication was clear: the schedule parameter—blocked vs. interleaved—determines whether the simple prediction error model holds. Researchers who used blocked schedules saw clean signals; those who used interleaved schedules saw messy ones. The parameter had been hiding in plain sight, unreported in methods sections, because no one thought it mattered.

The Hidden Variable That Rescued the Model

In 2012, Samuel Gershman and Nathaniel Daw published a computational model that resolved the discrepancy. They introduced the concept of latent state inference: the animal does not know which block it is in, but it can infer the block identity from the sequence of outcomes. The dopamine prediction error, they argued, is computed relative to the expected value of the inferred block, not relative to the true block. This introduces a free parameter: the hazard rate of block switches—the probability that the block changes on any given trial.

With this single additional parameter, the model could fit both blocked and interleaved data. In blocked schedules, the hazard rate is low, so the animal quickly converges on the true block identity and the prediction error behaves as Schultz described. In interleaved schedules, the hazard rate is high, so the animal's inference lags behind the actual block changes, producing the paradoxical signals that had been observed. The parameter that fractured the model also rescued it.

Gershman and Daw's model made several testable predictions. First, the prediction error should depend on the recent history of rewards, not just the current trial. Second, the effect should be stronger when the hazard rate is higher. Third, animals should show behavioral evidence of block inference, such as slower learning after a block switch. All three predictions were confirmed in subsequent experiments, including a 2014 study by Gershman and colleagues that used a two-armed bandit task with hidden block transitions.

Gershman and Daw's model did not discard Schultz's original insight but embedded it in a richer framework that acknowledged the brain's need to infer the structure of the environment. The prediction error signal remained central, but its reference point was no longer the true expected value; it was the expected value of the inferred latent state. This shift reconciled the blocked and interleaved data, and it highlighted the importance of a parameter that had been overlooked for over a decade.

The hazard rate parameter has practical consequences for experimental design. If a study uses an interleaved schedule without specifying the block length or the transition probability, the results may be uninterpretable. Two labs using different block lengths could obtain different dopamine responses even if the reward probabilities are identical. The parameter must be reported and controlled.

What This Means for Every Dopamine Experiment Since

Most rodent and primate studies of dopamine function continue to use blocked schedules, because they are easier to implement and produce clean signals. But the results from these studies may not generalize to the interleaved, volatile environments that animals encounter in the wild. The dopamine system evolved to handle change, not stationarity.

Human fMRI studies of reward prediction error often use interleaved or mixed designs, because they are more efficient for estimating blood-oxygen-level-dependent (BOLD) responses. These studies may be inadvertently measuring not just prediction error, but also block-inference signals. Re-analyses of older fMRI data could reveal that some of the reported effects are confounded by schedule structure. The field may need to revisit decades of work with this parameter in mind.

Pre-registration of schedule parameters is now recommended by several funding agencies. The Open Science Framework guidelines for neuroeconomic tasks include a checklist that asks researchers to specify the block length, the number of blocks, and the transition probability between blocks. This ensures that the hidden parameter is no longer hidden. It also allows meta-analyses to control for schedule effects across studies.

The lesson is not that the prediction error model was wrong, but that it was under-specified. Every model has hidden assumptions, and the reward schedule parameter was one of them. By making it explicit, the field has gained a more nuanced understanding of how dopamine supports learning in dynamic environments. The model is stronger for having been fractured and repaired.

Lessons for Designing Reproducible Neuroeconomic Tasks

First, specify the block length and transition probability upfront in the methods section. Even if the schedule is blocked, report the number of trials per block and the total number of blocks. This allows others to assess whether the schedule structure could influence the results.

Second, pilot both blocked and interleaved versions of the same task. If the prediction error signals differ between the two versions, that is valuable information. It tells you that the task engages latent state inference, and that your computational model should account for it. If the signals are similar, you have evidence that the schedule parameter does not matter for your specific task.

Third, report whether prediction error signals are schedule-dependent. In publications, include a supplementary analysis that compares blocked and interleaved conditions, or at least acknowledges the potential confound. This transparency will help the field accumulate evidence about when schedule structure matters and when it does not.

Fourth, consider using computational models that infer hidden states. The Gershman and Daw (2012) model is one example, but there are others. These models allow you to estimate the hazard rate parameter from the behavioral data, providing a direct test of whether animals are inferring block transitions. This approach can be more informative than simply comparing averaged firing rates across conditions.

Finally, the parameter list for any dopamine study must include the schedule. This seems obvious in retrospect, but it was not obvious in 1997. The history of science is full of such oversights. The reward schedule parameter fractured a model that had seemed unshakeable, but it also deepened our understanding of how the brain learns from a changing world. That is how science progresses: not by avoiding fractures, but by examining them closely and rebuilding.

Recommend Posts
Science

One Uncosted Mirror Alignment Jig Fractured a Billion-Pixel Sky Survey

By Alice Chen/Jul 9, 2026

How a single uncosted mirror alignment jig degraded a billion-pixel sky survey, costing half its resolution and years of delay. A tale of fixed-price contracts and corner-cutting in big science.
Science

One Unrecorded Atmospheric Seeing Monitor Drift Collapsed a Transiting Exoplanet Radius Measurement

By Renu Shah/Jul 9, 2026

A missing atmospheric seeing monitor inflated the radius of exoplanet WASP-76b by ~15%. This piece explores how such systematic errors creep into transit photometry and what the field is doing about it.
Science

One Unfrozen Atmospheric Reanalysis Grid Stretched a Decade of Storm Tracking

By Alice Chen/Jul 9, 2026

A subtle grid freeze in the ERA5 reanalysis led to systematic storm-count biases. Researchers found 11% fewer cyclones in one version, reversing trends and highlighting infrastructure fragility.
Science

One Uncaptured Laboratory Social Desirability Prompt Bent a Cooperation Game Replication

By Karim Osman/Jul 9, 2026

A single added sentence—'Please be honest'—may have inflated cooperation rates in a classic economic game replication from 50% to 80%, revealing how unnoticed wording changes can distort findings.
Science

One Misaligned fMRI Voxel Size Selection Fractured a Working Memory Localization Model

By Jonas Eriksen/Jul 9, 2026

How a seemingly trivial choice of fMRI voxel size—3 mm instead of 2 mm—obscured submillimeter functional columns in the prefrontal cortex, leading to a decade of conflicting results about working memory localization.
Science

How a Behavioral Nudge for Organ Donation Moved into Public Health Policy

By Jonas Eriksen/Jul 9, 2026

How a simple opt-out nudge for organ donation, rooted in behavioral science, moved from academic labs into public health policy worldwide, saving thousands of lives.
Science

One Unreported Holographic Grating Polarization Bias Skewed a Dark Energy Survey Shear Calibration

By Renu Shah/Jul 9, 2026

A subtle polarization bias from the Dark Energy Survey's holographic grating introduced a 0.5–1% shear calibration error, mimicking an additive signal. New corrections reduce the bias below 0.1%, with lessons for LSST and Euclid.
Science

One Unreported Crystal Growth Flux Ratio Bent a Topological Superconductor Gap Map

By Alice Chen/Jul 9, 2026

A hidden variable in crystal growth—the flux ratio—was found to bend the superconducting gap map of Sr2RuO4, reshaping the phase diagram and prompting new reporting standards.
Science

One Unreported Rodent Light-Dark Cycle Shift Inflated a Fear Conditioning Meta-Analysis

By Karim Osman/Jul 9, 2026

A single lab's accidental reversal of the light-dark cycle during rodent fear conditioning experiments inflated effect sizes in a meta-analysis, raising questions about circadian confounds in preclinical neuroscience.
Science

How a Fluid Dynamics Code Mapped Neural Activity Across a Mouse Visual Cortex

By Jonas Eriksen/Jul 9, 2026

A fluid dynamics code originally designed for pipe flow was repurposed to model neural activity in the mouse visual cortex, revealing traveling waves and feedback loops with 87% accuracy.
Science

One Missing Fringe-Phase Calibration Thread Bent a LIGO Noise Budget

By Jonas Eriksen/Jul 9, 2026

A single overlooked fringe-phase calibration thread bent LIGO's noise budget for two observing runs. The fix cost $2M and saved 15% of observing time, exposing deep flaws in how large-scale science funds noise debugging.
Science

One Unversioned Mesh Refinement Parameter Broke a Turbulence Simulation Replication

By Alice Chen/Jul 9, 2026

A 40% discrepancy in a turbulence simulation replication traced to a single unversioned mesh refinement parameter. The episode exposes gaps in computational reproducibility.
Science

One Unreported Rat Chow Selenium Lot Shift Inflated a Thyroid Hormone Study

By Karim Osman/Jul 9, 2026

A mid-experiment selenium lot shift in rat chow inflated a thyroid hormone study. The retraction exposes a blind spot in model organism infrastructure and the economics of replication.
Science

One Unversioned Solver Tolerance Parameter Bent a Climate Model Ensemble

By Renu Shah/Jul 9, 2026

A single unrecorded solver tolerance parameter shifted a climate ensemble's spread by 10-15%. This methodology piece traces the root cause and what it means for reproducible science.
Science

How a Fluid Dynamics Code Solved a Solid-State Electron Flow Mystery

By Alice Chen/Jul 9, 2026

A fluid dynamics algorithm originally built for turbulence now simulates electron transport in quantum dots with 5% error, revealing vortices and interference patterns that classical models missed.
Science

How One Underpowered Nudge Replication Fractured a Cooperation Theory

By Jonas Eriksen/Jul 9, 2026

A landmark 2008 study on eye-like cues boosting cooperation failed to replicate in a massive multi-lab project. The fracture exposed deep methodological flaws and reshaped behavioral science.
Science

One Uncorrected Attrition Log Split a Classic Social Belonging Intervention

By Alice Chen/Jul 9, 2026

How a single uncorrected attrition log fueled debate over a classic social belonging intervention, revealing deeper issues in handling dropouts in behavioral science.
Science

One Undocumented Spectrograph Temperature Drift Split a Galactic Archeology Collaboration

By Alice Chen/Jul 9, 2026

A few millikelvin of thermal drift in a spectrograph fiber feed caused two subgroups to disagree on correction methods, delaying a galactic archeology catalog and splitting the collaboration.
Science

One Unreported Reward Schedule Parameter Fractured a Dopamine Prediction Error Model

By Jonas Eriksen/Jul 9, 2026

How a single unreported parameter—whether reward probabilities were blocked or interleaved—fractured the canonical dopamine prediction error model, revealing hidden assumptions in decades of neuroscience research.
Science

One Unreported Quartz Sample Etch Protocol Split a Luminescence Dating Standard

By Karim Osman/Jul 9, 2026

A hidden variation in quartz etching protocols caused 15–20% age offsets across luminescence dating labs. The discovery reshaped how geochronologists document sample preparation.