One Uncorrected Attrition Log Split a Classic Social Belonging Intervention
In the early 2010s, a brief writing exercise designed to bolster a sense of social belonging among college students seemed to produce a remarkable result: African American students who completed the exercise earned higher grades than their peers who did not. The study, led by Gregory Walton at Stanford University, became a touchstone for a generation of psychological interventions aimed at closing achievement gaps. But as replication attempts accumulated over the following decade, the effect proved elusive. Then, in 2022, a meta-analysis in Nature Human Behaviour turned attention to a detail that had been hiding in plain sight: an attrition log. One contributing study had lost more control participants than treatment participants, and how that missing data was handled—or not handled—became a flashpoint for a broader debate about rigor in behavioral science.
A Classic Intervention, a Replication Crisis, and an Attrition Log
The social-belonging intervention, pioneered by Walton and his colleague Geoffrey Cohen, was deceptively simple. Incoming college students read survey results suggesting that most students worry about belonging during their first year, and that these worries fade with time. The students then wrote an essay describing how their own experiences mirrored this pattern, and recorded a video message for future students. The goal was to normalize adversity and reframe it as temporary and shared. The entire exercise typically took about one hour.
The original 2011 study, published in Science, reported that African American students who completed the exercise earned a grade point average roughly 0.25 points higher over the next three years compared to a control group. For a one-hour activity, that effect size was striking. The sample was modest—92 students at a selective university—but the result resonated widely. Schools across the United States adopted versions of the intervention.
Yet when other labs tried to replicate the finding, results were mixed. Some found smaller effects; others found none. A large-scale replication project, Many Labs 3, published in 2019, tested the intervention across 20 sites and more than 4,000 participants. The overall effect shrank to about 0.07 GPA points, and heterogeneity across sites was high. The intervention seemed to work in some contexts but not others, and the reasons were unclear.
Enter the attrition log. In 2022, a team led by Paul T. von Hippel at the University of Texas at Austin published a meta-analysis in Nature Human Behaviour that re-examined the original studies. They noticed that one of the contributing experiments—a replication conducted at a different university—had a striking pattern: 28% of control participants dropped out of the study, compared to only 12% in the treatment group. Differential attrition can bias results if the reasons for dropout are related to the outcome. If struggling students are more likely to drop out, and they are unevenly distributed across conditions, the apparent treatment effect may be inflated. The authors of the original work pushed back. In a correspondence published in PNAS, they argued that the attrition coding was flawed—some participants classified as dropouts had simply not completed follow-up surveys but still had GPA data available. The dispute highlighted a deeper issue: attrition logs are often incomplete, and the rules for handling missing data are rarely pre-registered or standardized.
The Original Effect Size and Its Contested Ground
The 0.25 GPA boost reported by Walton and Cohen in 2011 was large enough to attract attention—and skepticism. For a brief writing exercise to produce an effect comparable to reducing class size or increasing instructional time seemed implausible to some. The study's small sample size (92 students) meant that the confidence interval around the effect was wide, and the result was fragile: removing a few participants could change the conclusion.
Replication attempts often failed to match that magnitude. A 2014 study by Walton and colleagues at a different university found a smaller effect, and a 2016 replication at an elite private university found no significant difference. By 2019, the Many Labs 3 consortium had pooled data from over 4,000 students and found an average effect of roughly 0.07 GPA points—a third of the original estimate. The heterogeneity across sites was substantial, with some sites showing positive effects and others showing null or even negative trends.
Why the variation? One possibility is that the intervention works best in contexts where belonging concerns are acute—for instance, at highly selective institutions where minority students feel particularly isolated. Another is that the delivery method matters: some replications used online modules instead of in-person sessions, or different prompts. But a third possibility, raised by the attrition analysis, is that differential dropout artificially inflated the original estimates.
The debate over effect size is not merely academic. School administrators and policymakers have invested resources in belonging interventions based on the promise of large effects. If the true effect is smaller and more context-dependent, the cost-benefit calculus shifts. Understanding what drives the variation is essential for deciding where—and whether—to implement these programs.
Where the Attrition Log Entered the Argument
The 2022 meta-analysis by von Hippel and colleagues did not set out to target the belonging intervention specifically. The team was interested in the broader problem of attrition bias in randomized experiments. They compiled a dataset of 35 studies from the social-belonging literature and found that differential attrition was common: in several studies, dropout rates differed by more than 10 percentage points between conditions. When they applied statistical corrections for missing data, the overall effect estimate shrank.
One study stood out. A replication conducted at a large public university had a control-group dropout rate of 28% versus 12% in the treatment group. The authors of that study had not reported attrition by condition in their original paper; the discrepancy only emerged when von Hippel's team requested the raw data. The log showed that many control participants had stopped responding to surveys, but their GPA data—the primary outcome—was still available from university records. The question was whether to include those participants in the analysis.
The original authors argued that they should be included, because GPA data were available regardless of survey completion. Von Hippel's team countered that the missing survey data could indicate disengagement, which might correlate with academic performance. Without knowing why participants stopped responding, the safest approach was to treat the missing data as potentially informative and conduct sensitivity analyses.
The correspondence in PNAS in 2023 laid out the disagreement in detail. Walton and Cohen pointed out that the attrition classification used by von Hippel was not the same as the one used in the original studies, and that the effect remained significant under several alternative coding schemes. Von Hippel responded that the key issue was not the coding but the lack of transparency: without a pre-registered attrition plan, researchers could choose the coding that best supported their hypothesis.
How Researchers Disagree on Handling Dropouts
The attrition log controversy is a microcosm of a larger methodological debate. In randomized experiments, the gold standard is intent-to-treat (ITT) analysis, which includes all participants regardless of whether they completed the intervention. But ITT can be misleading if dropout is differential. Per-protocol analysis, which includes only those who completed the study, can also be biased if completers differ from dropouts.
Missing data can be classified into three types: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). MCAR is rare in practice. MAR is the assumption behind most multiple imputation methods. MNAR is the hardest to handle, because the missingness itself is related to the outcome—for example, if struggling students are more likely to drop out. In the belonging intervention, if control participants who were struggling academically were more likely to stop responding, their absence could inflate the control group's average GPA, making the treatment effect appear larger.
Sensitivity analyses can test how robust results are to different assumptions about missing data. For instance, researchers can impute missing outcomes under various scenarios—assuming dropouts have worse outcomes, or better outcomes—and see if the conclusion changes. But such analyses are rarely reported in original papers. A 2020 survey of psychology experiments found that fewer than 10% reported any sensitivity analysis for attrition.
Pre-registration of attrition rules is also uncommon. Without a pre-specified plan, researchers may be tempted to choose the analysis that yields the most favorable result. The belonging intervention debate has prompted calls for journals to require attrition flow diagrams, similar to those used in clinical trials, and for authors to pre-register their handling of missing data.
What Subsequent Large-Scale Replications Found
The Many Labs 3 project, published in 2019 in Social Psychological and Personality Science, was one of the largest replication efforts in psychology. It included the social-belonging intervention as one of its target studies, with 20 participating labs and over 4,000 participants. The overall effect on GPA was small—about 0.07 points—and not statistically significant after correcting for multiple comparisons. But the heterogeneity was striking: some sites found effects as large as 0.3 GPA points, while others found negative effects.
Why such variation? The Many Labs 3 team examined several moderators, including campus climate, demographic composition, and delivery format. None consistently explained the differences. One possibility is that the intervention's effectiveness depends on the specific concerns of the student population, which may vary across institutions and over time. Another is that the implementation fidelity varied: some sites may have delivered the exercise with more enthusiasm or in a more supportive setting.
A subsequent meta-analysis by the original authors, published in 2020 in Educational Researcher, included 14 studies and found an average effect of about 0.11 GPA points for underrepresented minority students. That estimate is smaller than the original 0.25 but still meaningful. However, the meta-analysis included several studies from the original lab, which may have introduced bias. Independent replications tended to show smaller effects.
The open science collaboration that produced Many Labs 3 also highlighted the importance of sharing raw data. Without access to the original data, the attrition log would never have been discovered. The project's data are publicly available, allowing other researchers to conduct their own analyses and test alternative assumptions.
Methodological Lessons for Belonging Interventions
The belonging intervention saga offers several lessons for behavioral science. First, effect sizes from small studies should be interpreted cautiously. The original 2011 study had a sample of 92 students, which is sufficient to detect large effects but not small or moderate ones. The confidence interval around the 0.25 estimate was wide, and the result was sensitive to the inclusion or exclusion of a few participants.
Second, context matters. The intervention may work well at elite, predominantly white institutions where minority students feel particularly isolated, but less so at diverse or less selective schools. Researchers should measure and report contextual factors—such as campus climate, demographic composition, and baseline belonging—to help understand when the intervention is likely to be effective.
Third, attrition logs should be public and coded blind to condition. Researchers should report flow diagrams showing how many participants were randomized, how many completed each follow-up, and why they dropped out. Ideally, attrition coding should be done by someone unaware of the condition assignment, to avoid unconscious bias.
Fourth, pre-registration of analysis plans should include rules for handling missing data. Researchers should specify whether they will use ITT, per-protocol, or some other approach, and what sensitivity analyses they will conduct. This prevents post-hoc decisions that could favor the desired result.
Finally, replication efforts must share raw data. The attrition log controversy would not have emerged without open data. Journals and funders should mandate data sharing as a condition of publication, with appropriate privacy protections.
Practical Takeaways for Researchers and Reviewers
For researchers designing randomized experiments, the first step is to minimize attrition. This means keeping follow-up surveys short, offering incentives, and maintaining contact with participants. But even with best efforts, some dropout is inevitable. For instance, in a typical longitudinal study in education, attrition rates of 10–20% per wave are common, and differential attrition of 5–10 percentage points can bias estimates if not addressed. The key is to document it thoroughly.
Report a flow diagram with reasons for dropout, as recommended by the CONSORT statement for clinical trials. Include the number of participants randomized, the number who completed each assessment, and the number with missing outcome data. If possible, compare dropouts and completers on baseline characteristics to assess whether attrition is selective.
Conduct multiple imputation to handle missing data under the MAR assumption, and report sensitivity analyses under MNAR assumptions. For example, researchers can assume that dropouts have outcomes that are 0.2 standard deviations worse than completers, and see if the conclusion changes. If the result is robust, confidence increases.
Editors and reviewers should require attrition disclosure as a condition of publication. Many journals already do, but enforcement is uneven. A simple checklist at submission could help: Did the authors report attrition by condition? Did they conduct sensitivity analyses? Did they pre-register their attrition plan?
Replication efforts must share raw data, as the Many Labs 3 project did. This allows other researchers to re-analyze the data with different assumptions and check for errors. The attrition log that sparked the debate over the belonging intervention was not a scandal—it was a diagnostic tool. It revealed a vulnerability in the evidence base that could be addressed through better methods.
Looking Ahead: Toward Better Attrition Reporting
The debate over the social-belonging intervention has not ended, but it has catalyzed important changes. Several journals now require attrition flow diagrams as part of their submission guidelines. Pre-registration templates often include a section for handling missing data. And funding agencies are increasingly mandating data sharing for large-scale projects.
Yet challenges remain. Many researchers still do not report attrition by condition, and sensitivity analyses are rare. The incentives for novelty and positive results can discourage thorough reporting. To shift norms, journals could adopt a badge system for transparency, similar to the Open Science Framework badges for data sharing and pre-registration. Reviewers could be trained to check for attrition disclosure during the review process.
Another promising development is the use of automated tools to detect attrition patterns. For instance, the “attrition-check” R package allows researchers to quickly identify differential attrition in their own data and generate sensitivity analyses. Such tools lower the barrier to good practice.
Ultimately, the story of the social-belonging intervention is not one of fraud or incompetence. It is a story of how science progresses through disagreement and self-correction. The attrition log, once uncorrected, is now a case study in the importance of transparency. As one researcher put it, “The log is not the enemy; the enemy is the assumption that missing data doesn’t matter.” Moving forward, the field must ensure that every study includes a clear, pre-registered plan for handling dropouts—so that the next controversy is averted before it begins.