Don't hate the replicator, hate the game
Planet Money: The Replication Games and the Replication Crisis
概览
This episode follows economist Abel Brodeur and his effort to address the replication crisis in social science. The central problem is that many published findings do not hold up when researchers rerun the code, inspect the data, or test whether the results survive reasonable alternative choices.
The story begins with Abel’s own experience as a graduate student, when he realized how easy it was to keep adjusting data analysis until a statistically significant result appeared. That experience led him to study publication incentives, p-hacking, and the pressure on academics to produce novel, significant findings.
The episode then moves to Abel’s solution: the Replication Games, a hackathon-style event where teams of researchers spend a day reproducing and stress-testing published papers. The larger idea is that even a small chance of being checked can change researchers’ behavior and push the field toward cleaner, more transparent work.
分段落总结
[00:17] A Trip to Montreal for the Replication Games
[事实] The hosts travel to Montreal to meet Abel Brodeur, an economics professor at the University of Ottawa. [事实] Abel is hosting an event called the Replication Games, where teams try to reproduce recently published social science papers. [事实] The event is framed as part of a broader response to the replication crisis, where older or published research often fails to produce the same results when checked again. [推测] The opening positions the event as both serious academic oversight and an unusually lively, social way to make that oversight happen.
[03:42] The Episode’s Main Question
[事实] The hosts describe science as being in an existential crisis over whether researchers know the things they think they know. [事实] They say the replication crisis began in psychology and spread to medicine and economics. [事实] The episode focuses on how Abel tried to identify what was broken in social science research and build a large, crowdsourced system to keep researchers accountable.
[05:17] Abel’s Smoking Ban Paper
[事实] In 2011, Abel was a master’s student studying whether smoking bans in restaurants and workplaces made people smoke less. [事实] He used public CDC data on smoking prevalence at the county level. [事实] Existing research suggested smoking bans were highly effective, but Abel’s analysis initially found no effect. [事实] Because null results are harder to publish, Abel kept changing his analysis until he found a statistically significant result.
[06:03] The Temptation of Statistical Significance
[事实] The episode explains that a statistically significant result is usually one that would appear by chance less than five percent of the time. [事实] Abel became excited when he finally found a significant result showing that smoking bans reduced smoking. [事实] He later became uncomfortable because the result came from repeatedly reshaping the data rather than from a straightforward test. [事实] He ultimately wrote the paper showing no effect instead of using the tortured result.
[08:13] P-Hacking and Academic Incentives
[事实] Abel learned from other students that adjusting analysis to find publishable results was treated as normal. [事实] The episode identifies this as an incentive problem: academics need journal publications, and journals prefer novel, statistically significant findings. [事实] The practice can cross into p-hacking, where researchers make many analytical choices until a result clears the significance threshold. [推测] The episode suggests that p-hacking can emerge from career pressure and gradual decision-making, not only from deliberate fraud.
[09:21] Researching the Research
[事实] Abel and colleagues scraped significance data from top academic journals. [事实] They found a noticeable cluster of results just above the five percent significance threshold. [事实] The episode says this could reflect researchers not submitting near-miss results, or researchers tweaking analyses until they barely became publishable. [事实] Their paper was initially rejected several times before being published in 2016 under the title “Star Wars: The Empirics Strike Back.”
[10:47] Existing Fixes Were Not Enough
[事实] Some top journals had begun requiring replication packages containing the data and code behind papers. [事实] Researchers were also beginning to pre-register hypotheses before doing research. [事实] Abel wanted to go beyond studying the problem and change research incentives at scale. [事实] He believed small-scale efforts would not shift academic norms.
[11:32] The Clean Apartment Analogy
[事实] Abel compares published research to someone presenting a cleaned-up version of themselves before a date. [事实] He argues that a published paper may look polished, but outsiders cannot easily see how messy the underlying data analysis is. [事实] The episode describes this as an information asymmetry between paper authors and readers. [推测] The analogy supports Abel’s view that transparency alone matters less unless researchers believe someone may inspect the underlying work.
[13:11] Creating the Institute for Replication
[事实] Abel wanted to obtain researchers’ code and data, but individual requests often received no response. [事实] In 2022, he created a website for the Institute for Replication. [事实] He recruited well-known economists to appear on the board, making the institution look legitimate. [事实] The Institute helped him get responses, data sets, and coding packages.
[14:12] The Need for Scale
[事实] Abel concluded that reproducing one paper at a time would not change the system. [事实] He wanted the academic community to feel that anyone’s work could be checked at any time. [事实] The episode compares this idea to an IRS for academia. [推测] The intended mechanism is deterrence: researchers may behave more carefully if they believe inspection is possible.
[14:44] The First Replication Game in Oslo
[事实] Abel was invited to seminars in Oslo and proposed a small workshop for his free day. [事实] After he posted about the workshop online, around 70 to 80 people registered quickly. [事实] He organized attendees into teams by field and collected papers for them to reproduce. [事实] The first Replication Game took place in October 2022.
[16:12] The First Big Error
[事实] In the first Oslo game, one team found major duplicate observations in a paper about inequality. [事实] Abel describes a data set where many entries appeared to be the same people duplicated. [事实] The episode says the underlying data had been merged improperly, like a large copy-and-paste error. [事实] By the end of the day, many other teams found that their papers were clean.
[17:21] Crowdsourced Academic Auditing
[事实] Abel realized the games could crowdsource a large academic auditing project at very low cost. [事实] The episode says enough games each year might pressure social science researchers to behave more carefully. [推测] The model works partly because participants receive training, collaboration, and authorship credit rather than direct payment.
[18:55] How the Montreal Game Works
[事实] The Montreal event is described as an all-day hackathon rather than a competition with winners or prizes. [事实] Teams include mostly economists and some psychologists. [事实] Each team has seven hours to inspect a paper’s replication package, check code, examine author decisions, and report findings. [事实] Results can range from no issue to a serious flaw.
[19:41] Participants and Their Papers
[事实] Education economist Jolene Hunt says PhD students often work in silos, so the event offers a rare chance to work together. [事实] Thibault Dupré’s team looks at a paper about pensions in different countries and considers whether country selection affects results. [事实] Other teams study topics including agriculture, negotiations, education, government trust, and cartel behavior in Mexico.
[21:20] Different Motivations for Replication
[事实] One agriculture researcher says she hopes the paper checks out and expresses sympathy for the original authors. [事实] Felix Fosu, a postdoc at Queen’s University, says his group definitely wants to find something. [事实] Felix argues replication matters because economics needs to know whether results really support their claims. [推测] The episode shows a tension between collegial sympathy and the professional value of finding consequential errors.
[23:03] Phase One: Reproducing the Code
[事实] The first phase is simple replication: teams run the authors’ original code using the replication package. [事实] They check whether the code runs and whether it produces the same tables and results reported in the paper. [事实] Possible problems include broken code, different outputs, or flawed raw data. [事实] The agriculture team successfully reproduces one published result about egg prices, matching the reported number.
[25:23] Phase Two: Robustness Checks
[事实] If the first phase succeeds, teams move to robustness checks. [事实] In this phase, they alter parts of the model or analysis to see whether the original conclusion still makes sense. [事实] The episode says this phase is less objective and requires judgment about what the paper’s authors did or did not test. [事实] Because time is limited, teams can only examine some possible alternative choices.
[26:40] Problems in the Government Trust Paper
[事实] One team examines a paper about whether people who trust government comply more readily with policies. [事实] The team finds that a folder labeled as raw data contains files labeled clean. [事实] After downloading data from the source and following the authors’ instructions, they encounter missing variables. [事实] The episode says the authors appeared to use data in their analysis that was not accounted for in the supposed raw data set.
[28:02] The Cartel Paper’s Fragile Result
[事实] Felix’s team studies a paper about whether Mexican cartels changed the types of crimes they committed after a government war on drugs. [事实] The team finds that removing one cartel makes the results insignificant. [事实] The replicators suspect the paper may not pass the robustness-check phase. [事实] A professor on the team says this is the kind of robustness check he would have tried if he had written the paper.
[29:37] End-of-Day Results in Montreal
[事实] At the end of the game, teams report their findings in short presentations. [事实] Some teams find no major issues, while others report missing variables, attrition, or results that fail robustness checks. [事实] The episode says 71 replicators participated in the Montreal game. [事实] Fourteen teams mainly double-checked published work, while two teams found more serious issues.
[30:47] How Abel Handles Serious Findings
[事实] Abel sends authors a standardized, neutral email explaining who the Institute is and what the replicators found. [事实] Authors are given a chance to respond, fix problems, and prepare a formal response before anything becomes public. [事实] Abel does not assume bad intentions. [事实] The Institute’s role gives junior replicators some insulation from direct conflict with senior researchers.
[31:41] The Cartel Authors Respond
[事实] The authors of the cartel paper were happy that their code replicated. [事实] They were less worried about the robustness criticism because they believed the replicators misunderstood the paper’s core hypothesis. [事实] The authors say they had Los Zetas, one large Mexican cartel, in mind from the start. [事实] They argue that removing Los Zetas is like removing the central object of the study.
[33:41] Ambiguity in the Original Paper
[事实] The hosts say the original cartel paper did not explicitly state that it focused only on Los Zetas. [事实] The Los Zetas data was grouped with several other new cartels. [事实] The episode says that if the authors meant to study only Los Zetas, that focus was not clearly spelled out. [推测] The disagreement shows that replication can expose not only coding errors but also unclear framing and interpretation.
[34:03] A System Failure, Not Just Individual Mistakes
[事实] Abel says that finding problems after publication is not necessarily a success, because a working scientific system should catch them before publication, citation, and dissemination. [事实] The episode notes that papers had already passed journal referees before being replicated. [事实] In the government trust case, journal referees apparently did not catch missing documentation or missing numbers. [事实] Abel says the rate of failures is higher than many people think.
[35:23] What the Games Have Found So Far
[事实] The episode says every Replication Game so far has found something. [事实] The findings have not included career-ending fraud, but have included major data errors, coding errors, and robustness failures. [事实] Abel has held more than 50 games and replicated about 300 papers. [事实] Some participants say the experience will change how they conduct their own research.
[36:00] Changing Behavior Through Enforcement Odds
[事实] Abel argues that people change behavior based less on the severity of punishment and more on the odds of actually being caught. [事实] The episode ends by returning to the clean-apartment analogy: the possibility that someone might inspect the work may be enough to make researchers keep it cleaner. [推测] The Replication Games are presented as a practical deterrence system rather than a complete solution to the replication crisis.
播客点评/总结
[推测] This episode is valuable because it turns an abstract methodological problem into a concrete story about incentives, career pressure, code, data, and institutional design. The strongest part is the reporting from inside the Montreal Replication Game, where the listener can hear how replication work actually unfolds.
[推测] The episode is also careful not to reduce the replication crisis to simple misconduct. It shows that errors, ambiguous hypotheses, missing documentation, and fragile results can emerge from ordinary academic processes, not only from intentional manipulation.
[推测] A limitation is that the episode focuses mainly on Abel’s project and selected cases from one event, so it does not fully compare the Replication Games with other reform efforts across science. Still, it is well suited for listeners interested in economics, social science, research integrity, statistics, and how institutions can change professional behavior.