Inside the mindset of AI safety workers

Why AI Workers Fear the Technology They Build—and Keep Building It

Episode guide Published Marketplace Tech 12 min

概览

This episode examines the apparent contradiction at the heart of the AI industry: some researchers sincerely believe advanced AI could radically disrupt society or even threaten humanity, yet they continue developing it. Anthropologist Spencer Kaplan argues that these beliefs are not merely public rhetoric; they can shape workers’ careers, finances, and daily lives.

Kaplan connects this mindset to a Bay Area culture that prizes intellectual debate, personal agency, and high-impact action. Within that culture, quitting publicly and remaining inside an AI organization can both be understood as attempts to produce the greatest positive effect.

The discussion also explores effective altruism, its influence on AI safety researchers, and its shift from addressing present-day problems to preventing catastrophic future risks. It concludes by explaining why AI workers assign precise probabilities to highly uncertain outcomes—and why outsiders should not dismiss their concerns as pure science fiction.

分段落总结

[01:05] An anthropological view of AI workers

[事实] Recent warnings about uncontrolled AI have come from people involved in building the technology, including Anthropic employee Jacob Coxon, who publicly resigned.

[事实] Yale PhD candidate Spencer Kaplan spent two years embedded in San Francisco’s AI research community.

[事实] Kaplan says many researchers seriously believe dramatic projections about AI’s future and organize parts of their lives around those beliefs.

[事实] Some workers report saving less money or declining to open retirement accounts because they expect AI to make current financial planning less relevant.

[02:41] Why concerned researchers continue working on AI

[事实] AI researchers themselves debate why they should continue developing a technology they consider potentially dangerous.

[事实] Kaplan identifies “agency” as a guiding value in the Bay Area technology community: people are expected to use their position and available resources to produce positive social impact.

[事实] This approach involves comparing possible actions by the magnitude and probability of their effects, then choosing the option with the highest expected value.

[事实] Some researchers, such as Coxon, leave publicly to influence politics, regulation, and public debate, while others believe staying in the industry offers a better path to impact.

[03:47] The intellectual and social world of AI research

[事实] Kaplan observed AI researchers online, visited their group houses, attended their parties, and participated in community activities during his fieldwork.

[事实] He describes the community as intensely intellectual, with reading groups and salon-style discussions functioning as social events.

[事实] Conversations frequently turn toward philosophical questions about how AI systems developed over the next few years could alter humanity’s future.

[06:23] Effective altruism and its evolution

[事实] Effective altruism began in the United Kingdom and later gained influence internationally, especially among Bay Area technologists and AI researchers.

[事实] Its central principle is to use available resources in ways that create the greatest possible positive impact.

[事实] Early initiatives included strategically funding mosquito nets to prevent malaria.

[事实] As the movement evolved, some participants shifted their attention toward preventing future catastrophes such as nuclear war, pandemics, and uncontrolled AI.

[08:12] How deeply effective altruism influences AI culture

[事实] Kaplan says many—and possibly most—AI safety researchers identify with effective altruism or have been influenced by it, while noting that he cannot speak for every researcher.

[事实] Broader values associated with the movement, including effectiveness, personal agency, and measurable impact, have spread through San Francisco’s technology community.

[事实] Kaplan repeatedly encountered the saying “you could just do things,” which expresses a belief that individuals can turn ideas into reality despite constraints such as institutional inertia.

[推测] Effective altruism’s influence extends beyond formal affiliation because its action-oriented language overlaps with the entrepreneurial norms of the technology sector.

[09:22] What outsiders misunderstand about AI workers

[事实] Kaplan argues that directly working with AI models prompts researchers to question intelligence, humanity, and the technology’s large-scale implications.

[事实] He characterizes the models as mystifying machines whose behavior pushes workers toward broader philosophical and existential questions.

[事实] In his view, AI risk narratives are not purely science fiction or products of internal conversation; they also arise from researchers’ firsthand experiences with the systems.

[事实] Kaplan nevertheless acknowledges that these beliefs can contain a narrative element.

[10:18] How researchers calculate “P(doom)”

[事实] Public anxiety can intensify when knowledgeable AI workers assign numerical probabilities—such as 10%—to human extinction.

[事实] Estimating probabilities for possible future outcomes is a common effective-altruist technique for choosing how to act.

[事实] During an eight-week introductory effective altruism course, Kaplan practiced assigning probabilities to extreme or highly unlikely events.

[事实] He describes these calculations as a form of back-of-the-envelope mathematics informed by substantial thought.

[事实] Kaplan declines to judge the accuracy of any specific extinction estimate but says assigning precise percentages to abstract outcomes is common within this community.

播客点评/总结

The episode’s main strength is its anthropological framing. Instead of deciding whether AI workers’ warnings are correct, it asks how their beliefs are formed, reinforced, and translated into personal and professional choices.

It also offers a useful explanation for an apparent contradiction: resignation and continued employment may represent different responses to the same ethic of maximizing impact. The discussion of effective altruism helps connect individual behavior to a broader intellectual culture.

[推测] The episode is especially useful for listeners seeking to understand the social world behind AI-risk discourse rather than the technical merits of particular forecasts. Its chief limitation is that it does not evaluate the evidence for specific extinction probabilities or include perspectives from researchers outside the community Kaplan studied.