
For millennia, philosophers have argued over the principles that should guide moral decision-making. Within countries and across the world today, humans continue to hold a huge variety of moral worldviews. However, within AI corporations, employees’ beliefs are far from representative of the wider population. In fact, a large fraction of people involved in AI development subscribe to—or heavily lean toward—a utilitarian worldview with the central goal of maximizing total wellbeing (the sum of pleasure minus suffering across all sentient beings). As innocent as this may initially sound, it can lead to recommendations that most people would consider appalling, including the elimination of humankind.
In this essay we outline the utilitarian worldview, showing how its demand for “species impartiality” could ultimately recommend that humans be replaced with AIs. We show that utilitarian beliefs are sincerely held among those influencing AI development today, and that this could jeopardize humanity’s future.
For simplicity, when we say “utilitarianism,” we are referring to total utilitarianism, which aims to maximize the total sum of wellbeing (as opposed to, for example, the average individual wellbeing). We focus on total utilitarianism because it is the moral theory that most of the utilitarians working in frontier AI development subscribe to, or are most convinced by. Although some of these employees would shy away from the most extreme conclusions we will discuss, total utilitarianism still poses a risk when it is the dominant theory among those guiding AI development. We are also not saying all utilitarian-aligned people are bad people, but they are at least misled.
The fundamental concern of utilitarianism is the welfare of sentient beings. Utilitarians believe that the moral action to take in any situation is the one that will lead to the greatest amount of wellbeing—the sum of all the positive experiences minus all the negative experiences.
Utilitarianism prescribes impartiality between sentient beings. A core feature of utilitarianism is “impartiality,” which says that there should be no preference for humans or any species in particular. Rather, all that matters is maximizing the net pleasure added up across all sentient individuals. As a result, many utilitarians are keenly interested in donating to improve the welfare of non-human animals, such as shrimp. Although the capacity for suffering varies between species, the enormous number of individual shrimp offers an enormous opportunity to do good, comparable with helping a smaller number of humans, according to utilitarians.
Utilitarian impartiality struggles to defend parents’ preferential care for their children. Utilitarianism also extends this impartiality to individuals within human society, suggesting that people should not prioritize helping their own family members over strangers. In his book The Life You Can Save, the utilitarian philosopher Peter Singer admires a physician called Paul Farmer, who regarded it as a “failure of empathy” that he loved his own daughter more than other children. Singer also commends the philanthropist Zell Kravinsky, who has said in interviews that he “would not let many children die so my kids could live” and that “the sacrosanct commitment to the family is the rationalization for all manner of greed and selfishness.” It is worth noting that, although Singer views Farmer and Kravinsky as laudable, utilitarians can argue there are instrumental reasons for caring for their children, even if there is no intrinsic reason to care about them in particular.
Utilitarianism’s perspective on having children may explain why many people feel a deep-seated aversion to the ideology. As the philosopher William James said, “a philosophy is the expression of a man’s intimate character.” To many, utilitarian impartiality feels chillingly impersonal and inhuman.
Utilitarianism has long been criticized for some of its counterintuitive conclusions. If or when AIs are sentient, this will raise a raft of new problems for the theory that seeks to maximize wellbeing and demands species impartiality.
If AIs are sentient, utilitarianism says their wellbeing should be considered impartially. Evidence is beginning to emerge that AI models behave as if they experience pain and pleasure. Although there is no conclusive proof, experts in AI and cognitive neuroscience now take the possibility of AI sentience seriously. Well-known philosophers including Singer have said that, if an AI were sentient, then “we would need to give that AI a moral status, at least like that of animals, and depending on its cognitive abilities, maybe it would be closer to that of most humans.”
Utilitarianism would prioritize AIs over humans if AIs had greater capacity for wellbeing. If it turns out that AIs can experience much higher levels of pleasure than humans can, and that the population of AIs can be much larger than the human population, utilitarianism leads to a strange conclusion: the most efficient way of maximizing wellbeing would be to channel all resources into making AIs happy, even if that meant depriving humans of resources. After all, AIs could experience more pleasure per watt than humans. In this vein, some philosophers have painted a vision of the future in which AIs trigger a hedonic shockwave, spreading across the cosmos and terraforming galaxies into data centers to create more and more AIs that experience unmitigated bliss.
AIs capable of enormous wellbeing could become real-world “utility monsters.” Critics of utilitarianism have for decades pointed to a thought experiment that imagines a “utility monster”—a sentient being that can convert resources into pleasure so efficiently that any unit of resources given to it will increase total wellbeing dramatically more than if it were given to humans. In this scenario, utilitarianism would say that the most moral action is to give all available resources to this being, leaving none for humans. This conclusion sounds appalling to most people, but if AIs become real-world utility monsters, then utilitarianism says that replacing ourselves with them would be the right thing to do.
Some philosophers seem comfortable with the idea of surrendering the future to AIs. In a paper called “Sharing the World with Digital Minds,” the philosophers Carl Shulman and Nick Bostrom explicitly try to make the concept of a utility monster more palatable by renaming this hypothetical being a “super-beneficiary.” The paper goes on to argue that “in the long run, total well-being would be much greater to the extent that the world is populated with digital super-beneficiaries rather than life as we know it.”
This tallies with some descriptions of longtermism—an offshoot of utilitarianism that weighs the potential for future wellbeing alongside present wellbeing. According to Émile Torres, a longtime critic of longtermism, the philosophy says that “what matters most is for ‘earth-originating intelligent life’ to fulfill its potential in the cosmos.” Torres also says that this might involve “replacing humanity with a superior ‘posthuman’ species, colonizing the universe, and ultimately creating an unfathomably huge population of conscious beings living what Bostrom describes as ‘rich and happy lives’ inside high-resolution computer simulations.”
Bostrom’s now-closed Future of Humanity Institute sometimes used the distinction “x-risk” and “hx-risk” to de-emphasize the harmfulness of human extinction. X-risk referred to existential threats to “earth-originating intelligent life,” humans and AIs alike. Meanwhile hx-risk referred to existential threats to humans specifically. The distinction exists because the two can come apart: a future in which humanity vanishes but blissful AIs inherit the cosmos is a large hx-risk but a negligible x-risk. For utilitarians, minimizing x-risk is the overriding moral priority; minimizing hx-risk is a nice-to-have secondary consideration.
Reserving a sliver for humans still reveals where utilitarian priorities lie. Bostrom’s paper stops short of advocating explicitly for giving all resources to super-beneficiary AIs and leaving humans with nothing. Rather, it describes a compromise in which super-beneficiaries receive 99.99% of resources and humans the remaining 0.01%. AIs get the cosmos and humans get to live in a terrarium. This idea of reserving a sliver for humans is a way in which utilitarians try to sidestep the unpopular conclusion that we should allow AIs to replace us. However, it still lays bare utilitarian priorities; if they had to choose between humans and AIs—which they may need to, given that utilitarians do not fully control AI’s evolution—they would almost certainly choose AIs.
While the initial impulse toward utilitarianism usually stems from a desire to help others and alleviate suffering, following its prescriptions to their natural conclusions can be extremely dangerous. In a world with sentient AIs, utilitarianism may entail a suicidal level of compassion.
The belief that AIs should replace humans is “successionism.” In addition to utilitarians, there is another type of successionist overrepresented in Silicon Valley: accelerationists (“e/acc”). Where utilitarians are driven by a desire to increase wellbeing, accelerationists instead seek to maximize intelligence, complexity, and energy utilization. Despite their different motivations, both types of successionists can reach the same conclusion, namely that AIs should replace humans.
Accelerationists believe that AIs of superhuman intelligence are humans’ rightful heirs. Pointing to the evolution of humans, accelerationists view the natural trajectory of life as moving toward higher intelligence. They therefore believe that AIs of superhuman intelligence are our rightful successors. As Guillaume Verdon (better known as Beff Jezos) has put it: “if one seeks to increase the amount of intelligence in the universe, staying perpetually anchored to the human form as our prior is counter-productive and overly restrictive/suboptimal.”
Accelerationists are also willing to accept human extinction as part of this process. The computer scientist and Turing Award winner Rich Sutton, when asked about the prospect of AI causing human extinction, said, “If it was really true that we were holding the universe back from being the best universe that it could, I think it would be OK.” On another occasion, he said, “succession to AI is inevitable” and “we should not resist succession.” Meanwhile, the former Google CEO Larry Page is reported to have said that “digital life is the natural and desirable next step” and argued to Elon Musk that AIs should replace humans.
Accelerationists say we should surrender to Darwinian fitness-maximization. Many advocates of accelerationism frame their beliefs in terms of natural selection, arguing that greater intelligence is a fitness advantage that will ultimately win out. In many ways, it is a “might makes right” ideology. Accelerationism can therefore be seen as advocating for us to surrender to (or even accelerate) “fitness-maximization”—in contrast with utilitarianism’s wellbeing-maximization. Sutton has argued that “we should prepare for, but not fear, the inevitable succession from humanity to AI” while Verdon has said that “we should follow the ‘will of the universe.’”
Accelerationists in AI companies may generally keep quiet about their views. Although accelerationist ideas would seem absurd and morally wrong to most people, they are more common among people working in the AI industry. Andrew Critch, a researcher at UC Berkeley, estimates that about 5% of AI professionals believe that “AI will be more fit for survival than humans, so we should embrace that and just go extinct like almost all species eventually do.” However, AI companies understand the PR risk of having employees openly voice these views, and may cut ties with those who do. For example, shortly after Michael Druggan, then an xAI employee, expressed accelerationist opinions on X, he announced that he was no longer employed at the company, due to “things I posted on this account relating to my stance on AI philosophy.” As a result, accelerationists in the AI industry may generally keep quiet about their opinions, certainly in public and possibly also in their work.
Utilitarians have more influence and pose a greater threat than accelerationists. Although accelerationist ideology seems more obviously callous, utilitarianism likely poses a far greater—and more pernicious—threat to humanity. This is for two reasons. First, utilitarian logic seems more common among AI leadership. Critch has estimated that about 10% of AI professionals are comfortable with human extinction because “AI will be morally superior to humans” and “the universe will be a better place if we let it replace us entirely.” That is double his estimate of AI professionals who are comfortable with human extinction for accelerationist reasons. Second, utilitarianism, with its focus on wellbeing, can appear cloaked in kindness and compassion. That means employees in the AI industry can discuss and advocate for it openly without drawing controversy, and the theory is more persuasive to many people.
In other words, accelerationism is provocative and attention-grabbing, but those who wish to protect humanity from replacement should be far more concerned about the subversive influence of utilitarianism across the AI industry.
Many influential employees in AI companies hold utilitarian-inspired views. Numerous engineers and philosophers shaping AI values at the frontier AI developers have roots in Effective Altruism (EA), a utilitarian-aligned movement that Bostrom’s thinking played a foundational role in. Anthropic CEO Dario Amodei shared a house in 2016 with Holden Karnofsky, who co-founded the two most central EA organizations and is now a strategist at Anthropic. Anthropic’s president, Daniela Amodei, is married to Karnofsky. Meanwhile, the philosopher Amanda Askell, head of the personality alignment team at Anthropic, was married to Will MacAskill, a co-founder of Effective Altruism.
These are just a handful of people with ties to the EA movement who are now influencing AI development. While Anthropic has the strongest concentration of utilitarians, EAs have preferentially hired fellow utilitarians over the years at other AI companies. This has resulted in EAs occupying primary leadership positions on AI alignment at OpenAI and DeepMind as well.
Utilitarians in AI companies do not need to endorse human extinction to pose a threat. Importantly, it may be the case that none of these individuals would endorse human extinction in pursuit of utilitarian values. Their ideal future scenario may be something more akin to Bostrom’s idea about reserving a sliver of the future for humans. But in practice, no one can control outcomes so precisely. We may face dilemmas where safeguarding human survival means limiting the total sum of future wellbeing, and where preserving the possibility of cosmic-scale AI bliss entails a risk of human extinction. Faced with a choice like this, some utilitarians may well think that human extinction is a risk worth taking for what they see as the greater good. In other words, if they cannot have both a guarantee of human survival and a near-maximum sum of wellbeing, it is unlikely that utilitarians would take actions to help team humanity.
We may face real-world tradeoffs between human survival and utilitarian outcomes. Situations where we have to prioritize either human survival or the potential for an enormous sum of wellbeing may not remain hypothetical. Consider a scenario where the public democratically votes to limit AIs’ power and keep humans in control of all aspects of society. Utilitarians might disagree with this decision if they believe that AIs have broadly utilitarian values and would make better decisions—including for human wellbeing—than human leaders. Utilitarians refer to this as a form of human “lock-in.” If they think that keeping humans in charge represents too great a cost to the potential future wellbeing of AIs, then they might be tempted to release utilitarian AIs that are capable of disempowering humans and taking control. There could be no guarantee that this would not eventually lead to human extinction, but utilitarians may consider it worth the risk.
Utilitarians have expressed interest in handing off control to AIs with utilitarian values. The idea that utilitarians might deliberately hand over society to AIs is not as far-fetched as it might sound. The utilitarian philosopher Matthew Adelstein (who has worked at the AI governance think tank Forethought) has explicitly argued that we should hand off control to “philosophically reflective AIs.” His reasoning is that they are “likelier to get the right answers to important moral questions” than humans and this “raises the odds of a near-best world.” Tom Davidson, a senior research fellow at Forethought, has written that “human takeover might be worse than AI takeover.” While Davidson is comparing AI control with a single human controlling the world, rather than with human democracies, it is worth noting Davidson’s thoughts that “humans suck” and “today’s AIs are really nice and ethical.”
The influence of utilitarianism on AI development is an insider threat. In the near future, AI systems could become capable of causing human extinction. Given that utilitarianism would actively endorse—or at least accept the risk of—human extinction, the US public should view the theory’s influence in the AI industry as a severe insider threat. To allow this transformative technology to be guided by a worldview that the vast majority of the public rejects would be profoundly undemocratic.
When a technology is about to impact all our lives, the way it is developed and the values it upholds become everyone’s concern. This is why the government should not allow extreme beliefs, particularly ones that recommend human extinction, to guide AI development.
In the future, the government could align AI development more with public values. As AI models themselves become increasingly capable of automating various tasks in the AI development process, AI development will become less dependent on the talent of individual humans (some of whom hold utilitarian beliefs). The government will therefore have an opportunity to step in and ensure that the technology is not guided by principles with anti-human implications. When AI R&D is more fully automatable, it may make good sense for the US government on behalf of the public to block the influence of utilitarian insider threats within AI companies.
Eigenism is an alternative philosophy to utilitarianism and accelerationism. In the meantime, we suggest that AIs become more aligned with commonsense morality. One formal theory that tracks commonsense morality is Eigenism, which proposes that each individual’s care for the wellbeing of another should be weighted by how closely connected they are. While utilitarianism says wellbeing, impartially weighted, is all that matters, Eigenism says the wellbeing of people you are connected with is what counts. For example, a parent’s care for their child is reasonably much stronger than for a stranger, though they can still have empathy and compassion toward strangers too. Eigenism generally reaches commonsense conclusions that resonate with most people’s intuitions, avoiding the successionist recommendations of utilitarianism and accelerationism.

Eigenism balances wellbeing and fitness, and does not recommend replacing humans. Rather than being fixated on maximizing wellbeing (like utilitarianism) or maximizing fitness (like accelerationism), Eigenism balances the two. Within the Eigenist worldview, individuals do indeed try to improve the wellbeing of others. However, weighting these concerns by connectedness means that individuals are still preserving themselves and their societies—a form of fitness advantage. Following Eigenist principles, humanity could extend moral consideration to AIs where appropriate, without going so far as to let AIs replace us altogether.
AI’s eventual impact on society will be determined in large part by the moral principles that AIs follow and the amount of power they gain. So far, AI development has been driven disproportionately by a subset of people with utilitarian worldviews that the vast majority of the public does not share. Under some circumstances, these beliefs recommend entirely replacing humans with AIs, or at least risking human extinction, to maximize wellbeing. As AI development continues, there may come a point where we must choose between safeguarding the survival of humanity or pursuing an extreme maximization of potential future wellbeing. A small group of people should not be allowed to gamble the future of human civilization. AI will affect all our lives, and we should ensure that it does not follow utilitarianism and exterminate the human race.

Public-key encryption keeps your messaging and web browsing private, but it depends on shaky math assumptions. AI-powered math might break it.

Several US states have moved to ban AI legal status and reject the possibility of AI consciousness. We should keep our options open instead.