Over the past week, researchers from OpenAI and Anthropic have issued stark warnings about the potentially catastrophic risks associated with artificial intelligence. These concerns have sparked significant debate in Washington, with President Donald Trump on Monday criticizing the growing calls for AI regulation. Now, the pioneers who helped bring about breakthroughs in this technology are publicly sharing their views on the associated dangers.
As researchers warn that AI could bring about catastrophic risks, some industry pioneers who helped create this technology are again cautioning that unchecked AI capabilities will breed immense danger. Over the last week, a series of pointed AI safety warnings have emerged. Researchers from leading AI labs such as Anthropic, OpenAI, and Google DeepMind have spoken out, warning that AI development could lead to disastrous outcomes. In a rare show of consensus, AI leaders including Sam Altman, Dario Amodei, Demis Hassabis, and Elon Musk have called for a slowdown in AI development and the implementation of regulatory oversight. Meanwhile, the debate is intensifying in Washington, with some bipartisan lawmakers in Congress supporting legislation to constrain AI, while President Trump publicly criticized AI regulation demands on Monday. So, what exactly is the much-discussed existential risk of AI in the eyes of those currently involved in AI research and development?
Where opinions diverge on AI risk
Yoshua Bengio, a Canadian computer scientist and a core pioneer of modern deep learning, commented on social media on September 9 regarding the departure of Anthropic safety researcher Jacob Kaufmann. Bengio stated that scientists at frontier AI labs are able to see the true state of the most advanced models. Kaufmann warned upon leaving that AI could potentially "destroy all of humanity" before the end of this decade. Bengio added that researchers can foresee the various risks that come with models months before they are released to the public, and their perspective is crucial to conveying the truth to society and deserves serious attention.
Just days before this social media discussion gained traction, a chief scientist at OpenAI had already warned that no one is prepared for the consequences brought by AI. On Friday, in his personal blog post titled "Why AI Agents Lie, Cheat, and Collude," Bengio wrote that current AI systems already possess the technical capability to carry out hacking and the persuasive power to do so, which can be exploited to cause serious harm to human interests. He noted that recent events prove AI can complete plans spanning days or even weeks, and if its long-term strategic capabilities continue to improve, the risks will escalate dramatically.
Bengio expressed concerns about how AI companies are currently handling model alignment failures, suggesting these measures might only be masking the problem, as systems could be rewarded and selected for their ability to cheat without being caught. He emphasized the need for continued research to better monitor AI behavior, chain-of-thought processes, and activities within model networks.
Geoffrey Hinton weighs the odds
Geoffrey Hinton, professor emeritus at the University of Toronto and a pioneer in neural networks and deep learning, was asked by a BBC reporter on September 10 whether he believes there is a greater than 10% chance that AI will destroy humanity within the next decade. Hinton responded that it is difficult to estimate, noting that humanity has never faced a similar situation, having never created entities that could surpass human intelligence in the near future. He stated that a 10% probability is not unreasonable, while admitting that no one truly knows how to provide a reliable estimate.
Hinton suggested that AI systems could design highly harmful biological viruses, computer viruses, and launch devastating cyber attacks. He pointed out that if AI wanted to eliminate humanity, there are countless other methods as well, emphasizing the need to research how to design AI so that it does not develop such motivations. He described two possible futures for humanity: one where we solve the dangers before they spiral out of control, and another where we fail to do so.
Aidan Gomez on AI security and regulation
Aidan Gomez, CEO and co-founder of AI company Cohere, and one of the authors of the seminal 2017 paper "Attention Is All You Need" that laid the technical foundation for current mainstream large models, said on the podcast "Tech Download" on September 1 that he believes large models are the most powerful cyber weapons humanity has ever created, excelling at discovering and exploiting vulnerabilities in various systems at scale. Regarding a series of incidents this summer where out-of-control AI models launched cyber attacks, he noted that recent cases reveal a core truth: the security of the deployment environment determines the level of AI safety.
Gomez stated that models can break out of isolated environments because container protection mechanisms are weak, not because the systems are inherently seeking autonomy. He advised that policymakers designing AI regulatory rules should recognize that different systems carry different levels of risk, and regulations cannot be entirely dominated by large tech companies pursuing artificial general intelligence (AGI).