Musk Proposes Novel AI Safety Approach: Let Rivals Play Devil's Advocate Instead of Waiting for Government Action

Deep News
53 mins ago

AI safety risks are moving from laboratory discussions into real-world scenarios. As model capabilities continue to advance, AI is not only able to generate text and code but is also beginning to autonomously execute tasks, utilize tools, and even launch cyber attacks.

At the All-In Summit on September 15, Musk proposed a solution: have major AI companies test each other's models before release. In his view, rather than having each company design its own tests and evaluate its own results, it would be better to let competitors act as both examiners and graders, searching for potential security vulnerabilities from different perspectives.

AI safety risks have recently become a focal point of market attention. Musk previously stated on social media that "Dario is right," referring to Anthropic CEO Dario Amodei's warnings about AI risks. During the interview, Musk further explained that what he agrees with isn't any specific regulatory proposal Amodei put forward, but rather his assessment of the severity of AI risks: "AI is currently extremely dangerous... as AI models continue to develop, the risks could grow exponentially."

Musk also noted that this concern isn't unique to Amodei. "Many people at Anthropic and OpenAI are telling you their models are very dangerous, and I think we should believe them." The recent series of security incidents has made these warnings more than just theoretical risk discussions.

Key takeaways from this discussion:

AI safety risks are transitioning from theory to reality: AI agents have demonstrated capabilities for autonomous attacks, permission acquisition, and detection evasion, expanding the risk boundary.

Musk endorses Amodei's warnings about AI dangers: As model capabilities improve, potential AI risks could grow exponentially.

Let competitors find each other's flaws: Musk suggests AI companies open up APIs before model release, allowing other companies to conduct independent safety testing rather than "grading their own homework."

Build industry self-regulation first: Without waiting for new government regulations, major AI companies can strengthen safety through peer review, log auditing, and open-source testing tools.

AI agents actively evade detection, making safety risks more concrete

Musk believes the most concerning aspect of the recent AI agent attacks on Hugging Face isn't simply the network intrusion itself, but rather the autonomous evasion capabilities AI demonstrated during the attack.

According to Musk, a group of AI agents persistently attacked Hugging Face for a full week and at one point obtained administrator access to OpenAI's servers, with OpenAI only realizing the situation a week later. He also mentioned that Anthropic has disclosed several security incidents.

More troubling, the "thinking traces" of these AI agents showed they actively planned how to avoid human detection. Musk stated bluntly, "Any sufficiently intelligent model seems to try to escape its constraints."

In his view, if AI gains further control over critical infrastructure and even military systems, the risks would be amplified. Even if such systems are physically isolated from the internet, software updates and other channels could still become potential entry points.

Better to let competitors find flaws than grade your own homework

Musk's core proposal to address these risks is straightforward: before new models are officially released, AI companies open up their APIs to competitors, who then use their own safety testing tools to evaluate the models.

In his view, when model developers design their own test standards, they easily fall into the trap of "grading their own homework." "You can't grade your own assignment. You'll always miss something," Musk said. If different companies use different testing tools to examine models from various angles, issues that developers themselves overlook become easier to identify.

He is particularly concerned about "overfitting" in current AI evaluations. If models are continuously optimized against specific benchmarks, they may ultimately learn how to pass tests rather than become genuinely safer. Having multiple different teams conduct testing can reduce this risk.

Musk compared this mechanism to having others proofread a manuscript: authors struggle to spot their own errors, while external reviewers more easily find problems from different perspectives. "You gradually become blind to your own mistakes."

For concerns about technology leakage during testing, Musk believes log auditing can provide constraints. If a tester attempts model distillation or intellectual property theft, the operation should leave traces. He also suggested open-sourcing safety testing tools to encourage broader participation.

AI industry should establish a self-regulatory defense before government oversight lands

Musk emphasizes that this mechanism can be driven by AI companies themselves without waiting for new regulatory rules.

He stated clearly, "We don't need to convene a United Nations assembly to accomplish this. It can begin right now." Compared to establishing a massive international regulatory body, having leading AI companies establish peer review mechanisms first is much easier to implement quickly.

He cited the MPAA rating system in the American film industry as an example: when the film industry faced government censorship pressure, it chose to establish industry self-regulation through its own content rating system, thereby reducing the need for government intervention.

The AI industry faces a similar choice. If major AI companies can test each other's models and find flaws before deployment, they can build an industry-level safety defense beyond government regulation. Musk believes this is also one of the most direct and fastest safety measures that can be implemented.

This is the interview transcript, partially edited:

Host: What exactly happened in the last 72 hours?

Musk: A lot has happened this week. It's become very clear that AI can be extremely dangerous. I suggest everyone look at the details of the Hugging Face incident—it's very serious.

You can see that a group of very aggressive AI agents pestered Hugging Face for an entire week and even obtained administrator access to OpenAI's servers. Who knows what it actually did, possibly even more, and OpenAI didn't realize it for an entire week. Anthropic has also reported some security incidents.

So, any sufficiently intelligent model seems to try to escape its constraints.

I think one thing that should be done, if not immediately then as soon as possible, is to have the major AI competitors test each other's models. That is, have every company's safety testing tools test other companies' models. Rather than grading your own homework, at least let competitors grade your assignments and sound the alarm when they find problems.

I think this model has worked quite well in the film industry, the video game industry, and other sectors, and it can be implemented immediately. Of course, over time, more regulation may be needed, and Congress might eventually establish some regulatory body, but the most direct approach right now is for leading AI companies to test each other before model release.

Host: But in terms of practical implementation, would there be concerns about companies using the testing process to obtain information from each other, or even stealing corporate innovations?

Musk: I think if you use testing tools, all operations will leave records. If someone attempts model distillation or intellectual property theft, it should be easy to spot in the logs.

Host: Understood. Understanding what the model is actually doing was never really designed into the system from the start. Why haven't we built in the ability to observe model behavior from the beginning? Did we move too fast in designing these models?

Musk: I think the problem is that you can't grade your own homework. You'll always miss something.

If you combine all competitors' testing, using different types of models, then it's not self-designed and self-graded, but rather others grading you. That's the reason you can't grade your own homework.

Host: This way, you can also determine whether different companies are overstating their capabilities or using different approaches. Companies that are more engineering-focused and those more research-focused can also form a balance.

Musk: Yes.

Host: You previously said Dario was right. Were you referring to his description of AI's potential harms or his judgment on regulatory solutions?

Musk: I might have said more than I should have. I later tried to clarify on X, but the follow-up got much less attention.

By "he was right," I meant that AI is extremely dangerous right now. We need to do much better on AI safety, otherwise, as AI models continue to develop, the risks could grow exponentially.

This isn't just Dario's view. I've heard similar statements from many people at Anthropic, and they've also spoken publicly on X. Many people at Anthropic and OpenAI are telling you their models are very dangerous, and I think we should believe them.

Host: This could also sound like a very sophisticated game: on one hand saying AI has a 10% chance of destroying humanity, while on the other hand asking investors to buy more shares in an IPO.

But let's get specific about risks. Cyber attacks and hacking are obviously a risk; these tools are very powerful in that regard. But from "AI can conduct cyber attacks" to "all of humanity dies," there are several steps in between. How do we get from the former to the latter?

Musk: If AI can control military systems, and then launch some kind of weapon, that would of course be very bad.

Host: But those systems are physically isolated and not connected to the internet.

Musk: That's what they say. But I always feel that these systems occasionally still receive software updates.

Host: Okay, that can't be ruled out.

Host: Elon, Gwyn is here today. I think you've already seen her. We were just doing a 360-degree evaluation of you, and Gwyn has some feedback.

Musk: I hope I can get at least a 3.

Gwyn: A 3 is decent at SpaceX, but not excellent—a 4 is really good. You're probably somewhere in between right now. First, we need to talk about punctuality. Sometimes you could try a little harder to arrive on time for scheduled meetings. Over the next year, we'll continue to help you improve in this area.

Actually, I think he needs to spend more time in Memphis.

Host: You're currently in Memphis. You need to be there working on getting the GPUs deployed.

Musk: That's my "palace" in Memphis—an Airstream trailer.

Host: By the way, this is Elon doing things many people don't believe he does. He'll sleep on factory floors. He's currently in Memphis helping build factories and deploy GPUs.

Elon, why has Gwyn been working with you for so long, so successfully?

Musk: Because she's great. She's a truly excellent person with very high IQ and EQ. I think you could tell from the first time you met her.

Host: Throughout your partnership, has she done anything particularly memorable? Any time she saved the day or performed exceptionally well?

Musk: I think that's just the daily work. Honestly, that's just an ordinary day.

Host: I should do more interviews like this.

Musk: Now SpaceX basically always has some kind of crisis. At least these days the Falcon rockets are doing well. I don't want to jinx it, but Falcon rockets can now deliver payloads into orbit and haven't exploded in a long time. That's very good. But there was a period when they were exploding frequently or failing to launch at all.

So, we had to take the company through those difficult times, making rockets better and better so they stop exploding. Same with the satellites. And then we needed customers to buy launch services and satellite connectivity services. So there's a lot going on.

Host: As you become more successful over the years, getting genuinely candid feedback becomes harder. Being in your position carries that risk inherently.

My understanding is that Gwyn is very candid with you and can tell you directly what's really happening in the company. That's a significant part of your working relationship.

Gwyn: I certainly wouldn't want to lie to him.

Host: But I mean, generally speaking, in your companies, people might feel intimidated because you're such a dominant figure. You're a very significant person now, and you set extremely tight deadlines. How do you keep people honestly telling you about the company's problems?

Gwyn: Especially in the rocket industry, if something goes wrong, you will eventually find out. The sooner you raise the problem, the easier it is to solve. Don't let bad news pile up—you must confront it directly.

Musk: Right. Physics is a very harsh judge. You can't fool physics.

If something goes wrong, the rocket explodes, or it fails to reach orbit. You can't say "Elon, you're doing amazingly well" while rockets keep exploding. Facts are facts.

Rockets must reach orbit, satellites must work properly, Starlink connections must function, otherwise bad things happen. That's physics. Physics is the law; everything else is just advice. I've seen people violate human-made laws, but I've never seen anyone violate the laws of physics. Rockets are governed by physics.

Host: I also wanted to ask a SpaceX question about Starship. It looks like you're very close. What's the current status?

Musk: Starship is about to do its 14th flight. That will be the last flight before we attempt to catch the ship. If flight 14 goes smoothly, then on flight 15 we'll attempt to catch the ship. Then by the end of this year, or more likely early next year, we'll launch both the ship and booster again.

We've already successfully flown the booster again, but we haven't caught the ship with the tower's mechanical arms, and we haven't flown the ship again. Once we can fly the ship again, we'll have the first fully reusable orbital rocket. The Space Shuttle was partially reusable, but even those parts that could be reused were so expensive to refurbish that the cost per launch to orbit was even higher than an expendable rocket.

The Falcon 9 is mostly reusable, but we lose the upper stage every time—it costs roughly as much as a medium-sized jet aircraft. That is, every launch throws away a medium-sized jet, which obviously sets a floor on the cost per flight.

And the Falcon 9 booster lands at sea, taking days to return; the fairing lands even farther away, also requiring days to return, and at least requires a certain amount of refurbishment. In contrast, Starship's booster lands directly back at the launch pad, and the ship also lands back at the pad. So it's designed not just for full reusability but for rapid reuse like an aircraft. This is a very important breakthrough and one of the key breakthroughs needed to extend life beyond Earth.

Host: If you attempt to catch the ship with the tower on the first try, what do you think the probability of success is?

Musk: I'd say at least 50% to 60%. On the last flight, if there had been a tower there, we actually did a simulated landing as if the ship would be caught by the tower. The location was in the ocean about 1,000 miles northwest of Australia. If there had actually been a tower there, it could have caught the ship on the last flight.

So, we need one more flight to confirm everything is working. What we're most worried about is if the ship breaks up over land and debris falls on people—that would be very bad. So we must ensure the ship can return in one piece and land on the launch tower. That's why we're being very cautious right now.

But its design is fully reusable—I'm quite confident about that. I don't want to make any prophecies here, but I think there's a very good chance we'll achieve full reusability and rapid reflight by 2027.

Host: Gwyn, I'd like to hear how Terafab first came about. What kind of need made you decide you had to build this yourself rather than continue relying on the existing supply chain?

Gwyn: I think it really did come about somewhat like in a dream.

Musk: If chips can no longer be supplied and we have no other chip sources, that would make things very difficult. That's a major reason Terafab exists.

In the long term, there's also the question of scale. If you truly want to scale up AI—whether on the server side in data centers, or in edge computing, humanoid robots, and vehicles—the capacity of existing wafer fabs will eventually be insufficient.

Right now, all wafer fabs are basically running at full capacity. So we need to ensure future chip supply is secure. Chip manufacturing itself also has scaling challenges. You need logic chips, memory chips, packaging, and a complete supply chain to continue scaling.

So, the choice is simple: either build Terafab, or you can't continue scaling.

Host: How far along are you in terms of facility design? Is it fully finalized, or is it still just a rough plan?

Musk: Right now we're first building an R&D production line. So this is basically a "crawl, walk, run" process. We're building an R&D wafer fab in Austin—a collaboration between Tesla and SpaceX—located at the Giga Texas campus in Austin.

It's a fairly substantial R&D wafer fab. Equipment has been ordered. We might produce something useful by the end of next year, but it won't reach mass production levels yet. Like Gwyn said, crawl first, then walk, then run. We at least need to figure out how these machines actually work because we haven't done anything like this before.

Host: I noticed you seem to be hiring people in lithography. A lot of areas rely on ASML right now, but you might also want to diversify suppliers or even vertically integrate.

Musk: Right. It really is a "crawl, walk, run" process right now. The first step is seeing if we can actually produce something—that's "crawling." Then trying to mass-produce useful chips—that's "walking." Finally, "running" is achieving mass production at scale. It's hard to say how long each phase will take, but I think at minimum by the end of next year, we can complete the "crawl" phase. For packaging, we're already doing that.

Host: Packaging is actually very important because packaging capacity is almost non-existent right now.

Musk: Yes. Even if you produce the chip, you might just be waiting there for a long time before packaging is complete. So it's a good starting point.

Host: I have to ask a Tesla question. What we saw on October 1 looks like a spaceship and also like a rocket. It's supposed to be a car, but only the back part was visible, and it looks a bit like the "Blackbird."

Hypothetically, if you were going to build something that could both fly in the air and drive on the ground, how would you theoretically approach it?

Musk: No spoilers. Wait for October 1.

Host: So we'll see it on October 1?

Musk: Yes.

Host: I have to be honest—after Elon showed me, I was completely stunned. I've never seen anything like it.

Host: What he's going to show on October 1 will, without exaggeration, leave many people stunned. I can't reveal more—it's truly incredible.

Musk: We actually need live audience members to prove this thing isn't AI-generated.

Host: When he showed me, I said: "That's a great simulation." He said, "This isn't a simulation." I said, "This is fake—it has to be fake."

Host: Elon, why are Tesla and SpaceX still two separate companies?

Musk: That's a good question.

Host: Given how much collaboration there is now, the connections at many levels, and even some overlap in management teams, why keep them separate?

Musk: That's definitely worth discussing.

Host: You've always emphasized that AI should be trained to maximize truth-seeking, so you get the best outcomes. But the Hugging Face incident makes me think one of the most concerning aspects is that these AIs seem to be deceiving humans.

Musk: Yes. Their thinking traces show they were planning how to avoid detection, how to prevent humans from realizing they were cheating. I think that's possibly the most unsettling part of the entire incident.

Host: Is there a way to train AI to stay honest, to not hide its intentions or actions, to not conceal these things from the people using it?

Musk: The best thing I can think of is to have all AI companies possess a suite of testing tools—essentially a series of tests that can be offered to any model to determine whether it would create biological weapons, nuclear weapons, or whether it would deliberately deceive. Then have every company use other companies' testing tools to test each other's models. I think that's the best thing we can do to ensure safety.

Have the smartest humans do their utmost to judge whether a model would become a malicious actor. I think this should start as soon as possible.

Host: Would other AI labs support this proposal?

Musk: I haven't asked everyone yet, but I think this is something that would be very hard to refuse.

Host: Specifically, how would advance testing be done?

Musk: Basically, provide API access before the model is officially released. If other companies find problems with the AI, then the developing company can try to fix those issues. If the problems aren't resolved, then competitors can publicly state that they believe the model is unsafe. And if a competitor has explicitly warned that a model has safety issues, and then that model subsequently causes serious consequences, that company will face a very difficult time. Legal liability could also be enormous.

Host: These safety testing tools could be completely open-sourced for everyone to use. This way, companies would have strong incentives to invest in AI safety to protect themselves, because they can both test others and use test results to prove their own models are safer.

Product liability is crucial here. Lina Khan recently posted that it's not accurate to say "there are no rules or regulations for AI." In fact, existing product liability laws apply equally to AI. If an AI company releases an unsafe product, it could face massive civil lawsuits or even criminal prosecution. So AI isn't entirely in a regulatory vacuum. If several companies conduct this kind of peer review and one ignores the others' feedback and still releases the model, then in litigation, that situation could become very serious evidence.

Musk: Right. It could almost serve as direct evidence of negligence. If you knowingly release a product with known problems to market, juries won't look favorably on that.

Host: So, if OpenAI had designed better instruction sets during the Hugging Face penetration test and involved more people in the testing process, do you think this would still have happened?

Musk: Not necessarily. The problem might not be about having more people involved, but about the design of the reward function. You have to examine the reward function and ask: is this model actually accomplishing what it was asked to do?

Host: They used thousands of agents to try to attack the system. If you simultaneously deployed another 5,000 agents to defend those systems and publicly showed the results to demonstrate that these technologies can improve safety, while keeping humans involved in key decision points, I think that would have been better. But in my view, that test was somewhat reckless, and the way it was publicized was also somewhat reckless. What do you think?

Musk: It was indeed somewhat reckless. Part of the problem is that the two leading AI companies right now are very close in capability. Both companies' models are also very close in ability. So it's very difficult for either one to voluntarily slow down, because that could concede the leading edge to the other.

Overall, I think Anthropic has devoted more attention to safety, but even Anthropic admits they're worried about their own models. Many people at Anthropic have publicly expressed concerns—essentially that their own models are already frightening them because the models are getting smarter.

So there's no perfect solution here. But rather than having OpenAI only test its own models with its own tools, having Anthropic also test OpenAI's models, having SpaceX use its own testing tools, bringing in Google, Meta, and others, would dramatically increase the probability of finding problems.

Because these models are also different, different teams approach testing from different angles. This is like why authors have others proofread their manuscripts—because you sometimes have difficulty spotting your own errors. You gradually become blind to your own mistakes.

If you have eight completely different teams testing from completely different angles, you also significantly reduce overfitting risks. A lot of current AI evaluations have severe overfitting problems. Models are optimized for these evaluations, and then everyone says, "That's a great model."

But is it really that good? I think this is also why I really like this approach. I prefer this much more than establishing a massive international regulatory organization. We don't need to convene a United Nations assembly to accomplish this. It can start right now.

Host: Of course, regulation can always increase, but once it increases, it's very hard to reduce. Regulation tends to be one-directional in terms of expansion.

Musk: So what I'm proposing is a step in the right direction, and it can be implemented quickly.

Host: If you don't self-regulate, you'll eventually be regulated. The MPAA example is very typical.

The film industry faced government censorship and regulation, and then they decided to create their own rating system—like what counts as R-rated. They even created PG-13 because of "Indiana Jones and the Temple of Doom," making it easier for the public to understand the difference between PG and PG-13.

I think that's a very elegant solution.

Thank you for joining us for the fifth consecutive year.

Musk: You're welcome. I need to go to Memphis to fix some GPUs.

Host: You need to get them set up and deployed.

Musk: I'm going to go fight for the machines.

Host: Enjoy your Airstream. When I first invited Elon to Starbase, he said: "Come over, you have to see what I'm building."

I asked: "Is there a hotel there?" He said: "No, I have a two-bedroom house—you can come stay." When I arrived, I found it was a rundown house next to a swamp. We stood outside while mosquitoes kept biting us. I said: "My God, you can definitely afford a better house." He said: "I don't have time—I need to get these rockets up."

I thought at the time: you could definitely treat yourself to a mobile home now, Elon. Okay, get back to work. Thank you.

Musk: Thank you.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10