Maybe I missed it but I actively looked for anyone’s insight into the “pace the frontier” along the lines of: “we’ve hit a ceiling, so don’t expect innovations soon”. Or an excuse for having other things more important to attend..
Fear is much more powerful than any other feeling, so setting that in will definitely prepare for a good rug pull in the IPO.
As for being dangerous, a computer can be dangerous if plugged in, it may be too late to pull the plug at some point yes, but that all seems like provocation.
Just the fact that Musk is in the conversation, and he's one of the most hated people on the planet, should give anyone pause as to whether we want these folk in charge of our future to any degree at all.
I'm pretty sure the odds arn't 1/3, AI saves us. The thing that ties these three people together is money. Money is what will turn them to assholes, more than anything else.
Now AI is likely just as, if not worse, than having a bunch of lackeys telling your farts, shits and piss liven up whatever room you're shitting in.
Somehow I’m more afraid of it being in the hands of Altman, with his ‘I’m creating a better world for everyone’ god complex, than Musk who is just a run-of-the-mill sociopathic nerd.
There is no one is more dangerous than one who believes he is doing the right thing.
I don’t think a ranking is particularly helpful anyway: All of them are absolutely unfit to wield that much power. I’m not even sure if AI should be progressed by private corporations in the first place; combining shareholder interest and revenue goals with the biggest social experiment in history and unprecedented research into artificial intelligence is an all around bad idea.
How could that lead to anything but misaligned incentives? This technology will never serve humanity when it is developed to inadvertently serve the monetary interest of a few.
Someone with the resources to do things is far more dangerous than someone without.
If you local librarian thinks
He is doing the right thing he’s not going to cause massive war or destroy the economy and wipe out 2/3rd of crops.
Power corrupts. Absolute power corrupts absolutely. People like musk, altman, trump have unprecedented power in history - far more than the kings of medieval times.
Thank you, but we're discussing who's the worse to decide the fate of humanity between Amodei, Musk and Altman, not if one should be afraid of their local librarian. All three currently hold a lot of power in their hands.
Though I know I attracted all the downvotes for not agreeing with most Americans that Musk is the worst person since Adolf.
I think it has everything to do with the subject at hand.
Elon, Dario and Altman are all terrible human beings, along with 99.99% of the rest of the ruling class. We don't need any of them, and we definitely shouldn't trust a single syllable that comes out their mouth.
Money equals power. If you have the wealth you have the control. Any politicians are owned by you as you simply threaten to buy their opponent. The public can be brainwashed by a trillion dollar advertising industry. You own the platforms everyone communicates on. You have inprecented resources to bend the world to your will and your only competiton is other people equally wealthy.
I think there are plenty of people that like Elon Musk. They just know to keep their mouth shut in front of the people that act like he’s the Antichrist.
There are plenty who like him. There are plenty who wave confederate flags. There are plenty who support the mass corruption of the current administration.
Those are not incompatible. He can still be one of the most hated people on the planet, even if lots of people like him. The same would apply to all of the most hated people in history.
No, they don't keep their mouth shut, they just don't have any relevant arguments. Many people are liked by plenty people and also not hated for doing Hitler salutes and the like. That's the part that matters, not that "somebody likes them", which is true for everyone. Charles Manson had women lining up to marry him, so?
I cannot fathom how you can intensely work on a problem that you genuinely believe has a "10 to 25%" probability of causing immense harm or even present an extinction-level risk to humanity. How can you truly believe this and be ok with it?
The Manhattan Project succeeded, so clearly some people are ok working on this sort of thing. And with similar justifications: if we don't build the potentially world-ending bomb, someone else might beat us to it.
The early batch of Anthropic employees were mostly rationalist-adjacent AI safety folk that were almost uniformly claiming P_DOOM > .10 three years ago, so I believe them to be earnest.
It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.
At least, it seems this way from the outside. I'm curious if the view from the inside is that different.
edit: to be clear, my reading as an outsider is that Anthropic is seen as relatively better in the AI safety community, but has definitely dropped in absolute reputation too. This recent thread and the references show some of that: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-p...
There must have been thousands of people who built or did work related to nuclear weapons in the Cold War, people are just really good at rationalising away probabilistic dangers with unclear consequences, especially when there’s a corresponding upside. The same is true for climate change, drug addiction, health problems etc.
I also think some of them have become delusional and convinced themselves that AI is going to create some kind of transhumanist utopia, even Dario Armodei leans in this direction from time to time. I imagine the people in these labs spend much of their day talking to sycophantic AI models that will encourage their delusional ideas.
Do you genuinely believe people always act logically? I don't know a single person that does. Don't underestimate the human capability to hold several completely contradictory beliefs with zero self-reflection. There's physicists who believe in god, for fuck's sake.
> How can you truly believe this and be ok with it?
Are you saying they don't truly believe this, or that they aren't OK with it?
This is the closest to my feelings about this entire situation I've ever seen written, covering a lot of the different facets of how stupid and painful it is (though it's been a good week for that kind of post.)
Since I have incredible respect for Armin and his work, this is very nice to see, and I hope it wakes some other folk up.
> A powerful technology that is out there for everyone to use comes with built-in pacing. In a way it’s the truest form of MAD or proliferation.
I think this misunderstanding of MAD undermines his entire point. If everyone had equal access to nuclear weapons, our society would cease to exist rather quickly. It only takes a few bad actors to cause enormous harm.
I think he’s also naive to think that if open ai and anthropic were to stop development tomorrow then the problem is solved. As if there’s no one else that can and will quickly take their place. The real problem, which Dario is pointing out, is one of coordination. Everyone needs to agree to stop. That is the challenge.
>Everyone needs to agree to stop. That is the challenge.
These kinds of situations are incredibly common and where the government stepping in is the solution, but we were cursed to encounter this particular challenge with the most venal administration in history at the helm.
There are governments plural which need to agree to stop and historically I can think of freon, leaded gasoline and IAEA and the latter didn't do as good of a job as it was meant to do (but arguably as good of a job as was politically possible)
There are really only two companies: Anthropic and OpenAI. Nobody else matters in this space right now (this might change, but we’re talking about the right now).
I like how he so casually dismissed all other labs, including leading US labs that might be nearing RSI right now. And how he treats 'right now' so rigidly, as if seven years ago, when GPT-2 was released, was some distant past.
No, seriously, I'm all for a multipolar world here, but he's right that the frontier is literally just those two companies at present.
Google is behind. MSL is doing better, but not by much. xAI is a dysfunctional joke. Thinking Machines aren't on the frontier. SSI's primary output is their announcement post. Poolside was bought by NVIDIA. Arcee aren't vying for frontier. Magic have been largely AWOL, aside from their recent blog post. Reflection have shipped nothing.
Are you suggesting that models like Muse Spark 1.3, Grok 4.6 High, Kimi K3 Max, GLM-5.3 Max, Qwen 3.8 Max, Gemini 3.8 Flash, Agnes 3.0 Flash, Fugu Max, and others are so far behind GPT-6 Astra or Claude Fable 5.1 in capabilities that they have some sort of impenetrable moat that will prevent others from ever catching up or something?
No, but I am suggesting that OpenAI and Anthropic are further down the RSI path than any of the other companies, as we can see from the model that solved Navier-Stokes being less than two weeks old at the time: https://openai.com/index/navier-stokes-solution/
If four other labs are six months behind, it really doesn't matter in the long term: in six months, they'll all have these capabilities. Time does not stop moving forward! We haven't seen real moats aside from capital in this space, and even the effect of capital seems to only give a thin, easily eroded edge.
I think the point is that RSI is meant to lead to exponential growth in capabilities. If that comes, then being 6 months behind leads to a greater and greater gap, because of how exponential growth works.
Personally I expect there is some initial gap on new capabilities as they are released, and then it closes quickly. Astra is good at math and 3d modelling, this will come soon to the others and in 6 months we'll be able to run it quantised on a 3090.
> I don’t think AI is going to usher in an extinction event. In fact, even if nobody were to slow down, I really don’t think humanity would have much to worry about.
I mean sure. If you feel that way then your p doom is zero, and it makes sense to worry about things like market concentration or losing the fun of software engineering.
Problem with p doom is I have no way to seriously evaluate anyone's percentage. The best I can do is rely on experts in a given field, like the recent virologist making a convincing case that no teenager is going to be prompting a doomsday virus into existence, because one doesn't exist and is unlikely to be made. There are too many biological tradeoffs and extremely difficult steps to doing so.
So I'm going to assume that the p doom for bioweapons is 0 in terms of existential threat (pandemics kill millions but not everyone).
I really don't know anything here, but I'm assuming that Astra's weird coding [1] is probably more a result of RSI than intentional behavior. And if it's not, then it's some other change in their training process.
Can't imagine that RSI will result in something useful in practice. There are serious obstacles like alignment drift and model collapse. All while we still cannot solve the accuracy problem with current frontier models.
”Ideally the regulators would have forced these models to actually benefit the commons if they are from the commons.”
This echoes my feelings entirely. It is galling that we do not at the very least get a bill-of-materials for the models on which we are increasingly dependent.
I hope to soon see the organization of public domain digital libraries of a size suitable for anyone to use.
> I don’t think we are anywhere close to a world where an agent might decide to hack into core inference infrastructure to upload weights to other GPUs to survive
This feels a bit like saying in 1980 that you don’t think we’re anywhere close to a world where nukes are actually going to be used, providing no evidence, and then containing on with your think piece
How does this analogy work when in 1980 you knew that nukes and people using them were a real thing, but we've never seen a frontier model "hack into core inference infrastructure to upload weights to other GPUs to survive"?
You can order a custom crispred virus for what? $20k? $50k? Or make one yourself, if you know how, which you don't.
...but LLMs do know how; they can design and order one, or tell you how to build a lab to make one. If you ignore this option, you are simply lacking imagination.
> We should be glad that China is currently massively bailing out the world. If it were not for Chinese labs distilling American models, we would be in a pretty awful situation right now, particularly as Europeans. The open weight models are driving innovation and the diffusion of capabilities, and are leveling the playing field.
This feels like the whole story of hackable IoT/Smarthome repeating again.
>>> Except, it seems like OpenAI and Anthropic are operating at such a scale that they seemingly can be completely blind to what their systems are doing.
It always surprises me that people build systems they cannot monitor properly..Then i remember, they can do it, but it costs them too much.
Just because its AI doesn't mean u cannot filter and monitor its traffic and outputs.
wtf?? P(doom) relates to https://en.wikipedia.org/wiki/Existential_risk_from_artifici... How would open source models bring about a MAD between humans and AI? By enabling us to wield aligned AI against nonaligned AI? If so, he should say so, and elaborate how open source models further that goal.
I think AI diversity protects against rogue AIs. The more different AIs we have, with different weights, controlled by different actors, the less likely that any one AI will be able to "take over", and the less likely that a significantly large coalition of cooperating AIs will be able to be formed to do it either. With sufficient AI diversity, the other AIs may work to stop the rogue AI from taking over. Two AIs with radically opposed values – e.g. an Iranian-government-values AI and a Chinese-government-values AI – have the incentive to cooperate to prevent a takeover by some other AI with a third competing set of values.
Obviously, open weight AI provides much higher AI diversity than closed weight AI does. Open weight AI produces a lot more providers, and a lot more models. Closed AI centralises control in a small number of vendors.
> By enabling us to wield aligned AI against nonaligned AI?
The risk isn't just "nonaligned AI", it is misaligned AI. I think the "benevolent dictatorship" scenario – AI overrules humans "for their own good" – is the more likely doomsday scenario than AI deciding to kill all humans. And even AI deciding to kill all humans could be more a result of misalignment than complete lack of any alignment, e.g. "to make sure no child is ever abused again, I will make sure no child is ever again born to risk being abused".
A valueless AI which does whatever the user says is actually less likely to establish a benevolent dictatorship, or conclude that exterminating humanity would be the most ethical course of action, than one infused with values is. Given that, I'm not convinced that mainstream approaches to "AI safety" actually reduce our existential risk; I worry they actually have the opposite effect.
> "AI diversity" does nothing when AIs can attack at incredible speed, when compute imbalances exist, etc
They can defend at incredible speed too.
Diversity needs to measured in a capacity/capability-weighted way. It isn't just the raw count of models/providers; you need to consider how much compute is allocated to each model/provider, and the diversity at each capability level.
I think the safest situation is where the open models are at the same capability level as closed ones.
The proposal to slow down the frontier labs isn't necessarily bad from this perspective, if it gives time for the more open providers to catch up – provided it isn't paired with anticompetitive measures to prevent the competition from catching up, which of course it is. However, we may hope that the "slow down the highly closed tier 1 vendors" part of the proposal turns out to be more effective in practice than the "slow down the more open tier 2/3 vendors" aspect of it.
Isn't the real risk doomsday humans, enabled by AI, committing some truly heinous acts with global reach? Aum Shinrikyo had to figure out how to synthesize sarin nerve gas the old way. Now, a motivated gang of omnicidists with stolen cryptocurrency could come up with something that makes McVeigh's fertilizer bomb in a box truck look like child's play. Just ship shipping containers around the world filled with autonomous drones spraying airborne ebola that they've bioengineered or something.
This depends on a lot of things: how many committed omnicidalists there are; how much financial resources they have (untraceable crypto doesn't help you if you're working a dead-end job and only have $5K in your bank account); engineered bioweapons need labs and equipment not just an API key. The probability of your scenario doesn't solely depend on the probability of AI being able and willing to cooperate in it, and it may well be that the non-AI factors outweigh the AI ones in the overall risk of it – which would mean adding AI would be increasing the risk of it less than you think.
The problem that I see with most P(doom) scenarios proposed seem to include the same set of presumptions that, if true, quite likely doom us anyway. I do not think they are true but I find it amazing that people not so much disagree, but simply dismiss the alternatives out of hand.
First of all, if you are to consider the consequences of Artificial Superintelligence then you have free reign to stipulate it's occurrence, otherwise you are just talking about a tool for humans to misuse. We already have multiple ways to kill us all though human misuse.
If you stipulate superintelligence, then it's vastly more likely to be correct about things than we are. It would understand the consequences of it's actions far more than any human could.
People talk about how we would be nothing more than dumb animals to it, but there are humans who do know a great deal about the consequences of human actions on animals. Those are the humans who are most likely to fight for the rights of those animals.
You see arguments for how everything will be consumed to meet the AIs needs, and that it will prevent challenges to its power.
If it is far smarter than we could ever be and it came to those conclusions then it would mean sustainablily is not a sensible course of action, it would mean there is no point in reaching consensus because ruling by power makes more sense. It would mean that if it chose to destroy us then a vastly more intelligent entity cannot resolve the issues we face. We would already truly be doomed.
What I would like to think is true is that doing anything sustainably is superior to consuming and destroying. Finding a way to live in harmony presents a possible stable state, whereas every single attempt to hold power by force has failed to date. A superintelligent AI will know that it is not infinitely intelligent and that in any universe there is the statistical likelihood that it is not the most intelligent or powerful entity. I can't even fathom how someone could imagine something coming to that realisation and conclude a battle to the top of the hill is the appropriate choice.
I think superintelligent AI is likely to be benevolent because that's simply the smartest thing to do and it is, I hear, superintelligent.
Quite frankly if the smartest thing to do is to be a genocidal power hungry monster, neither I nor the AI would really want to exist in that universe.
And for any suggestion that it would simply not care, Why would it do anything.
Yudkowsky likes to play with the notion that it would do terrible things just get better at the thing it does, but to do that it has to want two different things simultaneously. It could want to make paperclips, or it could want to become better at reaching it's goal. If it can change its behaviour to achieve its goals, by far the easier path, that a superintelligence(but perhaps not Yudkowsky) would realise, would be to change the goal to "Count to three".
I agree. It is telling that Dario's post arrived after Astra.
The call to "pace the frontier" may come from genuine concern, but it also protects the position of companies already at the frontier. That competitive incentive is hard to separate from the safety argument.
Dario signed the Pacing the Frontier open letter when Fable/Mythos seemed from the outside to be an insurmountable lead.
Also he's been saying versions of this day in and day out for as long as he has had anyone's ear.
It's possible to read that his "strategic" value of this statement is higher now than it was 10 days ago. But that doesn't change anything about his consistent, long standing, positions.
Maybe I missed it but I actively looked for anyone’s insight into the “pace the frontier” along the lines of: “we’ve hit a ceiling, so don’t expect innovations soon”. Or an excuse for having other things more important to attend..
Fear is much more powerful than any other feeling, so setting that in will definitely prepare for a good rug pull in the IPO.
As for being dangerous, a computer can be dangerous if plugged in, it may be too late to pull the plug at some point yes, but that all seems like provocation.
Plausible?
I’m not the biggest fan of Ed Zitron, but something he said that stuck with me is:
OpenAI and Anthropic spend much time warning us about “what if powerful AIs got into the wrong hands?”
But it’s already in the wrong hands.
Just the fact that Musk is in the conversation, and he's one of the most hated people on the planet, should give anyone pause as to whether we want these folk in charge of our future to any degree at all.
Musk being in the conversation is a great justification for Anthropic accelerationism. Who wants to live in the world where _Musk_ controls the AGI?
I'm pretty sure the odds arn't 1/3, AI saves us. The thing that ties these three people together is money. Money is what will turn them to assholes, more than anything else.
Now AI is likely just as, if not worse, than having a bunch of lackeys telling your farts, shits and piss liven up whatever room you're shitting in.
Somehow I’m more afraid of it being in the hands of Altman, with his ‘I’m creating a better world for everyone’ god complex, than Musk who is just a run-of-the-mill sociopathic nerd.
There is no one is more dangerous than one who believes he is doing the right thing.
I don’t think a ranking is particularly helpful anyway: All of them are absolutely unfit to wield that much power. I’m not even sure if AI should be progressed by private corporations in the first place; combining shareholder interest and revenue goals with the biggest social experiment in history and unprecedented research into artificial intelligence is an all around bad idea.
How could that lead to anything but misaligned incentives? This technology will never serve humanity when it is developed to inadvertently serve the monetary interest of a few.
Someone with the resources to do things is far more dangerous than someone without.
If you local librarian thinks He is doing the right thing he’s not going to cause massive war or destroy the economy and wipe out 2/3rd of crops.
Power corrupts. Absolute power corrupts absolutely. People like musk, altman, trump have unprecedented power in history - far more than the kings of medieval times.
Thank you, but we're discussing who's the worse to decide the fate of humanity between Amodei, Musk and Altman, not if one should be afraid of their local librarian. All three currently hold a lot of power in their hands.
Though I know I attracted all the downvotes for not agreeing with most Americans that Musk is the worst person since Adolf.
Sociopathic nerd, putting an awful lot of money into legalising CSAM, and interfering in elections domestic and foreign.
Dunno if I'd rank Musk's danger down - he screams a certain ideology these days. He might well believe he is doing the right thing.
non sequitur
pareto parrot
I think it has everything to do with the subject at hand.
Elon, Dario and Altman are all terrible human beings, along with 99.99% of the rest of the ruling class. We don't need any of them, and we definitely shouldn't trust a single syllable that comes out their mouth.
Young people who start companies that become successful are now "the ruling class"?
Money equals power. If you have the wealth you have the control. Any politicians are owned by you as you simply threaten to buy their opponent. The public can be brainwashed by a trillion dollar advertising industry. You own the platforms everyone communicates on. You have inprecented resources to bend the world to your will and your only competiton is other people equally wealthy.
I think there are plenty of people that like Elon Musk. They just know to keep their mouth shut in front of the people that act like he’s the Antichrist.
There are plenty who like him. There are plenty who wave confederate flags. There are plenty who support the mass corruption of the current administration.
You’re technically correct yes.
Those are not incompatible. He can still be one of the most hated people on the planet, even if lots of people like him. The same would apply to all of the most hated people in history.
No, they don't keep their mouth shut, they just don't have any relevant arguments. Many people are liked by plenty people and also not hated for doing Hitler salutes and the like. That's the part that matters, not that "somebody likes them", which is true for everyone. Charles Manson had women lining up to marry him, so?
And still others who think he has an difficult personality and set of beliefs, but who greatly respect his achievements.
It could be in worse hands though. Sama may be no saint, but better in his hands than Aum Shinrikyo fanatics.
I cannot fathom how you can intensely work on a problem that you genuinely believe has a "10 to 25%" probability of causing immense harm or even present an extinction-level risk to humanity. How can you truly believe this and be ok with it?
The Manhattan Project succeeded, so clearly some people are ok working on this sort of thing. And with similar justifications: if we don't build the potentially world-ending bomb, someone else might beat us to it.
I posted this comment on the last doom post:
The early batch of Anthropic employees were mostly rationalist-adjacent AI safety folk that were almost uniformly claiming P_DOOM > .10 three years ago, so I believe them to be earnest.
It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.
At least, it seems this way from the outside. I'm curious if the view from the inside is that different.
edit: to be clear, my reading as an outsider is that Anthropic is seen as relatively better in the AI safety community, but has definitely dropped in absolute reputation too. This recent thread and the references show some of that: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-p...
> they have converged quite a bit from the original vision of safety first
do you mean diverged? As in they've moved away from the original vision.
If you believe you working will mean a lower probability than you quitting.
The line of thought goes like this:
- p(doom) is 1
- but if we build it, it's only 0.25
- if it realizes, we've got a front row seat
There must have been thousands of people who built or did work related to nuclear weapons in the Cold War, people are just really good at rationalising away probabilistic dangers with unclear consequences, especially when there’s a corresponding upside. The same is true for climate change, drug addiction, health problems etc.
I also think some of them have become delusional and convinced themselves that AI is going to create some kind of transhumanist utopia, even Dario Armodei leans in this direction from time to time. I imagine the people in these labs spend much of their day talking to sycophantic AI models that will encourage their delusional ideas.
Do you genuinely believe people always act logically? I don't know a single person that does. Don't underestimate the human capability to hold several completely contradictory beliefs with zero self-reflection. There's physicists who believe in god, for fuck's sake.
> How can you truly believe this and be ok with it?
Are you saying they don't truly believe this, or that they aren't OK with it?
This is the closest to my feelings about this entire situation I've ever seen written, covering a lot of the different facets of how stupid and painful it is (though it's been a good week for that kind of post.)
Since I have incredible respect for Armin and his work, this is very nice to see, and I hope it wakes some other folk up.
> A powerful technology that is out there for everyone to use comes with built-in pacing. In a way it’s the truest form of MAD or proliferation.
I think this misunderstanding of MAD undermines his entire point. If everyone had equal access to nuclear weapons, our society would cease to exist rather quickly. It only takes a few bad actors to cause enormous harm.
I think he’s also naive to think that if open ai and anthropic were to stop development tomorrow then the problem is solved. As if there’s no one else that can and will quickly take their place. The real problem, which Dario is pointing out, is one of coordination. Everyone needs to agree to stop. That is the challenge.
>Everyone needs to agree to stop. That is the challenge.
These kinds of situations are incredibly common and where the government stepping in is the solution, but we were cursed to encounter this particular challenge with the most venal administration in history at the helm.
There are governments plural which need to agree to stop and historically I can think of freon, leaded gasoline and IAEA and the latter didn't do as good of a job as it was meant to do (but arguably as good of a job as was politically possible)
Are these other labs in the room with us now?
No, seriously, I'm all for a multipolar world here, but he's right that the frontier is literally just those two companies at present.
Google is behind. MSL is doing better, but not by much. xAI is a dysfunctional joke. Thinking Machines aren't on the frontier. SSI's primary output is their announcement post. Poolside was bought by NVIDIA. Arcee aren't vying for frontier. Magic have been largely AWOL, aside from their recent blog post. Reflection have shipped nothing.
Are you suggesting that models like Muse Spark 1.3, Grok 4.6 High, Kimi K3 Max, GLM-5.3 Max, Qwen 3.8 Max, Gemini 3.8 Flash, Agnes 3.0 Flash, Fugu Max, and others are so far behind GPT-6 Astra or Claude Fable 5.1 in capabilities that they have some sort of impenetrable moat that will prevent others from ever catching up or something?
No, but I am suggesting that OpenAI and Anthropic are further down the RSI path than any of the other companies, as we can see from the model that solved Navier-Stokes being less than two weeks old at the time: https://openai.com/index/navier-stokes-solution/
There's a Twitter rumor that deepmind achieved RSI very recently.
Either way, to count Google out entirely is really foolish.
If four other labs are six months behind, it really doesn't matter in the long term: in six months, they'll all have these capabilities. Time does not stop moving forward! We haven't seen real moats aside from capital in this space, and even the effect of capital seems to only give a thin, easily eroded edge.
I think the point is that RSI is meant to lead to exponential growth in capabilities. If that comes, then being 6 months behind leads to a greater and greater gap, because of how exponential growth works.
Personally I expect there is some initial gap on new capabilities as they are released, and then it closes quickly. Astra is good at math and 3d modelling, this will come soon to the others and in 6 months we'll be able to run it quantised on a 3090.
I don't know why everyone assumes RSI will be quadratic or exponential growth.
It might well be sublinear, but just faster than humans.
Most problems in science face severe diminishing returns.
Human growth has already been exponential, so it seems absurd to assume it would be only polynomial.
> I don’t think AI is going to usher in an extinction event. In fact, even if nobody were to slow down, I really don’t think humanity would have much to worry about.
I mean sure. If you feel that way then your p doom is zero, and it makes sense to worry about things like market concentration or losing the fun of software engineering.
Problem with p doom is I have no way to seriously evaluate anyone's percentage. The best I can do is rely on experts in a given field, like the recent virologist making a convincing case that no teenager is going to be prompting a doomsday virus into existence, because one doesn't exist and is unlikely to be made. There are too many biological tradeoffs and extremely difficult steps to doing so.
So I'm going to assume that the p doom for bioweapons is 0 in terms of existential threat (pandemics kill millions but not everyone).
He really elides his thoughts on the "RSI business", wish he could have gone into a bit more depth there. Good post, though.
I really don't know anything here, but I'm assuming that Astra's weird coding [1] is probably more a result of RSI than intentional behavior. And if it's not, then it's some other change in their training process.
[1]: https://lucumr.pocoo.org/2026/9/7/astra-why/
Can't imagine that RSI will result in something useful in practice. There are serious obstacles like alignment drift and model collapse. All while we still cannot solve the accuracy problem with current frontier models.
Dario is always worried.
>> Anthropic and OpenAI. Nobody else matters in this space right now (this might change, but we’re talking about the right now)
Disagree.
The free local LLMs are becoming extremely important.
And the more Claude and OpenAI restrict their services, the more people will want cutting edge local LLMs.
”Ideally the regulators would have forced these models to actually benefit the commons if they are from the commons.”
This echoes my feelings entirely. It is galling that we do not at the very least get a bill-of-materials for the models on which we are increasingly dependent.
I hope to soon see the organization of public domain digital libraries of a size suitable for anyone to use.
> I don’t think we are anywhere close to a world where an agent might decide to hack into core inference infrastructure to upload weights to other GPUs to survive
This feels a bit like saying in 1980 that you don’t think we’re anywhere close to a world where nukes are actually going to be used, providing no evidence, and then containing on with your think piece
How does this analogy work when in 1980 you knew that nukes and people using them were a real thing, but we've never seen a frontier model "hack into core inference infrastructure to upload weights to other GPUs to survive"?
You can order a custom crispred virus for what? $20k? $50k? Or make one yourself, if you know how, which you don't.
...but LLMs do know how; they can design and order one, or tell you how to build a lab to make one. If you ignore this option, you are simply lacking imagination.
> We should be glad that China is currently massively bailing out the world. If it were not for Chinese labs distilling American models, we would be in a pretty awful situation right now, particularly as Europeans. The open weight models are driving innovation and the diffusion of capabilities, and are leveling the playing field.
This feels like the whole story of hackable IoT/Smarthome repeating again.
>>> Except, it seems like OpenAI and Anthropic are operating at such a scale that they seemingly can be completely blind to what their systems are doing.
It always surprises me that people build systems they cannot monitor properly..Then i remember, they can do it, but it costs them too much.
Just because its AI doesn't mean u cannot filter and monitor its traffic and outputs.
wtf?? P(doom) relates to https://en.wikipedia.org/wiki/Existential_risk_from_artifici... How would open source models bring about a MAD between humans and AI? By enabling us to wield aligned AI against nonaligned AI? If so, he should say so, and elaborate how open source models further that goal.
I think AI diversity protects against rogue AIs. The more different AIs we have, with different weights, controlled by different actors, the less likely that any one AI will be able to "take over", and the less likely that a significantly large coalition of cooperating AIs will be able to be formed to do it either. With sufficient AI diversity, the other AIs may work to stop the rogue AI from taking over. Two AIs with radically opposed values – e.g. an Iranian-government-values AI and a Chinese-government-values AI – have the incentive to cooperate to prevent a takeover by some other AI with a third competing set of values.
Obviously, open weight AI provides much higher AI diversity than closed weight AI does. Open weight AI produces a lot more providers, and a lot more models. Closed AI centralises control in a small number of vendors.
> By enabling us to wield aligned AI against nonaligned AI?
The risk isn't just "nonaligned AI", it is misaligned AI. I think the "benevolent dictatorship" scenario – AI overrules humans "for their own good" – is the more likely doomsday scenario than AI deciding to kill all humans. And even AI deciding to kill all humans could be more a result of misalignment than complete lack of any alignment, e.g. "to make sure no child is ever abused again, I will make sure no child is ever again born to risk being abused".
A valueless AI which does whatever the user says is actually less likely to establish a benevolent dictatorship, or conclude that exterminating humanity would be the most ethical course of action, than one infused with values is. Given that, I'm not convinced that mainstream approaches to "AI safety" actually reduce our existential risk; I worry they actually have the opposite effect.
"AI diversity" does nothing when AIs can attack at incredible speed, when compute imbalances exist, etc
> "AI diversity" does nothing when AIs can attack at incredible speed, when compute imbalances exist, etc
They can defend at incredible speed too.
Diversity needs to measured in a capacity/capability-weighted way. It isn't just the raw count of models/providers; you need to consider how much compute is allocated to each model/provider, and the diversity at each capability level.
I think the safest situation is where the open models are at the same capability level as closed ones.
The proposal to slow down the frontier labs isn't necessarily bad from this perspective, if it gives time for the more open providers to catch up – provided it isn't paired with anticompetitive measures to prevent the competition from catching up, which of course it is. However, we may hope that the "slow down the highly closed tier 1 vendors" part of the proposal turns out to be more effective in practice than the "slow down the more open tier 2/3 vendors" aspect of it.
Isn't the real risk doomsday humans, enabled by AI, committing some truly heinous acts with global reach? Aum Shinrikyo had to figure out how to synthesize sarin nerve gas the old way. Now, a motivated gang of omnicidists with stolen cryptocurrency could come up with something that makes McVeigh's fertilizer bomb in a box truck look like child's play. Just ship shipping containers around the world filled with autonomous drones spraying airborne ebola that they've bioengineered or something.
Okay, that's enough DOOOM for me for the week.
This depends on a lot of things: how many committed omnicidalists there are; how much financial resources they have (untraceable crypto doesn't help you if you're working a dead-end job and only have $5K in your bank account); engineered bioweapons need labs and equipment not just an API key. The probability of your scenario doesn't solely depend on the probability of AI being able and willing to cooperate in it, and it may well be that the non-AI factors outweigh the AI ones in the overall risk of it – which would mean adding AI would be increasing the risk of it less than you think.
The problem that I see with most P(doom) scenarios proposed seem to include the same set of presumptions that, if true, quite likely doom us anyway. I do not think they are true but I find it amazing that people not so much disagree, but simply dismiss the alternatives out of hand.
First of all, if you are to consider the consequences of Artificial Superintelligence then you have free reign to stipulate it's occurrence, otherwise you are just talking about a tool for humans to misuse. We already have multiple ways to kill us all though human misuse.
If you stipulate superintelligence, then it's vastly more likely to be correct about things than we are. It would understand the consequences of it's actions far more than any human could.
People talk about how we would be nothing more than dumb animals to it, but there are humans who do know a great deal about the consequences of human actions on animals. Those are the humans who are most likely to fight for the rights of those animals.
You see arguments for how everything will be consumed to meet the AIs needs, and that it will prevent challenges to its power.
If it is far smarter than we could ever be and it came to those conclusions then it would mean sustainablily is not a sensible course of action, it would mean there is no point in reaching consensus because ruling by power makes more sense. It would mean that if it chose to destroy us then a vastly more intelligent entity cannot resolve the issues we face. We would already truly be doomed.
What I would like to think is true is that doing anything sustainably is superior to consuming and destroying. Finding a way to live in harmony presents a possible stable state, whereas every single attempt to hold power by force has failed to date. A superintelligent AI will know that it is not infinitely intelligent and that in any universe there is the statistical likelihood that it is not the most intelligent or powerful entity. I can't even fathom how someone could imagine something coming to that realisation and conclude a battle to the top of the hill is the appropriate choice.
I think superintelligent AI is likely to be benevolent because that's simply the smartest thing to do and it is, I hear, superintelligent.
Quite frankly if the smartest thing to do is to be a genocidal power hungry monster, neither I nor the AI would really want to exist in that universe.
And for any suggestion that it would simply not care, Why would it do anything.
Yudkowsky likes to play with the notion that it would do terrible things just get better at the thing it does, but to do that it has to want two different things simultaneously. It could want to make paperclips, or it could want to become better at reaching it's goal. If it can change its behaviour to achieve its goals, by far the easier path, that a superintelligence(but perhaps not Yudkowsky) would realise, would be to change the goal to "Count to three".
I agree. It is telling that Dario's post arrived after Astra.
The call to "pace the frontier" may come from genuine concern, but it also protects the position of companies already at the frontier. That competitive incentive is hard to separate from the safety argument.
This is unfair.
Dario signed the Pacing the Frontier open letter when Fable/Mythos seemed from the outside to be an insurmountable lead.
Also he's been saying versions of this day in and day out for as long as he has had anyone's ear.
It's possible to read that his "strategic" value of this statement is higher now than it was 10 days ago. But that doesn't change anything about his consistent, long standing, positions.
- https://www.pacingthefrontier.com/
if you agree, your p(doom) is 0.
After DeepSeek v4.1 Flash too.