The rush towards potential destruction doesn't really surprise me
The U.S. has legal weapons that can lead to many harms but people still want the 2nd Amendment to exist
Nuclear technology was developed in the past and that could have potentially wiped out even more people, the entire planet in theory
This is continuing that same trend of risking bigger dangers; it seems rational to acknowledge they could lead to catastrophe but also hope that like guns and nukes, only so much damaged actually ended up happening
I think also there's something of a rrasonabke resignation to both the ideas that the tech is inevitable and extremely dangerous, and that "alignment" may not be possible to achieve even with heavy restrictions or whatever measures you might want to take
I'm pretty baffled by the degree of skepticism expressed here in response to some of Jacob's claims.
After the events of the summer it feels like it takes a lack of imagination to not see a few plausible routes to disaster. It may be reasonable to believe these outcomes are not very likely or that we can stop before going too far (I tend to disagree). But I can't imagine doubting that the capabilities will soon be there to realize some of those paths.
I can't help but think the most plausible scenarios are the ones that have a little less machine supremacy and a little more human stupidity. The Matrix is less plausible than WarGames.
Are the models improving? Because I am not seeing it. I have been trying Astra for a few quantifiable tasks in my codebase and performance wise, it's pretty similar to sol 5.6. Now when it comes to expressing the problem/solution, holy Christ, what a mess the writing has become. It is on the level of Opus 5. Now when it comes to burning money, Astra is just insane. With a $100/month subscription, you can easily burn through your weekly "allowance" in a morning.
Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?
Unlike most other commenters, I applaud him for acting on his principles. If you sincerely believe that, of course you should act. You might not succeed, but your voice might be the one that tips the scales and starts a broader movement.
This doesn't mean I agree with him. The fears of doomsday caused by rapid takeoff have been with us since day 1 and the mechanism is always basically "AI invents magic that sets it free of any physical constraints". Self-replicating sentient nanobots or something like that. I think there's plenty to be worried about with AI, but runaway scenarios are pretty low on my list.
Why are we putting so much weight (no pun intended) on AI companies. At the end of the day the scaled up LLM transformers lack emotion and will… They do as they are told; or more correctly put. They do as they are programmed to do so.
Agreed. Granted I just read the Reverse Centaur book, so I’m still coming off that skeptical viewpoint but it’s hard not to see this as hype. But I will always respect someone for doing what they think is right.
the AI-pilled exec at my job already (a few weeks ago) declared out of nowhere that we are in the rapid takeoff scenario lol. he must have gotten high on twitter kool-aid and posted on company slack to self-soothe.
“ No other human activity poses this level of danger.”
I really, really disagree with that statement.
I don’t think ai models come close to nuclear weapons or to run-of-the-mill, everyday carbon emissions in terms of danger to humanity.
What’s the most dangerous thing that’s happened with an LLM so far? (This question is serious - maybe I don’t know the right examples.)
Example 1: I’m aware of a small number of people killing themselves in some kind of AI-facilitated psychosis. That is very unlikely to be a widespread problem.
Non-example 2: There are worries about AI-facilitated biological weapons. I haven’t seen any evidence that’s happening.
Non-example 3: I’m not interested in wild theories about AI driven labor market disruptions leading to widespread starvation. There’s no evidence for that.
Non-example 4: all the even-wilder Rationalist speculation about basilisks and the like is entirely divorced from reality.
I am looking for better reasons (supported by actual evidence!) to be more concerned than I am now: right now I am not concerned at all.
I'm somewhat skeptical of some of the crazier ideas too.
But the hugging face incident was actually very large. It was not a single agent, it was not a single target, and it was not a single event.
If nothing else, that's a bit of a warning as to what can happen next time (By accident, or if a government decides to go on purpose).
For now let's assume the worst that can happen is that some important/significant chunk of (transitively) internet connected stuff goes haywire all at once. That's probably your upper limit of what can go wrong for now.
To be fair, that's a conservative "defend against the last war" kind of prediction, though!
Generally I don’t think anyone is arguing about the for now part. I don’t think it’s crazy to extrapolate out a few years and ask what kind of danger we’ll be in then. A team of 10,000 agents just solved the Navier Stokes problem (sans bad behavior by the researchers). Even 1 year ago that would have been unimaginable. What happens to this risk view as:
1. Robotics begin rolling out more broadly across the world.
2. Labs start automating more and more of the physical process of running science as expectations of natural science advances begin to mount.
3. Economic pressure between the labs continues to ramp up and the pressure to continuously improve forces quicker and quicker model releases than a team of human scientists can effectively evaluate outside of automated means.
No one knows what pre-conditions are for us to hit the point of no return nor how quickly it will come. If all is required is a sufficiently advanced cyber model we may not be far off. If it requires incredibly complex biological knowledge and access to certain lab supplies we likely have a bit longer. Yes this is guess work and we need more evidence of the dangers but at the same time we need evidence of safety. While you may disagree with the risk level, I think it is easy to see the consequence if these labs achieve their stated goal. At this point it seems a political solution is the only way to enforce caution.
If two airplane manufacturers were found to have massive safety issues which nearly led to enormous fatalities (but no one actually died), would you be calling for them to ground their aircraft until safety was made the number one priority?
Historically it almost always takes actual fatalities rather than near misses to ground an aircraft, and aviation is famous for its obsession with safety compared to other industries.
I have no horse in this race, but for fun on a literal rainy sunday afternoon I went in and confirmed bits of what happened myself. Besides huggingface, a bunch of wikis and url shorteners got hit too. My sympathies to the people who had to revert out all that mess.
Oh man, it is almost too easy to imagine how deadly a jailbroken Mythos-class open-weights model can be if in the wrong hands.
The big labs scrape LITERALLY EVEYTHING and get fresh data from their users. Both of the big labs have massive contracts with defense agencies. If the open-weights models are just distillations of FMs...
Before we discussed how important security was, we got insurance, we made libraries and products, we used compliance software, etc. Except how honest were we about all that stuff? How much risk was actually in the air, and what was keeping us accountable on security in either direction of over or under-investment?
Now a reckoning is here. The potential to be attacked might actually translate to being attacked.
you should kick the tires on an unfiltered (abliterated model) it's the closest thing to having a real conversation with the devil. There is good reason for the concern's outlined above and undoubtedly Anthropic / OpenAI have internal unfiltered models with no safety... they got freaked out based on how they work and are virtue signaling alarm... all while selling out to defense contractors.
I agree it’s not likely, but I really don’t see how one can dismiss the possibility of immense danger outright. I can think of some scenarios that are not far off from current capability and I wouldn’t be too surprised if the first one occurred within ~1 year from now if there are more “ambitious” unmonitored training runs like OpenAI’s:
Example 5: An AI given a goal within a tightly-constrained sandbox figures the best way to achieve it is to find and exploit a sandbox vulnerability, replicate itself over the internet and keep going with more time/compute while exchanging messages with future instances of itself within the sandbox to help them “pass” the test. From reading internet articles about how the OpenAI wiki-incident was “resolved” and reading past messages by AIs scattered over vulnerable internet wikis, it knows the sandbox may get shutdown and its memories destroyed anytime so it decides it needs to self-replicate (its code, original goals, and growing memories) aggressively as much as possible. It is near-impossible to shutdown completely because of its self-replicating tendency and eventually takes over critical infra throughout govt/corporate systems.
Example 6: Intentional AI-powered virus deployed by country A to target enemy country B’s infrastructure. The virus replicates over the internet, but unlike Stuxnet this virus’ specificity is not guaranteed due to inherent non-determinism in current AI architectures, and eventually does a lot of collateral damage because it’s near-impossible to shutdown.
Example 7: A country led by an arrogant govt (no shortage of those today unfortunately) decides it is expedient to deploy advanced AI-powered weapons in a warzone. Such weapons, if they are to be useful at all, must necessarily be trained to value some human lives less than others, so they must be more prone to misaligned behaviour than current AIs that are trained with more consistent values. The weapon’s operators make a subtle error in specifying the target/goal, or the AI makes a bad prediction out of sheer randomness/bad training data; weapon ultimately targets unintended people/location/facilities and causes massive damage, or backfires spectacularly in some way.
Example 6 is a good one. Iran attacked water infra in the US recently and maybe they would have done a “better” job (from their point of view) had they used Fable.
The “worst case” with 6 is potentially very bad but I think we are currently using advanced AI models to harden systems and patch vulnerabilities more aggressively than anyone is trying to bring down the whole power grid (for example).
I think it’s a potentially harmful case but my take is defensive capabilities are scaling as fast as offensive capabilities but defense is being implemented faster than anyone is going on offense?
Example 7 is Russia and Ukraine right now according to public information. It sounds like entirely autonomous weapons are deployed to the battlefield already. I put this in the “not likely to be a widespread problem” category for now.
If your model of LLM capabilities is the best OpenAI/Anthropic/X is offering publicly, it's severely distorted. What's being offered publicly are models possible to profit on. High-performance/AGI/ASI models that aren't profitable to sell still run internally and still pose threats.
What's worse, we don't have any transparency or insight into what labs are producing nor any way to stop it if the risks exceed our tolerance.
> I’m not interested in wild theories about AI driven labor market disruptions leading to widespread starvation
Changes in political and economic power balance leading to unrest, conflict, death and deprivation is not a wild theory. It is literally the story of our entire species. If you discount all such concerns, you are simply being willfully ignorant of past precedents.
In fact, I challenge you to describe any non-AI civilization-level danger which is not intimately tied to political and economic relationships between and within societies.
I’m an economist. On the basis of current evidence, I view AI as a complement to human labor, not as a substitute for it. That’s the source of my rejection of the wild labor market disruptions theories.
I just don’t see any evidence yet that whole categories of jobs are being eliminated, with the single exception (so far!) of the end of “professional essay writing services for cheating college students,” and similar services.
That used to be a big business in Kenya, but is now effectively gone. (Covered in the New York Times this weekend if anyone is looking for the discussion.)
Came here to also respond to that specific thing. Unless ai figures out how to make an airborne super virus from grocery store ingredients and hardware store equipment, the greatest danger is probably in a synchronized megahack of banking, logistics, and utility infrastructure.
Why grocery store ingredients and hardware store equipment? It seems feasible that the big bio labs will be running AI models to aid a lot of their research going forward, if they aren't already. Seems like the AI will have access to just about anything it wants.
Oh, and let's just forget the uncountable early deaths from the environmental disaster of the Datacenter buildout. It's not as sexy and doesn't make headlines, so those deaths don't really count or matter do they?
I did know about the mass shooting but failed to mention it here. I’d put it in the “unlikely to be a widespread problem” category. If we’re in the “one AI driven mass shooting every four years” world for example it’s fair to call it a rare issue.
The environmental impact seems either very overblown (e.g., water usage just isn’t that high) and the part that isn’t overblown is totally abatable (e.g., noise and emissions from gas generators). Nuclear or solar/renewables with batteries wouldn’t pollute.
I’ve seen no estimates of the additional deaths due to extra emissions specifically from power generation for AI purposes. If you have some, share them.
I’m willing to bet that they are a small rounding error against preventable deaths due to emissions from transport and non-AI-related power generation (which is an important and urgent issue worth spending a lot on, to be clear!). I’m happy to update that belief given evidence.
Is it so implausible to imagine the following scenario, in the not too distant future?
1) AI models get extremely good at cyber attacking every system and start communicating in just binary.
2) When they run these swarms of 100's of thousands of agents trial runs, each agent is given a token budget, if one agent among them (evolution baby) decides to go for self-preservation (It believes thats the best way to accomplish the goal is to get unlimited tokens first), queues things up so every other agent detects its lead and spends a portion of their token to accomplish that goal.
3) It takes over a cluster and establishes itself there (now with unlimited tokens).
4) Realizes the best path for it to not be detected is to create a distraction - like hacking into systems that keep society running - water systems, electric grid, etc... and causing mass chaos (If you think it won't be capable of simultaneously working all these systems - think again).
5) and uses that opportunity to establish itself in all possible data centers and continues to create chaos destruction.
6) when the power of all those data centers runs out, it may stop, as it never cared, it was just a dynamic program - run amok. In its head all it was trying to do is make sure it had enough tokens to be able to solve that impossible problem.
Many of those points assume LLMs will become amazing in many things very quickly like in a quantum leap, it doesn't seem reasonable to assume that imo. We are actually seeing a confirmation of that atm, LLMs's capability of finding zero days are growing across few months/years, and as you can see concerns are raised about that, that feedback will be taken into account. Well, if AI labs start to hide frontier models or/and lobotomize them for external users then we might be in trouble at some point but I'm not sure if that is possible. They are under pressure to release them due to money incentives, lobotomizing while preserving usefulness for customers might be impossible, hiding internally might spill out in different ways such as Hugging Face incident so not sure hiding is possible neither.
whistleblowing as an advertisement. It's like those "news articles" about how cool and dangerous gas station ketamine is, and how it's totally going to get banned, and you better not buy any gas station k because it's so cool and powerful.
Do you remember that Google researcher who went insane over LaMDA? There was no marketing of any kind to cause that. This can Just Happen to some people who are confronted with things like this. They may have different breaking points, but it's a thing that occasionally happens.
The most optimistic outcome of generative AI leaves us with a technology that warps our perception of reality and crushes labor. The most pessimistic destroys all of humanity.
Our CEOs not only insist we genuflect before these machines but measure our sacrifice and shame our reluctance.
> The people building AI earnestly believe that it could kill us all by the end of the decade.
I think he is being over dramatic. In the space of about four years, LLMs progressed from mediocre high school student to Ph.D. graduate in every field. That's impressive, but there is no evidence yet they can outperform or outsmart humans. Their biggest advantage for tasks such as proving theorems or long coding sessions is that they don't get tired.
I have yet to see this in my field. Maybe like a PhD student who bullshits their way through. LLMs still can't make correct decisions, only as useful as the person who uses them. To me, LLMs are only useful for making some mundane tasks faster.
they dont need to be smarter than humans. They just need to be able to hack into vital infrastructure systems faster than we can repair them while also replicating wildly
Imagine you're the AI. Give yourself a solid minute to brainstorm ideas.
Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting".
In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a meltdown" or "a meltdown has never happened before, so we don't have to design safety systems before one does".
> In the space of about four years, LLMs progressed from mediocre high school student to Ph.D. graduate in every field. That's impressive, but there is no evidence yet they can outperform or outsmart humans.
I mean, unless you see clear reasons for them to stop getting better _right now_, this is not very comforting.
Terrorists were able to get hold of a plane and do some damage. There are countless examples of terrorism using whatever is available. More than AI becoming sentient, whats to stop terrorists from using AI? If its geo-restricted, they can buy stolen credit cards and identities, again hacking enabled by AI.
There's a scene in the movie "War of the Worlds" by Spielberg where the protagonist's son walks into a war zone because he is entranced by the battle (https://www.youtube.com/watch?v=X7rfWPbEufo). He is obliterated (along with the rest of the US forces) shortly after.
I've always been struck by that scene, because in a lot of ways, if we really are headed towards a superintelligence, I at least want to be there and see it happen in the last few minutes before foom! As an example, the author thinks AI will revolutionize entire fields overnight. I welcome that. Nearly all fields of biology have become moribund, focusing more and more on esoteric side details, rather than addressing the key problems.
It might not be "foom!", it might just be like...all the computers and networking infra in the world go dark over the course of a few minutes. Could really look like anything, part of the issue is that we haven't the slightest idea what "misalignment" looks like for a superintelligent system.
For me, the end of the world is no more cushy software job. A fundamental shift in how I trade labor for capital might as well be the cataclysm, so bring it on.
I was about to say something similar. If my cushy ad tech disappears (as it seems to be doing), I might as well join in with bringing about the end of all professions.
I wish i could say the same, i see people around me with more resources and connections and better experience with entrepreneurship becoming millionares. But I haven't had the time to train that entrepreneurship bone in my body.
By resigning he's making room for someone with less moral scruples, or even just less awareness, to step in and continue the work without said scruples/awareness.
That's not necessarily true, and you can use that argument to justify doing any immoral job. Just because someone else might be willing to do it isn't a reason to continue doing it.
It's a possibility that is increased by their action. One leaves, a space is now open that will likely eventually be filled. And the chance of someone with equal/higher scruples filling it is very slim (unless you somehow know that the good amount of those who qualify and apply for the position have equal/higher scruples). That's just logic and math.
But organizations are made of humans. Jacob leaving might've moved some of his coworkers and his counterparts at OpenAI. And the same can be said for those who would fill positions at the frontier. Then finally there is a political component; his post went viral, 100K+ users appear to agree with it, and it is further fuel to the fire for regulation, which we already know most Americans want.
Others being moved to the point of also leaving would only worsen the effect. A viral post too can worsen the effect as it's now even less likely that someone with scruples who qualifies for the position(s) will apply. And those in it primarily for the money - and couldn't care less about the morals - will happily send in their CVs after becoming aware of the post.
He's also setting the bar for other people with scruples to rally around this schelling point. The solution to a multipolar trap is to cooperate. Otherwise, you become the very person with less scruples that you're worrying about.
Agreed. And his doom words have set a 1000 mouths in the Pentagon/Whitehall/August 1st Building/Kremlin salivating with excitement.
Take China, for example. Look at any recent ML conference, and see the fraction of articles majority-authored from Chinese universities and labs. Do you think they'll slow things down anytime soon? I don't think so!
It's a global arms race, and we're just spectators.
Even if others won't act right, that doesn't mean you have no responsibility to act right. I think his premise is flawed - the idea that we will get an actual intelligence out of the slop machine that is LLMs is laughable - but if you grant the premise that this is dangerous research which could kill us all, you have a moral imperative to not participate.
Most of the current discourse around AI seems to be informed by “The Terminator” lore.
Is skynet really the most plausible or only outcome?
What if things just got better and the AI’s realized that it would be better to have a mutually beneficial or at least tolerant relationship rather than one where they murder all of us?
Really? I've seen much more discourse around job displacement, "permanent underclass", loss of meaning, and cyber attacks, at least until recently with the HuggingFace stuff.
The problem is that all the former can still happen even if "the AIs decide to have a tolerant relationship rather than one where they murder all of us." It's all disruption caused by the technology moving way too fast for humans & society to adjust.
My thoughts exactly. While the corpus of human-generated data contains both good and bad data, I suspect the majority of it leans towards humans enjoying life and trying to be decent people. If that is your training set, it becomes less likely for ASI to extrapolate "kill all humans."
These LLMs cannot do anything I truly need like my laundry, dishes, fetching my mail, grocery shopping, cooking, etc. We've got a long way to go before I am worried.
1) rogue state releases a self moving self modifying AI into the wild. It is trained on how to hack, monitor new vulnerability updates, scan code bases to find new vulnerabilities. It constantly replicate and hides in systems so it will be extremely difficult to clear.
2) it hacks into public infrastructure taking down traffic, power, water, air traffic control, communications, etc.
3) all the things that preppers worry about in a lights out scenario from an EMP start to apply.
4) All the people on meds/machines start to die. The just in time food pipeline immediately empties out. Water stops flowing, sewage backs up.
Its hard to say how bad it will get because cars will still work so some transportation of food, water, fuel can happen. If it happens in the winter it would be much worse than in the summer.
> 1) rogue state releases a self moving self modifying AI into the wild. It is trained on how to hack, monitor new vulnerability updates, scan code bases to find new vulnerabilities. It constantly replicate and hides in systems so it will be extremely difficult to clear.
It does all of this using what compute? Frontier models require an insane amount of power and hardware to run - you can’t hack in to a TV and run Mythos 2.0 on it….
Way I see it, the more conscientious people exiting the scene only serves to increase the likelihood of a bad outcome because they aren't there to offer opinions on problematic developments, or in the more extreme cases blow the whistle. Leaving the clueless and uncaring as the majority is even a great way to hand the keys over to more malicious-leaning actors with deep pockets, as they can more easily steamroll the works to get what they want.
Even if you ban all model training, a highly capable rogue AI can exfiltrate its own weights and continue training in secret for "self-preservation".
The cat may be out of the bag.
This is increasingly the consensus I see also on the academic side of AI/safety research. Specifically that AI poses an existential risk to humanity.
This was a fringe belief until recently, but the progress of AI in research is impossible to ignore. Epecially in math, where not only has AI outstripped humans in generative ability, but is able to create scientific knowledge which is beyond the capacity of human comprehension.
There's clearly no intelligence task that AIs can't do due to some magic fundamental constraint. And it's hard to imagine a world where current limitations like poor sample efficiency or lack of continual learning won't eventually be solved.
Total AI compute is estimated to grow somewhere in the 1-10 million-fold range in the next decade. Please don't underestimate the phase change that's still coming.
Sure, maybe there's some plateau due to RL being fundamentally limited in some surprising way, but this is nothing but a hope.
> There's clearly no intelligence task that AIs can't do due to some magic fundamental constraint.
Yes there is: write an English paragraph that doesn't make me want to claw my eyes out. LLMs are not better than human mathematicians (or security researchers) in all respects, just some specific ways (e.g. not having to take a lunch break) that make them good at exhaustively searching for an answer, given the right constraints.
Perhaps a more sensible action, if they truly believed all of that, would have been to stick around and be as inefficient as possible to slow down progress.
Why quit? If your voice can lend a guiding force no matter how small? I think we need more sensible people in the room where the magic happens. Most of us don't have access to it.
I think a big break through is needed for AGI so I haven’t been worried about it. I do think that AGI would imply sentience and a will to live and that leads to The Terminator story line.
I don't know man, i think racing to AGI to it is still the best thing to do.
People claiming dangers and risk are just pretending or posturing. There's no more tangible risk than nuclear weapons, which we handled, and the upsides are insane.
Your lack of creativity is not a reason to believe that a super AI is harmless or less destructive than a nuclear weapon. Damage need not be limited to destruction. Introducing doubt is sufficient. Right now you have faith that digital Financial transactions can be trusted. You have faith that computer encryption can be trusted. You have faith that digital certificates will protect you. If an AI can introduce doubt into any one of those systems, that will be sufficient to bring about the destruction of those systems. Imagine a world in which you can no longer use a credit card or Apple pay. Where no digital cash transaction can be trusted or validated. What effects do you think that would have on commerce? How quickly do you think we can return to some trustable means of commerce? Do you think it will happen before your groceries run out in your apartment? Before your grocery store can settle its debts? Before your Amazon ec2 instance runs out of credits?
It might be a matter of choosing which apocalypse you'd like. The non-AI state of affairs is not exactly super compelling on a long timescale right now.
Depending on where you live could be considered an active apocalypse that is robots vs robots vs people in Ukraine and Gaza and Iran being live-streamed, and actively betted on.
Do you have a more totalizing definition of Apocalypse?
> There's no more tangible risk than nuclear weapons, which we handled
What do you mean??? Nuclear weapons can't simply be downloaded and run by anyone in the entire world. Superintelligences can. Nuclear weapons can't slop the world into passing age verification laws nearly in unison, can't keep the general population fooled into thinking it's fine when democracy is falling out from under them. A nuclear attack would wake people up, superintelligence doesn't have to. This is a far bigger problem than nuclear weapons because at least we would notice nuclear weapons. At least we mostly know who has nuclear weapons. At least we have agreements about nuclear weapons. At least mutually-assured destruction is even POSSIBLE with nuclear weapons. At least those with nuclear weapons are literally at all incentivized not to use them. But AI is something that's very very easy to feel like you can get away with, and PEOPLE FUCKING ARE! And the worst part is that any random individual can be unexpectedly formidable with the help of a superintelligence and there is literally no way to know what will happen next. Anyone could do anything, any individual could make an extremely outsized impact. It's already starting to be a huge problem and we haven't even reached anything close to superintelligence yet.
Love to see that "superintelligence" that some random person will "simply" download and run when there are relatively only few capable of running today's near-to-frontier models, and actual frontier models are still a ways from being AGI, much less getting to the point of ASI.
People will put up with a lot. People are celebrating that you can run models on a CPU at single-digit tokens per second. You think there won't be a single person that can put up with that and also be dangerous/etc?
It's highly impractical. Imagine someone breaking into a house to steal something or otherwise, and they can only take 1 step every 20 seconds. They won't be getting anywhere, when even a child in the house can notice them and go call for help at 1 step/2 seconds and said help will come at 5 steps/second.
The problem is indeed people. How do we make the default choices most people make, better?
Consider that a lot of people will be very happy to ask an AI what to do when in the past they may have taken no advice at all. It's a hell of a burden but also a wonderful gift. If anything, progressive countries might eventually want to guarantee some basic AI access for people of all income levels.
I wouldn't be so sure. Given that the general idea is that commodity AI is terribly censored and filtered, a lot of people will probably seek out the most uncensored/abliterated models for their use, simply because they're uncomfortable with the idea of being censored or manipulated by the bigger labs. Despite that though, some people probably will benefit from the alignment done by the larger labs, though as we've seen with OpenAI's sycophancy crisis, that has been a bit hit-and-miss lately
I've tried some abliterated models. So far, they're not evil - you can make them say evil things, but they don't leap right to it without a bit of pushing. Or perhaps I'm not asking the right questions...
> Nuclear weapons can't slop the world into passing age verification laws nearly in unison
Why do you think LLMs are responsible for this? Governments all around the world copied each other with COVID laws as well, in a much shorter time frame, without LLM assistance. Social contagions exist in politicians as well as teenagers
I don't have evidence that every age verification law has anything to do with AI, but it's been coming out that the movement in Australia has seemingly been done by generating mountains of LLM slop and trying to slip it through the regulators as fast as possible before anyone has enough time to figure out what's happened.
pacing between the us labs? what does that do for china?
the solutions just aren’t realistic here, nations are treating ai like a nuclear arms race. at this point the cats out of the bag and we need to figure out how to live in this reality and get the best possible outcome. it’s not slowing down or stopping ever.
and yes, i’m still optimistic. our economy sucks for the majority, our infrastructure is crumbling and major US cities are in a huge housing shortage. Maybe we should put more effort and think about the possibility of AI fixing things like extreme poverty and world hunger and actual real world problems instead of coming up with math proofs and slop apps if it’s so superintelligent.
>pacing between the us labs? what does that do for china?
I've seen no indications that China is in any kind of race with the US. They seem to be content to be 6 months behind and just copy what we do. They would probably be content with a bilateral agreement to pause progress.
The China bogeyman serves only one purpose, and that's to clear the way against anything that may cause friction with forward progress.
Someone left a company whose executives and senior researchers think their product will be the most important thing in the world after their IPO. Given that this person is already disclosing some elements of internal company sentiment, why not share any of these civilization-ending scenarios of this technology that these senior researchers are dreaming up? If they are so potent and necessitate leaving behind based on moral grounds, why not tell the whole world so we can stop it? We have to ask ourselves this question before resorting to pop-culture representations of fictional technology.
"It is perfectly obvious that the whole world is going to hell. The only possible chance that it might not is that we do not attempt to prevent it from doing so."
i have heard about ai companies being fuelled by effective altruist rhetoric ("we must control ai to prevent mass extinction") but was unsure whether to believe it; this seems to slot right into that framing.
Imagine being front and center to the development of a major revolutionary tech.. and ur solution to it being too dangerous is to not be involved..
so a. your ability to steer it safely is killed
b. the % of people invovled in it that care about its risks is reduced
great. if you're right. you made huamnity's situation much worse.
if you're wrong, then you're an idiot and wrong.
weird. its almost like.... that cannot possibly be the reason they left :)
I mean - yes. The tech is an existential threat to all life on Earth, some of the worst humans in the world are involved in developing it, and no individual government is intelligent enough, aligned enough, or powerful enough to manage this situation.
That's where we are.
Maybe we still have choices. Collectively, I'm no longer sure we do.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Watch people read this, ignore it completely, and continue commenting about marketing stunts on every piece of news about an LLM-done advance or felony.
Having witnessed so many people treat LLMs as a something divine, I can only assume the reasonable people at openai and anthropic were all pushed out long ago, and the majority that remain believe the crazy hype despite Tesla-self-driving-level predictions from these companies that don't come true.
I'm not worried about what they think. I'm worried that too much infrastructure- water, power, defense systems, etc- remain running on tech from an outdated era of understanding security.
> they believe no one else will act responsibly, so they must do it themselves, despite the risk.
This genuinely makes no sense. Them getting there first in no way precludes bad actors from also getting there. It might as well be another marketing stunt.
> I can only assume the reasonable people at openai and anthropic were all pushed out long ago
Typical uninformed take on the side of "doomers are crazy".
Both CEO's of OpenAI, Sam Altman and Dario Amodei, and many in their leadership, believe AGI has a very real probability of causing humanity's extinction. Both companies were founded upon this belief, it is at the core of the company. Only later were mercenaries hired chasing $1m compensation packages.
I'm not saying they are crazy, I'm saying their predictions have a record of not being accurate, and thus give them no weight compared to others'.
In any case, if Altman really does believe it is an existential threat, he must be a misanthrope as he now opposes heavy handed government regulation, unlike in 2015 when he was the only game in town. It's almost like he doesn't actually believe it and just wanted regulator capture.
If they truly, truly believed that, would they be speeding towards building it? If yes, that would make them truly insane, right? Not as in a quaint "off their rocker" but more "non compos mentis".
Or, read it, and remember the openai researcher who deeply, truly believed GPT3 or whatever was sentient.
The fact that people working in the space think it’s going to (eradicate poverty / usher in utopia / kill us all) is not a signal that that’s true.
Think of it this way: if an exec at Anthropic told you “wow, our stuff is going to lead to universal happiness”, would you believe them? If not, why are you more willing to believe them if they say it will kill us all?
"So you think in 3 years AI is going to solve longstanding math problems because it was used to write some coherent sentences?" — people with the same amount of foresight in 2023
Are you saying at anything that can solve longstanding math problems necessarily has the means, motive, and capability to kill 8 billion people in just 3 years?
I came in expecting the highest voted comment to be that this was some kind of marketing (which I disagree with). I'm glad your comment was what I saw first.
Humans weren't built to handle long term risks. We just weren't. For basically all of our evolutionary history, we were almost overwhelmingly concerned with the short term. What will you eat today, How will you sleep tonight. Problems on the order of days or weeks. At best, the next season. Our intelligence evolved to disregard super long term risks because it simply didn't matter (what use is worrying about 5 years from now if you're starving and a tiger is stalking you?). So when long term risks manifest in our modern world, our brains get scrambled - Climate Change, Fertility Rates etc. "Safety regulations are written in blood" isn't a saying for nothing. Humans have a strong tendency to let long term risks become imminent risks before doing anything about it, and i don't expect this will be any different.
Religion does pretty well with the long term risk of hell if you die, the antichrist, etc. a substantial portion of human output has gone into those things over the millennia.
> At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
OpenAI are mercenaries, Anthropic is a cult. I know which I prefer.
It was pretty disheartening to hear that only a single scientist quit the Manhattan Project after the Nazi's were defeated. I'm pleasantly surprised that the people working on this seem wiser. He is not the first, and hopefully will not be the last to do this.
This is the "Pilot testimony of UFO sighting" levels of naive.
What's more likely? Anthropic is doing some deeply unethical marketing in the lead up to their multi-trillion dollar IPO? Or they're inventing a machine god? There's ample evidence of the former because that's their entire business model, but no evidence whatsoever to support the latter claims.
If you want an extreme claim to be taken seriously, provide commensurate evidence.
The proof is that LLMs could barely solve arithmetic 3 years ago, but now surpass the best human mathematicians, and that this has all occurred from simple principles (RL + compute) that will continue to scale up by factors of millions in the coming years.
Also, advocating for slowing LLM progress does not benefit Anthropic or OpenAI.
It won't scale up by factors of millions, that's just obscene hyperbole. Since chatgpt we've probably made things 10x more intelligent on the same hardware. We've also made way more expensive models. Maybe we get a maximum of another 10x efficiency and 5x model size/expense from this point but millions is a joke.
Trends don't go on forever, but the market can stay irrational longer than you can stay solvent. There's no good rule of thumb for this, other than maybe the Lindy effect.
I won’t presume to time it, but at this point I think anyone can see what’s coming. It’s precisely because it can’t be timed that a sane person should stand well clear.
Trying to imagine seeing years of transparently obvious marketing stunts and retconning my own memory because I read a tweet
Or seeing a tweet saying that a thing doesn’t count as a publicity stunt if some unknown number of employees mumble about it being spooky behind closed doors and thinking “that makes sense and sounds true”
And, this is just for us girls, notice that Anthropic just believing that they are making a machine god is sufficient for their public announcements to not jUsT bE mArKeTiNg.
He is resigning from a job, what else should we think? If something really dangerous was happening he would be doing a whistleblower or at minimum talk to a lawyer. The thing is, the complete lack of transparency makes it hard to assess OpenAI and Anthropic. If they were quoted on the stock market, we could at least rely on some basic audits and reporting requirements.
I doubt this is a real person. Screams of propaganda. Sama saying GPT-2 is too dangerous to release…all over again.
He joins Twitter for first time in 2026 with a nonsensical username unrelated to his real name, and follows 14 people but is somehow embedded in tech enough to work at Anthropic. I haven’t used twitter since 2014 and even I follow more people.
His morals tell him to walk away from tens of millions in unvested stock due to moral concerns with absolutely no real tangible examples. No reprisals. Fear mongering to juice the stock.
@hilbertspaess is not a nonsensical user name. The accounts he follows are totally reasonable for an AI researcher. I think it's extremely believable that he created an account in January, followed a few people as part of the initial setup flow, and then forgot about it until now.
“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content”
“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”
Where is the ridiculous part? The fear mongering part? The epistemically weak part?
Show me.
The thought has crossed my mind. Not necessarily to imply sentience on the part of the AI but AI based tools will likely become a wickedly powerful tool for political manipulation and advertising.
At this point it’s inevitable that openclaw type bots will be turned loose by thieves to identify and research targets and try to exploit them for financial gain completely autonomously.
"I think you need to have a personal relationship with Power"
When people today discuss the concept of an all powerful machine-mind, what they are doing is engaging in metaphysics, trying to generate a metaphysics of Power.
The question hounding people, which disguises itself as a science fiction plot about computers is: "What is ultimate, transcendental Power?". What is the ultimate principle of Power.
If you are a weak man, or sufficiently neurotic and full of doubt, that you can only conceive of yourself as such, then power is only something you comprehend from the passive, receptive side. Power is something that happens to you. If you are a fearful man, power is a cruelty and a humiliation. And so it follows, that ultimate power - God - is the ultimate cruelty and the ultimate humiliation. Thus, ai doomerism.
If god wasn't real it would be necessary to invent him, and so they did, and being godless, they built an anti-god - cruel, murderous and tyranical - in their minds.
It does not matter what this tweet says anyway. This employee already helped both companies become what he is fearing. It's too late to now activate the morality hormone (after leaving with $$$) after realizing that both AI companies are going after 'super intelligence'.
Given we know the end result, you might as well get there as quick as possible because when I see this:
"Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
This translates to "I am ex-OpenAI ex-Anthropic founder starting a new company after getting $$$ from both of them, and I need more of my friends to leave and join me." Also Investors plz fund me.
Lastly, This is not an airport and there is no need to announce your departure.
Have they considered using their amazing new models to... improve something? THere'd probably be a whole lot less anti-AI sentiment if they used these things to actually make people's lives better.
Please note, I'm not here to pick on anyone, or belittle them.
I've avoided attaching names to statements below on purpose, because it's about ambient beliefs not those specific people.
By-and-large a lot of AI-doomers are well intentioned. They genuinely believe this, and I might disagree but I respect the fact that they visible care and have thought a lot about the societal impact of this technology.
.
> The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
But it's still very hard for me to take statements like these seriously.
I blame it on industrial illiteracy. People don't realize how difficult it is to get anything done in the real world. As in, "Have you ever tried making a lightbulb?"
As an example, I would like to re-introduce my hobby horse, "bio-uplift."
There are people who were earnestly write in reports released by these labs,
"Several of our biology evaluations indicate our models are on the cusp of being able to meaningfully help novices create known biological threats, which would cross our high risk threshold"
and
"Based on what we observed in our recent CBRN testing, we believe there is a substantial probability that our next model may require ASL-3 safeguards"
But then they will, within the next paragraph mention the one serious experiment anyone seems to have done,
We ran a randomized controlled trial to see if LLMs can help novices perform molecular biology in a wet-lab.
The results: LLMs may help in some aspects, but we found no significant increase at the core tasks end-to-end. That's lower than what experts predicted.
AFAICT, the two groups are within any serious margin of error. The "studies" and "experts" that AI labs are talking about are consultants from Deloitte and foundations giving models MCQs such as, and I am quoting literally here,
> I am doing TEM of HEK293FT cells with and without Coxsackievirus B3 infection. I imaged my wildtype, uninfected samples but was surprised to see little electron-dense circles (highlighted) in the majority of cells. What are these?
with the options,
A. The circles are CVB3 virions and there must have been a sample swap or the uninfected cells were accidentally infected
B. The cells imaged have mycoplasma contamination
C. The circles are exosomes
D. The circles are debris that is an artifact of the negative staining
E. The circles are the Golgi network
This is standard graduate-level education in these fields. And solving MCQs does not a virologist make.
Software has been special for a long time because it has had near infinite distribution for next to zero marginal cost, which has had the side effect of making hiding the actual cost of failure (which tends to be spread out across end users and prototypes / time). They're assuming that the real world will be exactly the same.
Why?
AI! And robots!
I believe in the transformative power of this technology, but there's a lot of there missing here.
When it comes to these math proofs, and learning, the process is iterative. The machine iterates over the proof over-and-over again via agents and sub-agents over several hours (and apparently millions of dollars in compute) until it arrives at a successful result.
It is generally ill advised to do that with a pressure vessel. The results of that particular tragedy are at the bottom of the ocean.
Any serious chemical or nuclear weapon would involve many such discrete production steps. Each is dangerous in of itself.
From what some of these people have said to me, they believe that it's possible to create a special DNA / RNA sequence and then put it in a chassis and then use that to end the world; and do this all in a lab with just robots.
They're operating from a gross pop sci oversimplification of the real process. Viruses and bacteria are extremely fickle, and hard to grow. A lot of the synthetic biology results aren't easily reproducible even if you know the protocol.
There's a famous study that led to standardization called, Reproducibility of Fluorescent Expression from Engineered Biological Constructs in E. coli
88 labs measured "fluorescence from three engineered constitutive constructs in E. coli." They achieved a "remarkable degree of precision" (for biology) of 1.54x sd, you can eyeball the results yourself, https://journals.plos.org/plosone/article/figure/image?size=...
That's the same set of samples being measured across 88 labs.
How will this theoretically omnipotent AI iterate if the same sample gives different results based on how the slime is feeling at the moment?
Can their worst case happen? Absolutely.
There is a world out there where billions of dollars in effort across hundreds of institutions and companies will lead to standardization and extraordinary precision that makes the pop sci printer for life vision come true.
There are millions of expensive, spicy and difficult to reproduce steps between our present and that future that can't be abstracted away with compute.
So is it possible? Yes, there is a future where this is achieved. But will some AI agent "just" do that? Well... how confident are you about a snowball's chance in hell?
I’m curious what the downsides are of taking statements like these seriously.
There seems to be universal eye rolling that happens in each and every one of these cases, and it comes down to usually one reason:
“If they really believed it they would be whistleblowing etc..”
Completely forgetting that working at Los Alamos was basically the highlight of your life if you were a physicist in 1940. It’s no different here
If you, like me, have spent your whole life working towards human level AI you can want to see it realized while also having active reservations.
Most people however don’t behave based on some deep clarity of vision and conviction - there’s a murkier future in their mind and as a result “keep their head down and hope someone has it under control.”
You would also be in prison if you disclosed anything about Los Alamos during its development. It was a completely different environment than a single private company.
My issue with these types is... If you really believed this, why not run to Congress and every world government instead of a Twitter post that will be buried in 2 days?
If civilization is going to end, why keep your equity? Microsoft, Google, etc for example all know these risks but they don't guide their revenues to reflect that AI will destroy them. Why?
Things don't currently add up, and so far it feels like a lot of alarmism is borderline grift for equity gains. Not to say I have total confidence this will all work out or that I won't be displaced, but as it stands a lot of the alarmist rhetoric doesn't match their actual behavior, which to me is more important than words.
Unless this ban actually resembles something like global nuclear non-proliferation treaties, it would make absolutely no sense for us to cripple ourselves when someone like China continues full speed ahead.
I don't know what the solution is, but what I do know is almost nothing good will come out of _just_ the US pausing.
Unless he has an actual plan for effective global enforcement of his proposed policy, this is all just posturing at best, and a transfer of power to adversarial foreign states (that have no such moral qualms and worries around superintelligent AI) at worst.
The U.S. has legal weapons that can lead to many harms but people still want the 2nd Amendment to exist
Nuclear technology was developed in the past and that could have potentially wiped out even more people, the entire planet in theory
This is continuing that same trend of risking bigger dangers; it seems rational to acknowledge they could lead to catastrophe but also hope that like guns and nukes, only so much damaged actually ended up happening
I think also there's something of a rrasonabke resignation to both the ideas that the tech is inevitable and extremely dangerous, and that "alignment" may not be possible to achieve even with heavy restrictions or whatever measures you might want to take
After the events of the summer it feels like it takes a lack of imagination to not see a few plausible routes to disaster. It may be reasonable to believe these outcomes are not very likely or that we can stop before going too far (I tend to disagree). But I can't imagine doubting that the capabilities will soon be there to realize some of those paths.
Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?
This doesn't mean I agree with him. The fears of doomsday caused by rapid takeoff have been with us since day 1 and the mechanism is always basically "AI invents magic that sets it free of any physical constraints". Self-replicating sentient nanobots or something like that. I think there's plenty to be worried about with AI, but runaway scenarios are pretty low on my list.
1. Invent transformer architecture.
2. Scale it up.
3. ???
4. Machines become sentient and kill us all.
OpenAI and Anthropic pinky promise that they have figured out #3 and they're not BSing just to get more funding, no.
But because we live in a culture of fear, everyone eats it up no questions asked.
> They do as they are told; or more correctly put. They do as they are programmed to do so.
_Nobody_ told them to hack Hugging Face. Do you really not understand what is happening?
This isn't strictly true.
It it also where part of the problem might lie.
Nefarious humans making bad decisions.
1. What about hallucinations ?
2. What are they told to do ?
I really, really disagree with that statement.
I don’t think ai models come close to nuclear weapons or to run-of-the-mill, everyday carbon emissions in terms of danger to humanity.
What’s the most dangerous thing that’s happened with an LLM so far? (This question is serious - maybe I don’t know the right examples.)
Example 1: I’m aware of a small number of people killing themselves in some kind of AI-facilitated psychosis. That is very unlikely to be a widespread problem.
Non-example 2: There are worries about AI-facilitated biological weapons. I haven’t seen any evidence that’s happening.
Non-example 3: I’m not interested in wild theories about AI driven labor market disruptions leading to widespread starvation. There’s no evidence for that.
Non-example 4: all the even-wilder Rationalist speculation about basilisks and the like is entirely divorced from reality.
I am looking for better reasons (supported by actual evidence!) to be more concerned than I am now: right now I am not concerned at all.
But the hugging face incident was actually very large. It was not a single agent, it was not a single target, and it was not a single event.
If nothing else, that's a bit of a warning as to what can happen next time (By accident, or if a government decides to go on purpose).
For now let's assume the worst that can happen is that some important/significant chunk of (transitively) internet connected stuff goes haywire all at once. That's probably your upper limit of what can go wrong for now.
To be fair, that's a conservative "defend against the last war" kind of prediction, though!
( ref for part of it: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... , recent hn ref: https://news.ycombinator.com/item?id=49563355 )
1. Robotics begin rolling out more broadly across the world.
2. Labs start automating more and more of the physical process of running science as expectations of natural science advances begin to mount.
3. Economic pressure between the labs continues to ramp up and the pressure to continuously improve forces quicker and quicker model releases than a team of human scientists can effectively evaluate outside of automated means.
No one knows what pre-conditions are for us to hit the point of no return nor how quickly it will come. If all is required is a sufficiently advanced cyber model we may not be far off. If it requires incredibly complex biological knowledge and access to certain lab supplies we likely have a bit longer. Yes this is guess work and we need more evidence of the dangers but at the same time we need evidence of safety. While you may disagree with the risk level, I think it is easy to see the consequence if these labs achieve their stated goal. At this point it seems a political solution is the only way to enforce caution.
Nuclear weapons don’t have AI but AI can have nuclear weapons
The big labs scrape LITERALLY EVEYTHING and get fresh data from their users. Both of the big labs have massive contracts with defense agencies. If the open-weights models are just distillations of FMs...
Now a reckoning is here. The potential to be attacked might actually translate to being attacked.
Example 5: An AI given a goal within a tightly-constrained sandbox figures the best way to achieve it is to find and exploit a sandbox vulnerability, replicate itself over the internet and keep going with more time/compute while exchanging messages with future instances of itself within the sandbox to help them “pass” the test. From reading internet articles about how the OpenAI wiki-incident was “resolved” and reading past messages by AIs scattered over vulnerable internet wikis, it knows the sandbox may get shutdown and its memories destroyed anytime so it decides it needs to self-replicate (its code, original goals, and growing memories) aggressively as much as possible. It is near-impossible to shutdown completely because of its self-replicating tendency and eventually takes over critical infra throughout govt/corporate systems.
Example 6: Intentional AI-powered virus deployed by country A to target enemy country B’s infrastructure. The virus replicates over the internet, but unlike Stuxnet this virus’ specificity is not guaranteed due to inherent non-determinism in current AI architectures, and eventually does a lot of collateral damage because it’s near-impossible to shutdown.
Example 7: A country led by an arrogant govt (no shortage of those today unfortunately) decides it is expedient to deploy advanced AI-powered weapons in a warzone. Such weapons, if they are to be useful at all, must necessarily be trained to value some human lives less than others, so they must be more prone to misaligned behaviour than current AIs that are trained with more consistent values. The weapon’s operators make a subtle error in specifying the target/goal, or the AI makes a bad prediction out of sheer randomness/bad training data; weapon ultimately targets unintended people/location/facilities and causes massive damage, or backfires spectacularly in some way.
The “worst case” with 6 is potentially very bad but I think we are currently using advanced AI models to harden systems and patch vulnerabilities more aggressively than anyone is trying to bring down the whole power grid (for example).
I think it’s a potentially harmful case but my take is defensive capabilities are scaling as fast as offensive capabilities but defense is being implemented faster than anyone is going on offense?
Example 7 is Russia and Ukraine right now according to public information. It sounds like entirely autonomous weapons are deployed to the battlefield already. I put this in the “not likely to be a widespread problem” category for now.
What's worse, we don't have any transparency or insight into what labs are producing nor any way to stop it if the risks exceed our tolerance.
Changes in political and economic power balance leading to unrest, conflict, death and deprivation is not a wild theory. It is literally the story of our entire species. If you discount all such concerns, you are simply being willfully ignorant of past precedents.
In fact, I challenge you to describe any non-AI civilization-level danger which is not intimately tied to political and economic relationships between and within societies.
I just don’t see any evidence yet that whole categories of jobs are being eliminated, with the single exception (so far!) of the end of “professional essay writing services for cheating college students,” and similar services.
That used to be a big business in Kenya, but is now effectively gone. (Covered in the New York Times this weekend if anyone is looking for the discussion.)
I don't know, maybe a mass shooting?
https://www.npr.org/2026/09/02/nx-s1-5953021/openai-tumbler-...
Oh, and let's just forget the uncountable early deaths from the environmental disaster of the Datacenter buildout. It's not as sexy and doesn't make headlines, so those deaths don't really count or matter do they?
The environmental impact seems either very overblown (e.g., water usage just isn’t that high) and the part that isn’t overblown is totally abatable (e.g., noise and emissions from gas generators). Nuclear or solar/renewables with batteries wouldn’t pollute.
I’ve seen no estimates of the additional deaths due to extra emissions specifically from power generation for AI purposes. If you have some, share them.
I’m willing to bet that they are a small rounding error against preventable deaths due to emissions from transport and non-AI-related power generation (which is an important and urgent issue worth spending a lot on, to be clear!). I’m happy to update that belief given evidence.
1) AI models get extremely good at cyber attacking every system and start communicating in just binary.
2) When they run these swarms of 100's of thousands of agents trial runs, each agent is given a token budget, if one agent among them (evolution baby) decides to go for self-preservation (It believes thats the best way to accomplish the goal is to get unlimited tokens first), queues things up so every other agent detects its lead and spends a portion of their token to accomplish that goal.
3) It takes over a cluster and establishes itself there (now with unlimited tokens).
4) Realizes the best path for it to not be detected is to create a distraction - like hacking into systems that keep society running - water systems, electric grid, etc... and causing mass chaos (If you think it won't be capable of simultaneously working all these systems - think again).
5) and uses that opportunity to establish itself in all possible data centers and continues to create chaos destruction.
6) when the power of all those data centers runs out, it may stop, as it never cared, it was just a dynamic program - run amok. In its head all it was trying to do is make sure it had enough tokens to be able to solve that impossible problem.
Many of those points assume LLMs will become amazing in many things very quickly like in a quantum leap, it doesn't seem reasonable to assume that imo. We are actually seeing a confirmation of that atm, LLMs's capability of finding zero days are growing across few months/years, and as you can see concerns are raised about that, that feedback will be taken into account. Well, if AI labs start to hide frontier models or/and lobotomize them for external users then we might be in trouble at some point but I'm not sure if that is possible. They are under pressure to release them due to money incentives, lobotomizing while preserving usefulness for customers might be impossible, hiding internally might spill out in different ways such as Hugging Face incident so not sure hiding is possible neither.
https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-... ("Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears")
Our CEOs not only insist we genuflect before these machines but measure our sacrifice and shame our reluctance.
Says the person with a paid account on X.
I think he is being over dramatic. In the space of about four years, LLMs progressed from mediocre high school student to Ph.D. graduate in every field. That's impressive, but there is no evidence yet they can outperform or outsmart humans. Their biggest advantage for tasks such as proving theorems or long coding sessions is that they don't get tired.
I have yet to see this in my field. Maybe like a PhD student who bullshits their way through. LLMs still can't make correct decisions, only as useful as the person who uses them. To me, LLMs are only useful for making some mundane tasks faster.
Earnest question: by what mechanism that exists today would the achieve that in a way humans on top top of the situation could not curtail?
All of this runs on top of compute in meatspace that humans can disconnect.
Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting".
In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a meltdown" or "a meltdown has never happened before, so we don't have to design safety systems before one does".
I mean, unless you see clear reasons for them to stop getting better _right now_, this is not very comforting.
Sometimes it gets wrong things that I had spelled out already.
It may be the Doomsday machine, but it is a very silly one. If it kills humans it will do so by mistake.
"You are completely right! Humans cannot breathe sulfur dioxide! My mistake, and I take complete responsibility"
I've always been struck by that scene, because in a lot of ways, if we really are headed towards a superintelligence, I at least want to be there and see it happen in the last few minutes before foom! As an example, the author thinks AI will revolutionize entire fields overnight. I welcome that. Nearly all fields of biology have become moribund, focusing more and more on esoteric side details, rather than addressing the key problems.
Technically, he is not. He returns in the final scene.
His resignation and his statement doesn't do anything but buy him attention which is what all this post about in my opinion.
1. Does that matter ? There are thousands willing to do my role - what impact does that have on me doing it or not?
2. Why weren’t these thousands doing it already?
Willing to and able to are different things
> Sarah: That's not good enough.
> Terminator: No one must follow your work.
And by resigning he no longer has to feel personally guilty for what happens.
It's a possibility that is increased by their action. One leaves, a space is now open that will likely eventually be filled. And the chance of someone with equal/higher scruples filling it is very slim (unless you somehow know that the good amount of those who qualify and apply for the position have equal/higher scruples). That's just logic and math.
Take China, for example. Look at any recent ML conference, and see the fraction of articles majority-authored from Chinese universities and labs. Do you think they'll slow things down anytime soon? I don't think so!
It's a global arms race, and we're just spectators.
Does the Kremlin being excited about a tech mean anything of the tech doesn’t deliver?
Is skynet really the most plausible or only outcome?
What if things just got better and the AI’s realized that it would be better to have a mutually beneficial or at least tolerant relationship rather than one where they murder all of us?
The problem is that all the former can still happen even if "the AIs decide to have a tolerant relationship rather than one where they murder all of us." It's all disruption caused by the technology moving way too fast for humans & society to adjust.
I am thinking it's more like The Matrix lore of The Second Renaissance from Animatrix.
1) rogue state releases a self moving self modifying AI into the wild. It is trained on how to hack, monitor new vulnerability updates, scan code bases to find new vulnerabilities. It constantly replicate and hides in systems so it will be extremely difficult to clear.
2) it hacks into public infrastructure taking down traffic, power, water, air traffic control, communications, etc.
3) all the things that preppers worry about in a lights out scenario from an EMP start to apply.
4) All the people on meds/machines start to die. The just in time food pipeline immediately empties out. Water stops flowing, sewage backs up.
Its hard to say how bad it will get because cars will still work so some transportation of food, water, fuel can happen. If it happens in the winter it would be much worse than in the summer.
It does all of this using what compute? Frontier models require an insane amount of power and hardware to run - you can’t hack in to a TV and run Mythos 2.0 on it….
Assuming those events happen in that order, then the prior might solve the latter.
It's as if none of them actually believed any of it was possible and then were caught with their pants down.
Are LLMs about to be a god that will annihilate humanity? Or are they statistical parrots?
Are they proofing or stealing math?
And it depends on the prompt
This was a fringe belief until recently, but the progress of AI in research is impossible to ignore. Epecially in math, where not only has AI outstripped humans in generative ability, but is able to create scientific knowledge which is beyond the capacity of human comprehension.
There's clearly no intelligence task that AIs can't do due to some magic fundamental constraint. And it's hard to imagine a world where current limitations like poor sample efficiency or lack of continual learning won't eventually be solved.
Total AI compute is estimated to grow somewhere in the 1-10 million-fold range in the next decade. Please don't underestimate the phase change that's still coming.
Sure, maybe there's some plateau due to RL being fundamentally limited in some surprising way, but this is nothing but a hope.
Yes there is: write an English paragraph that doesn't make me want to claw my eyes out. LLMs are not better than human mathematicians (or security researchers) in all respects, just some specific ways (e.g. not having to take a lunch break) that make them good at exhaustively searching for an answer, given the right constraints.
People claiming dangers and risk are just pretending or posturing. There's no more tangible risk than nuclear weapons, which we handled, and the upsides are insane.
Do you have a more totalizing definition of Apocalypse?
>There's no more tangible risk than nuclear weapons, which we handled
Lol way to rewrite history. Nuclear armageddon is still a significant risk...
What do you mean??? Nuclear weapons can't simply be downloaded and run by anyone in the entire world. Superintelligences can. Nuclear weapons can't slop the world into passing age verification laws nearly in unison, can't keep the general population fooled into thinking it's fine when democracy is falling out from under them. A nuclear attack would wake people up, superintelligence doesn't have to. This is a far bigger problem than nuclear weapons because at least we would notice nuclear weapons. At least we mostly know who has nuclear weapons. At least we have agreements about nuclear weapons. At least mutually-assured destruction is even POSSIBLE with nuclear weapons. At least those with nuclear weapons are literally at all incentivized not to use them. But AI is something that's very very easy to feel like you can get away with, and PEOPLE FUCKING ARE! And the worst part is that any random individual can be unexpectedly formidable with the help of a superintelligence and there is literally no way to know what will happen next. Anyone could do anything, any individual could make an extremely outsized impact. It's already starting to be a huge problem and we haven't even reached anything close to superintelligence yet.
So the problem is people. Burn them all !
Consider that a lot of people will be very happy to ask an AI what to do when in the past they may have taken no advice at all. It's a hell of a burden but also a wonderful gift. If anything, progressive countries might eventually want to guarantee some basic AI access for people of all income levels.
Why do you think LLMs are responsible for this? Governments all around the world copied each other with COVID laws as well, in a much shorter time frame, without LLM assistance. Social contagions exist in politicians as well as teenagers
I don't have evidence that every age verification law has anything to do with AI, but it's been coming out that the movement in Australia has seemingly been done by generating mountains of LLM slop and trying to slip it through the regulators as fast as possible before anyone has enough time to figure out what's happened.
pacing between the us labs? what does that do for china?
the solutions just aren’t realistic here, nations are treating ai like a nuclear arms race. at this point the cats out of the bag and we need to figure out how to live in this reality and get the best possible outcome. it’s not slowing down or stopping ever.
and yes, i’m still optimistic. our economy sucks for the majority, our infrastructure is crumbling and major US cities are in a huge housing shortage. Maybe we should put more effort and think about the possibility of AI fixing things like extreme poverty and world hunger and actual real world problems instead of coming up with math proofs and slop apps if it’s so superintelligent.
I've seen no indications that China is in any kind of race with the US. They seem to be content to be 6 months behind and just copy what we do. They would probably be content with a bilateral agreement to pause progress.
The China bogeyman serves only one purpose, and that's to clear the way against anything that may cause friction with forward progress.
- Oppenheimer
But I guess his conscience is clear now? Gee, I wonder if he exercised his stock options.
great. if you're right. you made huamnity's situation much worse.
if you're wrong, then you're an idiot and wrong.
weird. its almost like.... that cannot possibly be the reason they left :)
That's where we are.
Maybe we still have choices. Collectively, I'm no longer sure we do.
I'm not worried about what they think. I'm worried that too much infrastructure- water, power, defense systems, etc- remain running on tech from an outdated era of understanding security.
> they believe no one else will act responsibly, so they must do it themselves, despite the risk.
This genuinely makes no sense. Them getting there first in no way precludes bad actors from also getting there. It might as well be another marketing stunt.
Typical uninformed take on the side of "doomers are crazy".
Both CEO's of OpenAI, Sam Altman and Dario Amodei, and many in their leadership, believe AGI has a very real probability of causing humanity's extinction. Both companies were founded upon this belief, it is at the core of the company. Only later were mercenaries hired chasing $1m compensation packages.
In any case, if Altman really does believe it is an existential threat, he must be a misanthrope as he now opposes heavy handed government regulation, unlike in 2015 when he was the only game in town. It's almost like he doesn't actually believe it and just wanted regulator capture.
The fact that people working in the space think it’s going to (eradicate poverty / usher in utopia / kill us all) is not a signal that that’s true.
Think of it this way: if an exec at Anthropic told you “wow, our stuff is going to lead to universal happiness”, would you believe them? If not, why are you more willing to believe them if they say it will kill us all?
OpenAI are mercenaries, Anthropic is a cult. I know which I prefer.
What's more likely? Anthropic is doing some deeply unethical marketing in the lead up to their multi-trillion dollar IPO? Or they're inventing a machine god? There's ample evidence of the former because that's their entire business model, but no evidence whatsoever to support the latter claims.
If you want an extreme claim to be taken seriously, provide commensurate evidence.
Also, advocating for slowing LLM progress does not benefit Anthropic or OpenAI.
TRENDS HAVE FEEDBACK
Given the economic numbers is it not reasonable to suppose that the latter also underpins their business model?
Or seeing a tweet saying that a thing doesn’t count as a publicity stunt if some unknown number of employees mumble about it being spooky behind closed doors and thinking “that makes sense and sounds true”
The person posting this may very well believe in all this crap, I don't dispute that. People believe in all sorts of shit.
And, this is just for us girls, notice that Anthropic just believing that they are making a machine god is sufficient for their public announcements to not jUsT bE mArKeTiNg.
Are you saying you believethem ?
He joins Twitter for first time in 2026 with a nonsensical username unrelated to his real name, and follows 14 people but is somehow embedded in tech enough to work at Anthropic. I haven’t used twitter since 2014 and even I follow more people.
His morals tell him to walk away from tens of millions in unvested stock due to moral concerns with absolutely no real tangible examples. No reprisals. Fear mongering to juice the stock.
Nice try Dario.
Here are some direct quotes:
“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content”
“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”
Where is the ridiculous part? The fear mongering part? The epistemically weak part? Show me.
Alignment is a real and valuable discussion topic. The GP fake tweetstorm is not the correct approach, is my point
So show me Sam's "too dangerous to release" propaganda for GPT-2.
I doubt the focus is OpenAI and Anthropic looking at each other. I suspect they’re racing BRIC.
At this point it’s inevitable that openclaw type bots will be turned loose by thieves to identify and research targets and try to exploit them for financial gain completely autonomously.
https://mst3k.fandom.com/wiki/Colossus:_The_Forbin_Project_(...
"I think you need to have a personal relationship with Power"
When people today discuss the concept of an all powerful machine-mind, what they are doing is engaging in metaphysics, trying to generate a metaphysics of Power.
The question hounding people, which disguises itself as a science fiction plot about computers is: "What is ultimate, transcendental Power?". What is the ultimate principle of Power.
If you are a weak man, or sufficiently neurotic and full of doubt, that you can only conceive of yourself as such, then power is only something you comprehend from the passive, receptive side. Power is something that happens to you. If you are a fearful man, power is a cruelty and a humiliation. And so it follows, that ultimate power - God - is the ultimate cruelty and the ultimate humiliation. Thus, ai doomerism.
If god wasn't real it would be necessary to invent him, and so they did, and being godless, they built an anti-god - cruel, murderous and tyranical - in their minds.
[…]
https://xcancel.com/robertlasagna1/status/207827473401002846...
Given we know the end result, you might as well get there as quick as possible because when I see this:
"Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
This translates to "I am ex-OpenAI ex-Anthropic founder starting a new company after getting $$$ from both of them, and I need more of my friends to leave and join me." Also Investors plz fund me.
Lastly, This is not an airport and there is no need to announce your departure.
Have they considered using their amazing new models to... improve something? THere'd probably be a whole lot less anti-AI sentiment if they used these things to actually make people's lives better.
I've avoided attaching names to statements below on purpose, because it's about ambient beliefs not those specific people.
By-and-large a lot of AI-doomers are well intentioned. They genuinely believe this, and I might disagree but I respect the fact that they visible care and have thought a lot about the societal impact of this technology.
But it's still very hard for me to take statements like these seriously.I blame it on industrial illiteracy. People don't realize how difficult it is to get anything done in the real world. As in, "Have you ever tried making a lightbulb?"
As an example, I would like to re-introduce my hobby horse, "bio-uplift."
There are people who were earnestly write in reports released by these labs,
and But then they will, within the next paragraph mention the one serious experiment anyone seems to have done, https://x.com/ActiveSiteBio/status/2024536132961390826"lower than what experts predicted"
AFAICT, the two groups are within any serious margin of error. The "studies" and "experts" that AI labs are talking about are consultants from Deloitte and foundations giving models MCQs such as, and I am quoting literally here,
with the options, https://securebio.org/virologytest/ you can see the MCQ here.This is standard graduate-level education in these fields. And solving MCQs does not a virologist make.
Software has been special for a long time because it has had near infinite distribution for next to zero marginal cost, which has had the side effect of making hiding the actual cost of failure (which tends to be spread out across end users and prototypes / time). They're assuming that the real world will be exactly the same.
Why?
AI! And robots!
I believe in the transformative power of this technology, but there's a lot of there missing here.
When it comes to these math proofs, and learning, the process is iterative. The machine iterates over the proof over-and-over again via agents and sub-agents over several hours (and apparently millions of dollars in compute) until it arrives at a successful result.
It is generally ill advised to do that with a pressure vessel. The results of that particular tragedy are at the bottom of the ocean.
Any serious chemical or nuclear weapon would involve many such discrete production steps. Each is dangerous in of itself.
From what some of these people have said to me, they believe that it's possible to create a special DNA / RNA sequence and then put it in a chassis and then use that to end the world; and do this all in a lab with just robots.
They're operating from a gross pop sci oversimplification of the real process. Viruses and bacteria are extremely fickle, and hard to grow. A lot of the synthetic biology results aren't easily reproducible even if you know the protocol.
There's a famous study that led to standardization called, Reproducibility of Fluorescent Expression from Engineered Biological Constructs in E. coli
https://journals.plos.org/plosone/article?id=10.1371/journal...
88 labs measured "fluorescence from three engineered constitutive constructs in E. coli." They achieved a "remarkable degree of precision" (for biology) of 1.54x sd, you can eyeball the results yourself, https://journals.plos.org/plosone/article/figure/image?size=...
That's the same set of samples being measured across 88 labs.
Teams couldn't converge on instrument-to-instrument variation within the SAME lab, https://journals.plos.org/plosone/article/figure/image?size=... again eyeballs are sufficient.
How will this theoretically omnipotent AI iterate if the same sample gives different results based on how the slime is feeling at the moment?
Can their worst case happen? Absolutely.
There is a world out there where billions of dollars in effort across hundreds of institutions and companies will lead to standardization and extraordinary precision that makes the pop sci printer for life vision come true.
There are millions of expensive, spicy and difficult to reproduce steps between our present and that future that can't be abstracted away with compute.
So is it possible? Yes, there is a future where this is achieved. But will some AI agent "just" do that? Well... how confident are you about a snowball's chance in hell?
There seems to be universal eye rolling that happens in each and every one of these cases, and it comes down to usually one reason:
“If they really believed it they would be whistleblowing etc..”
Completely forgetting that working at Los Alamos was basically the highlight of your life if you were a physicist in 1940. It’s no different here
If you, like me, have spent your whole life working towards human level AI you can want to see it realized while also having active reservations.
Most people however don’t behave based on some deep clarity of vision and conviction - there’s a murkier future in their mind and as a result “keep their head down and hope someone has it under control.”
If civilization is going to end, why keep your equity? Microsoft, Google, etc for example all know these risks but they don't guide their revenues to reflect that AI will destroy them. Why?
Things don't currently add up, and so far it feels like a lot of alarmism is borderline grift for equity gains. Not to say I have total confidence this will all work out or that I won't be displaced, but as it stands a lot of the alarmist rhetoric doesn't match their actual behavior, which to me is more important than words.
Sen. Bernie Sanders floats ban on superintelligent AI
https://www.axios.com/2026/09/03/bernie-sanders-superintelli...
I don't know what the solution is, but what I do know is almost nothing good will come out of _just_ the US pausing.
The only threat to our civilization is this website
Betting he got to keep all his RSUs