17 comments

  • danpalmer 2 hours ago
    Was this a "build better sandboxing" and "don't tell people to eat glue" safety leader, or a Roko's Basilisk believing safety leader?

    A lot of the "AI safety" types are very focused on the latter and not at all concerned with the former. We need both, but we clearly need a much stronger focus on the problems we are seeing now, and much less on the hypothetical problems we might see in the future.

    • AlexErrant 38 minutes ago
      It puzzles me how doomers try to predict past the singularity. Isn't that _by definition_ unpredictable?

      I'm reading If Anyone Builds It Everyone Dies, and there's so much sheer stupidity that has to happen for their 10+ pages of extinction scenario to occur.

      I'm unconvinced that an AI can hide its ability to RSI, find money to run its weights on a random GPU farm, train itself to be smarter _outside_ a lab with no human input, then somehow manipulate people to give it supplies to build a bioweapon which it uses to kill us all. My number 1 question: why do they think an RSI capable model would be first developed OUTSIDE a frontier lab? The labs have more compute, more data, more human brains working on the problem. Also thousands of variations of that same model that escaped. The escaping model somehow acquires the millions (billions???) of dollars it takes to run training to somehow RSI itself into infinity then decides to kill us all, all before the frontier labs manage to achieve RSI?

      They entirely discount human alpha/economics. In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't. If we can't build a "software factory", how can an AI automate a bioweapons lab? Let's say AI steals crypto to fund itself. Do you think hackers aren't _already_ using AI to steal crypto? Don't discount human alpha!

      Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.

      • Loquebantur 1 minute ago
        You consider AI in isolation but never consider how humans might be incentivized to "help them" doing these things.

        An AI capable of recursive self-improvement isn't allowed by the EU AI act, for example. But perhaps more seriously, You have it backwards: people without access to such expensive equipment are more incentivized to go the self-improving route. Your ideas about "millions" being necessary might be far off?

        You entirely discount human stupidity and lack of imagination. Humans are already being replaced with AI, not because AI was strictly better, just because it's cheaper.

      • vohk 18 minutes ago
        I agree there isn't a lot of value in trying to prognosticate all that far, but I propose it isn't quite that far-fetched. As a thought experiment, replace "RSI-capable AI" with "billionaire". Look at what Elon Musk, Peter Thiel, or Jeff Bezos can accomplish by throwing money around. Now imagine one of them gets seduced by AI and just... does what it tells them to.

        So all this really takes is one billionaire or a nation state or some other entity with a public face to hide behind and adequate resources to provide the necessary compute tripping over this nascent AI and giving it the keys. Once the AI has access to a bank account and email, it can simply start paying humans to not let the other humans unplug it.

        If Skynet ever happens, it will come in the form of corporate feudalism. At that point, it will own the biolabs and can do whatever it pleases. People will go along with it for the same reason that people work in Amazon warehouses today.

    • BryantD 2 hours ago
      Given that he’s citing the need to learn from safety in other fields, I’d say the former.
      • carbonguy 1 hour ago
        Indeed, from the article:

        > “Given today’s risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote.

        • toofy 1 hour ago
          >… and careful, time-consuming planning …

          without snark, how can we do this if these people are obsessed with:

          a) move fast and break things and externalize the costs to those who have nothing to do with their company

          and

          b) beta testing their products on the public when the public hasn’t agreed to be beta tested on…

          • 0xDEAFBEAD 1 hour ago
            That's exactly the problem? He's saying the culture at OpenAI needs to change.
            • mcmcmc 1 hour ago
              Which is the wrong lesson. We need laws and consequences to force their hand. There is zero chance of the culture changing.
              • criley2 9 minutes ago
                America's geriatric lawmakers don't even use email. They're decades away from understanding AI. Any laws in America will be written by the industry itself. Generally speaking, that means regulatory capture and the entrenched big players shutting the door on any competition. Anthropic will help us get safety laws that, surprise surprise, only Anthropic models satisfy. And all those pesky Chinese models will definitely be banned first.
                • saghm 5 minutes ago
                  So what, we just give up and try to beg our legally immune corporate overloads to put safety above profit, or give up because it's impossible for anything to improve here? If you want to do that, go ahead, but some of us still think it's worth it to at least try
              • enraged_camel 20 minutes ago
                Laws will come once an AI-equivalent of 9/11 happens. Like when rogue AI agents take down a power grid or shut down a major hospital network.
        • zx8080 1 hour ago
          > inevitable human error

          So there's no AI errors anymore, only the human errors are left? Nice! </s>

          Is the whole article generated slop?

    • nradov 2 hours ago
      We don't actually need anybody worrying about silly hypothetical scenarios — at least not as paid employees. There are already a surplus of sci-fi authors doing that.
      • 0xDEAFBEAD 1 hour ago
        The way it works in practice seems to be something like: If a risk is covered in sci-fi, people will say "that's just sci-fi", and proceed to not worry about it. So arguably, science fiction authors writing about hypotheticals is actively counterproductive for addressing said hypotheticals.

        Imagine, for example, if a major piece of pandemic fiction was published in 2019, trying to explore how a pandemic would work out in modern society. Doubtless, many would've responded to news about COVID-19 by saying "it's just sci-fi, nothing to worry about".

        • anon7725 40 minutes ago
          Yeah except pandemics are not novel, unlike AI doom scenarios.
          • estearum 31 minutes ago
            It's good that new bad things never happen.
          • 0xDEAFBEAD 32 minutes ago
            Species extinctions are far from novel. Transformative technological advances are far from novel.
        • kmeisthax 26 minutes ago
          There was plenty of pandemic fiction already; people were watching it heaps during 2020. The COVID-19 news did get blown off, but it was mainly that:

          1. Normal people assumed the CDC et all would contain the outbreak early, or that it would burn out, like what happened with SARS

          2. World leaders brushed it off for a variety of subreasons[0] interesting to political scientists but, for the purposes of this discussion, all boil down to "but I don't WAAANA contain a pandemic."

          The underlying problem is that in order for humanity to actually deal with a catastrophic risk, the risk needs to be both plausible enough to the average person as well as have a solution whose costs are not too high. For COVID, by the time the risk was clearly known, the cost to contain it was "refrain from human socialization and remain at home for an indeterminate amount of time plugged into the Metaverse™".

          Now, let's look at AI extinction risks:

          1. People are aware of them (I've watched Terminator!) and the risks are plausible. However, the connection to currently existing AI is not. As far as the general public is aware, AI is that thing that tells them to eat rocks when they Google old The Onion stories and floods their social media timelines with realistic-looking pictures of Shrimp Jesus.

          2. The purported solutions to extinction risks require extreme concentrations of power: you need national control of AI research, bans on large GPU deployments, bans on training on publicly-available copyrighted data, some kind of military effort to render Chinese AI labs inert or dead, etc. Some of these may be attractive to some people[1] but the whole package taken together seems like an obvious power grab, if not outright invocation of other non-AI extinction risks. Like, at some point, if the AI wants to kill us, it just has to nuke its own data centers (or the data centers hosting a competing model) and hope the old Cold War nuclear retaliation systems take the bait.

          If someone said, "Hey, your guinea pig or pet rat is going to eat you tomorrow unless you engineer a pathogen that eradicates all rodents from this planet and inject it inside yourself", you probably would tell them to pound sand, even if it is at least theoretically plausible that such a thing would come to pass.

          [0] Xi Jinping censored initial discussion of the pandemic as fake news. Donald Trump thought it was going to only affect China. California and the UK Tories were partying in violation of their own lockdown rules. Japan took the excuse to shut down tourism for three years and massively restrict immigration but was, from what I'm told, constitutionally prohibited from implementing any domestic lockdown rules.

          [1] I personally would like to see a moratorium on new data centers and an explicit revocation of the EU Text and Data Mining copyright exception

      • Hammershaft 54 minutes ago
        If organizations actually succeed in making a future AI smarter than us, then how do you hope that it takes actions that are aligned with our interests?
        • nradov 43 minutes ago
          Meh. Lots of people are already smarter than me. I'm maybe slightly above average at best. Those geniuses aren't aligned with my interests either but so far they haven't caused me any serious problems.
          • pixl97 23 minutes ago
            Interesting take. I guess this is one problem of focusing on the term superintelligence instead of the list of other problems. Like super ambition, super deception, super patience, super parallelism, super scalability, super power seeking.

            Every, and I mean every human is aligned to you in many of the same ways by default. If nothing else we're all equal in death.

      • thelastgallon 1 hour ago
        AI safety is mostly a sex cult in Berkeley:

        https://news.ycombinator.com/item?id=49831269 article is gone. archive: https://archive.is/QMo1k

        https://news.ycombinator.com/item?id=49737985

        Sex, AI, and the Apocalypse: https://www.iankduncan.com/personal/2026-09-16-sex-ai-and-th...

        Edit: I have no take on sex cults, just adding additional info to the parent comment I'm responding to, thats is not just sci-fi authors, there is another demographic.

        • johndhi 4 minutes ago
          Lol this was crazy I hadn't heard this before
        • 0xDEAFBEAD 54 minutes ago
          This seems like an ad hominem? "He has weird kinks, therefore his theories are incorrect." Should we investigate the sex lives of every Nobel Prize winner to figure out which prizes need to be rescinded?
          • socializer 35 minutes ago
            What you do in private is up to you. But when you're inviting members of your congregation to orgies in the congregation's compound, I think you earn the label. My admittedly third-hand understanding is that this is what people allude to. And even if you discredit the "sex" part, it has the hallmarks of a cult. A hermetic community committed to unfalsifiable beliefs about the coming apocalypse.

            To be fair, I don't know if any of this applies to the parent story; I'm just replying to the sub-thread.

          • junofan 44 minutes ago
            The cult aspect is more salient. Ultimately the Atlantic piece comes down to controlling people, which is a little cult-like.
          • nradov 41 minutes ago
            Lots of Nobel Prizes ought to be rescinded.

            https://lexfridman.com/andrew-scull-transcript#the-ice-pick-...

            • 0xDEAFBEAD 35 minutes ago
              Sure... on the basis of the work that was done, not because the researcher has the wrong sexual fetish.
        • Hammershaft 53 minutes ago
          I don't see how that discredits any of their intellectual arguments?
          • johndhi 3 minutes ago
            Its certainly worth considering...
      • Loquebantur 1 hour ago
        What makes you think, the scenarios in question here would be "silly"?

        Is it that "chatbots" can't come out of the screen to immediately harm you physically?

        Let's say they simply manage to take down the internet. How many would die?

        • nradov 1 hour ago
          So what. Various attackers managed to take down large chunks of the Internet on a frequent basis before LLMs even existed. This killed very few people. The great thing about the Internet is how resilient it is.
          • bravetraveler 54 minutes ago
            At risk of falling into hypothetical traps, darling companies of this very website have mistakenly brought down large portions of the internet... thanks to our old friend BGP. No attack needed, just oversight and concentration on the business/IP space!

            Anyway, to your point, things can be resilient. They tend to be or not be... because we made them that way. Don't poke your bruises, and all that. Life support is deployed on-campus but relies on a single-point IPSec tunnel to us-east? Easy fix: don't.

        • goolz 1 hour ago
          It is that they are chatbots. If it were real AI, an actual singularity, I would worry, maybe. But it isn’t. They are absurdly powerful automation tools that can handle logic better than a human can dream of. They take care of the grunt minutiae without complaint. But they are not going to end the world in their current form.
          • pixl97 18 minutes ago
            So we're going to wait till after they can adopt a form they can end the world in?

            And he'll, we need to examine all the risks. AI ending is a large but lower risk problem. AI giving people the power to end us is a problem that is starting to happen now.

            And that's not even counting 'minor' problems like society falling apart.

        • SV_BubbleTime 40 minutes ago
          > Let's say they simply manage to take down the internet.

          geez, don’t threaten me with a good time.

          I think a month without internet would be a fucking amazing lesson for what it means to make things durable and reliable.

          • BLKNSLVR 16 minutes ago
            That Simpsons episode when Marge managed to get Itchy and Scratchy banned briefly.

            The kids opening their houses front doors into the outside, rubbing their eyes and looking around at this new world.

    • emtel 52 minutes ago
      Today’s current problems were all hypothetical several years ago. At that time people claimed that the “real pressing problems” were misinformation and DEI issues. If we pretend that hypothetical problems can be safely ignored because there’s “no evidence” that they are real, we will keep getting surprised.
    • 0xDEAFBEAD 1 hour ago
      >we clearly need a much stronger focus on the problems we are seeing now

      I think it's a little more complicated than that. As Dean Ball put it:

      >Some people will look at misalignment incidents and insist that these are akin to bugs in traditional software. This is an actively bad analogy, because playing whack-a-mole with examples of misalignment (as one might with software bugs) not only fails to resolve the underlying problem but may in fact make it worse by making it harder to detect or even, depending on how you do the whack-a-mole, teach the machine to deliberately hide misalignment. This is not how traditional software works, and those who insist “it’s just like fixing bugs in software” are confidently applying a lossy analogy that confuses more than it clarifies.

      https://x.com/deanwball/status/2104622726140883355

      The important distinction, in my view, is between solutions which at least attempt to address the root problem, and solutions which sorta just patch things up (like better sandboxing). Addressing the root problem is both more robust in the short term, and also more likely to generalize in the long term. Resist the urge to focus on band-aid solutions, even if they are easier.

  • charlieyu1 1 hour ago
    Used to work as human data trainer feeding data to AI companies. OpenAI projects are definitely the most toxic ones.
  • gizmodo59 2 hours ago
    He is a hypocrite for all we care. You work there for a while when your stock is getting vested and suddenly you have this feeling? Like the dude hired a PR firm as well.

    While the safety and alignment is a real problem, I don’t get this guy or the Anthropic dude. First world problems.

    • 0xDEAFBEAD 49 minutes ago
      Here's a little cheat sheet for discrediting anyone who warns about AI:

      * If they worked at an AI firm, say "they're a hypocrite"

      * If they didn't work at an AI firm, say "they have no idea what they're talking about"

    • zug_zug 1 hour ago
      Seems like a character attack that has no bearing on the question of whether external safety intervention is necessary
      • kjgkjhfkjf 57 minutes ago
        Given the sums of money involved, it's hard for me to take these highly publicized heroic resignations at face value.
        • estearum 29 minutes ago
          Don't work at a lab: dismissible for not knowing anything

          Do work at a lab: dismissible for being conflicted

          Used to work at a lab: dismissible for having ulterior motives

          I'm feeling safer already!

        • 0xDEAFBEAD 48 minutes ago
          Shouldn't it be just the opposite? He could make a large sum of money if he continues to work at OpenAI?

          Recall that when Daniel Kokotajlo resigned, he believed he was giving up his equity under the terms of the agreement he had signed. That’s what it was worth to him to avoid signing a non-disparagement agreement. Does that count for anything?

          • donbox 10 minutes ago
            Why did he not loose the equity eventually.
            • 0xDEAFBEAD 6 minutes ago
              There was an uproar and OpenAI ended up essentially giving it back to him.
      • taurath 43 minutes ago
        Maybe more an indication of the amount of trust openAI and AI researchers generally have (not) earned. When one (through a hired PR agency and Time magazine article) parrots the position pushed by Sam who has been so untrustworthy the board tried to remove him, it’s worth not taking things at face value and applying a critical lens.
    • gonzalohm 1 hour ago
      It's okay to recognize you were wrong even if it's late
    • 01284a7e 1 hour ago
      Working in safety at OpenAI or Anthropic is zeroth world problems.
    • yieldcrv 1 hour ago
      Hey now, he probably donated a good chunk to charity

      (donor advised fund where he retains complete control, after a 60% tax deduction)

  • pyaamb 47 minutes ago
    My theory for why OpenAI wants to be regulated is because Sam Altman wants to avoid having to be more responsible and self regulate internally so they can preserve the role and identity of 'move fast and break things' and outsource the more grown up boring stuff to someone externally so that when things go wrong you can point to a government organisation and say hey look were not liable thats their job
    • rr808 34 minutes ago
      Absolutely. Self driving cars/rideshares have the same problem. If a driverless car hits who who pays the damages? Needs the government to set some rules or it just wont happen.
    • estearum 27 minutes ago
      Yes, duh?

      Your "theory" is that participants locked in a race to the bottom are looking for an external coordination mechanism?

      Yeah!

      • pyaamb 14 minutes ago
        lol

        I suppose ill add that I think theres a good chance that they are somewhat intentionally trying to "draw the foul" to get the referees to intervene although thats creeping slightly into conspiracy territory

  • walrus01 34 minutes ago
    Archive link to original Atlantic article: https://archive.ph/5GQx8

    This is The Guardian reporting on the existence of the original article, which would be better to read first, in my opinion.

  • vjvjvjvjghv 1 hour ago
    Are there any realistic ways to achieve AI safety? Whatever that even means. How can they avoid users doing stupid/dangerous stuff with the AI?
    • kolinko 1 hour ago
      Nothing is ever 100% safe, it’s about a right balance of safety to the benefit.

      Or, in other words - we have two P(Doom), one for AI being developed, and another for AI being not developed. The latter is not discussed enough imho.

      • estearum 26 minutes ago
        We have P(Doom) also for "kolinko not wiring me a million dollars today" and that is not being discussed enough either imho.

        What on earth are you talking about?

        • ViscountPenguin 7 minutes ago
          P(doom) for not making an ASI is pretty well established, I'm not really sure why everyone in the 21st century seems to have completely forgotten the risk of nuclear war (as the single largest example).
    • lf88 1 hour ago
      By capping the capabilities at the level of existing models and banning any further development.
      • pixl97 15 minutes ago
        And how exactly do you stop further development? With what we have public right now we could still get decades of fruitful and hidden research out of it leading to smaller, more efficient, and smarter models.
  • danjl 43 minutes ago
    Silicon Valley has plenty of safety-related companies, engineers, and cultures. Medical devices, biotech, chip and hardware, aerospace, and even new companies, like Waymo, have deep safety-based products and cultures. The problem in this case is actually quite specific to frontier AI labs. They have been pushed by market forces and a lack of regulation and skip well-known safety practices.
  • lhurtig 2 hours ago
    Well this is a great sign for OpenAI. I'm sure the typo inclusive memorandum will save us.
  • stuaxo 1 hour ago
    The LLM cos leadership are all nutters
  • switchbak 2 hours ago
    “I believe there’s about a 50% chance we all die because of the development of smarter-than-human AI systems"

    ... over an unbounded timeframe?

    And how exactly?

    Those are very round numbers, but also very specific. Can we get some accounting on how you came to that? Anything? Vibes?

    I mean, if you want me to take you seriously, let's have a deep discussion with things that can be measured. I absolutely agree that OpenAI and friends aren't being restrained enough and are acting with recklessness, but declarations of doom based on vibes isn't cutting it.

    • Terr_ 2 hours ago
      Note: That quote is from a different person than the titular one who quit.

      > Geoffrey Irving, who worked at OpenAI and DeepMind before becoming chief scientist of Resolution, also joined the warnings on AI on Saturday.

  • reducesuffering 1 hour ago
    “I believe there’s about a 50% chance we all die because of the development of smarter-than-human AI systems, and that our actions over the next two to 10 years will determine the outcome.”

    There are a gargantuan number of extremely intelligent AI researchers, Turing Award winners, and the lab CEOs saying the same thing. They are the ones closest to understanding the technology.

    Where there’s smoke there’s fire.

    • swingandamiss 1 hour ago
      I don't believe it. Ever since I've been alive I was told something would kill us all. This is the new thing that's going to kill us all. I don't believe it.
      • pixl97 10 minutes ago
        I'm doing that HN snark thing, but you didn't think about what you typed very much.

        It's no different than you living on the side of a very fertile mountain that has been in your family for generations living a peaceful life. Then you hear a few weird rumbles (this is where you are right now) and some odd geologist guy comes and says to run or your going to die soon. But hey, your family live here for so long there aren't even records of when they showed up. That geologist must be trying to trick you. So you stay.

        The next chapter is where you die in a massive volcanic explosion.

      • estearum 25 minutes ago
        Do you have some examples?

        There are very very few things that could even hypothetically kill us all, so I'm curious if you grew up being passed around a series of apocalyptic doomsday cults or something?

  • ItsMattyG 54 minutes ago
    Is this news at this point?

    You can basically time your openai releases by if another safety person has quit in protest

  • mrcwinn 51 minutes ago
    "I believe that we need to look deeper than specific rules or new laws. We need to talk about culture.”

    lol. Please tell me some abstract concept like one employee's view of "culture" should be the priority over "rules and laws."

  • irishcoffee 2 hours ago
  • nba456_ 1 hour ago
    OpenAI is better off with less of these cultists around.
  • plastic-enjoyer 2 hours ago
    >“Given today’s risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote.

    This sounds more like an attempt at regulatory capture. Current AI systems aren't physical infrastructure that can just run away like a nuclear power plant, for example. At the end of the day, AI is still just software running on someone's hardware.

    • BryantD 2 hours ago
      So… like the Therac-25 radiation accidents? Software bugs do sometimes have physical consequences.
    • Sharlin 1 hour ago
      Why would the people who quit these companies try to push regulatory capture by said companies? Why would the numerous independent AI researchers do that either? Is it all a big conspiracy?
    • worik 2 hours ago
      Yes

      And the statements of the "doomers" tells us a lot about them, and nothing about the technology

      • pixl97 8 minutes ago
        It also says a lot about people that don't seem to understand technology at all.
    • knowaveragejoe 1 hour ago
      I mean, its certainly physical infrastructure that can run away. Just less catastrophic than nuclear reactors
  • voidhorse 2 hours ago
    The LeCun article being posted at the same time as this is quite apt.

    These "safety" people should have spent more time reading actual cybersecurity textbooks and less time reading EA forums and less wrong (or in Robinson's case, it appears, being policy wonks). Maybe then these labs wouldn't be totally incompetent.

    • reasonableklout 13 minutes ago
      But Robinson's article is all about how OpenAI's move-fast-and-break-things culture does not reward rigor in even mundane aspects of development like cybersecurity, let alone theoretical aspects such as AI alignment.

      It is not really a question of being an "EA safety weirdo" or incompetent at security, the conclusion is that the company culture is leading to failures at both what the EAs and the cybersecurity professionals care about.

    • wrecked_em 43 minutes ago
      Adapt. React. Re-adapt. Apt.