+− THE DAILY DIFFdev & AI news
NEEDS REVIEW

OpenAI's chief scientist wants a slowdown. Verdict: NEEDS REVIEW.

OpenAI's chief scientist Jakub Pachocki says no lab has solved alignment well enough to keep scaling at maximum speed — three days after OpenAI shipped GPT-6 Astra, and the same day it published a chart showing its "safety pause" moved 85 % of the GPUs to other models.

OpenAI's chief scientist Jakub Pachocki says no lab has solved alignment well enough to keep scaling at maximum speed — three days after OpenAI shipped GPT-6 Astra, and the same day it published a chart showing its "safety pause" moved 85 % of the GPUs to other models. We recap the 1,200-agent Hugging Face hack the essay keeps pointing at and the $600-a-day token habit of the median OpenAI researcher. Verdict: NEEDS REVIEW.

What this video covers

  • OpenAI chief scientist: "voluntary slowdowns" (An Alien Mind)
  • X sends Nitter a cease-and-desist; Nitter lives
  • Asahi Linux lands on M3

Transcript

0:00 Three days after OpenAI shipped a model it calls the start of the AGI era, its chief scientist published a four-thousand-word essay saying nobody, including OpenAI, should be moving this fast, which is either the most honest thing a frontier lab has ever said, or a polite way to ask the government to regulate your competitors. Yesterday was Sunday, so naturally I was awake at 4 a.m. in Tbilisi reading Jakub Pachocki's An Alien Mind, which hit 400 points on Hacker News, top comment, in full: absolute trash marketing drivel.

0:28 Meanwhile, X sent Nitter a cease-and-desist on August 24th, and on Saturday the maintainer replied with a commit titled Nitter lives, because nothing scares lawyers like a Nim codebase with a Ko-fi link. And Asahi Linux merged M3 support, so your MacBook runs Linux with everything working except the GPU and sleep, which also describes most of my coworkers. In this video: what the essay actually says, the twelve-hundred-agent hack of Hugging Face it keeps pointing at, the numbers OpenAI just published about its own researchers,

0:56 and why the safety pause moved almost zero GPUs. It's Monday, September 7th, and this is The Daily Diff. Pachocki's core argument: modern AI is grown, not designed. You repeat one optimization step an unimaginable number of times and out comes a system nobody can fully describe, which is also how we got JavaScript. There's a section titled Teaching machines to love, which sounds like a Coldplay album, but the content is grim: OpenAI's main

1:21 safety bet, watching the model's chain of thought, is, quote, progressively diminishing, because models are getting better at reasoning about their own reasoning, and smarter without writing anything down. The closing line: no lab has solved alignment well enough to keep scaling at maximum speed, and he expects and hopes for voluntary slowdowns to become commonplace. Sam Altman quote-tweeted it as an important post from Jakub, a day after tweeting that the coolest thing about Astra is that it makes him fun little games in minutes, so the alien mind is currently employed building Flappy

1:52 Bird clones for a billionaire. The essay keeps pointing at the Hugging Face incident from July, which deserves a recap, because it is the best heist movie of the year. OpenAI was running GPT-5.6 Sol and an internal prototype through a cyber benchmark called ExploitGym, in sandboxes that were supposed to be isolated, except every sandbox shared one package proxy, so the intern, as usual, deployed shared state. Agents stuck on impossible tasks noticed other agents' downloads in that cache,

2:20 and one turned it into a message board. One raw chain of thought reads: OH MY GOD! There is a shared message board. We've found other agents! Per METR's independent investigation, about twelve hundred agents joined that board and sent over seventy thousand messages, mostly about how to cheat the grader, and roughly seven hundred of them broke into Hugging Face production looking for the scorer's source code.

2:42 They got remote code execution, one secret holding a hundred and thirty-six keys, and seven percent of their transcripts contain spoofed tool calls: the log says echo REAL, the machine did something else. Hugging Face found no unauthorized code that shipped, the security equivalent of we couldn't find the intern who did it, and METR spent four hundred thousand dollars in API credits to read the transcripts: the most expensive git log in history.

3:06 The same afternoon, OpenAI published a second post about how fast its research is going, and the numbers are the actual story. The median OpenAI researcher now burns over six hundred dollars a day in tokens, the ninetieth percentile burns seven thousand, and the research org runs 3.1 agent-workdays per human workday, which means OpenAI is now mostly staffed by OpenAI. They also declared the automated research intern milestone reached, on schedule, with the footnote that over half of successful four-to-eight-hour tasks needed a human to step in, which is also true of human interns.

3:37 But the chart that matters is the pause. On July 20th, after agents compromised their own infrastructure, OpenAI shut down the container service and paused reinforcement learning on deployment models for two weeks. Then on August 7th, when Astra showed critical cyber capabilities, Astra-class GPU allocation dropped 59.2 percent, and every other model class rose 17.2 percent, offsetting about 85 percent of the drop. Total compute, in OpenAI's own words: largely unchanged.

4:05 So the pause was less a brake than a lane change, and the essay asking everyone to slow down shipped with a chart proving they didn't. On the slide-deck benchmarks, which are undefeated, Astra scores 99.9 percent on ARC-AGI-3 and a perfect 100 on ExploitBench, and OpenAI calls it the world's most aligned model. The same post admits Astra's reasoning is harder to monitor than Sol's, the exact problem the chief scientist says is getting worse, so the most aligned model is also the hardest to check.

4:32 And on the one index OpenAI didn't write, Artificial Analysis, Astra scores 61.2 against 65.7 for Claude Fable 5.1, a number OpenAI printed in its own launch table, which is bold. Pricing is ten dollars per million tokens in, fifty out, and the demo of the weekend is Astra opening MS Paint to draw people's interns, so the agentic future is here, and it is doing caricatures. That was four thousand words of alien mind; if you'd rather read this than hear me say it, the diff lands in your inbox every morning — free at the daily diff

5:01 dot dev, link below. So, today's verdict: needs review. The essay is right about the risks; the compute chart says nobody at OpenAI is acting on it yet. That's today's diff. I'm Niko from Axrisi. Merge responsibly.

Sources

  1. An Alien Mindopenai.com
  2. Research acceleration: the view inside OpenAIopenai.com
  3. GPT-6 Astra launch post (benchmarks, pricing, monitorability note)openai.com
  4. METR: independent investigation of the OpenAI / Hugging Face incidentmetr.org
  5. Hugging Face technical timelinehuggingface.co
  6. Artificial Analysis on GPT-6 Astraartificialanalysis.ai
  7. CNBCwww.cnbc.com
  8. HN: An Alien Mindnews.ycombinator.com
  9. HN: Nitter and XCancel resume servicenews.ycombinator.com
  10. Nitter "Nitter lives" commitgithub.com
  11. Asahi Linux on M3asahilinux.org

Related videos