OpenAI publishes internal metrics on AI-agent-driven research automation and Pachocki essay warning that no lab has solved AI control/alignment
OpenAI published a blog post with internal usage metrics (agent inference spend, token output growth, agent-vs-human workday ratios, task success rates by difficulty) claiming it has met its self-declared "automated research intern" milestone, alongside a companion essay by chief scientist Jakub Pachocki ("An Alien Mind") warning that chain-of-thought monitoring is losing reliability and that no AI lab, including OpenAI, has sufficiently solved alignment/control for continued max-speed scaling.
Entities: OpenAI, Jakub Pachocki, GPT-6 Astra, GPT-5.6 Sol, Anthropic, Epoch AI
0 primary
What happened
OpenAI published a blog post with internal usage metrics, including median researcher inference spend of $600/day, 124x token output growth since December 2025, a 3.1 agent-workdays-per-human-workday ratio, and an 86% success rate on sub-15-minute tasks, to claim it has hit its self-declared "automated research intern" milestone. Alongside this, chief scientist Jakub Pachocki published an essay, "An Alien Mind," warning that chain-of-thought monitoring is becoming less reliable and that no AI lab, including OpenAI itself, has adequately solved alignment or control for the pace of scaling underway. The release came three days after the GPT-6 Astra announcement.
Why it matters
This gives regulators, competitors and enterprise buyers the first detailed internal trend line OpenAI has disclosed on agent-driven research productivity, which is more specific than typical lab messaging even if unverifiable. Pachocki's public admission that alignment and control are unsolved, paired with a call for binding, externally enforced scaling standards, is a notable signal because it comes from OpenAI's own chief scientist rather than an outside critic, and could feed into regulatory arguments for mandatory oversight. For most professionals, though, nothing here is deployable today: it is a disclosure and a warning, not a product or policy change.
What is noise
Every metric is self-reported and self-measured, including success rates judged by an AI classifier of unstated reliability, with no independent audit, so the "automated research intern milestone" framing is essentially unfalsifiable marketing dressed as data. The timing, three days after the GPT-6 Astra launch, and the pairing of an impressive-sounding capability claim with a safety warning reads as calculated positioning rather than a spontaneous disclosure, and coverage repeating the "recursive self-improvement" framing without noting the total absence of external verification overstates its evidentiary weight.
Watch next
- 01Whether independent researchers or a rival lab publish comparable agent-productivity metrics that can be cross-checked against OpenAI's self-reported figures
- 02Any concrete move toward Pachocki's proposed binding, externally enforced scaling standards, e.g. draft legislation, industry pact, or regulator response, versus this remaining a talking point
- 03OpenAI's actual progress toward the stated March 2028 fully automated AI researcher target, tracked via subsequent disclosures rather than restated milestones
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679