Signum
Feed
Useful signal7 Sept 2026medium confidence

OpenAI publishes internal metrics on AI-agent-driven research automation and Pachocki essay warning that no lab has solved AI control/alignment

OpenAI published a blog post with internal usage metrics (agent inference spend, token output growth, agent-vs-human workday ratios, task success rates by difficulty) claiming it has met its self-declared "automated research intern" milestone, alongside a companion essay by chief scientist Jakub Pachocki ("An Alien Mind") warning that chain-of-thought monitoring is losing reliability and that no AI lab, including OpenAI, has sufficiently solved alignment/control for continued max-speed scaling.

CapabilityLabourGovernanceInfrastructure

Entities: OpenAI, Jakub Pachocki, GPT-6 Astra, GPT-5.6 Sol, Anthropic, Epoch AI

61Useful signal
1 source
0 primary
Was this useful?
01

What happened

OpenAI published a blog post with internal usage metrics, including median researcher inference spend of $600/day, 124x token output growth since December 2025, a 3.1 agent-workdays-per-human-workday ratio, and an 86% success rate on sub-15-minute tasks, to claim it has hit its self-declared "automated research intern" milestone. Alongside this, chief scientist Jakub Pachocki published an essay, "An Alien Mind," warning that chain-of-thought monitoring is becoming less reliable and that no AI lab, including OpenAI itself, has adequately solved alignment or control for the pace of scaling underway. The release came three days after the GPT-6 Astra announcement.

02

Why it matters

This gives regulators, competitors and enterprise buyers the first detailed internal trend line OpenAI has disclosed on agent-driven research productivity, which is more specific than typical lab messaging even if unverifiable. Pachocki's public admission that alignment and control are unsolved, paired with a call for binding, externally enforced scaling standards, is a notable signal because it comes from OpenAI's own chief scientist rather than an outside critic, and could feed into regulatory arguments for mandatory oversight. For most professionals, though, nothing here is deployable today: it is a disclosure and a warning, not a product or policy change.

03

What is noise

Every metric is self-reported and self-measured, including success rates judged by an AI classifier of unstated reliability, with no independent audit, so the "automated research intern milestone" framing is essentially unfalsifiable marketing dressed as data. The timing, three days after the GPT-6 Astra launch, and the pairing of an impressive-sounding capability claim with a safety warning reads as calculated positioning rather than a spontaneous disclosure, and coverage repeating the "recursive self-improvement" framing without noting the total absence of external verification overstates its evidentiary weight.

04

Watch next

  1. 01Whether independent researchers or a rival lab publish comparable agent-productivity metrics that can be cross-checked against OpenAI's self-reported figures
  2. 02Any concrete move toward Pachocki's proposed binding, externally enforced scaling standards, e.g. draft legislation, industry pact, or regulator response, versus this remaining a talking point
  3. 03OpenAI's actual progress toward the stated March 2028 fully automated AI researcher target, tracked via subsequent disclosures rather than restated milestones

Coverage

1 story

More capability signals

Full feed →