Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Anthropic Says Automated System Beat Human Researchers on AI Safety Fixes, Cost $4 an Hour Versus $150

Anthropic published a paper this week claiming its automated research system outperformed human researchers at fixing specific AI safety problems, and did it for a fraction of the cost.
The paper, titled "Automated Researchers Can Reliably Mitigate Alignment Failures," was led by Anthropic fellow Chen Yueh-Han, according to TechCrunch. The system was tested against 10 benchmarks, each targeting a specific kind of misaligned model behavior. It improved performance on every single one without degrading the model's overall performance elsewhere.
Here is how it worked, per TechCrunch's reporting on the paper: the automated system, which Anthropic calls an Automated Alignment Researcher (AAR), searches existing research literature, proposes a fix, and trains the model on that fix for 30 minutes. It repeats this over multiple rounds, keeping what works and throwing out what doesn't. That's standard research methodology, just run by a machine instead of a grad student.
The headline numbers are the ones Anthropic itself chose to highlight. "The best AAR method beats what experienced humans propose, on average within six hours," the paper states, according to TechCrunch. And on cost: "An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers."
That's a 37-to-1 cost advantage, according to Anthropic's own comparison. If a company is paying human PhDs $150 an hour to run experiments a machine can run for $4, and the machine gets better results faster, that's not a small efficiency gain. That's a category shift in how research gets done, at least for this narrow task.
What This Actually Proves, and What It Doesn't
Anthropic's paper is upfront about the limits, and TechCrunch's coverage flagged them clearly. The automated system only works as well as the benchmarks it's optimizing against. If a benchmark doesn't actually capture the alignment goal researchers care about, the system will happily optimize for the wrong thing. Anthropic itself acknowledges there's "significant work to be done" in building and maintaining benchmarks that mean something.
This is the fundamental issue. An automated system that's faster and cheaper than humans at gaming a benchmark isn't automated alignment. It's automated benchmark-gaming, if the benchmarks are weak. Anthropic's paper doesn't claim to have solved that problem. It claims the automated approach works well on the 10 benchmarks it tested.
The paper also frames this as a step toward "recursive self-improvement," the idea that AI systems could eventually improve their own training processes with less and less human involvement. TechCrunch noted the paper is explicit about this ambition and about the implication that human AI researchers could eventually become less central to the process.
Whoever is skeptical of AI safety hype has a fair point here: a company publishing a paper showing its own automated system beats its own human researchers on benchmarks it built is not a neutral third-party audit. Anthropic has a commercial interest in demonstrating it can scale alignment work cheaply as it races against OpenAI, Google DeepMind, and Chinese labs like Z.AI. That doesn't mean the result is wrong. It means the result comes from an interested party grading its own homework, and it hasn't been independently replicated.
Where the Coverage Gets Sloppy
A report from Crypto Briefing covering the same paper claimed Anthropic has an "unreleased Model 2" that reportedly outperforms a "current Mythos 5 model," and framed the alignment paper as tied to prediction markets on Anthropic's IPO valuation and benchmark leadership. None of that appears in Anthropic's actual paper or in TechCrunch's reporting on it. There is no confirmation from Anthropic, in the material reviewed here, of models by those names, and no established link between this alignment research and Anthropic's corporate valuation or public-listing status. That framing should be treated as unverified.
A separate republication from PressBee added no new reporting and simply pointed readers back to the original TechCrunch piece.
What's actually confirmed: Anthropic ran an experiment, published the methodology, and got a result that favors automation on a narrow, self-defined task. The company hasn't said when or whether this approach will be used in production training of its Claude models. Anthropic also hasn't addressed, in the material published so far, how it plans to guard against the benchmark-gaming risk it flagged in its own paper. That's the open question worth watching as other labs, including OpenAI and Google DeepMind, decide whether to publish comparable experiments of their own.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.