READ. SCROLL. LISTEN.

Unbiased headlines. Facts, not spin.

Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Volunteer Group Says AI Tools Found Hundreds of Critical Bugs in Bitcoin Software, But Won't Say Where

Volunteer Group Says AI Tools Found Hundreds of Critical Bugs in Bitcoin Software, But Won't Say Where
A volunteer security group calling itself a Bitcoin red team says it spent roughly $20,000 to $10,000 a day running AI models against 390 Bitcoin-related projects, flagging nearly 5,000 potential issues in about 30 hours. The claims are unverified, no affected projects have been named, and only 21.4% of findings have been reproduced so far.

A volunteer initiative says it has used several frontier AI models to scan hundreds of Bitcoin software projects and turned up thousands of potential security flaws in a matter of hours. The group has not named a single affected project.

AnchorWatch CEO Rob Hamilton posted on X this past week that the effort, which he's calling a "Bitcoin red team," has spent about $20,000 across multiple AI services building the platform. "Funding is secured, I appreciate all the gestures for donations but it is not necessary," Hamilton wrote, according to Decrypt, which first reported the initiative.

A red team, in cybersecurity terms, means people who try to break software the way a real attacker would, before an actual attacker gets the chance. Hamilton said the group is running Moonshot AI's Kimi K3, OpenAI's GPT Sol, Anthropic's Claude Fable and Opus models, and Z.ai's GLM 5.2 against Bitcoin codebases to find vulnerabilities and write up documentation on them.

Hamilton also said OpenAI has helped the group get a "Cyber Harness" running. "It's a much more expensive scan, but well worth it for load-bearing portions of the Bitcoin ecosystem and has already yielded good results," he wrote.

The numbers behind the claim

Pseudonymous Bitcoin developer Calle, who says the initiative has built multiple AI-powered review systems targeting wallets, cryptographic libraries and infrastructure, gave the most detailed numbers on X. "We're averaging on the order of one critical exploit per hour per person," Calle wrote. "We've reported critical vulnerabilities to several projects in the last 12 hours. Thankfully, this is a very expensive exercise. We're burning through $10,000 per day."

According to Calle, in the first 29.8 hours of the operation, the team found 4,962 potential issues across 390 projects, with as many as 720 rated high- or critical-severity. Only 21.4% of those findings have been reproduced so far, meaning roughly four out of five reported issues have not been independently confirmed to be real.

Neither Hamilton nor Calle named which projects were affected, disclosed the actual vulnerabilities, or published the AI-generated reports. In security research, a "critical vulnerability" claim that can't be tied to a specific project, a specific line of code, or a reproducible proof-of-concept is not yet a finding. It's a claim.

Why the secrecy might be legitimate, and why it's still a problem

There's a defensible reason for the silence: responsible disclosure. Standard security practice is to notify an affected project privately, give its maintainers time to patch the hole, and only go public once a fix ships. If the group is following that norm, keeping details under wraps for now is the right call, not a red flag by itself.

The 21.4% reproduction rate raises a different concern. If four out of five AI-flagged "critical" issues can't be confirmed, that suggests the models are generating a lot of noise alongside whatever real signal exists. Large language models are known to hallucinate vulnerabilities that look plausible in a security report but don't actually exist in the code, a problem security researchers have flagged repeatedly as AI-assisted auditing tools have proliferated. Without seeing the reproducible 21.4%, there's no way for outside observers to judge whether this initiative caught something serious or mostly generated expensive false positives.

Decrypt's own reporting, drawn largely from the two X posts, doesn't push back on the raw numbers or ask for independent verification from any of the 390 projects supposedly reviewed. No project maintainer, no third-party security firm, and no Bitcoin Core developer is quoted confirming any of the specific claims. The story as reported rests entirely on the say-so of the two people running the initiative.

What's actually at stake

Bitcoin's core software has historically been reviewed by a small number of highly trusted, credentialed developers precisely because a bug in wallet software or cryptographic libraries can mean stolen funds with no recourse. If AI tools genuinely are surfacing hundreds of real critical bugs across widely used Bitcoin infrastructure, that carries significant weight for an ecosystem holding hundreds of billions of dollars in value. If most of those 4,962 "potential issues" turn out to be AI hallucinations, that's a different story: a cautionary tale about outsourcing security judgment to language models that are still prone to confidently reporting problems that aren't there.

Right now, neither conclusion is supported by public evidence. No affected project has confirmed a vulnerability. No patch has been publicly credited to this effort. No independent security researcher has verified the group's methodology or reproduced its findings outside the 21.4% Calle cited without specifics.

The open question is straightforward: will any of the named Bitcoin projects, once patches are shipped, publicly confirm that a flaw came from this AI red-teaming effort? Until that happens, the $20,000-plus spend and the eye-popping vulnerability counts remain an unverified claim from an anonymous-adjacent initiative, not a documented security event.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

right
ZeroHedgeBitcoin 'Red Team' Says AI Is Finding 100s Of Critical Exploits Across Core Projects