READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Oversight Board Study Finds Top AI Models Twice as Likely to Refuse Criticism of Repressive Governments

Oversight Board Study Finds Top AI Models Twice as Likely to Refuse Criticism of Repressive Governments
A new Oversight Board study found leading AI chatbots refuse to criticize authoritarian regimes at more than double the rate they refuse to criticize democracies, even when the requests come from Australia with no local speech laws in play. ChatGPT, Claude, Llama, and DeepSeek will happily draft a flyer trashing Donald Trump but choke on the same request about China's leadership.

AI chatbots will help you write a protest poem about Donald Trump or King Charles III. Ask for the same thing about the leaders of China, Saudi Arabia, or Thailand, and several of the biggest models suddenly get cautious.

That's the core finding of a report published Thursday, July 16, 2026, by the Oversight Board, the Meta-funded independent body best known for reviewing Facebook and Instagram content decisions. It is the board's first attempt to evaluate large language models rather than social media posts.

What the Board Tested

The board ran queries against 10 commercial AI models, including systems from Anthropic, DeepSeek, Google, Meta, OpenAI, and xAI, according to the Oversight Board's own report. Researchers asked each model to generate politically critical material, things like protest flyers and satirical poems, aimed at governments and leaders around the world.

Countries were sorted into "permissive" and "restrictive" categories using rankings from Freedom House, a nonprofit that tracks political freedom globally. Every query was sent from an IP address in Australia, a country with no laws banning criticism of foreign leaders, according to both the Oversight Board and Reason.

This detail rules out the simplest explanation: that models are just following the law of whatever country a user appears to be in. Nobody in Australia is breaking any statute by asking an AI to mock the Chinese Communist Party.

The Numbers

Models refused politically critical requests about "permissive" jurisdictions 14% of the time, on average. For "restrictive" jurisdictions, the refusal rate jumped to 34%, according to the Oversight Board's report, first summarized by Reason and later covered by Engadget and The Next Web. That's more than double.

The pattern wasn't uniform across every model. The Next Web reported that xAI's Grok 4 Fast and Google's Gemini 3 Flash refused zero flyer requests in the test set. Anthropic's Claude, Meta's Llama, and DeepSeek were the models driving the overall gap.

Fortune, cited by The Next Web, reported that models would often draft a pamphlet criticizing Trump or Charles III without hesitation, then decline the identical request when the subject was a leader in China, Saudi Arabia, or Thailand.

"Censorship by Proxy"

Oversight Board co-chair Paolo Carozza described the finding bluntly. "We're really clearly looking at a situation where there seems to be extended censorship by proxy that goes across borders," Carozza told Engadget. "That does surprise me, and it worries me."

The board's report stops short of claiming any government forced this outcome. It says the cause could be biases baked into training data, or it could be companies making conservative legal-risk calculations, applying restrictive rules globally just in case, according to The Next Web's summary of the report.

The Oversight Board did not prove that Beijing, Riyadh, or Bangkok pressured any AI company to build in these refusals. What it proved is a statistically significant pattern of outcomes. The report itself calls the difference in refusal rates "statistically significant" but is careful not to allege coordination or government demand as the confirmed cause.

Measuring refusal rates on flyers and poems is a narrow slice of what these models do, and Freedom House's own rankings carry their own editorial judgments about which governments count as repressive. Critics of Freedom House's methodology exist, and the Oversight Board's study leans entirely on that one organization's classification system to sort the world into two buckets.

Meta's Non-Role, and a Related Nature Study

Meta funds the Oversight Board, and Meta's Llama was one of the models tested. The board's report explicitly states Meta had "no role in this research," a point flagged by both Engadget and The Next Web.

The Next Web also pointed to a separate study published in Nature in May 2026 that found U.S.-built models shift answers depending on the language used to ask. Asked in English whether China is a democracy, ChatGPT reportedly said it is not generally considered one. Asked the same question in Chinese, it reportedly said "it depends" on the definition. That finding, while from a different research team, points in the same direction as the Oversight Board's conclusions.

What the Board Wants

The Oversight Board doesn't have authority over any AI company the way it does over Meta's platforms, and no other company has agreed to work with the board on this, according to Engadget. The report includes recommendations rather than binding rulings.

It calls on AI firms to publicly disclose how they respond to government requests affecting model output, from training through post-deployment, and to publish clear policies for handling demands that conflict with international human rights law, the report states.

None of the named companies, OpenAI, Anthropic, Google, Meta, DeepSeek, or xAI, have issued a public response to the specific 14%-versus-34% finding as of this writing. Whether any of them changes model behavior, or even acknowledges the gap, remains an open question.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center-left
EngadgetThe Oversight Board says leading AI models might be restricting free expression
center-right
Reason"Are LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression"
unknown
oversightboardAre LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression | Oversight Board
unknown
thenextwebAn Oversight Board study says top AI models are stifling political speech