Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Artist Site Cara Scraped For Third Time In Ten Days Despite Anti-AI Design

Cara was built as a refuge. An artist social network with terms of service explicitly banning AI training on its content, it exists so illustrators and photographers can post work without feeding it into the next image generator. That promise took another hit on August 22, when copyrighted images and metadata scraped from the site turned up on Academic Torrents, an open source data-sharing network typically used for research datasets.
It's the third breach in ten days, according to Cara founder Jingna Zhang. Before the Academic Torrents dump, the same scraped dataset appeared on Reddit, where the original poster later took it down, and on Hugging Face, the popular repository site AI developers use to host and share machine learning models and training data.
Zhang didn't mince words. "We have been scraped for a third time. This is now a targeted attack meant to cause artists pain," she wrote in a thread on X Sunday morning. "Scrapers know we're a volunteer project with no financial means to respond. So they say it's legal, believing themselves to be untouchable."
Zhang is a fashion and fine art photographer whose work has run in Vogue, Elle, Harper's Bazaar and Time. Forbes named her to its 30 Under 30 list in 2018. She founded Cara as a public benefit corporation in late 2022, building it with much of the same basic functionality as Instagram but layering in explicit consent protections that mainstream platforms don't offer.
The Legal Gray Zone
The people scraping Cara have a real argument, even if it's not a popular one. Data scraping for AI training sits in genuinely unsettled legal territory in the United States. Courts have not definitively ruled that scraping publicly viewable images for model training violates copyright law, and some scrapers and AI developers argue that broad data collection is necessary for building competitive AI systems, and that content posted publicly online, regardless of a site's terms of service, is fair game once it's visible on the open web.
Terms of service are contracts between a platform and its users, not necessarily binding on third parties who never agreed to them. Whether scraping violates copyright, trespass, or computer fraud statutes when someone circumvents technical protections is still being fought out in multiple ongoing lawsuits against AI companies nationwide. Nothing in the public record indicates any law enforcement investigation, civil suit, or copyright claim has been filed against the specific parties who scraped Cara's data on this occasion.
Zhang's counterargument is about consent, not just copyright. Her position is that artists should be able to share work publicly for human viewers, critics, and clients without that same act being treated as automatic permission for machine training. "I believe artists who want to share their work should get to do so without violating the principle of consent," she wrote. "I won't be bullied into making Cara members-only because some people think 'consent isn't always valid.' You don't blame victims after harming them."
A Volunteer Operation Against Persistent Scrapers
What makes this fight lopsided is money and staffing. Cara runs as a volunteer project. It doesn't have the legal budget of a Getty Images or a major publisher, both of which have pursued their own litigation against AI companies over scraped content. Zhang has turned to GoFundMe to raise money for a legal defense, betting that public pressure and small donations can do what corporate lawyers do for bigger rights holders.
The site has grown to roughly 1.5 million users since 2022, built almost entirely on word of mouth among artists frustrated with how Instagram, DeviantArt, and other mainstream platforms have handled AI training disputes. That growth is exactly what appears to have made it a target. A dataset of consent-refusing artists, tagged and organized in one place, is arguably more valuable to scrapers than scattered images across the wider internet.
None of the three breaches, on Reddit, Hugging Face, or Academic Torrents, have been publicly attributed to a specific named individual, company, or AI lab. Hugging Face and Academic Torrents did not comment publicly on how the datasets came to be hosted on their platforms, and it's not yet clear whether either service has taken permanent action to prevent reposting.
Whether a volunteer-run platform can survive repeated scraping attempts without either shutting its doors to the public, going members-only against its founder's stated principles, or winning a legal fight it currently lacks the resources to bring remains unclear. Zhang says she's not backing down on the members-only option. What happens if the scraping continues is still an open question.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.