READ. SCROLL. LISTEN.

Unbiased headlines. Facts, not spin.

Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Study Finds Human Reviewers Often No Better Than AI at Catching Facial-Recognition Mistakes, as New Research Revisits Eyewitness Memory Flaws

Study Finds Human Reviewers Often No Better Than AI at Catching Facial-Recognition Mistakes, as New Research Revisits Eyewitness Memory Flaws
Scientific American reports fresh nuance on how unreliable eyewitness memory really is in criminal cases, while a University of Colorado Boulder study found human reviewers of AI facial-recognition matches are only as good as their own face-recognition skills, which vary wildly. Put together, both point to the same uncomfortable fact: the human safeguards courts and police rely on aren't automatically reliable just because they're human.

Human memory has sent innocent people to prison. More than 70 percent of convictions overturned by DNA evidence between 1989 and 2019 involved mistaken eyewitness identification, according to Scientific American. New research is now adding nuance to that picture. A study out of the University of Colorado Boulder raises a related question for modern policing: can a human reviewer actually catch an AI's facial-recognition mistake?

Memory Isn't a Recording, It's a Reconstruction

Each time someone recalls an event, the memory doesn't get pulled from storage and put back unchanged. It goes through a process called reconsolidation, according to Scientific American, and that process lets new information bleed into the original memory.

Elizabeth Loftus, a memory researcher at the University of California, Irvine, demonstrated this back in the early 1970s. She found that simply changing a word in a question about a car crash changed what witnesses remembered. Ask if two cars "smashed" instead of "bumped," and witnesses were more likely to falsely recall broken glass that was never there.

The same distortion shows up in what researchers call "flashbulb" memories, the vivid recollections people have of major events like September 11, 2001. People are confident they'll never forget where they were that day. Scientific American reports they still get roughly 40 percent of the details wrong.

Distance from a crime, lighting conditions, and whether a witness and a suspect are the same race all affect how accurate an identification is likely to be, according to Scientific American. Newer research, including work by Maastricht University forensic psychologist Lilian Kloft-Heller on how intoxication affects memory, is still catching up to how many variables actually matter.

Eyewitness testimony is not worthless. Laura Mickes, a memory researcher at the University of Bristol, told Scientific American that memory should be treated "like any other piece of forensic evidence." Her point: if it's contaminated, it won't hold up, but that doesn't mean uncontaminated memory evidence has no value. The implication is that procedure, not blanket distrust, is what separates a useful identification from a wrongful conviction.

The Same Problem Shows Up With AI

A study published in February 2026 in the Journal of Applied Research in Memory and Cognition found a parallel weak spot in a technology increasingly used by police departments: facial-recognition software.

Most of these systems work with a "human in the loop." The algorithm flags a potential match, and a person signs off before any action is taken. David Dobolyi, lead author of the study and an assistant professor at the University of Colorado Boulder's Leeds School of Business, tested how well human judgment lines up with what the algorithms produce.

The result: agreement between humans and AI was highest among people who scored well on standardized face-recognition tests, sometimes called "super recognizers." Nearly 4,000 participants took the test as part of the research. Dobolyi's team found the AI systems themselves performed comparably to those higher-scoring humans.

"I wouldn't say that having a human in the system solves the problem," Dobolyi said. "If you take the average human, that person may not be better than a specific model." He added that treating a human reviewer as an automatic fail-safe misreads what's actually happening. "People have their own biases. These systems have their own biases. It's not about one being better than the other."

The study also compared six commercial and open-source facial-recognition systems and found they don't all see faces the same way. Agreement with human judgment sometimes differed depending on whether the comparison involved Black or white faces, though those differences varied by which system was tested, according to the study.

Police departments and prosecutors have a reasonable counterargument: eyewitness identification and facial-recognition matches are rarely used as the sole piece of evidence in a case. Courts already require corroboration in most jurisdictions, and departments that use facial recognition typically treat a match as an investigative lead, not a conviction. That's a fair point, and nothing in either study proves these tools are useless when paired with other evidence.

What the research does establish is narrower and harder to dismiss: the assumption that a human reviewer will reliably catch what an algorithm gets wrong isn't automatically true. Dobolyi's data shows it depends entirely on how good that specific reviewer is at recognizing faces, a skill that varies enormously across the general population and one departments generally don't test for.

Dobolyi points to the free UNSW Face Test, developed by researchers at the University of New South Wales, as one way individuals or agencies could measure that ability. No source indicates any police department currently screens facial-recognition reviewers this way. Whether departments start vetting who's allowed to sign off on an AI match, the same way courts have slowly tightened lineup procedures after decades of wrongful-conviction data, remains an open question.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center-left
Scientific AmericanWhat eyewitness accounts tell us about the science of memory
unknown
colorado.eduThink a human reviewer can catch an AI mistake? Facial-recognition study suggests it's complicated
unknown
PositronWhat eyewitness accounts tell us about the science of memory