ROR Labs cover: 10–100× false hits
|

Facial Recognition Errors: What the Evidence Shows

Key takeaways · 13 min read

  • In NIST’s 2019 test of 189 algorithms, false positive rates often differed by a factor of 10 to beyond 100 between demographic groups, but the most accurate algorithms showed the smallest differences.
  • The best systems improved fast: NIST measured search failures falling from 4% in 2014 to 0.2% in 2018. How often a system errs, and for whom, depends heavily on the threshold its operator sets.
  • In the UK’s 2023 test of the Met’s live system, a threshold of 0.6 gave about 1 false alert in 6,000 passers-by and no significant demographic gap in detection; at 0.56 and 0.58 more Black people were falsely flagged.
  • The ACLU counts 14 Americans wrongfully arrested after a facial recognition match. In the cases reviewed, the failure was usually an investigation that treated a lead as proof.

In January 2020 Detroit police arrested Robert Williams at his home in Farmington Hills, in front of his wife and two young children. He was accused of stealing watches from a Detroit shop, and the lead had come from a facial recognition search of a surveillance image. He was not the man in the footage. In June 2024 the city settled his lawsuit, agreed to pay him $300,000 and adopted what the American Civil Liberties Union called the strongest police policy on the technology in the country.

Williams is the best known of at least 14 Americans the ACLU has counted as wrongfully arrested after a facial recognition match. Over the same years, testing by the US National Institute of Standards and Technology (NIST) has shown the most accurate algorithms becoming remarkably good at their job. Both statements are true. The gap between them is where the risk sits.

This article sets out what the most quoted error figures actually measured, why the threshold an operator chooses matters as much as the software, how far the best systems have improved, what went wrong in the known arrests, how well people check the machine’s suggestions, what live deployments in London show, and the rules now taking shape in the United States and Europe.

A pointillist illustration: an empty city pavement at dusk lined with tall lamp posts, a single small camera mounted high on one pole, with one small amber point of light at the camera lens.
A camera sees a face. A threshold decides what it means.

What the famous numbers measured

The study that put the issue on the public agenda was Gender Shades, published in 2018 by Joy Buolamwini and Timnit Gebru. They found two widely used benchmark collections were 79.6% and 86.2% lighter-skinned, then built a balanced set of photographs of parliamentarians and tested three commercial systems. Darker-skinned women were misclassified at rates of up to 34.7%. The highest error rate for lighter-skinned men was 0.8%.

It matters what those systems were doing. Gender Shades tested gender classification, guessing whether a face is male or female, not matching a face to a person’s identity. That is a different task from the one police use. The audit also had an effect. When Inioluwa Deborah Raji and Buolamwini re-ran it in August 2018, all three companies had released new versions within seven months, and the gap between their best and worst subgroups had narrowed sharply. Amazon’s system, not in the original study, still misclassified 31.4% of darker-skinned women.

Gender Shades: the gap before and after the audit

Difference in error rate between the best and worst classified subgroups, in percentage points. Gender classification, not identity matching.

IBM, May 201734.4
IBM, August 201816.7
Face++, May 201733.7
Face++, August 20183.6
Microsoft, May 201720.8
Microsoft, August 20181.5

Sources: Buolamwini J, Gebru T, Proceedings of Machine Learning Research 81:77–91 (2018); Raji ID, Buolamwini J, Communications of the ACM (2022), doi:10.1145/3571151.

The identity-matching question was answered at scale in December 2019, when NIST published its report on demographic effects. It ran 189 mostly commercial algorithms from 99 developers on 18.27 million images of 8.49 million people, drawn from State Department, Homeland Security and FBI databases. False positive rates, where two different people are wrongly declared the same, often varied by a factor of 10 to beyond 100 between groups. They were two to five times higher in women than in men. In a search of 1.6 million FBI mugshots, they were higher for African American women.

What NIST’s 2019 demographic study found

One-to-one and one-to-many tests across 189 algorithms. Results varied widely by algorithm.

Error typeFinding
False positives, by race and originOften 10 to beyond 100 times higher in some groups; highest in West and East African, East Asian and American Indian people
False positives, by sexTwo to five times higher in women
False negativesMore algorithm-specific; differences often below a factor of three
Algorithms developed in ChinaSome showed no dramatic gap between East Asian and white faces

Source: Grother P, Ngan M, Hanaoka K, NISTIR 8280, Face Recognition Vendor Test Part 3: Demographic Effects (2019).

NIST itself warned against summarising it in one sentence. “While it is usually incorrect to make statements across algorithms,” said its lead author, Patrick Grother, “we found empirical evidence for demographic differentials in the majority studied.” Majority is not all, and the range between products was enormous.

The threshold decides the error

A face recognition system does not say yes or no. It produces a similarity score, and someone decides how high the score must be to count as a match. Set the threshold low and the system finds more true matches but flags more strangers. Set it high and false alerts fall but real matches are missed. NIST notes that the vast majority of systems use one fixed threshold for every comparison. If impostor scores run higher for one group, the same setting produces more false matches for that group.

That means a single error rate, or a single “bias” figure, is only meaningful at a stated threshold. A 2019 study by Jacqueline Cavazos and colleagues tested four algorithms on East Asian and white faces and found that for all four the degree of race bias changed with the threshold; to reach equal false accept rates, East Asian faces needed a higher threshold. An earlier NIST study found the threshold needed to hold false positives steady shifted with the demographic mix of the people being compared.

So two police forces using the same software can get different error patterns. One that lowers the threshold to get more candidates is choosing more false leads. Whether they cause harm depends on what happens next.

How much better the best systems got

The technology changed dramatically in the 2010s. In 2018 NIST tested 127 algorithms from 39 developers against a database of 26.6 million photographs. The best search algorithm failed to find a matching photo 0.2% of the time, against 4% in 2014 and 5% in 2010. NIST described the software as 20 times better at database search in four years, while warning that there remained “a very wide spread of capability across the industry.”

Best search failure rate in NIST tests

Share of searches in which the leading algorithm failed to find the matching photo in a large database.

20105%
20144%
20180.2%

Source: NIST, “NIST Evaluation Shows Advance in Face Recognition Software’s Capabilities” (November 2018); 2018 test of 127 algorithms against 26.6 million photos.

Accuracy and fairness turned out to be linked. In congressional testimony in February 2020, NIST’s Charles Romine said the most accurate algorithms failed in about a quarter of one percent of searches, and that “demographic effects are smaller with more accurate algorithms.” Some of the most accurate one-to-many algorithms gave similar false positive rates across groups.

A 2022 NIST follow-up separated the two kinds of error. Differences in missed matches, it said, were largely due to poor photography of some groups, including under-exposure of people with dark skin, and could be reduced with better cameras and lighting. The much larger differences in false positives, which appear even in high-quality photographs, had to be fixed by the developers. A 2023 study by Gabriella Pangelinan and colleagues similarly found that how much of the face is visible in the image best explained gaps in accuracy.

Laboratory results describe an algorithm on good photographs, not a grainy still from a shop camera. A 2025 preprint using degraded synthetic faces found blur and low resolution raised errors, more for women and Black faces, though accuracy stayed above many traditional forensic methods.

Fourteen known wrongful arrests

A pointillist illustration: an empty wooden chair beside a small bare table in a pale room, a closed door behind it, with one amber point of light resting on the table.
Where a lead became an arrest.

In April 2026 the ACLU listed 14 people in the United States known to have been wrongfully arrested after a facial recognition match. Three were arrested by Detroit police, and all three were Black. The list is a floor: it counts only cases that came to light.

Porcha Woodruff’s case shows how a match turns into evidence. In February 2023, eight months pregnant, she was arrested over a carjacking. Detroit police had run a petrol station video through facial recognition, put her file photo in a lineup, and the victim picked her. The charges were dropped. In August 2025 a federal judge dismissed her lawsuit, finding probable cause independent of the match, although the software had chosen who was in the lineup.

The newest case is a Tennessee woman arrested in July 2025 on North Dakota bank fraud charges that were dismissed about six months later. Her lawsuit alleges the match was made against a photo from the suspect’s fake ID, not the surveillance footage.

A Washington Post investigation in January 2025 found at least eight wrongful arrests, seven of Black people, in which officers had described results as a “100% match” or said the software had identified a suspect “immediately and unquestionably.” It found 15 departments in 12 states whose use of matches departed from their own standards.

What investigators skipped in eight wrongful arrests

Number of the eight cases examined by the Washington Post in which officers missed each basic step.

Did not check the suspect’s alibi6 of 8
Did not collect key evidence5 of 8
Ignored differences in physical features3 of 8
Disregarded contradicting evidence2 of 8

Source: Washington Post investigation (January 2025), as summarised by Fortune, 16 January 2025.

A study of 1,136 American cities, published in 2022, found that police use of facial recognition was associated with a larger gap between Black and white arrest rates. It was observational and cannot show cause, but it fits the pattern in the individual cases: the people harmed were overwhelmingly Black, and the failures were in the investigation as much as the software.

The human in the loop

Every police policy on facial recognition says the same thing: a match is a lead, to be checked by a person. The research on how well people check faces they do not know is sobering. In a 2014 study, 30 passport officers matching live people to photo IDs made errors 10% of the time and wrongly accepted 14% of fraudulent photos. Years of experience made no difference.

Specialists do much better. A 2018 study led by Jonathon Phillips gave the same difficult face pairs to 57 forensic facial examiners, other specialists, fingerprint examiners and students, and to four algorithms. The newest algorithm scored slightly above the median examiner. The single best result came from combining one examiner with that algorithm, which beat two examiners working together. But every specialist group included someone who scored below the median student.

People and algorithms on a hard face-matching test

Median accuracy (area under the curve; 1.0 means no errors, 0.5 is chance).

Students0.68
Fingerprint examiners0.76
Super-recognisers0.83
Forensic facial reviewers0.87
Forensic facial examiners0.93
Best algorithm (2017)0.96
One examiner plus best algorithm1.0

Source: Phillips PJ et al., PNAS (2018), doi:10.1073/pnas.1721355115. Twenty challenging image pairs.

The lesson is narrower than “humans are bad” or “machines are good”. Review helps when the reviewer is trained and the check is independent. A detective glancing at two photos, or a witness shown a lineup built around the software’s suggestion, is not an independent check at all.

Live cameras on London streets

A pointillist illustration: a plain white van parked on a wide, empty pedestrian street under an overcast sky, a short mast of cameras on its roof, with one amber point of light at the top of the mast.
Thousands of faces, a handful of alerts.

Live facial recognition works differently from a search after a crime. Cameras compare every passing face to a watchlist in real time, and officers stop people the system flags. In 2023 the UK’s National Physical Laboratory published a police-commissioned study of the NEC system the Met uses. It placed 405 volunteers among crowds of 7,000 to 38,000 people at five deployments and used a watchlist of nearly 180,000 images, about 20 times larger than any used operationally at the time.

At the Met’s default threshold of 0.6, the system found 89% of the volunteers on the watchlist and falsely flagged about 1 in 6,000 passers-by. Differences in detection by sex and ethnicity were not statistically significant. At lower settings of 0.56 and 0.58, false alerts showed a significant imbalance, with more Black people wrongly flagged. At 0.64 and above there were none.

One system, four settings

How the threshold changed the Met’s live system in the 2023 National Physical Laboratory test.

ThresholdWhat the test found
0.56 and 0.58False alerts showed a statistically significant imbalance, with more Black people wrongly flagged
0.60 (default)89% of watchlisted volunteers found; about 1 false alert in 6,000; no significant demographic difference in detection
0.62A single false alert during testing
0.64 and aboveNo false alerts

Source: National Physical Laboratory, NPL Report MS 43, Facial Recognition Technology in Law Enforcement: Equitability Study (March 2023). NEC NeoFace V4 algorithm.

Small rates still add up. At 1 in 6,000, a crowd of 38,000 would produce about six false alerts. In its figures for September 2024 to September 2025, the Met reported 203 deployments, 3,147,436 faces scanned, 2,077 alerts, 962 arrests and 10 false alerts, roughly one per 315,000 faces. Eight of the ten people wrongly flagged were Black, too few to calculate a rate but enough to keep the question open.

Rules in Detroit, Brussels and The Hague

Detroit’s 2024 settlement with Robert Williams set out the most specific rules in the United States. Police may not arrest anyone based solely on a facial recognition result, or on a photo lineup that follows directly from one without independent evidence. Officers must be trained on the technology’s higher misidentification rates for people of colour and women, and the city must audit every case since 2017 in which a search led to an arrest warrant. The court keeps jurisdiction for four years. The ACLU says more than 20 US cities and other jurisdictions have banned police use of the technology altogether.

The European Union has gone further on live use. Article 5 of the AI Act, applying since 2 February 2025, prohibits real-time remote biometric identification in publicly accessible spaces for law enforcement, except where strictly necessary to search for specific victims of abduction or trafficking, to prevent a specific and imminent threat to life, or to locate suspects in serious crimes. Each use needs prior authorisation from a judge or independent authority. The Act also bans building facial recognition databases by untargeted scraping of images from the internet or CCTV.

That last rule describes Clearview AI’s business model. In September 2024 the Dutch data protection authority announced a €30.5 million fine against the company for building a database of more than 30 billion photographs without consent, with further penalties of up to €5.1 million for continued breaches. Other European regulators had already fined it. “Facial recognition is a highly intrusive technology, that you cannot simply unleash on anyone in the world,” said the authority’s chairman, Aleid Wolfsen.

What the evidence suggests

The evidence supports three points more firmly than the rest. First, error rates are properties of a particular algorithm, image and threshold, not of “facial recognition” in general; the best algorithms are very accurate, and their demographic gaps are smallest. Second, false positives differed between groups far more than missed matches did, and in the known wrongful arrests the people harmed were overwhelmingly Black. Third, those arrests followed investigations that treated a candidate as a suspect without independent corroboration. Better software reduces the first problem. Only procedure fixes the third.

Questions people ask

Is facial recognition biased against Black people and women?

Many algorithms NIST tested in 2019 had higher false positive rates for some groups, including Black and Asian people and women. The most accurate algorithms showed much smaller differences.

How accurate is police facial recognition?

It depends on the algorithm, the photo and the threshold. The best NIST-tested algorithms missed a match in about 0.2% of searches of a large database in 2018, but poor surveillance images reduce accuracy.

How many people have been wrongly arrested because of it?

The ACLU counted 14 known cases in the United States by April 2026. That is a minimum, because not every misidentification becomes public.

Is live facial recognition legal in Europe?

The EU AI Act prohibits real-time identification in public places by police except in narrowly defined serious cases with prior authorisation. The UK is not bound by the Act, and the Met continues to deploy live systems.

What should I do if I think I was misidentified?

Ask whether facial recognition was used and what independent evidence exists, and speak to a lawyer. In the Washington Post’s review, alibis went unchecked in six of eight wrongful arrests.

The short version

  • In NIST’s 2019 test of 189 algorithms, false positive rates often differed by a factor of 10 to beyond 100 between demographic groups, but the most accurate algorithms showed the smallest differences.
  • The best systems improved fast: NIST measured search failures falling from 4% in 2014 to 0.2% in 2018. How often a system errs, and for whom, depends heavily on the threshold its operator sets.
  • In the UK’s 2023 test of the Met’s live system, a threshold of 0.6 gave about 1 false alert in 6,000 passers-by and no significant demographic gap in detection; at 0.56 and 0.58 more Black people were falsely flagged.
  • The ACLU counts 14 Americans wrongfully arrested after a facial recognition match. In the cases reviewed, the failure was usually an investigation that treated a lead as proof.
  • Trained people are not a reliable safety net: passport officers accepted 14% of fraudulent photos. Detroit now bans arrests based solely on a match, and the EU restricts live identification in public places.

This article summarises published research, official evaluations and news reports for general information. It is not legal advice. If you are affected by a police identification or have concerns about how your image is used, contact a qualified lawyer or your national data protection authority.

Further reading: Grother, Ngan and Hanaoka, NISTIR 8280 (2019), for the full demographic results by algorithm. National Physical Laboratory, NPL Report MS 43 (2023), for how a live system behaves at different thresholds. Phillips et al., PNAS (2018), for how people and machines compare at matching faces.

Three books
  • Unmasking AI, Joy Buolamwini (2023). The Gender Shades researcher’s own account of how the audit began and what followed. Clear on how unbalanced test sets hide errors; it is a memoir and an advocate’s case, not a neutral review.
  • Your Face Belongs to Us, Kashmir Hill (2023). A reporter’s history of Clearview AI and the scraping of billions of photos. Strong on how the database was built and used; lighter on accuracy testing.
  • Race After Technology, Ruha Benjamin (2019). A sociologist on how new technologies can reproduce old inequalities. Useful framing for the arrest cases; it argues a position and predates the newer NIST results.

Sources

Buolamwini J, Gebru T. Gender Shades. Proceedings of Machine Learning Research 81:77–91 (2018). — Raji ID, Buolamwini J. Actionable Auditing Revisited. Communications of the ACM (2022), doi:10.1145/3571151. — Grother P, Ngan M, Hanaoka K. NISTIR 8280, Face Recognition Vendor Test Part 3: Demographic Effects (2019). — NIST, “NIST Study Evaluates Effects of Race, Age, Sex on Face Recognition Software” (December 2019). — Cavazos JG, Phillips PJ, Castillo CD, O’Toole AJ. arXiv:1912.07398 (2019). — O’Toole AJ et al. NISTIR 7757 (2011), doi:10.6028/NIST.IR.7757. — NIST, “NIST Evaluation Shows Advance in Face Recognition Software’s Capabilities” (November 2018). — Romine CH, NIST congressional testimony on facial recognition technology (6 February 2020). — Grother P. NISTIR 8429, FRVT Part 8 (2022), doi:10.6028/NIST.IR.8429.ipd. — Pangelinan G et al. arXiv:2304.07175 (2023). — Cuellar M, To HK, Mehrotra A. arXiv:2505.14320 (2025, preprint). — ACLU, “More than a Dozen Wrongful Arrests Due to Police Reliance on Facial Recognition Technology” (14 April 2026). — ACLU of Michigan, press release on the Williams v. City of Detroit settlement (28 June 2024). — Fortune, Detroit settlement (29 June 2024) and Washington Post investigation summary (16 January 2025). — CBS Detroit, Woodruff ruling (August 2025). — Valley News Live, Fargo wrongful-arrest lawsuit (15 September 2026). — Johnson TL et al. Government Information Quarterly (2022), doi:10.1016/j.giq.2022.101753. — White D, Kemp RI, Jenkins R, Matheson M, Burton AM. PLoS One (2014), doi:10.1371/journal.pone.0103510. — Phillips PJ et al. PNAS (2018), doi:10.1073/pnas.1721355115. — National Physical Laboratory, NPL Report MS 43 (March 2023). — The Register, Met live facial recognition figures (3 November 2025). — EU Artificial Intelligence Act, Article 5 (prohibitions applying from 2 February 2025). — Library of Congress Global Legal Monitor and NL Times on the Dutch DPA Clearview fine (2024).

Similar Posts