AI bodycam transcripts preserve words but confuse speakers, researchers find
An October 9 report highlights errors in assigning police body-camera dialogue to speakers. Researchers recommend human correction before transcripts are used where attribution matters.
Northeastern University researchers warn that AI transcripts of Kansas City, Missouri, police body-camera recordings can capture much of the dialogue while confusing who said it, according to an October 9 report. The findings highlight why human review matters before transcripts inform criminal cases or officer accountability.
The report, published by Northeastern Global News and republished by Phys.org, brings expert commentary to research published on August 6 in the Journal of Experimental Criminology. It does not announce a new policing system or demonstrate that transcription mistakes changed the outcome of a criminal case.
Fewer speakers and missing words in bodycam transcripts
Criminology professor Eric Piza and doctoral student Savannah Reid examined transcripts from 176 body-camera videos covering 73 incidents and approximately 23 hours. The study used a convenience sample of Kansas City encounters recorded between January 1 and March 5, 2024.
Researchers compared Rev.ai transcripts with versions edited by human transcribers at Rev. Those editors worked from AI drafts rather than independently transcribing the recordings. That distinction matters: the comparison measured differences between automated and corrected text, with both versions sharing a starting point.
According to the October report, automated transcripts identified an average of 3.3 participants, compared with 4.8 in human-edited versions. The automated versions identified at most seven participants; human editors identified as many as 16. AI also assigned longer passages to individual speakers, reducing the apparent back-and-forth.
The automated transcripts contained about 25,000 fewer words across the sample, the report said. Missing dialogue could remove an explanation, warning or response relevant to understanding an encounter. That is a potential consequence identified in the report, rather than a demonstrated effect on a prosecution or disciplinary decision.
One example shows why preserving words is insufficient. The automated transcript grouped a civilian's statements about stopping drug sales together with an officer's response about imprisonment, assigning them to one speaker. Much of the wording remained, but the transcript changed who appeared to be saying it.
Searchable recordings offer benefits with limits
The workload helps explain interest in automation. The researchers estimated that Kansas City police generate more than 350 hours of body-camera footage in a typical day, far beyond the footage examined in this study. Reid said recording encounters is much easier than finding time to watch and analyse them.
Logan Seacrest, a criminal justice and civil liberties fellow at the R Street Institute, told Northeastern Global News that searchable transcripts could help public defenders find a short relevant passage within hours of video. The study did not measure that potential time saving.
“The catch is that AI puts a layer of interpretation between the recording and the record,” Seacrest said. He also raised automation bias—the tendency to give machine-generated material undue credibility. Reid noted that body-camera recordings are made in uncontrolled conditions where even people present may struggle to distinguish speakers.
Shellie Solomon, chief executive of research and consulting firm Justice & Security Strategies, stressed that humans must interpret the details AI gathers. Calli Schroeder of the Electronic Privacy Information Center warned that transcripts can omit context needed to assess encounters. Solomon also noted that the camera itself records only a limited perspective.
Earlier speech-recognition research adds context
A separate PNAS study by Allison Koenecke and colleagues, published in March 2020, illustrates why transcription performance needs scrutiny across speakers. It tested speech-recognition systems from Amazon, Apple, Google, IBM and Microsoft using 19.8 hours of interviews with 73 Black and 42 white speakers.
Average word-error rates were 0.35 for Black speakers and 0.19 for white speakers, with disparities across all five systems. Those historical results do not establish current performance. Nor did that study test Rev.ai, the Kansas City recordings or speaker attribution; it cannot establish racial disparities in the body-camera sample.
What the body-camera study cannot establish
The body-camera paper covers one department and one vendor. Human-transcriber reliability was unmeasured, and transcripts were not independently validated against the source audio. Its text-based metrics also exclude tone, emotion and nonverbal behaviour, limiting what the comparison can say about the full meaning of an encounter.
The authors recommend human correction for attribution-sensitive uses and further testing across vendors, agencies and languages. They identify searching and aggregate text analysis as more promising applications. Reid also urged scrutiny of encounters that automated review leaves out, because those recordings might otherwise receive no further attention.
Kansas City Police Department did not immediately respond to Northeastern Global News's request for comment, the outlet reported. Piza expressed cautious optimism about future tools combining video-pattern analysis with audio transcription; the report does not establish improved criminal-case outcomes from such tools.
Sources and context
- AI turns bodycam footage into text. Can it keep the speakers straight?Phys.org; republished from Northeastern Global News
- AI vs. human transcription: evaluating accuracy and meaning in police body-worn camera footageJournal of Experimental Criminology / Springer Nature
- Racial disparities in automated speech recognitionProceedings of the National Academy of Sciences
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.