HomeAILLMs Identifying Vulnerability Systematically Over-Flag UK Police Cases

LLMs Identifying Vulnerability Systematically Over-Flag UK Police Cases

Artificial intelligence can scan thousands of police records in the time it takes a human analyst to read a handful. But a new study published on arXiv by researcher Sam Relins reveals just how carefully that speed needs to be managed — particularly when the records involve some of the most vulnerable people in society. The research examines how well LLMs identifying vulnerability indicators in UK police incident logs actually perform, and the findings are more complicated than either AI optimists or skeptics might expect.

Key takeaways

  • The study analyzed nearly 3,000 de-identified incident logs from a UK police force to estimate the prevalence of four vulnerability indicators.
  • Mental ill health was flagged in approximately one in five incidents — the most common of the four indicators measured.
  • Single-pass LLM classifications were found to be unstable and to systematically over-assign vulnerability indicators compared to human judgment.
  • A locally hosted open-weight LLM was used throughout, reflecting the strict data security requirements of police environments.
  • Population-level estimates are achievable but require significant human review and statistical correction, limiting their practical scalability.

Methodology and Data Sources

The study draws on nearly 3,000 de-identified incident logs from a UK police force — a dataset large enough to produce statistically meaningful patterns, but also one that carries the real-world complexity of frontline police writing. These are not clean survey responses; they are narrative accounts written under pressure, full of shorthand, ambiguity, and inconsistency.

Adapting a US pipeline for UK police data

The classification pipeline at the heart of the research was originally developed using open-source US police data and then adapted for the UK context. That adaptation matters. Policing vocabulary, service structures, and recording conventions differ enough between the two countries that a direct transfer would be unreliable. The study does not claim the adaptation is seamless, and some of the accuracy challenges observed may partly reflect that cross-jurisdictional tension.

Running on a locally hosted model

One of the more practically significant design decisions was the choice to run the entire pipeline on a locally hosted open-weight LLM. This wasn’t an academic preference — it reflects the legal and operational reality that police forces cannot send sensitive incident data to external cloud services. Any AI system that wants to operate inside a policing environment has to work within those constraints, and this study was built around them from the start.

Vulnerability Indicators and LLM Performance

The four indicators the pipeline targets are mental ill health, substance misuse, alcohol dependence, and homelessness. Each represents a dimension of vulnerability that policing increasingly has to account for — not because police are social workers, but because vulnerable people interact with the justice system at disproportionate rates, and understanding that interaction requires data.

Mental health dominates the picture

Mental ill health indicators appeared in approximately one in five incidents — roughly 20% of the dataset. The other three indicators were less frequent, though the study does not give individual breakdowns for substance misuse, alcohol dependence, or homelessness beyond noting their lower prevalence. That one-in-five figure is striking: it suggests that a significant share of routine police work already involves people experiencing mental health difficulties, with major implications for resourcing, training, and multi-agency response planning.

The over-assignment problem

Here is where the technology runs into trouble. When the LLM was given a single pass at classifying each record, its outputs were both unstable and systematically biased. Run the same text through the model twice and you may get different classifications. Aggregate those outputs and the model consistently over-assigns indicators — flagging more cases as involving mental ill health, substance misuse, or homelessness than a human reviewer would. That kind of systematic inflation is not just a minor calibration issue; it would materially distort any policy conclusion drawn from the raw numbers.

This finding reflects a broader challenge with deploying large language models on real-world administrative text. Police logs are not written to be machine-readable. They contain implied context, idiomatic language, and gaps that a human reader fills in automatically but that an LLM may misinterpret or over-interpret. The model’s tendency to over-assign suggests it is picking up on linguistic cues that loosely correlate with vulnerability but don’t confirm it.

Challenges and Limitations of LLM Classification

The study’s methodology addresses the over-assignment problem through a multi-stage process combining repeated model inference, label aggregation, structured human review, and statistical correction. In practice, this means running the model multiple times, comparing outputs, and then having human reviewers assess cases where the model was inconsistent or where aggregated scores were borderline. Statistical adjustment was then applied to bring the final prevalence estimates closer to what human judgment would produce.

The human review burden

That process works — but it is expensive. Relins notes that correcting the biases in raw LLM output required substantial human input and statistical adjustment, and that even after correction, considerable uncertainty remained. For a research project, that’s an acceptable trade-off. For an operational policing environment with constrained analytical resources, it raises serious questions about whether the benefits justify the investment at scale.

Not suitable for individual decisions

The study is explicit on one critical point: LLM outputs cannot be used as valid measurements for individual operational decisions. Errors at the record level remain frequent and unpredictable. A system that is wrong about an individual’s mental health status or housing situation in an operationally consequential context is not just inaccurate — it could lead to harmful outcomes for the people it misclassifies. The pipeline is designed and validated for aggregate, population-level analysis only.

Resource-intensive path to defensible estimates

At the population level, the study concludes that defensible prevalence estimates are achievable — but only with the full methodological apparatus in place. Strip out the human review layer or skip the statistical correction, and the outputs revert to the inflated, unstable classifications of naive deployment. That ceiling on easy automation is one of the paper’s most practically important findings.

What This Means for Policing Practice

The broader ambition behind this research is worth stating clearly. If police forces could reliably estimate how often their officers encounter people experiencing mental ill health, homelessness, or substance dependence, that data could inform everything from officer training programmes to multi-agency referral pathways and budget allocations. Administrative data — the incident logs officers file routinely — could become a source of strategic insight rather than a passive archive.

The study shows that machine learning classification can move that ambition from theoretical to achievable, but not cheaply and not without human oversight baked into the process. The pipeline Relins developed demonstrates a viable architecture: locally hosted for security, multi-pass for stability, human-reviewed for accuracy. What it does not provide is a shortcut. The implication for any police force considering similar tools is that the investment required — in model setup, human review capacity, and statistical expertise — needs to be weighed honestly against the insights generated.

There is also a question of what happens as these tools mature. The current limitations around single-pass instability and systematic over-assignment are partly model-specific and may improve as open-weight LLMs become more capable and better calibrated on administrative text. But the fundamental challenge of applying probabilistic AI outputs to decisions about individual vulnerable people is unlikely to disappear with the next model generation. That tension — between aggregate utility and individual reliability — will remain the defining constraint on LLMs identifying vulnerability in operational policing contexts for the foreseeable future.

FAQ

What vulnerability indicators were targeted in the study of UK police incident logs?

The study targeted four indicators: mental ill health, substance misuse, alcohol dependence, and homelessness. Mental ill health was the most prevalent, appearing in approximately one in five incidents analyzed.

Can fine-tuned LLMs reliably classify vulnerability indicators in UK police data for individual cases?

No. Single-pass LLM classifications were found to be unstable and tend to over-assign indicators relative to human judgment, resulting in frequent errors that make them unsuitable for individual operational decisions.

How does the classification pipeline ensure data security when processing UK police incident logs?

The pipeline operates on a locally hosted open-weight LLM, which means incident data is never sent to external cloud services — a requirement for compliance with secure police data environment standards.

Are population-level estimates of vulnerability feasible using LLMs?

Yes, but only with significant methodological support. Defensible population-level estimates are achievable when the pipeline incorporates repeated model inference, structured human review, and statistical correction — a process the study describes as resource-intensive.

Article produced with the assistance of artificial intelligence and reviewed by the editorial team.

Francesco Antonio Russo
Web 3.0 entrepreneur for over 4 years, expert in Cryptocurrencies and Artificial Intelligence. He uses his cross-functional skills for functional and trend-following Social Media Management.
RELATED ARTICLES

Stay updated on all the news about cryptocurrencies and the entire world of blockchain.

Featured video

LATEST