TL;DR — Cases that reach a jury are statistical outliers where the parties failed to value the claim. Anchoring settlement negotiations on public verdicts ignores selection bias and post-trial reductions, artificially inflating your financial exposure.
A claims manager requests a valuation on a severe bodily injury file. The analyst pulls recent comparable cases from a legal database. The search returns a cluster of eight-figure verdicts. The reserve is immediately adjusted upward. The defense shifts to a reactive posture. The entire negotiation is now anchored to an illusion.
This workflow is standard practice across the insurance industry. It is also a fundamental mathematical error. Relying on public verdicts to estimate settlement exposure ignores the underlying mechanics of civil litigation. Public verdict data is not a representative sample of case value. It is a record of cases where the normal valuation process failed.
The selection mechanics of a trial
George Priest and Benjamin Klein formalized this dynamic decades ago. The Priest-Klein hypothesis demonstrates that the disputes selected for trial are not drawn randomly from the pool of all claims. If the facts are clear and the damages are easily quantified, the plaintiff and the defense will arrive at similar estimates of the expected trial outcome. They will settle to avoid the deadweight loss of litigation costs. The cases that actually reach a jury are the outliers. They proceed to trial precisely because the two sides hold mutually exclusive, highly divergent expectations about liability or the magnitude of damages.
When you query a database for comparable verdicts, you are intentionally sampling the extreme tail of a distribution. You are looking exclusively at files where at least one party was severely miscalibrated. Using this biased sample to set a reserve on a new, unlitigated claim guarantees an inflated valuation. You are treating the exception as the baseline. The availability heuristic compounds this error. Adjusters and defense counsel remember the nuclear verdicts because they dominate industry media. The quiet, rational settlements disappear into the archives. To build a valid predictive model, you have to triangulate the missing data. You must construct honest uncertainty bands around the sparse settlement data you do have, rather than filling the gaps with the loudest verdicts.
The gap between the verdict and the check
Even if you accept the biased sample of cases that reach a verdict, the numbers reported on the jury form do not represent actual financial exposure. A highly publicized fifty-million-dollar verdict is rarely followed by a fifty-million-dollar wire transfer. The stated verdict is merely a data point in an ongoing negotiation. Post-trial mechanics systematically compress extreme jury awards.
Statutory caps on noneconomic damages automatically truncate the top end of the distribution. Uncollectability and strict policy limits force settlements for fractions of the awarded amount. High-low agreements, struck while the jury is deliberating, establish hard floors and ceilings that render the final verdict largely irrelevant to the actual payout. The asymmetry of the appeals process heavily favors the defense in reducing massive awards. The plaintiff faces years of delay and the risk of a complete reversal. The rational choice is to accept a substantial, immediate reduction on the verdict in exchange for guaranteed payment.
David Hyman and his co-authors documented this extensively in their research on medical malpractice payouts. The gap between the jury award and the actual amount paid is vast, particularly at the upper extremes. A predictive model trained strictly on the raw text of public jury verdicts will learn the wrong target variable. It will predict the theatrical number, not the economic reality. If an insurer allocates capital based on that theatrical number, they lock up funds that should be deployed elsewhere. This drives unnecessary reserve volatility across the entire portfolio.
Separating generation from calibrated prediction
This is why building a forecasting platform requires a strict architectural boundary between reading documents and predicting outcomes. Generative AI is exceptionally useful for the first task. We use large language models to read thousands of pages of pleadings, medical records, and correspondence. The LLM extracts the entities, builds the timeline, and maps the specific injury characteristics into a neural-symbolic structure. It converts unstructured text into a computable state. The LLM does the reading and structuring. It does not make the prediction.
Language models are text generators. They predict the next token in a sequence based on training data distributions. They do not calculate distance in a geometric space, and they have no mathematical concept of calibration. Passing a raw case file to an LLM and asking for a settlement value is a dangerous misuse of the technology. To predict a settlement, we pass the structured facts to separate, mathematical machine-learning models. These geometric models operate on numerical representations of the claim, calculating probabilities across large numbers of resolved cases where the actual financial outcome is known. They map the claim against a true baseline, accounting for the probability of liability and the field-specific priors of the jurisdiction.
Conformal prediction solves a critical problem in litigation forecasting. Point estimates in litigation are mathematically dishonest. If an algorithm outputs a settlement value of exactly four hundred thousand dollars, it is lying about its own precision. The system must produce a calibrated settlement range, defined by conformal prediction bands. This provides a mathematical guarantee that the true value falls within the specified range a given percentage of the time. We quantify the uncertainty. If the medical damages are highly variable or the jurisdiction has sparse data, the confidence interval organically widens.
The platform identifies the probability of escalation, calculates a reserve delta against the current numbers, and isolates the specific drivers behind the valuation traceable directly to the source documents. Crucially, we keep verdict-heavy comparable cases in a distinct module, separated from the settlement-anchored baseline. You can view the worst-case scenario without letting it poison the expected value of the claim.
Valuing a claim requires confronting the reality of the data. Social inflation, rising nuclear verdicts, and third-party litigation funding are pushing demands higher. This makes accurate, day-one reserves critical for allocating defense spend and detecting escalation early. If you build your strategy around the most visible, extreme trial outcomes, you will negotiate against yourself before the plaintiff even makes a demand. The goal of a predictive platform is to calculate the clearing price of a dispute. You cannot find that price by studying the files where the parties refused to clear.
Related articles.
When AI Should Say It Doesn't Know
Language models are built to generate plausible text, not calculate risk. When an AI offers a single dollar figure for a complex claim without quantifying its own doubt, it is not predicting. It is guessing.
Training on Known Outcomes
Large language models are built to talk, not to calculate risk. Relying on them to predict claims outcomes conflates reading comprehension with mathematical forecasting.
Benchmarking Litigation Outcome Prediction
A point prediction for a complex liability claim is mathematically meaningless. True litigation forecasting requires separating the extraction of text from the calculation of risk, delivering calibrated ranges rather than brittle guesses.
Want to talk to an executive?
Press, partners, investors, candidates — the inbox is monitored. Tell us who you are and we'll route it to the right person within two business days.