Blog • Product

Policy Limits in the Comparable Set

Policy limits artificially cap historical settlement data, creating a massive blind spot for predictive models that fail to track collectability.

TL;DR — Mixing limit-constrained settlements with unconstrained claims destroys forecasting accuracy. To output reliable reserve ranges, a forecasting system must read the policy declarations and geometrically isolate censored data from the comparable set.

A traumatic brain injury settles for $50,000. The medical bills alone are five times that amount. To a naive clustering algorithm, this looks like a cheap venue or a weak plaintiff. To a claims adjuster, it is obviously a minimum-limits auto policy. The settlement value was dictated entirely by collectability, not the underlying severity of the injury. Insurance settlement data is heavily censored. Policy limits cap the financial recovery regardless of the actual economic damage. If you train a predictive model on raw settlement numbers without factoring in the policy limits of the historical cases, your system will systematically under-predict severity. It learns the artificial ceiling instead of the true exposure. Building a forecasting platform for litigation requires treating this truncation as a core architectural constraint. You cannot build a reliable reserve model on top of censored historicals.

At Canotera, we separate the reading from the math. The generative AI layer processes the unstructured case file. This is not a trivial ingestion task. We process raw PDFs, faxed medical chronologies, encrypted correspondence, and complex policy declarations. The AI structures this text. It extracts the specific coverage limits, attachments, umbrella layers, and exclusions. It maps the entities and timelines. We execute this extraction before any prediction occurs. The generative models do not forecast the outcome. They simply structure the reality of the file into a machine-readable format. We establish the contractual ceiling before we calculate the floor. This pipeline runs with strict latency constraints, moving the structured payload to the modeling layer rapidly.

Censored Data in the Geometric Space

When an adjuster evaluates a new claim, they need comparable resolved cases to anchor their negotiation. Finding comps based purely on injury type and venue is dangerous. If you pull a historical comp that settled for $250,000, you need to know the context. Did it settle there because $250,000 was the fair value of the claim, or because $250,000 was the absolute limit of the policy? Supplying an adjuster with limit-constrained comps to evaluate an unconstrained commercial claim sets them up for failure. Our mathematical models map these historical cases in a high-dimensional space. We map the severity indicators from the medical records against the liability facts, the venue data, and the plaintiff firm history. A case constrained by collectability occupies a fundamentally different geometric region than an unconstrained settlement. We explicitly segregate limit-capped comps from the unconstrained comparable set unless the new claim shares that exact limit profile. Blending unconstrained settlements with limit-capped settlements creates a mathematically invalid average. When settlement data for a specific injury in a specific venue is sparse, we rely on triangulation. We analyze adjacent venues or similar injury profiles to build the reference class, applying field-specific priors to anchor the distribution.

We output calibrated settlement ranges, not single point guesses. When the underlying historical data is heavily censored by policy limits, the uncertainty bands widen. The math must reflect honest uncertainty. We do not invent certainty where the data is truncated. The API returns the settlement range, the escalation probability, and the specific drivers behind the numbers. This allows the claims team to set realistic reserves on day one and allocate defense spend logically.

This geometric separation intersects directly with selection bias. Claims approaching the policy limit frequently settle early to avoid bad faith exposure. Claims where liability is highly contested go to trial. This dynamic distorts the available public and private data. We isolate verdict-heavy comparisons from settlement-anchored data. Mixing them contaminates the settlement prediction with jury awards that reflect a different risk profile. A jury award is not the final paid amount. We keep the modules strictly separate to prevent appeals asymmetry and post-trial reductions from skewing the settlement baseline.

Traceability as an Engineering Baseline

Engineering a forecasting system for claims requires user trust. A black box that returns a large reserve recommendation is useless to an adjuster who must justify that number to a supervisor. The interface must show the work. Traceability is a baseline engineering requirement. It dictates how we design the entire pipeline, from the extraction logs to the final API response. The payload returned to the client contains the reserve delta, the escalation probability, and the exact historical cases that drove the calculation.

When Canotera displays comparable resolved cases, it surfaces the specific drivers behind the historical settlement. If a comp was limit-constrained, we flag the constraint. The user sees the original policy limit directly alongside the settlement amount. They can click through the interface to view the extracted text from the historical file. Every variable that influenced the mathematical model is traceable back to the source documents. Latency and security shape this architecture. Processing thousands of pages of medical records and policy documents requires significant compute. We optimize the ingestion pipeline to return structured data and calibrated predictions fast enough to fit into the actual workflow of a claims professional. Sensitive records never leak across tenant boundaries. The security model isolates each client's historical data while allowing the mathematical models to learn the geometric relationships between injury severity and settlement value.

You cannot negotiate from a gut feeling. You need data. Bad data is worse than no data. Feeding an adjuster a list of artificially low comparable cases creates a false sense of security. It leads to inadequate reserves on day one. It masks the risk of escalation. When the plaintiff lawyer, backed by third-party litigation funding, finds a deep pocket or an umbrella policy, the resulting nuclear verdict appears to come out of nowhere. Driven by social inflation and reserve volatility, these verdicts do not come out of nowhere. They come from a fundamental failure to account for the constraints on the historical data. The math only works if it respects the contract.

Want to talk to an executive?

Press, partners, investors, candidates — the inbox is monitored. Tell us who you are and we'll route it to the right person within two business days.