Challenge
A regional P&C carrier facing a DOI audit inherited loss run data from multiple prior carriers spanning a decade. The files used inconsistent field definitions, mixed paid losses with total incurred values, trapped critical cause-of-loss codes inside free-text adjuster notes, and contained duplicate claim IDs and mismatched valuation dates.
Approach
Outlierr ingested the raw PDFs, flat files, and mixed schemas into an isolated optical-parsing pipeline. Claim fields were isolated using text-boundary mapping, carrier coding structures were normalized into a single NAIC-aligned template, and reserve values were verified against actual payment histories. The final dataset included an Audit Shield Ledger detailing every duplicate removal and field correction.
Outcome
- 180,000 claim rows structured and deduplicated
- 60-hour turnaround for actuarial-ready export
- Passed carrier audit on first submission
