Not legal advice; educational only. These are public-records research sources; their output is leads that require records-level review before anything is called a finding.
Every recovery in this section, the cardiac-device case, the PPP (Paycheck Protection Program) affiliation suits, the Medicare outlier screen, started in the same place: a free, official, public dataset. There are no credentials, no vendor platform, and no FOIA (Freedom of Information Act) request required to reach most of them. The limiting factor is not access. It is knowing which file carries which signal, and how they join.
This is the reference shelf. Healthcare first, because that is where the dollars are, then federal contracting, then the enforcement layer you validate against.
Healthcare
| Dataset | What it carries | Fraud signal |
|---|---|---|
| CMS (Centers for Medicare & Medicaid Services) Medicare Physician & Other Practitioners | Per-provider, per-procedure (HCPCS, the Healthcare Common Procedure Coding System) volumes, charges, and Medicare-paid amounts | The single richest outlier file, “who bills this code far more than peers” |
| CMS Medicare Part D Prescriber | Per-prescriber drug-level scripts, day-supply, beneficiary counts | Pill-mill, opioid-overprescribing, phantom-prescribing patterns |
| CMS DMEPOS (durable medical equipment) | Supplier-level equipment billing | A classic high-fraud category, braces, catheters, test strips |
| CMS Open Payments (Sunshine Act) | Every pharma/device payment to physicians and teaching hospitals | The kickback / conflict-of-interest signal, cross-reference against prescribing |
| HHS-OIG LEIE (the Department of Health and Human Services (HHS) Office of Inspector General’s List of Excluded Individuals/Entities, the exclusions list) | Everyone barred from federal health programs, with exclusion dates | The labeling key, join by NPI (National Provider Identifier) to flag billing after exclusion |
| NPPES (National Plan and Provider Enumeration System) NPI Registry | Identity for every provider, individual vs. organization, addresses, officials | The join key, and the entity-clustering backbone |
| CMS PECOS (Provider Enrollment, Chain, and Ownership System) enrollment | Who is actively approved to bill, with ownership and reassignment chains | Reveals controllers and reassignment-of-benefits structures |
The two most important fields in the whole shelf are humble: the NPI (the National Provider Identifier) is the key that joins identity, billing, and exclusions together; and an organization’s authorized official is what tells you how many entities one person stands behind. Identity resolves through those two, not through names, common names are dangerous.
Federal contracts and grants
| Dataset | What it carries | Fraud signal |
|---|---|---|
| USAspending.gov (+ free API) | Every federal contract and grant: recipient, agency, amount, dates, modifications | Sole-source clustering, end-of-year spend spikes, split awards under thresholds |
| SAM.gov (System for Award Management), Entity & Exclusions | Registered entities + the federal debarment/exclusion list | Shell-company indicators, shared addresses, banned entities still touching awards |
| SBA (Small Business Administration) PPP / pandemic-relief loan data | Borrower, amount, entity, ownership | The affiliation and eligibility-certification frauds (see the PPP cases) |
The enforcement layer you validate against
A screen is only as honest as the ground truth you test it against. HHS-OIG enforcement actions and DOJ False Claims Act settlements tell you whether your outlier flags actually map to real, charged conduct, the academic standard (the Florida Atlantic University and NBER (National Bureau of Economic Research) work on Medicare-fraud detection) labels models exactly this way, by joining provider data to known exclusions and enforcement outcomes.
What the toolkit can and cannot do
Be precise about the boundary, because overstating it is how a screen becomes a smear. These files are aggregated and run roughly two years behind. They carry no diagnoses, no medical necessity, no operative notes, no modifiers, no beneficiary identity, and no intent. They are superb at producing a ranked, peer-relative, explainable list of outliers, and incapable of proving a knowing false claim. The academic methods are published and the public scaffolding is more mature than most people assume; the gap between “this provider is an outlier” and “this provider knowingly submitted false claims” is the entire ballgame, and no dataset closes it.
Which is the same lesson as everywhere else in this section. The toolkit manufactures leads. Turning a lead into a case, the legal hurdles, the qui tam process, the records and the proof, is the forensic work the data cannot do for you. Start with the screen, and bring the discipline.
How the files join, and what they’ve produced
The toolkit is only as powerful as the joins between the files, and they all hinge on one key: the NPI. Identity (NPPES) joins to billing (the CMS provider files) joins to labels (the LEIE exclusions) joins to conflicts (Open Payments), all on that single provider identifier. The published academic methods (the Florida Atlantic University and NBER work on Medicare-fraud detection) formalize exactly this: label providers by joining the claims data to known exclusions, then score the rest for how closely they resemble the caught ones, an approach the published work reports substantially outperforms random case selection at surfacing likely-fraudulent providers.
And the public scaffolding is not theoretical. Every recovery in this section was built on these free files:
| Built on free public data | Recovery |
|---|---|
| Cardiac-device (ICD, implantable cardioverter defibrillator) nationwide sweep | \$250 million+ |
| PPP affiliation (Sidesolve / Empire Roofing) | \$9 million |
| Vascular-procedure case (Lincoln Analytics) | \$6.73 million |
| FY2025 FCA recoveries, all sources (record) | >\$6.8 billion |
Sources: DOJ FY2025 recoveries; per-case citations in the section’s case studies.
The data is free, the methods are published, and the track record is real, which is precisely why the boundary below matters so much. The toolkit manufactures leads at a scale and quality that did not exist a decade ago; it still cannot manufacture proof.
By Noah Green CPA CFE, for Sheepdog Prosperity Partners. Educational only; not legal advice. All sources are public; code and methods described elsewhere in DD Tech Lab are illustrative.
Primary sources & tools: CMS data · Open Payments · HHS-OIG Exclusions (LEIE) · NPPES · USAspending.gov · SAM.gov · DOJ, False Claims Act · ACFE (Association of Certified Fraud Examiners), Report to the Nations
