leeky.sh

Data & Fraud ·

I Vibecoded a Medicare Fraud Detector in a Weekend to See How Hard It Actually Is

2025 and 2026 turned government fraud into a headline argument with numbers ranging from tens of millions to nine billion for the same programs. The data behind it is public, so I built a screening system over 47 million rows to find out what it can actually show.

DuckDBFastAPIfraud detectionopen datavibecoding

For about eighteen months, finding fraud in federal health programs has been an argument conducted almost entirely in press releases.

A federal prosecutor said fraud in Minnesota-run Medicaid services likely exceeds $9 billion. The state’s own human services department said it had seen tens of millions. That is a gap of three orders of magnitude, about the same programs, in the same month.

I wanted to know what the underlying data actually supports. Not to referee the argument, but because the datasets everyone is arguing about are published, downloadable, and free. So I spent a weekend building a screening system over them, mostly by prompting a model and correcting it, which is a fair description of what people now call vibecoding. The question was narrow: how hard is it, in 2026, for one person to build something that surfaces genuine anomalies in Medicare claims data?

The answer turned out to be interesting in both directions.

The numbers everyone is shouting are not the same number

Before any code, the definitional problem, because almost every public claim in this space trips over it.

Improper payments are not fraud. CMS estimated improper payments in its FY2025 fact sheet at $28.83B in Medicare Fee-for-Service (6.55%), $23.67B in Medicare Part C (6.09%), $4.23B in Part D (4.00%), $37.39B in Medicaid (6.12%), and $1.37B in CHIP (7.05%). Roughly $95.5 billion across those five programs. An improper payment is one that should not have been made or was made in the wrong amount, and the category is dominated by documentation and eligibility errors. A missing signature on a form produces an improper payment. So does a fabricated patient. The estimate does not separate them, and anyone citing $95 billion as stolen money is misreading the source document.

Intended loss is not actual loss. The June 2025 National Health Care Fraud Takedown was the largest in US history, charging 324 defendants including 96 doctors, nurse practitioners, and other licensed professionals across 50 federal districts, in schemes involving over $14.6 billion in intended loss. Intended loss means fraudulently billed, not paid out. CMS reported preventing over $4 billion in fraudulent payments in connection with the takedown, and law enforcement seized over $245 million in assets. The seizure figure is the only hard dollar number in that set. These are also charges rather than convictions.

The famous percentages are not measurements. The National Health Care Anti-Fraud Association, a trade association of insurers and investigative agencies, says on an undated page that a conservative estimate of fraud losses is 3% of total health expenditures, with some agencies placing it as high as 10%. The FBI carried a nearly identical 3-to-10-percent line in its Financial Crimes Report from at least FY2006 through FY2010-11, never sourced it, and has since dropped it. Both are estimates of unknown provenance that have been recycled for two decades.

Once you hold those three distinctions, most of the public argument becomes much easier to read.

What DOGE claimed, and what was verifiable

The Department of Government Efficiency made fraud and waste discovery its public identity, so its track record is directly relevant to the question of what detection actually requires.

Its “wall of receipts” ended at a claimed $215 billion in estimated savings, frozen on doge.gov in January 2026 and left there until the site went dark on 4 July 2026. That was against an original public target of $2 trillion, halved to $1 trillion in January 2025. Roughly 30% of the claimed total was itemized with supporting documentation, and the figure mixes contract cancellations, projected annualized savings, lease terminations, workforce reductions, and previously identified improper payments, categories that cannot be summed without double counting.

Several specific claims were checked against the public federal record and did not survive.

The single largest claimed saving, an $8 billion ICE contract, was actually worth $8 million, an error of three orders of magnitude that news organizations caught by reading the procurement record. Correcting it roughly halved DOGE’s then-itemized savings. DOGE subsequently deleted its five largest entries after reporters documented accounting errors, incorrect assumptions, and outdated data, including one contract it claimed to have cancelled that had already been cancelled in November 2024 under the previous administration. Separately it counted the maximum ceiling value of indefinite-delivery contracts as realized savings, overstating by as much as $1.96 billion.

On USAID specifically, the White House said DOGE found $50 million about to be spent on condoms for Gaza. CNN checked it against USAID’s own obligation data: total USAID condom aid worldwide came to roughly $8 million in FY2023, with none going to the Middle East in FY2021 through FY2023. Musk publicly walked it back, saying “some of the things that I say will be incorrect, and should be corrected.” The fallback explanation, that Gaza Province in Mozambique was meant, does not hold either, since federal data showed no condom shipments there. A second claim, that USAID paid Politico more than $8 million, resolved to about $44,000 in USAID subscriptions to E&E News; the $8.2 million figure was total Politico Pro subscription spending across the entire federal government, purchased by agencies rather than granted to the publication.

The one agency-confirmed dollar figure reporters could corroborate was $1.9 billion at HUD, and procurement law specialists described it as a routine deobligation of already-identified funds rather than a discovery. Jessica Tillipman of George Washington University put it on the record: “Nothing they have identified is, to my knowledge, evidence of fraud or corruption. Fraud and corruption are crimes.”

No criminal prosecutions or law enforcement referrals have been publicly documented as arising directly from DOGE’s work. That is an absence-of-evidence finding rather than proof of a negative, since sealed matters would not appear in the public record. DOGE was disbanded as a centralized entity in late 2025, formally ended around its statutory expiration on 4 July 2026, and OMB confirmed no after-action report would be produced, which means its savings claims will likely never be independently reconciled.

The fraud that was actually proven came from somewhere else

Set against that, the cases that produced convictions.

USAID. In June 2025 a former USAID contracting officer, Roderick Watson, and three corporate executives pleaded guilty to a bribery conspiracy running from 2013 to 2022 that steered at least 14 prime contracts worth over $550 million. Bribes came as cash, laptops, NBA suite tickets, a country club wedding, mortgage down payments, and jobs for relatives. Two companies entered deferred prosecution agreements. The $550 million is the value of contracts touched, not money stolen. As of September 2025 USAID’s Office of Inspector General reported 149 active investigations, 35 accepted for criminal prosecution and 16 as civil fraud matters. The case predates DOGE entirely and came from career investigators at OIG and DOJ.

Feeding Our Future. More than $250 million in federal child nutrition funds stolen, with about $50 million recovered. Over 70 indictments and 66 convictions. Ringleader Aimee Bock was convicted on all counts in March 2025 and sentenced in May 2026 to 500 months, nearly 42 years, and ordered to repay almost $243 million. Prosecutors called it the largest pandemic relief fraud case in the country. This one is fully adjudicated, which makes it the load-bearing fact when anyone asks whether large-scale fraud in these programs is real. It plainly is.

Minnesota autism services. The state’s Early Intensive Developmental and Behavioral Intervention program billed a little over $600,000 in 2018 and more than $400 million by 2025. Roughly a 650-fold increase in seven years. DOJ has charged what it called the largest Medicaid autism fraud case it has ever brought, a roughly $46.6 million scheme, following a first $14 million case charged in September 2025. Those are charges, not convictions.

Housing Stabilization Services. Minnesota terminated the program on 31 October 2025 after its OIG suspended payments to 77 providers over credible fraud allegations. In July 2026 four men pleaded guilty to defrauding it of about $2.2 million by enrolling roughly 350 people using AI-generated fake records. That last detail is worth sitting with: it is the first well-documented Medicaid fraud case where generative AI manufactured the supporting documentation that integrity reviews are built to check.

Notice what found these. Claims data, contract records, and investigators reading them. Not a dashboard, and not a public savings tracker.

The Nick Shirley episode is a case study in why detection is hard

In December 2025 a 23-year-old independent YouTube journalist named Nick Shirley posted a 42-minute video alleging that federally supported child care centers around Minneapolis, many Somali-run, were billing taxpayers while serving no children. It reportedly reached about 140 million views across platforms, a self-reported figure, and was amplified by Elon Musk and Vice President JD Vance, who wrote that Shirley “has done far more useful journalism than any of the winners of the 2024 Pulitzer prizes.”

The specific allegations did not hold up well. Minnesota’s Department of Children, Youth and Families said a state inspector had visited each of the day cares shown in the video within the prior six months and found children present. None of the featured centers had formal fraud allegations against them at the time. Most had other violations covering safety, cleanliness, and staff training, which are not fraud. Fact-checkers found the video conflated the Child Nutrition Program, which was the Feeding Our Future vector, with the Child Care Assistance Program, which is structured so that a phantom day care cannot bill the same way. Shirley disputes the fact-checks.

There is a genuine counterweight. KARE 11 found that the centers featured in the video had collectively received $6.3 million from Feeding Our Future, and one closed after the video aired.

What followed was substantial. HHS paused child care payments to Minnesota, the SBA suspended funding citing $430 million in suspected statewide PPP fraud, the state paused payments across 14 high-risk Medicaid programs, and the governor ordered a third-party audit of DHS Medicaid billing. Shirley testified before a House Judiciary subcommittee in January 2026. Causation is entangled, though: the US Attorney’s statement came eight days before the video posted, and Housing Stabilization Services had already been terminated in October. The video accelerated an investigation that was already running.

The $9 billion figure came from First Assistant US Attorney Joseph Thompson, reasoning that 14 high-risk programs had billed $18 billion since 2018 and that half or more could be fraudulent. Governor Walz called it sensationalism. Minnesota DHS said it had seen tens of millions. Thompson later characterized $9 billion as an early estimate rather than a firm number, and charged Minnesota Medicaid fraud to date totals roughly $90 to $150 million. The derivation was an assumption applied to total program billings, not a finding.

So: a real fraud problem, a viral investigation whose specific claims mostly did not check out, an official estimate built from an assumption, and a state response that was already underway. Every participant was pointing at something real and almost nobody’s number was right.

That is the actual problem. Not that fraud is hidden, but that the gap between “this looks wrong” and “this is provably wrong” is enormous, and public argument lives entirely in the first half.

So I built the boring version

The premise of the weekend was that the second half is a data engineering problem, and that the inputs are already public.

DatasetSourceRows
Medicare Provider Utilizationdata.cms.gov~10M
CMS Open Payments (Sunshine Act)openpaymentsdata.cms.gov~12M
Medicare Part D Prescriberdata.cms.gov~25M
OIG LEIE Exclusion Listoig.hhs.gov~75K

Forty-seven million rows. Everything CMS and OIG publish about who billed what, who paid which physician, who prescribed what, and who is barred from federal health programs entirely. Nobody joins them, and the join is where the signal is.

DuckDB made it a laptop project. Embedded, columnar, file-based, no server and no cluster. Analytical queries over 25 million Part D rows come back in seconds because a columnar layout only touches the columns a query names. Postgres would work with tuning and a server. Spark would work with a cluster, to answer questions about a few gigabytes. Around it: FastAPI and Python, Next.js 16 with React 19, twenty-one tables, fifteen routers.

Five detection algorithms, each targeting a different way of stealing, each producing a 0 to 100 score.

Excluded provider billing, weighted 30%, cross-references the LEIE against active billing. Exact NPI matching misses most of it, because excluded providers do not keep billing under the same identifier, so matching runs four tiers with explicit confidence: exact NPI at 1.0, name plus first initial plus state at 0.8, a soundex phonetic match at 0.5 for transliteration variants, and a zip cluster match at 0.3 for someone operating under a practice name at the same address.

Upcoding, 20%, examines the distribution of evaluation and management codes 99211 through 99215, where 99215 pays substantially more. The trap is the peer group: compare a hospital-based physician to office-based peers and you flag the entire hospital. Peer groups are stratified by specialty and place of service, with a fallback to specialty-only when a stratified group is too small to mean anything.

Impossible days, 20%, compares estimated daily service volume against specialty-specific plausibility limits with a 1.5x multiplier for facility settings, since a surgeon in an operating room legitimately does more procedures than one in a clinic. A second dimension uses beneficiary counts, because billing more services per patient and billing for patients who were never there are different frauds that produce different signatures.

Billing anomalies, 15%, does procedure-level outlier detection per HCPCS code against same-specialty peers, plus services-per-beneficiary intensity and an HHI concentration measure over the provider’s procedure mix. Concentration is the sharper signal. One lucrative code at extreme volume beats high volume spread across a normal portfolio.

Kickback correlation, 15%, is the join that justifies the whole exercise. Open Payments records what manufacturers paid physicians. Part D records what those physicians prescribed. Cross them and you can ask whether a provider prescribes their payers’ drugs well above peer rates. Payments alone prove nothing, since speaker fees and consulting are legal and common. The signal is correlation: payment burden relative to prescribing revenue, brand versus generic preference, and a boost only when payment and prescribing anomalies agree.

The composite applies a weighted sum, then a correlation bonus from 1.10x to 1.50x as two through five algorithms converge on the same provider, a persistence bonus of 1.05x to 1.15x for providers flagged across multiple runs, small additive bumps for analyst watchlisting and active investigation, and separate tracking of score velocity, because a provider climbing from 40 to 75 is more interesting than one parked at 78 for a year. Severity is the maximum across algorithms rather than the average, so one critical finding does not get diluted by four clean ones.

Then the part that makes it usable rather than a demo: provider detail pages, geographic drill-down, a watchlist, dismissal tracking, an audit log, and a case builder that exports a CSV or PDF evidence package. Qui tam actions under the False Claims Act start with exactly that kind of package.

Note the shape of the EIDBI story against this. A program going from $600,000 to $400 million in seven years is a volume anomaly against its own history and against every peer program. That is precisely the class of signal these algorithms are built to surface, and it is visible in billing data without any investigator setting foot anywhere.

What the weekend actually proved

The technology is not the bottleneck, and has not been for a while. One person, public data, an embedded database, and a model doing most of the typing produces a working five-algorithm screening system over 47 million rows in a weekend. Ten years ago this was a funded project with a data warehouse. That change is real and it is the strongest argument for why detection capacity should be far ahead of where it is.

Vibecoding got me to a working system and could not get me to a correct one. The model wrote the ingestion, the SQL, the scoring, and the React without much fight. What it could not do was know that comparing a facility-based physician to office-based peers invalidates the entire upcoding algorithm, or that soundex and exact NPI matches are not the same claim and must not collapse into one boolean. Those are domain judgments, they are where all the precision lives, and every one of them came from me reading the algorithm output and finding it wrong. Peer group definition alone was the single largest accuracy improvement in the project.

A screening tool produces leads, not findings. Everything above outputs “this provider is statistically unusual.” Unusual is where an investigation starts. The distance from there to a charge is subpoenas, records, and interviews, which is exactly the distance the $9 billion estimate skipped and the reason the charged total sits near $90 to $150 million.

Which is the real lesson of the last eighteen months. The loudest fraud claims came from people extrapolating from totals, and they mostly did not survive contact with the underlying records. The convictions came from people reading claims data and contracts carefully and slowly. The tooling to do the first part of that job well is now nearly free. The second part, the part that turns an anomaly into a case, has not gotten any cheaper at all.

On sourcing. Figures here are attributed to whoever published them. Where a claim is contested I have said so and given both the claim and the correction, and where something is an estimate rather than a measurement I have labeled it. Charges are not convictions, intended loss is not actual loss, and improper payments are not fraud.