Model-based synthesis of a patient record
Point Rockfish at a real electronic health record and get back a synthetic one: same tables, same relationships, statistically faithful, ready to build on, with no real patient inside. One model learns the whole record; the identifiers are masked before it ever sees them.
a patient record is radioactive
Real patient records live under HIPAA. You can't drop them into a test environment, hand them to a vendor, or give them to a model team without an IRB, a data-use agreement, and months of review, and even a de-identified extract can still leak. So the teams who most need realistic health data, to build readmission and sepsis models, test pipelines, and run analytics, usually can't get it. A synthetic replica carries no real patient and is shareable on day one.
mask, learn the whole record, generate every table
Rockfish masks the identifiers, learns the patient record as one model, and samples brand-new patients.
Patients and their linked encounters, conditions, medications, and labs.
Identifiers detected and replaced; they never enter the model.
One Rockfish model learns each patient as a whole, all their tables jointly.
Sample new patients, each a coherent clinical history.
The same five tables, foreign keys intact, no real patient.
one patient, many linked tables
A real EHR is not one table. Each patient has a stream of hospital encounters, and every condition, medication, and observation (lab and vital) hangs off the encounter where it happened. Rockfish generates this clinical history and reuses the shared reference data, the hospitals, providers, and medical code catalogs, unchanged. So every code and provider a synthetic patient points to already exists, and every reference resolves.
| Group | Tables | How |
|---|---|---|
| Modeled | patients, encounters, conditions, medications, observations | Generated by the model; referential integrity guaranteed |
| Carried through | code catalogs (SNOMED / RxNorm / LOINC), organizations, providers, payers | Reused unchanged; codes and lab values sampled at their real-world frequencies |
the healthcare data everyone recognizes
The record is Synthea,
MITRE's open patient generator, in plain CSV, the same shape as a real hospital export:
demographics with the full set of identifiers, plus diagnoses (SNOMED), prescriptions (RxNorm),
observations (LOINC), and encounters over time. Specifically we use its
COVID-19 sample set (10k_synthea_covid19_csv, 12,352 patients): a realistic
longitudinal cohort with a strong clinical signal to model against, the pandemic's testing,
diagnoses, hospitalizations, and readmissions running through each patient's history. It stands in
for a customer's real data; in production you point Rockfish at your own.
detected, masked, and never fed to the model
Rockfish scans every column of the patients table and flags the direct identifiers on its own, with no column list handed to it: SSN, driver's license, passport, names, birth and death dates, street address, ZIP, and geo-coordinates. Each is replaced on the source with a realistic fake of the same type, and none of them are ever fed to the model, which learns only from the demographics and clinical behavior. So the synthetic patients carry no identifier that traces to a real person: no name, no SSN, no address.
Each identifier is masked by the strategy its type calls for, and the same value always maps to the same fake, so joins across tables still resolve. Shown for one real patient:
| Field | Detected as | Real value | Masked / mapped to |
|---|---|---|---|
FIRST, LAST | PERSON | Jayson Fadel | Megan Richard (fake) |
SSN | US_SSN | 999-27-3385 | 999XXXXXXXX (masked) |
DRIVERS | US_DRIVER_LICENSE | S99971451 | V45139155 (fake) |
PASSPORT | US_PASSPORT | X53218815X | W32815055 (fake) |
BIRTHDATE, DEATHDATE | DATE_TIME | 1992-06-30 | 1992-04-24 (shifted, distribution kept) |
ADDRESS | STREET_ADDRESS | 1056 Harris Lane | 15213 Thomas Squares (fake) |
ZIP | US_ZIP | 01001 | XXXXX (masked) |
LAT, LON | GEO_COORDINATE | 42.18, -72.61 | 40.66, -93.09 (fake) |
the distributions a health team would expect, and every key resolves
The synthetic record reproduces the real one field by field and, harder and more telling, in the structure across fields. Overall distribution fidelity is 0.96, and the cross-field structure holds: numeric correlations match at 0.99, so a synthetic diabetic looks diabetic across conditions, drugs, and labs. The aggregate numbers a health-data team actually queries match too, real vs. synthetic:
| Measure | Real | Synthetic |
|---|---|---|
| 30-day readmission rate | 3.6% | 3.4% |
| Hypoxemia prevalence | 15.1% | 15.6% |
| Encounters per patient | 21.6 | 21.9 |
| Conditions per patient | 9.3 | 9.7 |
| Medications per patient | 26.4 | 26.6 |
| Inpatient share of encounters | 3.6% | 3.2% |
The population's age structure is reproduced across every decade.
Light and heavy utilizers alike, with the full spread preserved.
The COVID clinical profile, real vs. synthetic, lines up across the top diagnoses.
And the structure holds exactly. Every synthetic
condition, medication, and observation hangs off a generated
encounter, and every encounter off a generated patient. Referential integrity is guaranteed by
construction:
proven useful, proven private
Fidelity is the means; the value is what a team can build. Take a 30-day readmission model that flags high-risk patients before discharge. We built exactly that model on the synthetic record, and put it to the two tests that decide its worth: it has to work on real patients, and it must not expose one.
01 · Train on synthetic, test on real
We trained a readmission model on synthetic patients alone, then tested it on real patients it had never seen. It scored 0.951, against 0.961 for the same model trained on the real data.
99%as accurate as training
on the real record
Readmission AUC, tested on real patients
02 · Find a rare group
Readmitters are rare. With only 5 real ones to learn from, a model is effectively blind. We added synthetic readmitters to the training set, retrained, and tested on the same real patients: a working model where there was none.
67%of at-risk patients caught,
that the real-only model missed
Real readmitters the model caught
03 · Membership-inference attack
We gave an attacker the synthetic data and one real patient's full record, and asked whether that patient was used to train the model. Patients who were look no closer to the synthetic data than patients who were not.
0.515membership-inference attack,
no better than a coin flip
Membership attack accuracy
The attack comes up empty because the model never copies a real patient in the first place, and the whole batch confirms it:
Not one of the 12,352 synthetic records reproduces a real patient. Every synthetic patient is newly generated.
The unlock
The model builders, analysts, and outside partners who could never touch the real record can build on this one on day one. They ship a readmission model that works on real patients, learns the rare cases the real data was too thin to teach, and carries nothing that traces back to anyone.
Safe to share, and ready to build on.
The pair healthcare has never
been able to have at once.
the next step
This is one worked example on a public Synthea COVID-19 cohort. Rockfish learns your own electronic health record and generates a synthetic replica that keeps the tables and relationships intact and carries no real patient. To see it on your data, get in touch with the Rockfish team.