RockfishRockfish · synthetic data Real EHR to synthetic replica

Model-based synthesis of a patient record

A synthetic patient record you can hand over.

Point Rockfish at a real electronic health record and get back a synthetic one: same tables, same relationships, statistically faithful, ready to build on, with no real patient inside. One model learns the whole record; the identifiers are masked before it ever sees them.

12,352
synthetic patients, a full record each
99%
of real-data accuracy, on a model trained only on synthetic
0
synthetic rows that copy a real patient
0.515
membership-inference attack, no better than a coin flip

Why this is hard to get

a patient record is radioactive

Real patient records live under HIPAA. You can't drop them into a test environment, hand them to a vendor, or give them to a model team without an IRB, a data-use agreement, and months of review, and even a de-identified extract can still leak. So the teams who most need realistic health data, to build readmission and sepsis models, test pipelines, and run analytics, usually can't get it. A synthetic replica carries no real patient and is shareable on day one.

What Rockfish does

mask, learn the whole record, generate every table

Rockfish masks the identifiers, learns the patient record as one model, and samples brand-new patients.

01

Real EHR

Patients and their linked encounters, conditions, medications, and labs.

→
02

Mask PHI

Identifiers detected and replaced; they never enter the model.

→
03

Train

One Rockfish model learns each patient as a whole, all their tables jointly.

→
04

Generate

Sample new patients, each a coherent clinical history.

→
05

Synthetic EHR

The same five tables, foreign keys intact, no real patient.

The data model

one patient, many linked tables

A real EHR is not one table. Each patient has a stream of hospital encounters, and every condition, medication, and observation (lab and vital) hangs off the encounter where it happened. Rockfish generates this clinical history and reuses the shared reference data, the hospitals, providers, and medical code catalogs, unchanged. So every code and provider a synthetic patient points to already exists, and every reference resolves.

GroupTablesHow
Modeledpatients, encounters, conditions, medications, observationsGenerated by the model; referential integrity guaranteed
Carried throughcode catalogs (SNOMED / RxNorm / LOINC), organizations, providers, payersReused unchanged; codes and lab values sampled at their real-world frequencies

Built on real standards

the healthcare data everyone recognizes

The record is Synthea, MITRE's open patient generator, in plain CSV, the same shape as a real hospital export: demographics with the full set of identifiers, plus diagnoses (SNOMED), prescriptions (RxNorm), observations (LOINC), and encounters over time. Specifically we use its COVID-19 sample set (10k_synthea_covid19_csv, 12,352 patients): a realistic longitudinal cohort with a strong clinical signal to model against, the pandemic's testing, diagnoses, hospitalizations, and readmissions running through each patient's history. It stands in for a customer's real data; in production you point Rockfish at your own.

Personal data is detected and masked

detected, masked, and never fed to the model

Rockfish scans every column of the patients table and flags the direct identifiers on its own, with no column list handed to it: SSN, driver's license, passport, names, birth and death dates, street address, ZIP, and geo-coordinates. Each is replaced on the source with a realistic fake of the same type, and none of them are ever fed to the model, which learns only from the demographics and clinical behavior. So the synthetic patients carry no identifier that traces to a real person: no name, no SSN, no address.

Each identifier is masked by the strategy its type calls for, and the same value always maps to the same fake, so joins across tables still resolve. Shown for one real patient:

FieldDetected asReal valueMasked / mapped to
FIRST, LASTPERSONJayson FadelMegan Richard (fake)
SSNUS_SSN999-27-3385999XXXXXXXX (masked)
DRIVERSUS_DRIVER_LICENSES99971451V45139155 (fake)
PASSPORTUS_PASSPORTX53218815XW32815055 (fake)
BIRTHDATE, DEATHDATEDATE_TIME1992-06-301992-04-24 (shifted, distribution kept)
ADDRESSSTREET_ADDRESS1056 Harris Lane15213 Thomas Squares (fake)
ZIPUS_ZIP01001XXXXX (masked)
LAT, LONGEO_COORDINATE42.18, -72.6140.66, -93.09 (fake)

Faithful to the real record

the distributions a health team would expect, and every key resolves

The synthetic record reproduces the real one field by field and, harder and more telling, in the structure across fields. Overall distribution fidelity is 0.96, and the cross-field structure holds: numeric correlations match at 0.99, so a synthetic diabetic looks diabetic across conditions, drugs, and labs. The aggregate numbers a health-data team actually queries match too, real vs. synthetic:

MeasureRealSynthetic
30-day readmission rate3.6%3.4%
Hypoxemia prevalence15.1%15.6%
Encounters per patient21.621.9
Conditions per patient9.39.7
Medications per patient26.426.6
Inpatient share of encounters3.6%3.2%
Real Synthetic

Patients by birth decade

The population's age structure is reproduced across every decade.

Encounters per patient

Light and heavy utilizers alike, with the full spread preserved.

Most common diagnoses, share of patients

The COVID clinical profile, real vs. synthetic, lines up across the top diagnoses.

And the structure holds exactly. Every synthetic condition, medication, and observation hangs off a generated encounter, and every encounter off a generated patient. Referential integrity is guaranteed by construction:

0
orphaned rows: nothing points to a patient or encounter that doesn't exist
1.39M
observation rows, each tied to a generated encounter
12,352
patients, every one a complete, linked record

Useful to build on, safe to share

proven useful, proven private

Fidelity is the means; the value is what a team can build. Take a 30-day readmission model that flags high-risk patients before discharge. We built exactly that model on the synthetic record, and put it to the two tests that decide its worth: it has to work on real patients, and it must not expose one.

01 · Train on synthetic, test on real

A model trained only on synthetic patients works on real ones.

We trained a readmission model on synthetic patients alone, then tested it on real patients it had never seen. It scored 0.951, against 0.961 for the same model trained on the real data.

99%as accurate as training
on the real record

Readmission AUC, tested on real patients

Trained on real data0.961
Trained on synthetic0.951

02 · Find a rare group

Synthetic data trains a model the real data alone cannot.

Readmitters are rare. With only 5 real ones to learn from, a model is effectively blind. We added synthetic readmitters to the training set, retrained, and tested on the same real patients: a working model where there was none.

67%of at-risk patients caught,
that the real-only model missed

Real readmitters the model caught

5 real cases only0.5%
With synthetic added67%

03 · Membership-inference attack

No one can tell which patients the synthetic data was built from.

We gave an attacker the synthetic data and one real patient's full record, and asked whether that patient was used to train the model. Patients who were look no closer to the synthetic data than patients who were not.

0.515membership-inference attack,
no better than a coin flip

Membership attack accuracy

0.5 · random1.0 · fully exposed

The attack comes up empty because the model never copies a real patient in the first place, and the whole batch confirms it:

0memorization

No synthetic patient is a copy

Not one of the 12,352 synthetic records reproduces a real patient. Every synthetic patient is newly generated.


The unlock

The model builders, analysts, and outside partners who could never touch the real record can build on this one on day one. They ship a readmission model that works on real patients, learns the rare cases the real data was too thin to teach, and carries nothing that traces back to anyone.

Safe to share, and ready to build on.
The pair healthcare has never been able to have at once.

the next step

See it on your EHR

This is one worked example on a public Synthea COVID-19 cohort. Rockfish learns your own electronic health record and generates a synthetic replica that keeps the tables and relationships intact and carries no real patient. To see it on your data, get in touch with the Rockfish team.

Try Rockfish →