Patient De-identification Viewer
Explore a synthetic patient sample and see how many records share the same identifying characteristics.
De-identified record
| Field | Value |
|---|
Shaded fields are the four quasi-identifiers used in this k check. Values are shown as exported from Colab.
Data and method
This educational viewer uses the cleaned 50-record Synthea sample exported from James Seegel's Lab 2 notebook. The records are synthetic. The uploaded sample excludes names, original patient IDs, full dates, full street addresses, full ZIP codes, city, county, and FIPS fields.
For each record, k is the number of records in this sample with the same AGE_BAND, GENDER, RACE, and ZIP3, including the selected record itself. The earlier _k values from the larger Colab table are not reused. If a grouping field is missing, its missing values are grouped together.
Joins and de-identification were performed in the course's Python notebook. This alternative viewer uses HTML and JavaScript for Hugging Face's free Static hosting option; it does not run the notebook's Gradio server. ChatGPT assisted with adapting the viewer and checking its sample group counts.