Guide for doctors

De-identifying a clinical case before you share it

A field guide for removing direct and indirect identifiers while preserving the chronology and reasoning that make a case educational.

By SameCase Editorial9 min read

Think beyond names

Removing a name is necessary but not sufficient. A phone number, email, record number, exact address, facility, clinician name, precise date, image label, device identifier, or link can point directly to a person. Indirect details can also combine into a recognizable story: an unusual occupation, a rare event in a small town, exact age, exact admission dates, and a distinctive family history may identify someone even when each detail seems harmless by itself.

The right question is not only “did I remove the obvious fields?” but “could an intended reader connect the remaining story with information reasonably available elsewhere?” United States HHS guidance describes Safe Harbor and Expert Determination as two formal routes under HIPAA, and explicitly notes that residual risk is not literally zero. Other countries and institutions use different legal frameworks. Follow the law, consent rules, and governance requirements that apply to your work; a platform checklist is not legal clearance.

Inventory every place identifiers hide

Review the title, narrative, pasted reports, comments, filenames, image pixels, image metadata, and any quoted correspondence. Free text deserves special attention because identifiers may be embedded in a sentence that looks clinical. Headers and footers from laboratory or imaging reports often contain names, dates, accession numbers, barcodes, or facility details. Cropping a visible label may not remove metadata stored inside an image file.

Do not include names of relatives, employers, household members, or staff when they make the patient traceable. Avoid exact appointment, admission, procedure, discharge, and birth dates unless a properly governed method permits them. Exact geography below a suitably broad region and very specific workplace or event descriptions often add risk without improving the teaching point.

  • Direct identifiers: names, contact details, addresses, account and record numbers, URLs, and device identifiers.
  • Narrative identifiers: exact dates, exact age, rare occupation, named facility, public incident, or distinctive family detail.
  • Media identifiers: faces, tattoos, room labels, wristbands, scans with burned-in text, filenames, and embedded metadata.
  • Cross-case identifiers: details that become revealing when several posts by the same author are read together.

Generalize without breaking the clinical logic

A useful case needs sequence, not calendar precision. Replace exact dates with relative intervals such as “three weeks after,” “on the second hospital day,” or “over six months.” Use age bands when exact age is not essential, broaden geography, and remove unnecessary occupational or family details. If an uncommon exposure is the teaching clue, describe the category at the least specific level that preserves its relevance.

Do not change a fact in a way that creates false medicine. Shifting every date by an arbitrary amount can distort intervals; swapping sex, anatomy, dose, or laboratory values can change the differential. Prefer suppression, categorization, ranges, and relative time. If a rare combination is itself identifying, the responsible choice may be not to publish, or to create a clearly labeled synthetic teaching example rather than imply that an altered record is a real case.

Separate drafting from privacy review

Write the educational narrative, then review it again with only privacy in mind. A second person or institutional privacy process may catch details the author overlooks because they already know the patient. Read the case as a local colleague, a family member, and a search engine might read it. Ask whether a phrase copied into a web search could lead back to a news report, fundraising page, conference slide, or social post.

Automated scanners are useful backstops, not proof of de-identification. They can flag common patterns but miss context and unusual identifiers. SameCase runs an identifier scan before submission and a server-side scrub, yet contributors remain responsible for excluding identifying material. If identification risk cannot be reduced without losing the case’s meaning, do not submit it.

Use a final pre-publication check

Confirm that the case contains only what is needed to understand the presentation, reasoning, workup, resolution, and learning point. Open every image at full resolution, rename files generically, and remove metadata through an approved workflow. Recheck the title, because a rare diagnosis plus location or occupation can be more revealing than the body. Make sure consent or institutional approval has been handled where required.

After publication, treat privacy reports as urgent. A case that may identify a patient should be hidden while it is reviewed, not left visible while the author debates whether the clue is decisive. De-identification is a continuing responsibility because outside information and re-identification techniques change over time.

Sources and further reading

Keep learning