A Single Data Point Tripped a $1.5 Million Settlement
In 2018, a health plan called Anthem paid $16 million to the Office for Civil Rights after a data breach exposed names, Social Security numbers, and medical IDs of nearly 79 million people. Every single exposed element was a recognized PHI identifier under the HIPAA Privacy Rule. But here's the question that trips up even experienced compliance officers: which is not considered a PHI identifier under HIPAA?
I get this question constantly during workforce training sessions. Clinicians, billing staff, and IT teams all assume they know the answer — and most of them get it wrong. They confuse clinical data, general demographics, and actual identifiers, which creates risk your organization can't afford.
This post breaks down the 18 PHI identifiers defined by HHS, explains what does not make the list, and gives you a framework to apply this knowledge every day. If you handle patient data in any capacity, this is foundational knowledge.
The 18 PHI Identifiers: The Exact List HHS Publishes
The HIPAA Privacy Rule at 45 CFR §164.514(b)(2) defines exactly 18 types of identifiers that, when linked to health information, make that data protected health information (PHI). Here they are:
- Names
- Geographic data smaller than a state (street address, city, ZIP code, etc.)
- All dates directly related to an individual (birth date, admission date, discharge date, date of death) — and all ages over 89
- Telephone numbers
- Fax numbers
- Email addresses
- Social Security numbers
- Medical record numbers
- Health plan beneficiary numbers
- Account numbers
- Certificate/license numbers
- Vehicle identifiers and serial numbers (including license plates)
- Device identifiers and serial numbers
- Web URLs
- IP addresses
- Biometric identifiers (fingerprints, voiceprints)
- Full-face photographs and comparable images
- Any other unique identifying number, characteristic, or code
That last one is the catch-all. HHS intentionally left it open-ended to prevent organizations from finding creative loopholes.
So Which Is Not Considered a PHI Identifier Under HIPAA?
Here's the direct answer: a diagnosis code, treatment information, lab result, vital sign, or general demographic category like gender, race, or ethnicity is not considered a PHI identifier under HIPAA. These are health information or demographic attributes — but they are not identifiers by themselves.
This surprises people. A blood pressure reading of 180/110 is clinical data. A diagnosis of Type 2 diabetes is clinical data. Neither one, standing alone, can identify a specific individual. They only become PHI when they're linked to one or more of the 18 identifiers listed above.
Similarly, a patient's age (if under 90) presented without any other identifier is not PHI. Saying "a 42-year-old male was treated for a fracture" involves health information, but without a name, ZIP code, date, or other identifier, it doesn't qualify as PHI under the Privacy Rule's de-identification standard.
Common Exam Answers That Fool People
If you've taken a HIPAA certification exam or seen this as a test question, here are the typical answer choices and why they trip people up:
- Diagnosis codes (ICD-10): Not an identifier. Clinical data only.
- Social Security numbers: Definitely an identifier. Number 7 on the list.
- Telephone numbers: Identifier. Number 4.
- Email addresses: Identifier. Number 6.
The correct answer is always the clinical or demographic element that cannot, by itself, point to a specific person. Diagnosis codes are the textbook example.
Why This Distinction Matters for Your Covered Entity
Understanding which data elements are identifiers and which aren't has direct operational consequences. Here are three I see constantly in the field.
1. De-Identification Projects Go Sideways
If your organization shares data for research, quality improvement, or analytics, you need to de-identify it properly. The HHS guidance on de-identification describes two methods: the Safe Harbor method (strip all 18 identifiers) and the Expert Determination method (a statistician certifies re-identification risk is very small).
I've seen organizations strip names and Social Security numbers but leave ZIP codes and dates of service intact — then claim the data is de-identified. It isn't. Under Safe Harbor, you must remove or generalize all 18 identifier types. Miss one, and your "de-identified" dataset is still PHI, still subject to the Privacy Rule, and still a breach waiting to happen.
2. Workforce Training Gets the Basics Wrong
When your staff can't distinguish identifiers from clinical data, they either over-restrict information sharing (hurting patient care) or under-protect it (creating breach risk). Both outcomes cost money and trust.
I worked with a hospital system where nurses were redacting diagnosis codes from interdepartmental referrals, thinking they were "protecting PHI." They were actually delaying treatment. The data they should have been protecting — patient names and medical record numbers traveling in unencrypted emails — was sailing through inboxes unchecked.
Good training eliminates this confusion. A program like the HIPAA training course for physicians and clinical environments walks through exactly these distinctions in scenarios that mirror real clinical workflows.
3. Breach Risk Assessments Need Precision
When you conduct a breach risk assessment under the Breach Notification Rule, you evaluate whether PHI was actually compromised. If the exposed data contained only clinical information without any of the 18 identifiers, it may not qualify as a reportable breach. That assessment has to be precise and documented.
OCR doesn't accept vague reasoning. They want to see that your organization knows the 18 identifiers, applied them to the facts, and reached a defensible conclusion.
The $5.55 Million Reminder: When Identifier Confusion Gets Expensive
In 2017, Memorial Healthcare System paid $5.55 million to settle with OCR after employees accessed PHI — including names, Social Security numbers, and dates of birth — of 115,143 individuals without authorization. The settlement highlighted failures in access controls and workforce training. Every compromised element was a recognized identifier.
Had the exposed data contained only aggregate clinical statistics with no identifiers attached, the story would have been very different. The distinction between identifiers and non-identifiers isn't academic. It's the line between a reportable breach and a non-event.
You can review real OCR enforcement actions and resolution agreements on the HHS HIPAA Enforcement page.
A Quick Framework You Can Use Today
When you're evaluating whether a data element is a PHI identifier, run it through this three-part test:
- Does it appear on the list of 18? If yes, it's an identifier. Protect it accordingly.
- Can it reasonably be used to identify an individual — alone or combined with other available data? If yes, treat it as an identifier under the catch-all category (#18).
- Is it purely clinical, statistical, or demographic without any link to a specific person? Then it's health information or a general attribute — not an identifier by itself.
Post this framework in your break rooms, include it in your onboarding materials, and reinforce it annually. It takes five minutes to teach and prevents years of confusion.
What About ePHI? Same Rules, Extra Safeguards
Electronic protected health information (ePHI) follows the same identifier definitions. The difference is that the HIPAA Security Rule adds technical, administrative, and physical safeguards specifically for ePHI. So a diagnosis code stored in an EHR alongside a patient's name and medical record number is ePHI — not because the diagnosis code is an identifier, but because it's health information stored electronically alongside identifiers.
Strip the identifiers properly, and that same diagnosis code in a research database is no longer ePHI. The Security Rule's safeguard requirements no longer apply to it. This is why de-identification isn't just a privacy exercise — it directly affects your security obligations and audit scope.
Train Your Workforce Before OCR Trains Them for You
Every compliance program I've audited that struggled with PHI handling had the same root cause: staff couldn't define what PHI actually is. They knew it was "patient information" but couldn't distinguish an identifier from a lab value.
Fixing this starts with targeted training. Not a generic slide deck, but scenario-based education that forces your team to classify real data elements. The HIPAA training catalog at HIPAACertify includes courses built around exactly these challenges — designed for clinical staff who need practical answers, not legal abstractions.
Your organization's next OCR audit, breach investigation, or research data request will test whether your people actually understand identifiers. Make sure they can answer the question correctly before the stakes are real.
The Bottom Line
Diagnosis codes, lab results, vital signs, and general demographic categories like gender or ethnicity are not PHI identifiers under HIPAA. The 18 identifiers defined by HHS at 45 CFR §164.514(b)(2) are specific data elements that can point to an individual person. Everything else is health information or demographic data — important to protect in context, but not an identifier on its own.
Know the 18. Teach the 18. And build your policies around the distinction. That's where compliant organizations separate themselves from the ones writing seven-figure settlement checks.