A Lab Director's $100,000 Mistake Over a Spreadsheet Column

A lab director at a mid-size hospital once told me she'd been sharing de-identified data with a research partner for two years. She was confident she'd stripped out every identifier. She hadn't. Buried in one spreadsheet column were admission dates — one of the 18 identifiers under the HIPAA Privacy Rule. That single oversight triggered an OCR investigation and months of corrective action.

The question of which is not considered an identifier under the Privacy Rule comes up constantly in workforce training sessions. And the answer trips people up more than you'd expect. Because the list of what is an identifier is precise, but the list of what isn't lives in the negative space — and that's where organizations get into trouble.

If you handle protected health information (PHI) in any capacity, you need to know both sides of this equation cold. Here's the breakdown I walk through with every covered entity I consult for.

The 18 Identifiers the Privacy Rule Actually Lists

Before you can understand what's excluded, you have to know what's included. The HIPAA Privacy Rule, codified at 45 CFR §164.514(b)(2), defines exactly 18 types of identifiers that must be removed for health information to be considered de-identified under the Safe Harbor method.

Here they are, no fluff:

  • Names
  • Geographic data smaller than a state (street address, city, county, zip code, equivalent geocodes)
  • All dates directly related to an individual (birth date, admission date, discharge date, death date) — and all ages over 89
  • Telephone numbers
  • Fax numbers
  • Email addresses
  • Social Security numbers
  • Medical record numbers
  • Health plan beneficiary numbers
  • Account numbers
  • Certificate/license numbers
  • Vehicle identifiers and serial numbers (including license plates)
  • Device identifiers and serial numbers
  • Web URLs
  • IP addresses
  • Biometric identifiers (fingerprints, voiceprints)
  • Full-face photographs and comparable images
  • Any other unique identifying number, characteristic, or code

That 18th category is the catch-all. It's deliberately broad. HHS designed it to close loopholes.

Why That 18th Identifier Catches People Off Guard

I've reviewed breach reports where organizations stripped out the first 17 identifiers and left in a proprietary patient tracking code. They assumed it wasn't an identifier because it wasn't a name or SSN. Wrong. That code could be linked back to a specific individual, which makes it an identifier under the Privacy Rule's Safe Harbor standard.

The 18th identifier means any code, characteristic, or number that could identify an individual counts — unless it's specifically assigned by the covered entity for re-identification purposes under a documented system that meets 45 CFR §164.514(c).

Which Is Not Considered an Identifier Under the Privacy Rule: The Direct Answer

Here's the short answer that captures what most people search for: aggregate statistical data, general diagnoses, treatment codes, lab values, vital signs, and clinical observations are not considered identifiers under the Privacy Rule.

A blood pressure reading of 140/90 is not an identifier. A diagnosis of Type 2 diabetes is not an identifier. A cholesterol level of 220 mg/dL is not an identifier. An aggregate statistic — like "34% of patients in this cohort had hypertension" — is not an identifier.

These are clinical data elements. They describe health conditions or treatment, but they don't point to a specific person on their own. That's the critical distinction.

The Line Between Clinical Data and PHI

Here's where it gets nuanced. A diagnosis code by itself isn't an identifier. But combine that diagnosis code with a zip code and a date of birth, and you've created a dataset that could identify someone. The Privacy Rule doesn't just protect individual identifiers — it protects combinations that could lead to re-identification.

This is why workforce training matters so much. Your staff might understand that names and SSNs are protected. But do they understand that a discharge date combined with a five-digit zip code and a rare diagnosis could identify a patient in a small community? In my experience, most don't — until they're trained on it specifically.

Our HIPAA training catalog covers this exact scenario in detail, walking your team through real-world de-identification exercises.

The $3 Million Problem With Getting De-Identification Wrong

In 2023, OCR settled with Yakima Valley Memorial Hospital for $240,000 after 23 security guards accessed patient ePHI without authorization. While that case centered on access controls rather than de-identification failures, it underscores a broader truth: OCR investigates when organizations demonstrate systemic gaps in understanding what constitutes PHI.

De-identification failures tend to surface during breach investigations. OCR examines whether the data involved could identify individuals. If your team incorrectly classified identifiable data as de-identified, that's not just a breach — it's evidence of inadequate training and policies.

What OCR Actually Examines in De-Identification Cases

OCR looks at two things when evaluating whether data was properly de-identified:

  • Safe Harbor Method: Were all 18 identifiers removed? Does the covered entity have no actual knowledge the remaining data could identify someone?
  • Expert Determination Method: Did a qualified statistical expert apply accepted methods and certify the risk of identification is "very small"?

Most small and mid-size organizations use Safe Harbor because hiring a statistical expert is expensive and complex. That makes knowing the 18 identifiers — and what falls outside them — operationally critical.

Common Items People Wrongly Assume Are Identifiers

Let me save you some confusion. These are not identifiers under the Privacy Rule's Safe Harbor list:

  • Gender (male, female, non-binary)
  • Race or ethnicity categories
  • General age ranges (e.g., "patient in their 40s") — as long as specific dates and ages over 89 are removed
  • Diagnosis or procedure codes (ICD-10, CPT)
  • Lab results and vital signs
  • Medication names and dosages
  • Smoking status or other behavioral health indicators in aggregate

None of these, standing alone, can identify a specific individual. They become problematic only when paired with actual identifiers or when the dataset is small enough that combinations create re-identification risk.

Why Your Workforce Gets This Wrong (and How to Fix It)

I've conducted HIPAA training for organizations ranging from 12-person dental practices to 4,000-employee health systems. The pattern is always the same: staff over-classify or under-classify data as identifiable.

Over-classifiers refuse to share any health data for research, quality improvement, or public health — even when it's properly de-identified. This creates operational bottlenecks and slows down legitimate work.

Under-classifiers share data they think is de-identified when it isn't. They strip names and SSNs but leave in dates, zip codes, or unique account numbers. This creates breach risk.

The Fix: Scenario-Based Training That Sticks

Abstract lectures about the 18 identifiers don't change behavior. What works is putting your staff in front of realistic scenarios: "Is this dataset de-identified? What's missing? What would you remove?"

That's exactly the approach we take in our HIPAA workforce training programs. Each module presents actual data scenarios — not just definitions — so your team builds the judgment to handle these decisions in real time.

The Expert Determination Alternative

If Safe Harbor feels too restrictive for your use case — especially for research or population health analytics — the Expert Determination method under 45 CFR §164.514(b)(1) offers more flexibility.

Under this method, a person with appropriate knowledge of statistical and scientific principles applies accepted analytical techniques and determines that the risk of identifying any individual is "very small." The expert must document their methods and results.

HHS published detailed guidance on de-identification methods that walks through both approaches: HHS Guidance on De-Identification of PHI. I recommend every privacy officer read it at least once a year.

Building a De-Identification Checklist Your Team Will Actually Use

Here's what I recommend to every covered entity and business associate I work with:

  • Create a written de-identification policy that specifies which method (Safe Harbor or Expert Determination) your organization uses
  • Build a checklist of the 18 identifiers and require staff to verify removal before any data leaves your systems
  • Assign a de-identification review role — one person or team responsible for final sign-off on any dataset shared externally
  • Train annually on what qualifies as an identifier and what doesn't, using scenario-based exercises
  • Document everything — OCR wants to see that you had a process, followed it, and can prove it

Your organization doesn't need to be perfect. It needs to be systematic. OCR distinguishes between organizations that made a good-faith effort and those that never built a process at all. The penalties reflect that distinction.

What This Means for Your Organization in 2026

Data sharing is accelerating. Interoperability rules, research partnerships, AI-driven analytics — all of these require moving health data between systems and organizations. Every one of those data flows demands a clear answer to this question: is this data de-identified or not?

Knowing which is not considered an identifier under the Privacy Rule isn't academic. It's the operational foundation for every data-sharing decision your organization makes. Get it wrong, and you're looking at breach notifications, OCR investigations, and penalties that can reach into the millions.

Get it right, and you unlock the ability to share data for legitimate purposes — research, quality improvement, public health — without exposing your patients or your organization.

Start with the fundamentals. Explore our complete HIPAA training catalog to make sure every member of your workforce understands the line between identifiable and de-identified data.