Skip to main content
Cory Trimm
Model Apache-2.0 English

phi-redact

Finds patient names, MRNs, dates, phone numbers and 26 other kinds of PHI and PII in English text - on the device, so it never reaches your servers, your logs or an LLM provider.

11 million parameters, 11 MB as int8. Small enough to run in the browser tab you're reading this in, which is exactly what the demo below does.

Redact as you type

Pick a sample, paste your own text, or dictate it. Every sample is invented - but anything you paste or say stays in this tab either way.

Loading model…

How well it works

Strict span-level scoring: a prediction counts only if both the character boundaries and the label match.

0.860

Micro F1 on Nemotron-PII

13,134 entities, strict span match

86.1%

Recall, 5-benchmark average

Desert Ant Redact (EN): 86.5%

11.1M

Parameters

BERT-mini, 4 layers

11.3 MB

int8 ONNX in your browser

~5.6 MB est. as 4-bit Core ML

Against Desert Ant Redact v0.4.0 (English)

Desert Ant figures are their published English results on the same datasets.

phi-redact v7 Redact (EN)
WikiANN name recall 84.3% 69.5%
MultiNERD name recall 94.3% 94.9%
Composite recall (5 benchmarks) 86.1% 86.5%
Parameters 11.13M 23M
Est. 4-bit on-device size ~5.6 MB 11.6 MB
Languages English 27
License Apache-2.0 Source-available

F1 by entity, Nemotron-PII

2,000 test documents. PROFESSION over-fires on job titles in ordinary prose.

  • MRN 0.951
  • ADDRESS 0.949
  • NAME_PATIENT 0.944
  • NAME_FAMILY 0.937
  • LICENSE 0.920
  • COUNTRY 0.919
  • SSN 0.908
  • IDNUM 0.908
  • EMAIL 0.905
  • CREDITCARD 0.905
  • CITY 0.904
  • ROUTING_NUMBER 0.876
  • DATE 0.859
  • URL 0.828
  • DEVICE 0.791
  • PROFESSION 0.494

Where it fits

Server-side redaction means the raw text has already crossed the network to get redacted. Doing it on the device removes that hop.

Before the LLM call

Swap names, MRNs and phone numbers for placeholders like [NAME_PATIENT_1], send the prompt, then put the originals back in the reply. The model provider never sees the patient.

Intake and case forms

Scrub free-text fields in the browser before submit, so the notes box on a benefits or intake form doesn't quietly become the place PHI ends up in your database.

Logs and analytics

Run it at ingest on support transcripts, error reports and session notes. Nothing leaves the machine to be redacted, so there's no second system holding the raw text.

Test data from real data

Mask a production sample into something a developer can safely use to reproduce a bug - consistent placeholders keep the same person the same person across a record.

30 entity types

Grouped the way the demo colours them.

Names

NAME_PATIENTNAME_FAMILYNAME_PROVIDER

Contact

PHONEFAXEMAILURLIP

Location

ADDRESSCITYSTATEZIPCOUNTRY

Dates & age

DATEAGE

Medical

MRNHEALTHPLANHOSPITALDEVICE

Government IDs

SSNLICENSEPASSPORTIDNUMVIN

Financial

ACCOUNTCREDITCARDIBANROUTING_NUMBER

Work

EMPLOYERPROFESSION

Get started

The same weights ship as safetensors, ONNX, Core ML and LiteRT in the model repo.

redact-core.js is the tokenizer, windowing and span decoder this page runs - one dependency-free file. Get it on GitHub ↗

npm install onnxruntime-web
import * as ort from "onnxruntime-web";
import { loadRedactor, redact, restore } from "./redact-core.js";

const redactor = await loadRedactor({
  ort,
  baseUrl: "https://huggingface.co/CDT5058/phi-redact/resolve/main/",
});

const input = "Patient Maria Garcia, DOB 1985-03-22, MRN 4471829.";
const spans = await redactor.detect(input);
// [{ start: 8, end: 20, label: "NAME_PATIENT", text: "Maria Garcia", score: 0.9999 }, ...]

const { text, map } = redact(input, spans);
// "Patient [NAME_PATIENT_1], DOB [DATE_1], [MRN_1]."

const reply = await callYourLLM(text);   // the provider only sees placeholders
console.log(restore(reply, map));        // originals put back locally

Specs

Task
Token classification, BIOES tags (121 labels: 30 entity types × 4 + O)
Architecture
BERT-mini (google/bert_uncased_L-4_H-256_A-4), 4 layers, 256 hidden, 4 heads
Training
Distilled (α = 0.7) from a 66M-parameter DistilBERT teacher on ~157k commercial-safe documents - no real patient data
Context
256 tokens per pass; longer text runs as overlapping windows
Language
English only - other languages miss silently rather than degrade safely
Formats
safetensors (44.5 MB fp32) · ONNX fp32 (44.6 MB) and int8 (11.3 MB) · Core ML · LiteRT
License
Apache-2.0 weights; training data CC-BY 4.0 / CC-BY-SA

Know the limits

Not a certified de-identification system.

It doesn't satisfy HIPAA Expert Determination or Safe Harbor. Treat it as a detection aid alongside human review.

  • English only. Other languages produce silent misses.
  • Trained on synthetic and public data - typos, OCR noise and unusual formatting will cost accuracy.
  • URL recall drops to 0.318 in informal first-person prose.
  • Roughly 1 in 7 entities is missed on Nemotron-PII. Don't treat a clean output as proof there's nothing left.

Questions

Does the text I type here get sent anywhere?

No. The page downloads the model once (11.3 MB, served from this site) and ONNX Runtime - fetched from jsDelivr - runs it in WebAssembly in your tab. There's no inference endpoint - you can open the network tab and watch nothing go out while you type.

Where does dictation happen?

Also on the device. The first time you press Dictate, the page downloads Whisper tiny (about 41 MB) and transformers.js; after that your recording is transcribed in WebAssembly and handed straight to the redactor. The audio is never uploaded - it's the same arrangement as the text.

Is this HIPAA de-identification?

No. It isn't certified under Expert Determination or Safe Harbor. It's a detection aid: use it to reduce exposure and to pre-mark text for review, and evaluate it on your own documents before relying on it anywhere compliance-sensitive.

How does it compare to Desert Ant's Redact?

On English it's roughly tied - 86.1% vs 86.5% average recall across five benchmarks - at half the parameters and about half the on-device size. Redact covers 27 languages; phi-redact is English only and adds clinical types like MRN, health plan and provider names.

Why are some labels wrong in the demo?

The model sometimes gets the right span with the wrong type - an insurer tagged as EMPLOYER, or the last four digits of a card as a ZIP code. For redaction the span matters more than the label, but it's worth knowing. PROFESSION is the weakest type (F1 0.494) because it fires on job titles in ordinary prose; switch it off if you don't need it.

What happens with long documents?

The model reads 256 tokens at a time. The demo slides overlapping windows across the text and keeps each token's prediction from the window where it sits closest to the middle, so entities at a boundary still get full context.

Putting redaction in front of an LLM, or working with PHI in a regulated system? Reach out.