HomeVideos

OpenAI's "Privacy Filter": The Open-Weights Release Everyone Missed

Now Playing

OpenAI's "Privacy Filter": The Open-Weights Release Everyone Missed

Transcript

239 segments

0:00

This is a new iteration of somebody I

0:01

built where users can upload their

0:03

documents, redact information, and then

0:05

send it off to AI. So, I just have this

0:07

fake medical document I made, and

0:08

because it's content aware, I also tried

0:10

to trick it up. Let's just drop it in

0:12

here, upload it, and we're going to

0:13

press add and redact. So, OpenAI had a

0:15

pretty crazy week this week. They

0:17

released GPT 5.5, they released GPT

0:19

image 2, they updated Codex, they've

0:21

been pushing that very hard. But, one of

0:23

their releases kind of dropped under the

0:24

radar that I found interesting, and that

0:26

is the open weights PII redaction model.

0:28

This is a tiny classification model made

0:30

for detecting and redacting PII. And

0:32

because it's open weights, you can run

0:34

it on your device right now. Now, for

0:35

context, the reason this is so

0:37

interesting to me is over the last few

0:38

years I've had to go over, upload, and

0:40

translate a lot of medical

0:41

documentation. But, from a

0:42

privacy-conscious perspective, I didn't

0:44

just upload it to ChatGPT or to Claude.

0:46

I would spend time redacting it. Names,

0:49

emails, phone numbers, ID numbers,

0:50

addresses. There's a lot of it, and it

0:52

shows up in a bunch of different ways.

0:53

So, this becomes a very tedious process.

0:55

So, about 2 years ago, I made this

0:57

application that you can input a

0:58

document, be it a text file, a markdown

1:01

file, a docx file, even some PDF files,

1:04

it would break it down into chunks, use

1:05

a bit of OCR, and then use pattern

1:07

matching, regex, and a bunch of

1:09

different deterministic rules to figure

1:11

out if certain strings were actually

1:13

PII. And in any case, I'd always go over

1:15

it one more time manually, redact the

1:17

information that I missed, and then be

1:18

able to upload it or translate it or

1:20

share it with somebody if I needed to.

1:22

And I did this for hundreds of

1:23

documents. And as the AI started picking

1:25

up, other models came out. This is one I

1:27

used it for a while, the Piranha V1. And

1:29

I spent probably the most amount of time

1:31

on that specific feature, the personal

1:33

information detection and redaction

1:35

feature. And the truth of the matter is,

1:37

it broke all the time. And I'm just one

1:39

guy trying to figure this out with a

1:40

bunch of other things going on in life.

1:41

Then, this week, OpenAI dropped their

1:43

own privacy filter, and it's a lot

1:45

better than anything I made or anything

1:47

else that's available outside right now.

1:48

Privacy Filter is a small model with

1:50

frontier personal data detection

1:52

capability. It is designed to be able to

1:54

perform context-aware detection of PII

1:57

in unstructured text. It can run

1:59

locally, which means that PII can be

2:01

masked or redacted without ever leaving

2:03

your machine. And it can process long

2:05

inputs. I think I saw it was 128,000

2:07

tokens, whereas Piranha, I think it was

2:09

like a fraction of that. Here's OpenAI's

2:11

Privacy Filter. We can test it right

2:12

here in Spaces. I'm just going to put in

2:14

a bunch of text. My name is Steve Stark.

2:16

I live at Bangkok, San Francisco, 145

2:18

Pennsylvania Street, California 98760.

2:21

My email is captaintaco@bankrupt.com.

2:23

Social Security 123684432.

2:26

Let's just see if it redacts it. Great.

2:28

So, it picked up everything. Got the

2:29

name, got the address, got the email,

2:31

and it got the social security. They say

2:33

privacy detection depends on more than

2:35

pattern matching. Traditional detection

2:37

tools often rely on deterministic rules

2:39

for formats like phone numbers and email

2:41

addresses. They work well for narrow

2:43

cases, but they often miss more subtle

2:45

personal information and struggle with

2:46

context. And this is the bigger problem

2:48

that I faced in my tool. Privacy Filter

2:50

is built with deeper language and

2:52

context awareness for more nuanced

2:54

performance. By combining language

2:56

understanding with privacy-specific

2:58

labeling system, it can detect a wider

3:00

range of PII in unstructured text,

3:03

including cases where the right decision

3:05

depends on context. It can better

3:06

distinguish between information that

3:08

should be preserved because it's public

3:10

and information that should be masked

3:12

and redacted because it relates to a

3:13

private individual. So, here are all the

3:15

different data types that it can detect.

3:17

They say account number can include

3:19

banking information, credit card

3:20

numbers, bank account numbers. That's

3:22

what it also counted the social security

3:24

number as, and secrets like an API key

3:26

or a password. So, important for that me

3:28

to point out, this isn't a model you

3:29

could download and run on your computer

3:31

locally like you would with Gemma or

3:33

some other open-source LLM. This is a

3:35

classification model. So, what this does

3:37

is allows us to build this into our

3:38

software or build this into our

3:40

workflows. OpenAI is essentially giving

3:42

this to us, and in no means is it

3:44

perfect, but it's actually a huge deal

3:45

because it lowers the threshold or the

3:47

barriers for actually implementing

3:49

redaction tech. Cuz when implemented, it

3:51

can run on your computer or your

3:52

company's infrastructure to redact long

3:55

documents before sharing them with third

3:57

parties. Now, before I show you what I

3:58

built, from a privacy perspective, it's

4:00

very important for me to point out

4:01

Privacy Filter and tools like it are not

4:03

a one-stop solution. You have to

4:05

implement privacy by design. And they

4:07

say here, it's not an anonymization

4:09

tool, it's not a compliance certificate,

4:11

and it's not a substitute for policy

4:13

review. This is a tool that will help

4:14

you along the way. But, if you not only

4:16

want to be compliant with regulations,

4:18

but also avoid data breaches or just

4:20

avoid uploading your data or someone

4:22

else's data to third-party servers, you

4:25

have to implement privacy by design.

4:27

Because the truth of the matter is, the

4:28

minute you upload anything to a

4:30

third-party server, you lose control of

4:31

it. You don't know what's going to

4:32

happen. Doesn't matter what they say, it

4:34

can and will and probably be leaked or

4:36

breached in some way. So, if you

4:37

practice good data hygiene and you

4:39

redact and you implement privacy by

4:41

design from the get-go, you will be a

4:43

lot better off. In this Privacy Filter

4:45

Hugging Face page, they give us a lot of

4:47

information, including how to install it

4:49

using Transformers and PyTorch. What I

4:52

ended up doing was rebuilding my whole

4:54

tool with GPT 5.5. Had it focused just

4:56

on uploading a document, PDF, TXT, docx,

5:00

and markdown file, parse it, and then be

5:02

able to run the redaction on it. I just

5:03

want to show you an example cuz I think

5:05

it's really cool to see what it's able

5:06

to do with your own data. Okay, so here

5:08

we are, Privacy Cabinet. This is a new

5:10

iteration of something I built where

5:12

users can upload their documents, redact

5:14

information, and then send it off to AI

5:16

to get better information without

5:17

putting all their personal information

5:18

on someone else's cloud to later be

5:20

trained on. So, I just have this fake

5:22

medical document I made. Let's open it

5:24

up. This is a fake document, fake

5:25

information, fake people, but I want it

5:27

to look like a real medical document.

5:29

And what I tried to do here is give it a

5:30

lot of personal information, and because

5:32

it's content aware, I also tried to

5:33

trick it up. So, we have the name of a

5:35

clinic, we have the address, we have its

5:36

phone number, we have the name of the

5:38

doctor, doctor's phone number, the

5:40

doctor's email, patient's name,

5:42

patient's birthday, social security,

5:44

pretty much everything. And by the way,

5:45

I also put a medication that I think

5:48

kind of looks like an address just to

5:50

see if Privacy Filter will mistake that

5:52

for personal information like a private

5:54

address or it will know that it's

5:56

medication. Olanzol, I don't think

5:58

that's real. Okay, so let's just drop it

5:59

in here, upload it, and we're going to

6:00

press add and redact. And forgive the

6:02

UI, this is just a basic implementation.

6:04

Now, we see the fake medical document.

6:06

This is an RTF file, so it uses text

6:08

util to properly parse it. We see the

6:10

same document here. I know it doesn't

6:11

look great, but you see the same thing,

6:12

the address, name of the doctor, the

6:15

patient, the information, and it already

6:17

ran the redaction. And by the way, I

6:19

also made it possible to manually redact

6:21

things if it fails, but hopefully it'll

6:23

be fine. So, we see here that the first

6:25

thing that I put in here, the clinic,

6:28

the name of the doctor's office, none of

6:29

it was redacted. And that's because this

6:31

is public information. Doesn't need to

6:33

be redacted. But then, when we get down

6:34

to the doctor, I'll redact it. Private

6:36

name, private phone number, private

6:38

email, private credential. Same thing

6:40

with patient. Private person, private

6:42

date, private number. Everything that's

6:44

private information has been redacted

6:46

here. And if we go down, it also didn't

6:48

trip up on the fake medication that kind

6:51

of looks like an address. Overall, I

6:53

think this is really cool, especially

6:54

because of my approach to privacy and

6:56

also my background in trying to build a

6:57

similar solution. I think it's really

6:59

cool that OpenAI released this. It's

7:00

very overlooked, but very powerful. And

7:02

once people start building into their

7:04

workflows, I think we're going to see

7:05

some very cool and very useful tools.

7:07

So, I hope you found this video helpful

7:09

and insightful. If you have any feedback

7:10

or questions, drop in the comments

7:11

below. If you haven't done so already,

7:13

subscribe to the channel. It really

7:15

helps me grow. Thank you guys for

7:16

watching and have a great night.

Interactive Summary

The video discusses OpenAI's release of the Privacy Filter, a classification model designed for detecting and redacting personal identifiable information (PII) in unstructured text. The creator highlights the model's context-aware capabilities, which offer a significant improvement over traditional, deterministic regex-based redaction tools. The creator also demonstrates how they have integrated this technology into their own document processing application to securely handle medical files, emphasizing the importance of privacy-by-design principles when using AI.

Suggested questions

3 ready-made prompts