HomeVideos

Claude watermarks your code now

Now Playing

Claude watermarks your code now

Transcript

926 segments

0:00

One of the most frustrating things about

0:01

AI is that it's hard to know if someone

0:03

used it or not. Sure, there's all the

0:05

common tells like the M dashes and the

0:07

you're absolutely right and the it's not

0:09

X, it's Y's and all these types of

0:11

patterns, but it's generally pretty hard

0:13

to actually know if a given piece of

0:15

text or line of code was really written

0:17

with AI. And this has scared a lot of

0:19

people and for reasons that are good. If

0:21

it's impossible to know what text was

0:22

written with AI and what was written by

0:24

a human, the world will devolve into a

0:26

pile of slop really, really fast. The

0:28

European Union agrees and that's why

0:30

they just introduced the AI act which

0:32

has a specific goal of making it easier

0:33

to know and label AI generated media and

0:36

text and anthropic is complying with it.

0:39

Normally these types of watermarks are

0:41

things that only apply to media gen

0:43

things like images, videos, audio, etc.

0:45

because there's a lot of data there

0:46

that's relatively easy to sneak

0:48

watermarks into. But when it comes to

0:50

text, it's a lot harder. Which is why

0:53

it's pretty crazy that Anthropics

0:55

already updated their docs on this,

0:57

saying that they plan to do marking of

1:00

all AI generated content across all

1:02

platforms. Cloud models launched in the

1:04

EU on or after August 2nd, 2026 will

1:07

support machine readable markings at

1:10

launch. Generated text will carry

1:12

embedded watermarks and generated files

1:13

will include digitally signed Providence

1:15

metadata where supported. Oh boy. This

1:18

means our friends at Anthropic will be

1:20

marking not just the text they output,

1:22

but the code they output as well. This

1:24

is a very scary change and I have a

1:27

feeling it's not going to do what is

1:30

intended. As always, there's a lot of

1:32

layers to this one from the relatively

1:34

good intent of this EU act to the

1:36

questionable implementations we're going

1:38

to be seeing all across the field and

1:40

also the impact this will have on us

1:41

actually using the models. I'll talk

1:43

about all of that, the workarounds, and

1:44

more after a real quick break for

1:46

today's sponsor. Picking the right

1:47

database has never been more important,

1:49

but it's also never been harder. You can

1:51

go with industry standards like

1:52

Postgress for your application or

1:54

ClickHouse for your data warehouse. But

1:55

how do you pick the right thing if you

1:57

don't even know what you're building

1:58

yet? Wouldn't it be nice if there was a

1:59

great solution that combine the best

2:01

both of those? You know, like if

2:03

Postgress was just managed by

2:04

Clickhouse. Yeah, they're actually doing

2:06

this now. And it couldn't be cooler. You

2:08

get enterprisecale Postgress with all

2:10

the benefits of Clickhouse, too. The

2:11

Postgress database is what you should

2:13

use for your apps. It's super fast and

2:15

built on top of real NVMe drives on the

2:17

same machine resulting in absurd

2:19

performance up to 10 times faster for IO

2:21

heavy workloads. Where things get magic

2:23

is the ClickHouse integration. All of

2:25

the changes in Postgress get synced to

2:27

your ClickHouse. So you can do massive

2:29

analytics calls without slowing down

2:30

your production DB. And if you're

2:31

worried about latency, don't be. Your

2:33

data is replicated within seconds. And

2:34

if you want to access that analytics

2:36

data from the Postgress DB in your app,

2:38

you can with PG Click House. This would

2:39

have saved my butt when I was building

2:41

the T3 chat wrapped feature last year

2:43

and I had to query data from all sorts

2:44

of different places in order to pull it

2:46

together. Setup couldn't be easier. You

2:48

click start, you add a new database, you

2:50

get a connection string and now you're

2:51

good to go with a real Postgress

2:53

database with all the power of

2:54

ClickHouse behind it. Stop compromising

2:56

on your database today at

2:57

soy.link/clickhouse.

2:59

The best place to start on this chaos is

3:01

probably the official code that the EU

3:03

has put out. This is the code of

3:05

practice on transparency of AI generated

3:08

content. There's two different sections

3:10

in this code. There's section one for

3:11

the providers which is rules for marking

3:13

and detection of AI generated and

3:15

manipulated content and section two is

3:16

for deployers rules for labeling of deep

3:18

fakes and AI generated and manipulated

3:20

text. This code was drawn up by

3:21

independent experts in a

3:22

multistakeholder process facilitated by

3:24

the AI office of the EU. It helps

3:26

providers and deployers of generative AI

3:28

systems to comply with the AI acts

3:30

obligations for labeling and marking of

3:32

AI generated content. Even though

3:34

adherence to the code is voluntary, the

3:36

transparency requirements under article

3:39

50 are legal obligations. AI needs to be

3:42

deployed in ways where users know

3:44

they're talking to AI. That is totally

3:46

fair. Providers of AI systems generation

3:49

shall ensure the outputs of AI systems

3:51

are marked in a machine readable format

3:53

in detectable as artificially generated

3:55

or manipulated. The inclusion of text

3:57

content here is a bit concerning.

3:59

Providers shall ensure that their

4:00

technical solutions are effective,

4:01

interoperable, robust, and reliable as

4:04

far as this is technically feasible,

4:06

taking into account the specificities

4:07

and limitations of various types of

4:09

content, the costs of implementation,

4:12

and the generally acknowledged

4:13

state-of-the-art as may be reflected in

4:15

relevant technical standards. So, this

4:18

is the only section that is obligatory

4:20

for companies in this ruling. The rest

4:24

of the code is more suggestions. But

4:26

that one piece there, that uh section

4:30

two here, the providers must ensure that

4:33

outputs are marked in a machine readable

4:35

format. That's where this all starts to

4:37

become a show. There's an

4:38

interesting call at the end here of the

4:40

section two. This obligation shall not

4:42

apply to the extent the AI systems

4:45

perform an assistive function for

4:46

standard editing or do not substantially

4:48

alter the input data provided by the

4:50

deployer of semantics thereof or where

4:53

authorized by the law to detect,

4:55

prevent, investigate or prosecute

4:57

criminal offenses. What this means is if

4:59

you have an AI system that is like

5:01

editing or autocorrecting or things for

5:03

existing documents, images, whatever,

5:05

that does not need to have the same

5:06

level of like transparency and encoded

5:10

metadata proving that it was AI. So if I

5:12

use AI to like do a small correction

5:14

like to remove a zit from my face, that

5:16

doesn't need to be marked the same way.

5:18

If I'm writing an essay and I ask for

5:20

the agent to find little things that I

5:22

could improve in the essay or like clean

5:24

up the grammar and stuff, it doesn't

5:25

need to apply the same way there. But

5:26

the actual output from the LLM, it does.

5:29

Very interesting. And with that, we now

5:31

need to talk a little bit more about

5:33

watermarks. Going to start from a weird

5:35

point. So, hear me out. This image could

5:39

be AI generated. It might not be. It's

5:42

hard to know for sure. I don't know

5:44

where this image would have been taken,

5:46

but Maria posted it online on Twitter a

5:49

few days ago. In fact, when Maria posted

5:52

it, there was no marking here that said

5:55

it was generated with AI. So, clearly

5:57

this photo is real, right? Let me go a

6:00

step further. These two images are the

6:03

same, right? Like obviously these are

6:05

the exact same image with no differences

6:08

between them. What if I told you there

6:10

was not a single bite shared between

6:13

these two images? One of these is a PNG

6:18

and the other is a JPEG. This is what

6:20

makes watermarking for media both really

6:22

easy and really hard. That image we were

6:25

just looking at, it's not huge, but it's

6:28

not small either. It's 623

6:31

kilob. And the relationship between the

6:34

size of the file and what we see not

6:37

that simple. When you have this much

6:39

data in a file, there's a lot of room to

6:43

stuff things into it. a lot of which

6:45

won't even be perceivable to humans.

6:48

Let's say I want to add some special

6:50

encoding to make sure that I know this

6:52

image is the one I have here on my

6:55

computer. I could put it in the image

6:56

metadata and things like that, but I

6:58

want it to be harder to get rid of.

7:00

Let's just zoom into some of the like

7:02

noise and the gradient here and pick a

7:04

corner. Let's say let's just start from

7:07

the top left. Why not? I'll take a

7:09

random color from here. We'll take this

7:11

darker purple. And now I'm going to put

7:13

in a little encoding. I'm T3. So what

7:16

we'll do is every three dots I will fill

7:20

one with this darker color. One, two,

7:21

three. Do it. One, two, three. Do it.

7:25

One, two, three. Do it. Okay. I just

7:28

added a marking on this image in the top

7:31

left corner where I changed one of the

7:34

colors in specific pixels every three

7:37

pixels. So three are left normal, then

7:39

one I change it slightly. Can you tell

7:41

any difference in this image? If you

7:44

think you can, you're lying because the

7:45

compression in the video will not allow

7:48

such a thing. And this is me changing

7:51

the pixel to a civic color. So, we do

7:53

three normal, one change, three normal,

7:55

one changed, but I was changing it to a

7:57

fixed color. Imagine we get a little

7:59

sneakier with it. Imagine instead we

8:02

take the color of a given pixel here and

8:05

we knock it down or up one in any of the

8:09

different spectrums you have control of

8:10

with color because you have RGB values

8:14

or HSL and you can change any of these

8:17

at any point to or from anything. So we

8:19

could have a pattern where every four

8:21

pixels we bump one of these by one and

8:24

it wouldn't change anything visually for

8:26

any normal person. If we just go from 30

8:27

to 31 every four pixels, that is

8:30

relatively easy to encode in an image.

8:33

And it's in a way that will not be

8:34

noticeable to any human, but does

8:37

accurately with the actual pixels encode

8:41

real data. But in order for that type of

8:44

granular encoding to work where every

8:46

four pixels we go up or down one value,

8:49

we need to make sure we don't lose the

8:51

pixel values. And it turns out the

8:53

compression we use for media is directly

8:56

in opposition of the types of watermarks

8:58

we're talking about here. One of the

9:00

things most image compression algorithms

9:02

will do is they take areas where colors

9:05

are similar and they flatten them to be

9:07

the same so you don't have to encode as

9:08

much data. It'll take a grid every four

9:10

pixels and say okay these four have

9:13

these colors and these brightness

9:14

levels. Can we normalize that so we only

9:16

have to encode it once instead of four

9:18

times? That means that if we have this

9:20

pattern where every four pixels we

9:22

change the color slightly, that's going

9:24

to be flattened by pretty much every

9:26

compression algorithm ever. And all you

9:29

have to do to kill the watermark is turn

9:31

your PNG into a JPEG. It is seriously

9:34

that easy. And I have been able to evade

9:36

pretty much every single fancy type of

9:39

media detection algorithm for generated

9:41

AI image detection with something as

9:44

simple as reexporting as a JPEG or just

9:47

increasing the compression or scale the

9:49

image up and then back down. There's a

9:51

lot of little things you can do that

9:53

just destroy these watermarking

9:55

techniques. And the real reason is this

9:58

here. There is so much data available

10:01

here that leaving watermarks in it is

10:03

easy. But all the algorithms are

10:05

optimizing for what humans actually see.

10:08

And the difference between what data

10:10

exists and what humans perceive is the

10:14

problem here. With something like an

10:16

image, there's a wide range of data that

10:19

can be changed before a human notices a

10:21

difference. This means it's easy to

10:22

embed this type of metadata in an image,

10:24

but it's also hard to do it in a

10:26

resilient way because the compression

10:29

algorithms are working around this fact

10:31

that humans are bad at seeing the actual

10:34

ones and zeros in the bites. So there's

10:36

a lot of hacks we can do to compress the

10:39

files in order to save storage, but it

10:42

also in that process would kill a lot of

10:44

these hidden pieces of metadata. This is

10:46

why many people who post their AI

10:48

generated images on X get tagged because

10:50

they're just copy pasting the image

10:52

straight from chat GPT. And when you

10:54

copy from chat GPT or Gemini or wherever

10:56

else, they have these IDs in them.

10:58

Google's went as far as creating this

11:00

whole synth ID platform, which is a tool

11:02

that watermarks and identifies AI

11:04

generated content, mostly images, video

11:06

stuff that they've been talking about

11:07

forever. And no matter how hard I push

11:10

them, they will not let me in. And I

11:12

only get like three test runs a day

11:14

using the official Gemini site. Synth

11:17

ID's tagging is hilariously simple. They

11:19

apply this noise pattern on top of the

11:22

image. GBT image 2 has a similar one

11:24

where this pattern gets applied on top

11:27

and it looks that same deep fried way

11:30

that GPT images look. I'm not sure if

11:33

this is the entirety of how they do this

11:35

check and how they do these tests, but I

11:37

would guess it is a significant portion

11:39

of it. These patterns are what are being

11:42

used for a lot of the identification.

11:43

And again, the right compression, the

11:46

right application, the right layers

11:47

added to it, a very very slight filter

11:50

even, which is trivial for me to add. If

11:52

I just go here, filter, I can do like

11:56

the slightest sharpening pass where I

11:59

add like a 1% sharpen.

12:02

no perceivable difference. But now every

12:04

single bite in this image has been

12:06

re-encoded and any patterns that were

12:08

applied on top are fully destroyed. I

12:10

can even demo that trivially by copying

12:12

the noise pattern and then doing a quick

12:15

anything with this. If I apply that 1%

12:18

sharpen doesn't look any different, but

12:20

the bites that they're actually looking

12:22

for here have just changed entirely. If

12:24

I zoom in real deep and show you

12:27

see that you can literally see the

12:28

difference with like 2%. And I'm going

12:31

to apply. Now it's a very subtle change.

12:36

Zero is nothing. And then five. You can

12:38

see a slight change too. This is for

12:40

like a visible noise pattern. Any of

12:42

these are indetectable to humans. None

12:45

of them are indetectable to AI. Like any

12:47

one of those changes will entirely

12:48

destroy the patterns it's looking for

12:50

here. It's so easy to kill these things.

12:53

It's hilarious. It's even easier if

12:56

you're okay with a slightly blurriier

12:57

image or you do the lazy thing, which is

13:00

to resize the canvas and see how the

13:03

pattern got way blurrier. How do you

13:05

think that will work if it's also

13:07

applied at 1% opacity on top of

13:09

something else? Trivial to destroy these

13:12

patterns. There's no world in which they

13:14

will persist well across any of these

13:16

types of things, especially if you apply

13:18

them the way they're actually applied

13:20

IRL, which is something like this. You

13:24

have it on the image as like a 0.1%

13:26

opacity, maybe 1%. Yeah, you don't see

13:30

any difference here. And if I blur,

13:32

unblur, resize, or do anything like

13:34

that, the pattern gets destroyed very

13:37

fast. You could even do the stupidest

13:38

possible thing, which is to just take

13:40

that original image and do that with it.

13:44

I just change where every pixel lands

13:47

just by shortening it slightly. It's so

13:50

easy to kill these types of watermarks

13:52

just because there's so much wiggle room

13:54

in the data. When you have 623

13:56

kilobytes, I can delete random and

13:58

no human will notice, but the computers

14:00

no longer have access to that pattern.

14:01

All of this is to say this image is

14:03

clearly real and we would have no way to

14:05

know otherwise. So again, the problem is

14:08

this gap between the amount of data and

14:11

what humans perceive. And when you have

14:13

images that are millions of bits, it's

14:16

easy to flip things around in all

14:18

directions. But what happens when you go

14:20

from 623 kilobytes to I don't know 10

14:23

bytes. Each of these bytes is a letter

14:26

or a chunk for a word. Like once you're

14:28

responding with plain text that is

14:30

small, your ability to hide things in it

14:33

and also transform it goes down

14:36

massively. If I have a four pixel image

14:38

and I change one pixel, you can tell. If

14:40

I have a 4 million pixel image and I

14:42

change one pixel, you can't tell. If I

14:43

have a whole book and I change one word,

14:45

you're not going to be able to tell. But

14:47

if I have one sentence and I change a

14:48

word, you can absolutely tell. And

14:51

that's where this new text watermarking

14:53

comes in. I already read this section

14:55

about how Anthropic is going to be

14:57

following the rules. They simply call

14:59

out that the marks are going to work

15:00

everywhere that Claude is used, whether

15:02

that is through the API on the platform,

15:04

Claude, Claude Code, Claude Co-work, and

15:06

Claude tag. They also want to help us

15:07

detect Claude's marks, too. They'll

15:09

support users and other third parties to

15:11

detect Claude's marks as the code

15:12

requires. We'll share details in

15:14

forthcoming documentation. They are

15:16

actually planning on adding marking to

15:17

their old models as well, which is

15:19

interesting to me too. As AI generated

15:21

content becomes commonplace, greater

15:22

transparency and signals about where

15:24

content comes from can give people

15:26

useful context about the information

15:28

they consume. Support transparency and

15:30

comply with our legal obligations.

15:31

Anthropic is working to include machine

15:33

readable marks in content that Claude

15:35

generates. Embedded watermarks will

15:36

apply to all generated text and

15:38

Providence metadata will apply where

15:40

Claude supports processing files. This

15:42

will work with all the cloud platforms

15:43

too. Whether you're hosting on AWS,

15:45

Google Cloud or Foundry. So let's see

15:46

how they actually will be marking the

15:48

content. Cloud uses two complimentary

15:50

techniques to mark content generated and

15:52

processed by AI or by cloud

15:53

specifically. Watermarks embedded in

15:55

text and signed providence metadata

15:56

attached to files. First we have the

15:58

text embeds. When a supported cloud

16:00

model generates text, it weaves an

16:01

imperceptible watermark directly into

16:04

the text itself. You won't see it. It

16:06

doesn't change the meaning, quality, or

16:08

readability of clause responses. Because

16:10

the watermark is part of the text, it

16:12

will travel with the text when it's

16:14

copied and pasted elsewhere and may

16:16

persist through some editing.

16:17

Watermarking will be applied at the

16:19

model level, which means it will be

16:20

present no matter which cloud product or

16:22

surface the text comes from. This is a

16:24

very interesting call out specifically

16:26

may persist through some editing. Then

16:29

there's the signed providence metadata

16:31

which is that when they touch or create

16:33

things like SVGs, PGs or JPEGs, it will

16:35

provide signed providence metadata.

16:37

Metadata follows the coalition for

16:39

content providence and authenticity. The

16:41

CGPA open standard which is used across

16:43

the industry to record information about

16:45

content providence. If the metadata is

16:47

present, it signals that a file was

16:48

processed with claude and lets you

16:49

detect whether the file has been

16:50

tampered with. Thankfully, they call out

16:52

limitations underneath here. Detected

16:54

mark provides a signal that content was

16:55

processed by Claude, but it's not fully

16:57

conclusive. Detecting a Claude mark

16:59

tells you that content may have been

17:00

processed by Claude. It does not on its

17:02

own confirm the full providence of the

17:04

content. For example, Claude may not be

17:06

the original author because people often

17:07

use Claude to proofread, translate,

17:10

summarize, or convert files. The output

17:12

can carry a Claude mark even if the

17:14

underlying ideas, text, or data

17:16

originated from another source. The

17:17

content also may have been changed after

17:19

Claude processed it. Mark content may be

17:21

modified, exerted, or combined with

17:22

other materials after Claude is

17:24

processed. There's also the fact that a

17:25

lack of a detected mark doesn't mean the

17:26

content wasn't AI generated or

17:28

processed. There's also the fact that it

17:30

doesn't mean anything if it says no, we

17:32

didn't process that because there's so

17:34

many cases where it won't flag. Like if

17:37

the model generated text before this was

17:39

a thing, if the text has been heavily

17:41

edited, paraphrased, translated, or

17:42

mixed into other writing. If the passage

17:44

was too short, so there wasn't enough

17:45

text to encode a watermark. And for the

17:47

file stuff, if the file's been

17:49

converted, resaved, screenshotted, or

17:51

other things, again, the examples I was

17:52

giving here, then it won't work. And if

17:54

it was produced through a platform

17:55

feature or file type where a particular

17:57

marking type wasn't supported, very,

17:59

very fun. The harsh reality here is that

18:03

this is only ever going to catch the

18:05

lowest effort people because if you just

18:08

copy paste the text straight out of

18:10

Claude or you have a bot that auto

18:12

replies on Twitter for you and things

18:13

like that, then yeah, this will catch

18:15

those things. But if you put any effort

18:17

in or do even the most basic of

18:20

grammatical pass or changes to the text

18:23

that comes out, this will be trivial to

18:26

break. It is worth noting that other

18:28

labs will likely be doing this soon too.

18:29

Thoric called this out. Other labs will

18:31

be adding similar watermarking. It's

18:33

hard to identify AI generated text and

18:35

this gives people better tools to do it.

18:37

They'll be shipping a text detection API

18:39

as well that we can use ourselves. All

18:40

claw generated text will have this

18:42

embedded watermarking. For example, you

18:43

can check if a PR was generated by cloud

18:45

code. That said, it does have

18:46

limitations. He just called out the same

18:48

section we just read. We also got this

18:50

wonderful example of the watermark from

18:53

Nat here. Wow, the new claw text

18:55

watermarking is wild. Here we can see it

18:57

making a list of totally normal, not

18:59

suspicious text, but search engine

19:00

optimization, but if you look real

19:03

closely and you use those thinking

19:04

glasses,

19:06

Sam Alman sucks. Wow, it's really

19:11

subtle, isn't it? Back to the real world

19:12

for a sec though because Sean Godc wrote

19:14

yet another awesome piece about this

19:16

because when the act was officially

19:18

announced he had concerns. Article 50 is

19:21

the thing we're talking about here. If

19:23

LM providers want to do business in the

19:24

EU, they'll have to apply watermarks to

19:26

their outputs. Some hidden signatures

19:28

that can be used to identify AI content.

19:30

LM text watermarking is a fascinating

19:32

problem. Like the best engineering

19:33

problems, it's theoretically hard to

19:34

solve perfectly, but has multiple

19:36

partial solutions like synth ID and some

19:38

quiet unicode trickery from OpenAI and

19:40

anthropic. It'll be interesting to see

19:42

how these AI labs navigate these

19:43

trade-offs before the end of the year.

19:45

So, why is this so hard? It's easy to

19:47

watermark an image because digital

19:48

images contain lots of noise the human

19:50

eyes cannot see. For instance, you could

19:52

apply a watermark like these 20 pixels

19:54

in these exact spots will always share a

19:56

color. Text is much much harder. Unlike

19:58

images, text is very compressed already.

20:00

You cannot make any changes to a

20:02

sentence that a human won't notice with

20:03

one exception which we'll get to later.

20:05

So, how are you supposed to watermark

20:06

it? It's basically a textonography

20:08

problem concealing a secret code made

20:10

more difficult because the plain text

20:11

cannot be arbitrarily manipulated. Any

20:14

changes you make to apply the watermark

20:16

will compromise the quality of the

20:17

output. For instance, every fifth letter

20:20

is an E would be a good watermark, but

20:22

applied naively, it would make the AI

20:23

output full of typos. Could you just let

20:25

the model figure out how to fit the

20:27

watermark? Strong AI models are smart

20:29

enough to juggle this kind of

20:30

constraint, but it still consume

20:32

reasoning time that we better spend on

20:34

the user's problem and make the model

20:35

sound much less capable than it is. Yep,

20:38

sorry. It's steganography. My bad.

20:41

Steganography and stenography being so

20:43

similar of words is uh humans are not

20:46

good at text either. Do we even need

20:48

watermarks to detect AI content? If

20:50

you're enthropic and you're required to

20:52

be able to verify whether your models

20:54

produce a particular block of text,

20:55

can't you simply run the text through

20:56

each model, measuring as you go how

20:58

closely the models predicted tokens

21:00

match each token from the text? Not

21:02

really. The space of all possible cla

21:05

answers to a question is way larger than

21:07

the space of all possible watermarked

21:09

answers to a question. In other words,

21:11

you'd get too many false positives for

21:13

human text that reads like it was AI

21:14

written. I know a couple people,

21:16

especially YouTubers, who are hit really

21:18

hard with this where like the way they

21:19

naturally speak is similar to the LLM

21:21

because those models were trained on

21:22

them and now whenever they talk was

21:25

like, "Oh, that's AI generated, but in

21:26

reality, they just talk in the way that

21:28

the models liked and copied from them."

21:30

It's way more likely for a human to

21:31

accidentally write like Claude than it

21:33

is for a human to accidentally reproduce

21:34

a watermark. It would also be

21:36

prohibitively expensive to run every

21:37

anthropic model against a piece of text

21:39

in order to watermark it. The EU AI act

21:41

will eventually require labs like

21:42

Enthroic to offer free watermarking

21:44

services to every EU citizen. You

21:46

couldn't do that with a run this model

21:48

based approach. Yep. Apparently, synth

21:50

ID is being used for text. Like we need

21:52

more reasons for Gemini to be bad. When

21:54

an LLM generates text, it's generating a

21:56

series of tokens. At each step, the

21:58

model itself does not output a single

22:00

token, but instead outputs a full list

22:02

of 100,000 tokens in its vocabulary,

22:04

each annotated with the probability that

22:05

that token will be the next one. Tools

22:07

like chatbt and claude code will pick

22:09

semi-randomly from the most likely

22:11

output options in order to get their

22:13

outputs. The semi-random sampling

22:15

process can be influenced in a

22:16

detectable way. For instance, we could

22:18

choose a sampling strategy like we pick

22:20

the second most likely token, then the

22:21

first, then the second, then the first,

22:23

and so on to still produce high quality

22:25

output, but you'd be able to rerun the

22:26

model against the generated text and

22:28

verify that the pattern holds. But this

22:30

makes verification really expensive, and

22:32

any slight tweaks to the output would

22:33

break it. also changes to the input

22:36

because again the whole context window

22:37

is necessary for this. So synth ID does

22:40

this instead by assigning a score for

22:42

each token based on the previous tokens.

22:44

To apply this watermark, the model

22:46

adopts a sampling strategy like out of

22:48

the top five most likely tokens, pick

22:50

the ones with the top synth ID score.

22:52

The watermark can then be detected by

22:54

calculating the aggregate synth ID score

22:56

of a block of text. It's suspiciously

22:58

high. It's very likely to be AI

23:00

generated. This is similar to noticing a

23:02

lot of m dashes, but instead of a list

23:04

of keywords, it's subtle math

23:06

relationships between things that humans

23:08

can't identify as well. And since it's

23:10

trivial to sign the scores, it's trivial

23:11

to run the watermark detection. The

23:13

other strat people have been using are

23:15

these different space characters.

23:18

There's a lot of different ways to

23:19

encode a space in a thing that are not

23:21

just the normal space bar press. humans

23:24

will not see it. But if you encode these

23:26

strategically inside of something that

23:28

an AI model writes, you can detect which

23:31

ones are where and use this to get real

23:33

useful data. I remember Verscell a while

23:35

back was working on a feature where you

23:37

could edit content on a blog or

23:39

something via the Verscell toolbar and

23:42

the really clever strategy for linking

23:43

different pieces of text in the UI to

23:46

the official like CRM backing it for

23:48

your blog or whatever was to append

23:50

different white code text at the end of

23:53

each paragraph so that when you selected

23:55

that text in the UI and you made

23:56

changes, it'd be able to map it to a

23:59

different place in the actual database

24:02

simply by detecting which row it does

24:04

with an identifier using these different

24:06

unicode space characters to hide

24:08

messages inside of the UI. It's very

24:11

clever and really cool, but also like

24:14

easy to trim out, too. If you were to

24:16

use different space characters to apply

24:17

watermarks, it would be very cheap to

24:19

detect, but also easy to filter out too.

24:23

Cloud code is apparently doing this in

24:25

the past to tag suspicious requests from

24:27

Chinese users. In the last few years,

24:29

they've noticed that when they copy

24:30

blocks of text from chatb and paste in

24:32

the VS code, sometimes VS Code marks

24:33

some all of the spaces as unusual unic

24:35

code characters. Are openthropic using

24:38

homoglyphs as AI generated watermarks?

24:40

The author's not sure, but they're

24:41

definitely using them. Interesting. But

24:44

the problem here isn't whether or not

24:46

they're easy to add. It's how easy are

24:47

they to remove. Because all text

24:50

watermarks can be trivially removed. If

24:52

you use these Unicode homoglyphs, it's

24:54

trivial to just replace them with the

24:56

original character. Now you're out of

24:58

it. Watermark gone. If you have access

25:00

to even a relatively weak unwatermarked

25:03

LLM, you can strip out the synth ID

25:05

watermarks by asking the LLM to

25:07

paraphrase the text content. Because the

25:09

watermark is inherent to subtle

25:10

vocabulary choices, rewarding the

25:12

content will remove that watermark. You

25:14

could even do it by hand, although at

25:16

that point it's not really AI generated

25:17

content anymore. Since there will be

25:19

some kind of free public watermark

25:20

testing tool, you can just keep tweaking

25:22

until it comes back negative. Yep,

25:25

that's the other problem. If it's easy

25:26

for us to check if text is generated or

25:28

not, that works on both sides. Both for

25:30

us reading the text and also for people

25:32

trying to generate AI text and get it

25:34

out there looking like normal text.

25:36

Moreover, the AI act requires

25:38

watermarking techniques to be

25:39

interoperable as far as it is

25:41

technically feasible. That means AI

25:43

writers will have to publish their

25:44

watermarking process and potentially

25:46

even attempt to standardize on applying

25:48

the same kinds of watermarks. The author

25:50

does not see how this is compatible with

25:51

the kind of security by obscurity that

25:53

LM text watermarking depends on. You

25:56

can't be obscure in how you implement it

25:58

to hide it, but also give all of these

26:00

tools. Like there's a have your cake and

26:02

eat it too. By the way, the cake's

26:03

invisible and all made up type thing

26:04

here. Then we get to C2PA, which is

26:06

honestly one of the stronger things in

26:09

all of this, but not because of how it

26:12

makes it easy to verify AI content.

26:14

We'll come back to C2PA in just a

26:16

moment. The AI acting code of practice

26:19

talks a lot about digitally signed

26:20

media. The idea here is that you can

26:22

include an AI disclosure in the files

26:24

metadata itself. Ideally in a way that

26:26

cannot be tampered with like signing a

26:28

hash of the files contents. The sign

26:30

metadata process is basically the CGPA

26:32

content credentials. While you can

26:33

remove CGPA metadata, you can't fake it.

26:36

So a file with created by human metadata

26:38

can be trusted and a file with no

26:39

metadata at all can be held in

26:41

suspicion. Okay, this is the part that I

26:44

actually want to bring up here. The

26:46

thing that makes CGPA strong isn't that

26:48

it's flagging AI gen content. It's that

26:51

it provides standards to identify human

26:53

generated content. And that's the only

26:55

direction that makes sense here,

26:57

especially for media. It's a lot easier

27:00

to sign something as really generated by

27:03

the chip on your camera with a photo you

27:05

took than it is to detect all things

27:07

that weren't taken on cameras. And the

27:09

more ways we can find to sign human

27:11

generated media, the easier this gets.

27:13

Because I really do think long-term

27:15

that's the only solution that's going to

27:17

work. Sadly, the CTPA created by human

27:20

markers are not going to solve anything

27:22

in text. It only really applies to

27:24

files. In the words of the code of

27:26

practice, that's quote a data format

27:28

that supports attaching metadata like

27:31

audio, image, video, or containerized

27:34

text. The output of chat tools in most

27:36

of the output of AI agents is not

27:38

containerized text but plain old regular

27:40

text so it cannot be signed. What would

27:42

it even look like to sign a chatgbt

27:44

output? There's no artifact to pass

27:45

around. I think it's fascinating

27:47

question whether cloud code has to CGPA

27:49

sign any HTML files or PDFs it generates

27:51

for you seems kind of tricky to get

27:53

right. But in any case, the Act also

27:55

mandates some kind of actual

27:56

watermarking as well. And the harsh

27:58

conclusion here is that no matter what

28:00

happens, technical users will be able to

28:02

strip out the watermark at will. There

28:04

will be a plethora of tools for

28:05

nontechnical users that will do the same

28:07

thing. We already have a bunch of repos

28:09

that are doing watermark removal here.

28:11

Unicode text hygiene, statistical

28:13

rewrite hooks, and CGPA metadata removal

28:15

from PNG, JPEG, SVG, PDF, DOCEX, HTML,

28:18

and MD. This is a real repo that already

28:21

exists to remove watermarks from AI

28:23

generated things. It's an agent skill

28:25

and a standard Python script to strip

28:27

multi- vendor AI providence marks from

28:29

text and files for privacy and hygiene

28:30

on content that you own. They have a

28:32

skill you can install if you want to

28:34

remove AI marks from things that are

28:36

generated by the same agent in your

28:38

codebase, by the way. So, you can

28:39

literally ask Claude to remove the

28:41

watermark from Claude's outputs, which

28:43

is hilarious. Yeah. To conclude on my

28:45

end, the issue here is that this is a

28:47

really shitty cat and mouse game where

28:50

the mouse has all of the advantages.

28:53

It's way too easy to work around these

28:55

things and it's way way too hard to do

28:59

this type of watermarking in a way that

29:00

isn't trivial to bypass. Like I I just

29:04

cannot fathom any method where you can

29:06

watermark text that isn't either

29:09

incredibly expensive to detect or

29:12

incredibly cheap to work around. And

29:14

that catch 22 is pretty brutal and I

29:17

just don't ever see it working out. All

29:20

this will ever do is catch the lowest

29:22

effort spammers. People setting up bots

29:25

that just use Claude or other models to

29:28

spam people on Twitter and whatnot. And

29:30

they can all move to cheap openweight

29:31

models or add really basic checks in

29:34

front. But this is not going to stop

29:36

anything. This is one of those things

29:38

where it like sounds really good and to

29:40

normal people, to voters, to

29:41

politicians, to all the people involved

29:43

in making these processes, it sounds

29:45

like a really good idea. We should be

29:47

able to detect what text was generated

29:49

by models. We should be able to detect

29:50

when an image was edited by an LLM. Like

29:52

obviously we want these things, but we

29:55

literally cannot do it. We have to think

29:58

back on what our goal is here. Is our

30:01

goal to make it slightly easier to

30:02

notice when a random person copy pastes

30:04

an output from chat GPT? Is our goal to

30:07

make it harder for students in high

30:09

school to cheat on essays and have chat

30:11

GBT generate them? Or is our goal to get

30:13

in front of people using AI to do

30:17

destructive campaigns across society,

30:19

like a government using it to spread

30:22

propaganda and generate a bunch of

30:24

messaging that aligns with them? I think

30:26

this will help a little bit with the

30:28

stopping high school cheaters, but I

30:30

don't think it's going to help with

30:31

almost anything else here because any

30:34

medium or even like slightly higher than

30:36

zeroeffort attempt to mislead people

30:39

with AI generated stuff will not be

30:41

stopped at all by this. And even the

30:43

high schoolers are going to find

30:44

workarounds if we're being real here.

30:46

So, I just I don't see this being

30:48

useful. I see this being one of those

30:50

things that sounds really good,

30:51

especially to normal people that are

30:53

getting tired of AI slop destroying

30:55

everything around them, but I just don't

30:57

perceive this as a thing that will work

31:01

for any real use cases. And I think that

31:04

we'll get over this watermarking phase

31:06

relatively quickly. What we're going to

31:08

have to do is start verifying when

31:10

images are real instead of fake and

31:12

educating the public about the fact that

31:14

there are many things that look real and

31:16

aren't. Unlike the image I have on my

31:19

screen now, which is clearly a real

31:20

image of my face on a leg. But generally

31:24

speaking, the education side and

31:27

flagging of human content is what we

31:30

need long term. That's all I have to say

31:32

on this one. I honestly wish that the

31:33

effort of these really smart people was

31:35

put into things that would actually

31:36

work. I know this is a good faith effort

31:38

to try and make it easier for the public

31:40

to adjust to a world with more and more

31:42

AI generated stuff, but these efforts

31:44

just aren't going to work. Even a very

31:46

basic cursory understanding of

31:48

engineering and watermarking techniques

31:49

makes that obviously the case. How do

31:52

you guys feel though? Is this a net

31:53

positive or is it just a waste of time?

31:55

Curious how you all feel about it. And

31:56

until next time, peace nerds.

Interactive Summary

The video discusses the EU's AI Act and its requirements for AI labs to implement watermarking and provenance metadata in AI-generated content. The speaker explores the technical challenges and limitations of these requirements, arguing that watermarking text and media is fundamentally flawed because it is trivial for users to strip or circumvent these markers using simple edits, compression, or paraphrasing. The speaker concludes that while the intent is to increase transparency and combat 'AI slop,' these technical solutions will likely only deter low-effort users and will not effectively stop bad actors or sophisticated misuse, suggesting that focusing on verifying human-generated content is a more viable long-term strategy.

Suggested questions

3 ready-made prompts