HomeVideos

Claude's Invisible Watermark - Everything You Need to Know (in ~7 mins)

Now Playing

Claude's Invisible Watermark - Everything You Need to Know (in ~7 mins)

Transcript

243 segments

0:00

Claude is about to watermark the text

0:01

you generate with AI. So today I'll

0:03

share with you everything you need to

0:04

know and what you can do about it. If

0:06

you're new, my name is Jay. I've been in

0:07

AI since my masters in data science and

0:09

now I'm running my own AI business and

0:11

one of the largest communities in the

0:12

space globally. Let's get straight to

0:14

it. So here's the four things that I

0:15

want to cover. When will this watermark

0:17

start to happen? Who will it affect? How

0:19

will the watermark actually be

0:21

implemented? And what you should be

0:23

doing about it. So when is this

0:24

happening? Well, the short of it is that

0:26

this watermark will be applied by models

0:28

launched after August 2. So since

0:30

Entropic hasn't launched a model yet

0:32

since that date, that means it's not yet

0:34

live. But say when Fable 5.1 or Opus 5.1

0:37

is launched, then they will most likely

0:39

start applying that watermark. Now, just

0:41

a note regarding the older models, as

0:43

per Entropics official documentation,

0:45

they did mention that they have plans to

0:47

rework the older models so that they

0:49

will also add watermarking for them as

0:51

well. But as of right now, that is also

0:54

not yet live. Now, who is this going to

0:55

affect? Now, because Entropic is

0:57

applying the watermark because of the EU

0:59

AI act, some people might think that

1:00

this will only apply to Europe. But

1:02

remember, since the watermark will be

1:04

applied at the model level, it's much

1:05

more likely that everyone globally will

1:08

have this watermark. Now, apart from the

1:09

scope, I think what's also important to

1:11

realize is that it is not just entropic

1:13

who signed this EU AI act. So if you

1:15

look at the European Commission's

1:16

official article here published around

1:18

the end of July, they mention here that

1:20

by the end of that month around 190

1:22

organizations already signed this act

1:24

and some big names here include of

1:26

course Entropic, Google, Meta, Microsoft

1:28

and OpenAI. So right now Claude is

1:31

making the news with regard to this

1:32

watermark. But pretty soon it's likely

1:34

that all of these AI labs will implement

1:35

some sort of watermarking for the text

1:37

that they generate as well. Now how

1:39

exactly is this invisible watermark

1:41

going to be applied? Now, this can get

1:42

technical depending on how much detail

1:44

you want, but I do think it's still

1:45

useful to get an idea of how it works.

1:48

And the best reference for this is this

1:49

official FAQ page from Entropic, which

1:51

I've also read true. And what they made

1:53

sure to clarify here is that it's not

1:55

going to be a simple hidden character

1:57

watermark that some people might think.

1:59

So, it's definitely not special

2:00

characters or extra spaces because

2:02

obviously that would be pretty easy to

2:04

strip as a watermark if they were using

2:06

that. So rather the method that they're

2:07

using is more similar to a pattern-based

2:10

watermarking is how I like to describe

2:12

it. Because if you go back to Entropic's

2:13

documentation here, they're saying their

2:15

exact methodology is a version of the

2:17

scent ID text approach which is

2:19

originally from Google. And this scent

2:21

ID watermark is actually simpler to

2:23

understand when you draw a parallel on

2:25

how it's applied for images. Because for

2:27

images, let's say you have this picture

2:28

of a landscape. All that image really is

2:31

is a collection of pixels, right? So

2:33

really small squares that if you zoom in

2:35

you'll be able to see them. So what the

2:36

synth ID watermark does is that it

2:38

nudges some of those pixels so that

2:40

their color is slightly altered but the

2:43

change in the color is so tiny that only

2:45

computers can really detect it but the

2:47

human eye cannot. But since the

2:48

watermark is embedded in the pixels

2:50

themselves even if let's say you copy

2:52

this image and you paste it elsewhere

2:54

then that watermark will travel with it.

2:56

And so the send ID for text actually

2:58

follows a similar principle where let's

3:00

say you have an essay written by Claude.

3:02

What it will do is it will nudge some of

3:04

its word choices towards a certain

3:06

direction to apply the pattern-based

3:08

watermark. And a good example of this is

3:10

this illustration by Tariq, who's one of

3:12

the more well-known engineers at

3:13

Entropic. And he's basically showing

3:15

here how the text will differ, where the

3:17

watermark is applied versus text where

3:19

it isn't implemented. And if we take one

3:21

example there just to make this as

3:22

simple as it can be. Whenever you send a

3:24

prompt to claude like this question on

3:26

what is Isaac Newton's most important

3:27

book when it writes a response like this

3:29

what it actually does under the hood is

3:31

just list down the most probable words

3:33

in response to that question. And in

3:35

this illustration the response without

3:37

the watermark says that Isaac Newton's

3:39

best known work is the Principia while

3:41

the one with the watermark says Isaac

3:43

Newton's most famous work is the

3:44

Principia. So practically means the same

3:46

thing. But the difference here in the

3:48

watermark version is how those words are

3:50

selected. And when it comes to word

3:51

selection, Entropic mentions this useful

3:53

analogy where if you imagine you're

3:55

playing a game like Monopoly where your

3:57

next turn or in this case, your next

3:59

word choice is decided by the rolling of

4:01

a dice. They're saying that instead of

4:03

rolling the dice to get this randomness,

4:05

we decided to just use the digits of pi.

4:07

So if you go back to this example, when

4:08

it came to this word choice, let's say

4:10

that dice landed on a one and that

4:12

corresponded to the word best known. And

4:14

so that was what was used in the final

4:16

response. But in this response with the

4:18

watermark, the die will look something

4:20

more like this where right now we're

4:21

just using pi as an example. But

4:23

obviously we don't know the exact key

4:25

that entropic will be using. But what

4:27

entropic is saying is if you randomly

4:29

select from these numbers and say it

4:31

lands on five and that corresponds to

4:34

most famous as the phrase which is the

4:36

one that was used in the final watermark

4:38

text. Then for all intents and purposes

4:40

as per them, it's still random. The

4:42

meaning is still supposedly maintained,

4:44

but the difference is you can reverse

4:46

engineer if the randomizer that was used

4:48

was just this ordinary die or if it was

4:51

Entropics watermark key. So because of

4:53

this methodology for the watermark,

4:55

there are some valid questions that

4:57

people are asking. And a big one is will

4:59

this affect word and code quality? Now

5:01

the real answer to that is we don't yet

5:03

know for sure until it is live. But it

5:05

is useful to know that in the original

5:07

synth ID text paper which is the

5:09

inspiration of entropic when it comes to

5:11

this method that Google already tested

5:12

this with their users and they found

5:14

that people weren't really able to

5:16

notice the difference. In that same

5:18

article, they also included this section

5:20

around code. And what Entropic is saying

5:22

here is that the watermarking takes

5:23

advantage of decisions where either

5:26

choice of a word would be equally good

5:27

so that the meaning of the whole thing

5:29

doesn't change. But where an exact

5:30

output is required, which is more common

5:32

when it comes to coding, then the

5:34

watermark isn't applied. So for example,

5:36

if the model has written 2 plus 2

5:38

equals, then it will just say four as

5:40

the answer and the nudge of the

5:41

watermark wouldn't be applied for this

5:43

case. And so for the same reason code

5:45

which in many cases has to be exact has

5:47

generally less watermarking than other

5:49

forms of text. But again the watermark

5:51

is not yet live and so we're just basing

5:53

this from the article that entropic has

5:55

published. Another scenario is let's say

5:56

if you write an essay and then you pass

5:58

it along to AI to edit it is it now then

6:01

watermark. So for this one it really

6:02

depends on how much change you ask the

6:04

AI to make. Because for example, if you

6:05

just have it add punctuations or just

6:07

fix the formatting of your essay to

6:09

capitalize some letters, then claude

6:11

won't really have the space to change

6:12

the words, right? And apply those

6:14

watermarks. But if you ask it to

6:15

paraphrase the whole thing, then it can

6:17

change the words and it will be able to

6:18

apply that pattern-based watermark that

6:20

we talked about. And then on a similar

6:22

note, if you want to know how to remove

6:24

the watermark, the basic principle is

6:26

the more that you paraphrase a text that

6:28

was given to you by AI, then the more

6:30

likely it is that you are erasing quote

6:32

unquote the watermark. So what is it now

6:34

that we should do? Well, first of all,

6:36

watching this video and just being aware

6:37

of this is already a good step for you.

6:39

I probably wouldn't generally advise

6:41

people to switch away from cloud just

6:43

for this one exact reason. You can

6:45

switch away from cloud because of other

6:46

reasons, but because other AI labs will

6:48

likely implement something like this,

6:50

then overhauling your systems away from

6:52

cloud just because of this one news is

6:54

probably not going to be the best move

6:55

for you. It's also important to note

6:57

that no cloud model as of the time of

6:59

this recording has the watermark yet.

7:01

And there's also no tool yet to detect

7:03

if a given set of text has that

7:04

watermark or not. Although Entropic did

7:06

mention that they will launch a tool

7:07

like that sometime in the near future.

7:09

But if it's really important for you not

7:11

to be accused of using AI generated text

7:14

for your work, then one of the best

7:16

things you can do is to review your

7:18

agents work, which is probably good

7:19

practice regardless of the watermark

7:21

existing or not. I hope that was

7:23

informative and if it is, then consider

7:24

subscribing because that helps me a lot

7:26

to put out more educational content like

7:27

this. And I'll see you all next time.

7:28

Thanks.

Interactive Summary

Claude is preparing to implement an invisible, pattern-based watermarking system for text generated by its AI models, following the EU AI Act. This method, inspired by Google's SynthID, subtly alters word choices without changing the meaning of the content, making it detectable by computers but not by humans. The watermark will be applied to models released after August 2, with plans to update older models. While this ensures accountability for AI-generated content, it does not apply to scenarios requiring exact output, such as coding, and does not impact users' ability to paraphrase text to remove the watermark. Users are advised that this shift will likely be adopted across the industry and that reviewing AI-generated work remains the best practice.

Suggested questions

3 ready-made prompts