HomeVideos

Paste This Into Claude Code, Never Run Out Of Tokens Again

Now Playing

Paste This Into Claude Code, Never Run Out Of Tokens Again

Transcript

456 segments

0:00

Claude just told me to come back in 5

0:02

hours and I still have a lot of work to

0:05

do today. I didn't have it do anything

0:08

massive. I didn't even run it for 8

0:10

hours straight. So, I set myself out on

0:14

a journey to understand why I kept

0:16

reaching my session limit and how I can

0:18

stop that from happening ever again.

0:20

[music] And the crazy part is almost

0:23

none of my usage limit was because of

0:25

what I typed. 0.01% [music]

0:29

of it to be exact. That is just how

0:32

these tools work. And once you see the

0:34

mechanism behind it, you can cut most of

0:37

your costs down. So, in this video, I'm

0:40

going to give you a prompt you can paste

0:41

[music] into Claude code that audits

0:44

your setup, then the seven fixes that

0:47

actually move your token consumption

0:48

[music]

0:49

down. Starting with the one that costs

0:51

nothing and ending with a trap that is

0:53

probably [music] doubling your bill

0:55

right now. Let's get started. Here is

0:57

the best way for me to explain token

1:00

consumption. These models have no

1:02

memory. None. Absolutely none. So, every

1:05

time you hit enter, your entire

1:08

conversation gets wrapped up and sent

1:10

again from the top. Your first message

1:13

costs what you typed. Your second

1:15

message costs what you typed plus the

1:18

answer plus the first message. By

1:20

message 20, the thing you just wrote is

1:24

that silver box at the top. It's every

1:27

single thing here. Everything under it

1:30

is stuff you already paid for being paid

1:33

for again. That is why it compounds the

1:36

way that it does. A 3,000 token file

1:39

your agent reads at turn four of a

1:41

40-turn session does not cost you 3,000

1:45

tokens. It costs you 3,000 tokens 37

1:49

more times. I mean, look at this

1:51

messages line right here. That is not

1:54

what you asked. That is all of this

1:57

chat's history compounding over and over

2:01

and over again. And the reason that

2:03

matters is that your setup is not the

2:05

same as mine. So, let me show you how to

2:08

find your own version of this. This is

2:11

the prompt. I've included it in the

2:14

description for free. You don't have to

2:16

give me your email for it. Paste it into

2:18

Claude code and it will audit your

2:21

actual configuration. It reads your

2:23

context breakdown, checks whether tool

2:25

deferral is on, measures your memory

2:28

files, looks at your cache hit ratio,

2:31

and flags scheduled tasks that are

2:34

firing while you sleep. You can

2:35

screenshot this video right now and send

2:38

this to your AI agent. This will let you

2:40

know what is consuming the most amount

2:43

of tokens for you and which of these

2:45

next seven fixes matter the most for

2:48

you. Mine came back with these problems

2:52

and I'm fixing them in order of what

2:54

actually saves me the most. The highest

2:56

leverage thing you can do to cut costs,

3:00

costs you nothing and is five letters

3:02

long. When you finish a job and start

3:04

with a different one, use {slash}clear.

3:08

Do not keep going in the same thread

3:09

because it is it is convenient. That old

3:12

conversation is not sitting there

3:14

quietly. It is being resent on every

3:17

message you send until the session ends.

3:20

Watch this. Messages are at 80,000

3:22

tokens and if I

3:25

write {slash}clear

3:27

and hit enter and boom, just like that,

3:29

messages are back to 0%. And Anthropic's

3:33

own docs say it plainly. When you want a

3:36

fresh start instead of continuity,

3:39

{slash}clear costs you nothing. Here's

3:42

why this beats everything else in this

3:44

video. Every other fix I'm about to show

3:46

you reduces one component of your

3:49

context. Clearing resets the entire base

3:53

that all of those components are a

3:55

fraction of. 96%

3:58

of my spend was rereading history. This

4:01

is the one tool that deletes the

4:03

history. One thing before you go

4:05

clearing everything, use {slash} rename

4:08

inside your session first, so you can

4:11

use {slash} resume later, which allows

4:14

you to restore the session in case you

4:16

ever urgently needed it. You are not

4:19

throwing the work away, you're stopping

4:21

the next job from carrying it. And that

4:24

fix is free. The next one is the

4:26

opposite, because it is something you

4:28

are probably already doing on purpose to

4:30

save money. This is the one I want you

4:33

to actually remember. When you're

4:35

running low, what do you do? You

4:37

probably switch to a cheaper model. You

4:39

hit {slash} model, drop from Opus to

4:42

Sonnet, and you feel great about it.

4:44

That switch is the most expensive thing

4:47

you can do. Here's why. Your

4:49

conversation is cached, and cache reads

4:52

cost 1/10 of normal output. That is why

4:56

long sessions do not bankrupt you. But,

4:59

the model is part of the cache key.

5:02

Change the model, and none of your

5:03

history matches the cache anymore. So,

5:06

the entire conversation gets reprocessed

5:09

at full price. On Opus 5, at 200,000

5:14

tokens of context, that turns a 10-cent

5:17

turn into a 1-dollar turn. 10 times more

5:21

expensive. It's absolutely invisible,

5:23

and you did it to yourself while trying

5:25

to save money. Effort level does the

5:28

same thing, by the way. Fast mode does

5:30

the same thing, as well. So, things that

5:32

break it, switching models, changing

5:35

effort, turning on fast mode, connecting

5:38

or disconnecting an MCP server if your

5:40

tools load up front, enabling a plugin

5:43

that ships an MCP server, and

5:45

compacting. Also, and this one is nasty,

5:49

upgrading cloud code and then resuming a

5:51

long session. Anthropic's docs literally

5:54

call that the most expensive request you

5:57

will send. Things that are safe, editing

5:59

files in your repo, editing your memory

6:02

file, changing output style, changing

6:04

permission mode, invoking skills and

6:06

commands, recaps, rewinds, and spawning

6:09

a sub agent. So, the rule is simple.

6:12

Pick your model and your effort at the

6:14

start of the session and then leave

6:16

those settings alone. If you want to be

6:18

on a cheaper model, start there. So,

6:21

that is the expensive keystroke. The

6:23

next three fixes, however, aren't

6:25

something you've ever typed. Yet, they

6:27

burn away at your context. Picture this,

6:30

you type in install remotion for me.

6:33

Your agent goes off and runs the actual

6:35

commands and 800 lines come back.

6:38

Package names, version numbers,

6:40

warnings, a funding message. You did not

6:43

read a single one of them. You just

6:45

wanted to know that it actually worked

6:47

and it installed the thing you asked it

6:49

for. But, your agent does not get to

6:51

skim. All 800 lines went into your

6:55

conversation and you pay for them again

6:57

on every message until you clear. So,

7:01

put a filter in front of it and you do

7:04

not write it this next part. You asked

7:07

for it. You can give it a prompt like

7:09

the one I'm showing on the screen as I'm

7:11

talking right now. It will create one

7:13

small file that sits between your agent

7:16

and the command and it cuts the output

7:18

down before your agent ever sees it.

7:20

Your agent writes it, your agent

7:22

installs it, and from then on it just

7:25

works. Anthropic ships a working version

7:28

of this so your agent has something to

7:30

copy from. Their online on it is

7:33

reducing context from tens of thousands

7:36

of tokens to hundreds. You do this once

7:39

and it works on every session after

7:41

that. And this fix is output coming in.

7:44

The next fix is what is already sitting

7:47

in your context before you have ever

7:49

typed anything. You connect Gmail, then

7:52

Notion, then Slack, and each takes you a

7:55

single prompt to install, and it feels

7:58

free, but it's not free. Every tool you

8:01

connect comes with an with an

8:02

instruction manual, what it can do, what

8:05

to send, what comes back, and your agent

8:07

has to read that manual before allowed

8:10

to touch the tool. GitHub on its own,

8:13

for example, costs you 26,000 tokens.

8:16

Slack is 21,000 tokens. All of that gets

8:20

loaded into every single session before

8:22

you type a single word. Now, here's the

8:25

good news. Claude has released an update

8:27

that makes sure your agent does not read

8:29

every manual anymore. It loads the

8:32

contents page, and it only opens the

8:34

section it actually needs when it needs

8:37

it. You still have access to the same

8:39

tools, but that cost has went down 85%.

8:43

That part is handled for you. It is on

8:45

by default, and you do not have to do

8:47

anything. But, you are still paying for

8:50

the contents page. And the contents page

8:53

grows every single time you connect

8:55

something new. So, here's the fix, and

8:58

it takes about 30 seconds. Type in

9:01

And just with that, you will see a panel

9:04

of every single tool you have ever

9:06

connected in one list

9:10

with a switch next to each one. Go down

9:14

in that list and turn off anything you

9:16

have not used in the last like month.

9:18

You're not deleting it, it stays set up,

9:21

it just stops loading. One last thing to

9:23

check while you are here, run /context,

9:29

and find the tools line when it loads

9:33

up. So, we'll go ahead and expand, and

9:35

there we go. We have system tools and

9:38

system tools right here. If it says

9:41

deferred, you are on the new behavior

9:44

and the manuals are staying shut when

9:46

you don't use them. So, as you can see,

9:47

it's only using up 17,000

9:50

context tokens and 20,000 of the active

9:53

tokens of skills I am currently using in

9:56

this session right now. And the good

9:58

news is turning tools off mid-session

10:01

does not cost you anything. As long as

10:03

that line says deferred, connecting and

10:06

disconnecting just appends. It does not

10:08

rebuild your cache the way switching

10:10

models does. And the next thing

10:12

everybody tells you to do is delegate to

10:15

sub-agents. So, I want to be honest

10:17

about that because what you have been

10:20

told is sort of half true.

10:23

Everybody says sub-agents save tokens,

10:27

but they do not. They move them. Here

10:29

are Anthropic's own numbers from their

10:32

documentation. So, over here we have a

10:34

simulation of how context windows work.

10:37

And if I go ahead and go through the

10:40

simulation, we get a prompt like use a

10:42

sub-agent to research this session. And

10:44

if you look at the context window, when

10:46

we hit send, boom, just like that, the

10:49

sub-agent has spent very few tokens

10:52

because what we see is a sub-agent reads

10:55

about 6,000 tokens of files and what

10:57

comes back to your main context is a 420

11:01

token summary. That looks like a massive

11:04

win. And in your main window, it is.

11:06

But, do the math on the whole thing.

11:08

That sub-agent loaded its own system

11:11

prompt, its own copy of your memory

11:13

file, its own tools, and then did the

11:16

reading. It burned roughly 9,800

11:19

tokens to save you 5,700.

11:22

In isolation, you lost. And Anthropic is

11:25

blunt about this elsewhere. Their

11:28

multi-agent research post says agents

11:31

use around four times more tokens than

11:34

chat and multi-agent systems about 15

11:38

times more tokens than chat. But, when

11:42

is using sub-agents actually worth it?

11:44

The answer is when three things are

11:47

true. The output is high volume, you

11:50

will not need the detail again, and the

11:52

session is going to continue for many

11:54

more turns. And that third one is the

11:57

whole game because those 5,700 tokens

12:00

you avoided would have been resent on

12:02

every remaining turn. But, if you

12:05

delegate and then immediately end the

12:07

session, you just paid extra for

12:09

nothing. One free upgrade if you do

12:11

this, set the sub-agents model to Haiku.

12:14

That is a five times reduction on the

12:17

isolated work and it does not touch your

12:19

main session's cache. And that brings me

12:22

to picking models properly, where

12:24

there's a myth I want to kill. Half of

12:27

what you ask for is small. Rename these

12:29

files, write me a commit message, clean

12:32

up this list. The useful heuristic is to

12:34

use the dumbest model that will still

12:37

finish the job. But, remember fix two.

12:39

Pick the right model at the start of the

12:42

session. Switching mid-session costs you

12:45

more than the saving. The better way to

12:47

do this is per skill and per sub-agent.

12:49

So, you can run Haiku for the grunt work

12:52

without ever touching your main session

12:54

model and without invalidating anything.

12:57

Every fix so far assumes you are sitting

13:00

at the keyboard. The next one is what

13:02

happens when you're not. This is the

13:05

trap I promised you at the start. A

13:07

scheduled task fires on its interval

13:09

whenever you are there or not. And every

13:13

time it fires, it sends your full

13:15

context, not a bit of it, all of it. So,

13:18

if that task is attached to a bloated

13:21

session, you're paying for that entire

13:23

context on every fire forever at 3:00 in

13:27

the morning, while you're asleep. Now,

13:29

here is the part that turns it from

13:31

expensive to painful. Your cache

13:33

expires. On a subscription, it lasts 1

13:36

hour. So, if your task runs less often

13:39

than once an hour, every single fire

13:42

misses the cache and reprocesses your

13:44

whole context at full price, instead of

13:47

the 1/10 cache price. 10 times the cost

13:51

on a schedule, forever. So, how often

13:55

your task runs is a real cost setting.

13:58

If your task can run every 45 minutes

14:00

instead of every hour or 2 hours, it is

14:03

cheaper to run it more often. And while

14:06

I'm here, a correction on the opposite

14:08

claim. People say leaving Claude code

14:10

open in the background burns your

14:12

limits. Anthropic's documents background

14:15

usage at under 4 cents a session. That

14:18

is not your problem. Your scheduled

14:20

tasks are your problem, and your live

14:23

agent teams, because each one keeps

14:25

consuming until it exits. That is the

14:28

one that doubles your bill without you

14:30

touching the keyboard even. And those

14:33

are the seven that burn your token

14:35

consumption number. So, let me clear out

14:38

the advice that does not, because some

14:40

of it is actively wrong. All right,

14:42

let's talk about what does not work.

14:44

Number one, writing shorter prompts. In

14:47

my logs, everything I actually type came

14:50

to 0.01%

14:53

of the bill. Your prompt length is a

14:55

rounding error. Vague prompts do cost

14:58

you, but through the file reads and the

15:01

rework they trigger, not through length.

15:04

Number two, compacting to save tokens.

15:06

This one is backwards. To write you a

15:08

summary, it has to send your entire

15:11

conversation one more time. So, the

15:13

thing you did to save money is the

15:15

single most expensive message of the

15:18

session. And then it wipes your cache on

15:20

purpose, because the conversation is

15:22

summarized is no longer exists. Clearing

15:25

is free. Compaction buys you continuity,

15:29

not savings. And if you only want to

15:32

undo a few back turns, use {slash}

15:34

rewind instead. Rewind takes you back to

15:37

a point your cash already knows, so

15:39

nothing has to be reread. Number three,

15:42

screenshotting text to save tokens. A

15:44

picture is not cheaper than the words in

15:46

it. On Opus 5, one screenshot of your

15:50

screen costs you about 2,700 tokens. A

15:53

4K one is nearly 5,000. That is a lot of

15:58

text. Paste the text instead. It is

16:00

cheaper and your agent can actually edit

16:03

it. It cannot edit a picture though. And

16:05

finally, PDFs, while I'm on format.

16:08

Every page of a PDF costs you between

16:11

1,500 and 3,000 tokens just for the

16:15

words on it. And then your agent also

16:18

takes a picture of that page and you pay

16:20

for that picture, too. You basically pay

16:23

for every page twice. Ask your agent to

16:26

turn it into a plain text file first and

16:29

the same document costs you about a

16:31

quarter as much. And there is one last

16:34

thing you need if any of this is going

16:36

to stick. Four things, three of them are

16:38

already sitting in your terminal and you

16:40

have probably never opened them. {slash}

16:43

context. This is the one you're going to

16:46

use the most. It shows you what is in

16:48

your window right now, line by line, so

16:51

you can see exactly what is taking up

16:53

the room. Then we have {slash} usage.

16:57

This one is underrated. It shows you how

17:00

much of your plan you have burned

17:02

through. And then it tells you what

17:04

burned it, not roughly. It names the

17:06

specific skill, the specific tool, the

17:08

specific agent. So if one thing on your

17:11

machine is quietly eating your limits,

17:14

this is where it confesses. Then inside

17:16

of that, you have what is this one

17:19

session costing you? The cost or slash

17:21

cost. And how much of it was your agent

17:25

re-reading history versus doing new

17:28

work. And the fourth one is the burn

17:31

rate meter that is in the corner of your

17:33

cloth code screen. Having a number

17:36

moving while you work changes your

17:38

behavior more than any rule I have given

17:41

you. And if you want the number I opened

17:43

this video with, it is sitting on your

17:45

machine right now. Every session you

17:47

have ever run is logged in a folder. And

17:50

every single reply in there records what

17:53

it costs. So, ask your sub agent to go

17:55

read them and work out your own

17:57

percentage. That is literally what I

17:59

did. And then, you have your number

18:02

instead of mine. You did not run out of

18:04

tokens because you asked too much. You

18:07

ran out because almost everything you

18:09

paid for was your agent re-reading

18:12

things you already sent, and nobody ever

18:14

showed you where the switch was. So, now

18:17

you know. Clear between jobs because it

18:19

is free and it resets the base

18:21

everything else is a fraction of. Pick

18:23

your model and effort once and leave

18:26

them alone because the thing you do when

18:28

you are trying to save money is the most

18:30

expensive move on the board. Filter your

18:33

tool output before it lands. Disconnect

18:35

the tools you never use and check that

18:38

line in context stays deferred. Delegate

18:41

when the session has a long way left to

18:43

run and not when it does not. And go

18:46

look at what your scheduled tasks are

18:48

doing at 3:00 in the morning. One honest

18:51

note to finish on, the labs are not

18:54

going to fix this for you. Not because

18:56

they're being difficult, but because

18:58

they're not graded on how few tokens you

19:02

use. It is your desk. You have to keep

19:05

it clean. The token usage audit prompt

19:07

is in the description and I would run it

19:10

once a week or whenever you feel you've

19:12

been burning through a lot of your

19:14

tokens fast. Things drift, you add a

19:16

server, you install a plugin, you change

19:18

a setting, and 6 weeks later you are

19:21

back where you started wondering why you

19:23

are hitting limits again. Comment below

19:25

with what shocked you the most in this

19:27

video. And if you enjoyed this video,

19:29

make sure to leave a like. And if you're

19:31

new to my channel, then subscribe

19:33

because I have a ton more content like

19:35

this coming your way. And oh, what would

19:38

you know? Over here, the algorithm gods

19:41

seem to think that you will really enjoy

19:43

this video. So, click it, and I'll see

19:45

you there.

Interactive Summary

Este vídeo explica por qué los usuarios de Claude Code agotan rápidamente sus límites de tokens, revelando que la mayor parte del consumo proviene de la relectura constante del historial de chat, no de lo que el usuario escribe. El autor proporciona una guía práctica para auditar la configuración de Claude Code y detalla siete estrategias clave para optimizar el consumo de tokens, incluyendo el uso frecuente de /clear, mantener configuraciones constantes para no romper el caché, filtrar salidas de herramientas innecesarias, gestionar los conectores de herramientas, usar sub-agentes solo cuando sea necesario, y optimizar las tareas programadas que se ejecutan en segundo plano.

Suggested questions

4 ready-made prompts