HomeVideos

I spent $1,486 on Fable tokens so you don't have to

Now Playing

I spent $1,486 on Fable tokens so you don't have to

Transcript

472 segments

0:00

So, as you know, Fable has been

0:01

miraculously returned to us. Well, in

0:03

the first 4 hours since its release, I

0:05

spent over $1,400. In the last 24 hours,

0:07

I spent another around $1,000. And I

0:10

didn't do this just because I love

0:11

pissing money away, but because I wanted

0:13

to figure out what the best and optimal

0:16

token reduction and usage strategies

0:18

were because I anticipate Fable and

0:20

other extremely high intelligence models

0:22

are likely to continue being computed

0:24

for at least some period of time. What I

0:26

want to do in this video is I just want

0:27

to show you all of those significant

0:29

token reduction methods. These token

0:31

reduction methods have next to zero

0:33

impact on your quality, but can reduce

0:35

the total token and usage consumption by

0:37

at least 50% if not more in some cases.

0:40

The first is a strategy called RTK or

0:42

Rust token killer. What this does is it

0:44

takes all of the tool inputs and outputs

0:46

of cloud code and it minifies and

0:49

reduces anything that is not explicitly

0:51

necessary to the models function. Now,

0:53

if that means nothing to you, never

0:54

fear. I have a demo. I want to show you

0:56

guys what a tool call like Claude would

0:58

be making internally looks like with RTK

1:00

versus without RTK. For those of you

1:02

guys that don't know, Claude typically

1:03

has its own little internal terminal and

1:05

uses that terminal in order to send and

1:07

receive requests. And so this is just a

1:09

brief demo of what a request would look

1:10

like without RTK, sort of like the

1:13

vanilla, and then with RTK down here.

1:16

Okay, so without RTK, pretending this is

1:18

just some internal tool call. You can

1:19

see we're repeating a lot of

1:20

information. standard out, standard out,

1:22

standard out, standard out, setting up

1:24

fixtures, hydrating mocks, right? We're

1:26

just saying the same thing over and over

1:28

and over again. Unfortunately, Claude

1:29

doesn't know the difference and neither

1:30

does your token build. So, when you send

1:32

these uh internal tool calls, which

1:34

Claude constantly has to do, it's

1:35

sending hundreds if not thousands in a

1:37

typical session. It typically has to

1:38

read a lot of irrelevant same

1:40

information. And obviously, that's an

1:41

inefficiency that we can prune out. So

1:43

what RTK does is it takes all of this

1:45

data and then it applies a new format to

1:48

it which gives you all of the same

1:50

information which Cloud can use in order

1:52

to do whatever the heck you want it to

1:53

do but in significantly fewer wasted

1:55

tokens. And so instead of spending let's

1:58

say 612 lines on this tool call okay cuz

2:01

this is truncated we are only spending

2:03

four lines. Instead of spending 36,700

2:07

characters we're only spending 177. And

2:10

the difference pre and post RTK is

2:12

literally a 99% reduction in token usage

2:15

for the same thing. Now, not all tool

2:17

calls are going to look like this. Some

2:18

of them are not inefficient, but a vast

2:20

majority of them are less efficient than

2:22

they could be. And so, realistically,

2:23

you're probably going to save somewhere

2:24

between 30 to 50%. The second is called

2:27

semantic compression, which is

2:28

essentially just taking a sentence,

2:30

okay, and then rewriting that sentence

2:32

in as few words as humanly possible

2:34

without removing the meaning. If that

2:36

still doesn't make any sense to you

2:38

guys, I want to show you guys a demo

2:39

claude.mmd. Over here, we have just some

2:42

system prompt that somebody's created.

2:44

Maybe they voice noted it in, so it's a

2:46

little less efficient than it normally

2:47

would be. Project instructions and

2:49

guidelines for the AI assistant. Hello,

2:52

thank you so much for helping out with

2:53

this project. We really appreciate all

2:54

the work that you do. This document

2:56

contains all of the important

2:57

instructions, guidelines, rules, and

2:59

conventions that we would like you to

3:00

please follow whenever you are working

3:01

on any part of the codebase. Now, this

3:04

is a learned skill, but essentially

3:06

everything up here has zeroformational

3:08

value. Um, if you wanted to, you could

3:11

rewrite this exact same paragraph and

3:13

instead of saying hello, thank you, and

3:15

so on and so on and so forth, you could

3:17

probably just say project instructions

3:19

and guidelines for the AI system, and it

3:21

would have the same semantic value. And

3:23

so, what this is is this is essentially

3:25

taking all of your system prompts, all

3:27

of the memory files, and everything else

3:29

in your cloud. MD and then system

3:32

context and just compressing it by

3:34

removing things that are superfluous,

3:35

things that are unnecessary. Again, if

3:37

you guys want to see what this looks

3:38

like in practice, I have a brief demo up

3:40

here where I'm actually going to spin up

3:41

a cloud instance. Okay, this cloud

3:44

instance is running sonnet 5. What it's

3:46

doing is it's going to read through the

3:47

cloud that I just showed you and then

3:49

rewrite it with maximum information

3:51

density, which is essentially semantic

3:53

compression. It's then going to show you

3:55

the before and the after of the word

3:57

count. And you can see that before we

3:59

had 865 words, most of them junk. After

4:02

we had 211 estimated token-wise, we went

4:05

from 1,125

4:06

to 274. And you can apply this approach

4:10

all across your projects to

4:12

significantly reduce both token

4:13

consumption, but then also improve

4:15

brevity and then improve the quality of

4:16

the output. I do this on every single

4:18

project. Also, if you guys want any of

4:19

these, just check the description down

4:21

below. I'm giving them all to you for

4:22

free. The third is logs to SQL Lite. Um,

4:25

this doesn't apply for every single use

4:26

case because Claude will not always be

4:28

reading logs for a lot of your own

4:30

projects, but it's a useful hack in

4:31

situations where Claude is for whatever

4:33

reason trying to read a log file. And

4:35

the logic here is similar to Rust token

4:37

killer. Instead of actually having to

4:38

pour through a massive log file to look

4:40

for what you want, we just use a highly

4:42

compressed form of that SQL light, which

4:45

is a database which contains like

4:47

location and place information and a

4:48

bunch of stuff. Then we abstract away

4:50

the the search function for a simple

4:53

command that will send to the database

4:54

that will do it for us. If I open up

4:56

this brief demo over here, which is

4:57

actually logging claude code, you can

5:00

see that a customer says checkout failed

5:01

this morning. Find the root cause in the

5:04

app.log. Important, it's 5,000 lines. Do

5:07

not read that log directly. Instead, use

5:10

scripts/logq.py

5:12

with no arguments to see usage. Report

5:14

the root cause of the affected line

5:16

order and even cite the line numbers.

5:18

Okay. And so previously we would have

5:19

literally had to treat this like a giant

5:21

piece of text and read every single

5:22

line. Instead, what this is doing is

5:24

it's just calling our database which is

5:26

set up in SQLite. This is extremely easy

5:29

and extremely simple. And you can see we

5:31

actually have if I just zoom in here and

5:33

uh maybe move my fat head a little bit

5:34

out of the way, you can see the

5:36

timeline, the specific line that every

5:38

one of these logs is in. Um and then all

5:40

the information that you need. You can

5:41

apply the same logic to any sort of

5:43

database, whether it's a Google sheet,

5:45

whether it's a CSV, anything like that.

5:47

Make sure that you're not actually

5:48

searching through it just using text.

5:50

Make sure that you're using some sort of

5:51

database function to do the filtering

5:53

for you. The next strategy is to block

5:54

huge reads. To make a long story short,

5:56

there are some resources that are just

5:58

very long and not all resources need to

6:00

be read start to finish. And so instead

6:02

of actually doing the entire read,

6:04

similar to what we did with SQLite, um

6:06

we're passing it off to a search

6:07

function which will do most of the

6:09

procedural heavy lifting for us and then

6:10

just rely on claude or you know open ais

6:13

or codeex's intelligence to set up the

6:16

filters correctly. Here's a brief

6:17

example of what that would look like. If

6:19

I set up Claude and then tell it to read

6:21

a massive dump of data and telling me if

6:24

there's anything wrong with it, notice

6:26

how it'll now say the file is 618

6:29

kilobytes or 20,000 lines, too big to

6:31

read in one shot safely. Let me sample

6:33

it to understand its structure first.

6:35

then instead of actually reading the

6:36

entire thing, okay, which is right over

6:38

here, what it'll do is it will just read

6:40

the beginning section, maybe the end

6:42

section, and then sort of apply its own

6:44

logic to see, okay, can I match any of

6:45

these patterns to potentially search

6:47

through this data way more efficiently?

6:49

Then if we scroll down here, rather

6:51

than, you know, creating this massive

6:53

massive burden on us token-wise, instead

6:55

of actually reading all 20,000 lines,

6:57

what we do is we essentially just like

6:59

make one small function sed- and then we

7:02

actually give it the exact index of

7:04

where it wants to look and then it saves

7:06

us all of that. So, you know, instead of

7:07

reading 20,000 lines, maybe we only read

7:09

20 or 30. This is massive efficiency

7:11

savings on tool calls that are extremely

7:13

long resources. You're not going to do

7:15

this like every single tool call of

7:17

course, but in the tool calls where you

7:18

do read large resources, you will say

7:19

like 99% of that. The next is pretty

7:21

simple and most of you will already be

7:23

doing this, but I'd highly recommend

7:24

prompting in English. English is just

7:26

significantly more of a information

7:27

dense language than something like

7:29

Japanese, French, German. And by doing

7:32

so, you'll typically get somewhere

7:33

between a 20 to potentially 80%

7:35

reduction in total token usage or more.

7:38

Should also note that most of these

7:39

models are trained on English. Um,

7:41

interestingly enough, if you're using

7:42

the Chinese classes of models, those um,

7:45

also work really well in English. But,

7:47

for instance, Mandarin is actually

7:48

extremely efficient token-wise because

7:50

most of the language is symbolic. If I

7:52

were to show you guys a brief example of

7:54

the total number of tokens used for a

7:56

simple query in English versus Italian

7:59

versus German and versus Japanese, this

8:00

is what we got. As you can see, English,

8:03

you know, making some simple query only

8:05

consumes 51 tokens. In Italian, it's 87.

8:08

In German, it's 118. 2.31x and in

8:11

Japanese it's something like 74 1.45x.

8:14

So this one's a little bit less uh

8:16

prescriptiony. Obviously most of you

8:17

guys will already be doing it in

8:18

English, but I just want you guys to

8:19

keep that in mind if you do end up doing

8:20

projects in other languages. The next

8:22

hack is embedding some form of context

8:24

frugality into your system prompt. We

8:26

already showed you guys how to optimize

8:27

system prompts by semantic compression.

8:30

Taking a sentence that means x and then

8:32

making the sentence itself shorter but

8:34

keeping the meaning as x. This is

8:36

similar. It's just we actually are

8:37

inserting a highle rule where our model

8:40

will only ever look to a resource if we

8:42

explicitly specify it. Now I will note

8:44

that this is one of the uh tactics or

8:46

tweaks that can reduce quality

8:47

significantly. You know Fable and these

8:49

other models already have pretty solid

8:51

built-in look schemes. They sort of

8:53

understand where to look through

8:54

statistical pattern matching. Um so what

8:56

this will do is it'll basically make it

8:57

harder for it to find the thing that you

8:59

want it to unless you're very specific.

9:01

But I'm imagining or assuming you guys

9:03

are willing to be more specific for

9:05

tradeoff of better usage. So I have a

9:07

demo over here called frugal-claw MD. If

9:09

I zoom way in, you could see that we say

9:11

customers are reporting that a 10%

9:13

coupon takes off 10 times too much at

9:15

checkout. Find the bug and propose the

9:17

oneline fix. Now I should note that

9:19

inside of our context window, we've

9:22

actually inserted a claw.md that says

9:24

read only files directly relevant to the

9:26

task. Ask before expanding the scope

9:28

beyond three files. prefer glob or GP to

9:31

locate then read the region not the

9:33

entire directory and summarize findings

9:36

so far before deciding to read more just

9:38

in case we you know change our mind

9:40

halfway through or something. Also never

9:42

generated files, block files or

9:44

fixtures. Those should not be read

9:46

unless explicitly asked. If we go back

9:48

to the demo, you can see that

9:49

essentially what we've done is we've

9:51

just applied GP. We've applied targeted

9:53

searches as opposed to again reading

9:55

through the entire thing. And you'll

9:56

find that a lot of the tactics and

9:58

tweaks just really revolve around, you

9:59

know, instead of just reading through

10:01

the entire thing line by line, use your

10:03

intelligence to build a filtering

10:05

mechanism or a script that does so

10:06

instead. Fable and other models are

10:08

already okay at this, but this heavily

10:10

biases it towards doing that. Any

10:12

advanced users here will probably scoff

10:14

at this, but you should be periodically

10:16

slashcontexting to see where things are

10:18

taking up your context without telling

10:20

you. Um, I as I was using Fable just the

10:23

other day, realized that I was running

10:25

like a dozen Chrome MCP instances

10:27

simultaneously, every single one loaded

10:29

with its full context. So, I ended up

10:31

spending way more money than I actually

10:32

had to. In addition, Claude was bloated.

10:34

It got really confused about which one

10:36

to use. Um, this occurs even to seasoned

10:38

veterans. And for those of you guys that

10:39

don't know, this is more or less what

10:41

that looks like. Um, we see the total

10:43

context window of the model. Sonnet 5 in

10:45

this case has almost 1 million token

10:47

window, which is pretty sweet. The

10:49

system prompt as it is consumes about

10:50

10,000 or 1%. The tools that it has

10:53

access to, the GP, the Glob, the SEDD,

10:56

and so on and so forth. Um, this gives

10:58

it 16,400 tokens. My memory files, which

11:02

are my own, are around 10,700 tokens.

11:05

The skills are around 7,000 tokens. And

11:08

then the messages that it and I have

11:10

sent back uh and forth with each other,

11:12

which I think it counts each one of

11:13

these as one message, is eight tokens.

11:15

So, collectively speaking, I basically

11:16

already used 8% of that budget. A quick

11:18

and easy hack to do this is just like

11:20

set up another cloud instance somewhere

11:21

on a different desktop that you're not

11:23

using. Um, and then have it loop. Okay,

11:25

so type /loop and then set up a watcher

11:28

that just runs once every day to tell

11:30

you if there's anything in your clawed

11:32

context that wasn't there 24 hours ago.

11:35

It'll take a quick snapshot of all the

11:36

files and everything currently in your

11:38

context. And in this way, you can just

11:40

get periodically notified so that if

11:41

something does creep its way in, like an

11:43

MCP or, I don't know, a big set of

11:45

skills or a bunch of other data, um,

11:46

you'll know immediately and you'll

11:48

always be able to uh, ensure that that

11:49

context is really small because at the

11:51

end of the day, all this token

11:52

management stuff is context management.

11:54

The last major tweak is to cap your

11:55

thinking. Now, what I mean by this is

11:57

Claude has the ability to think at

11:59

multiple levels. Thinking is just

12:01

feeding in tokens through its loop over

12:03

and over and over again to become clear

12:05

and clear about a problem and come up

12:07

with a better solution. Similar to how

12:08

if you were to brainstorm something on a

12:10

piece of paper for a thousand words,

12:11

you'd probably be clearer than if you

12:12

only got to brainstorm it for 10 words.

12:15

The issue with Claude's adaptive

12:16

thinking, which is where it sets its own

12:18

thinking budget for you, is it tends to

12:20

just use way more than you actually need

12:21

to. And with really intelligent models

12:23

like Fable, you don't actually need to

12:25

be thinking all that much in order to

12:27

answer questions and solve problems. And

12:28

so rather than default to super high

12:31

thinking like most of you are doing, you

12:32

should default to super low thinking.

12:34

You should basically opt uh in only when

12:37

you need really big thinking budgets.

12:39

And this is actually massive. This is

12:40

going to improve like 30 to 40%. Let me

12:42

show you what I mean. I have a little

12:44

demo over here which is basically going

12:45

to ramp up the effort over two different

12:48

thinking modes. The first is low. And

12:50

the task that I'm giving Claude here um

12:52

is just to find a bug somewhere in a

12:55

codebase. It's a simple bug. Nothing

12:58

super crazy. Um all it's doing is it's

13:00

basically going top to bottom and it's

13:01

reading the thing. The first pass is on

13:03

low effort. This took seven turns, cost

13:07

1,028

13:09

tokens, consumed 18 seconds of my time,

13:12

and cost uh 16 cents in total. Okay, so

13:15

it looped back on itself 7 times. That's

13:17

about the cap for low. Contrast that

13:19

with X high, which is basically the

13:21

smartest mode that we have. It took nine

13:23

turns, spent 1363 output tokens,

13:26

consumed 21 seconds, and cost something

13:28

like 3 cents more. Um, I should note

13:30

that clot has the ability to mediate

13:32

again how much thinking it gives to

13:35

different problems. Despite the fact

13:36

that it had the ability to mediate this

13:38

to go fewer turns, it didn't. It used

13:40

more turns obviously to achieve the

13:42

exact same result. Um, so, you know,

13:45

same bug found either way. Extra high

13:47

spent around 1.3x the output tokens. It

13:50

was nine turns versus seven. You know,

13:52

if I ran this another dozen times, I bet

13:53

you the delta would actually be a little

13:55

bit more than that. In my testing, it's

13:56

usually closer to 1.5x. So, the real

13:58

takeaway there is just set it to low.

14:00

Only set it to high if you explicitly

14:01

need it for a specific business

14:03

function. Um, don't just trust the

14:04

adaptive thinking because of course the

14:06

if you think about it, that's what

14:07

Anthropic makes money off of, right?

14:08

They make money off of the total number

14:10

of tokens you use. While adaptive

14:12

thinking is great and all, they're

14:13

probably going to trend up a little bit

14:14

versus what you would have done, uh, you

14:16

know, looking at the glaring hole in

14:17

your wallet. Hopefully, you guys

14:19

appreciated that video. Had a lot of fun

14:20

putting it together for you. You can get

14:21

all of the resources as mentioned down

14:23

below in the description. It's a free

14:24

download, free sign up. Go ahead and uh

14:26

take what you need. If you guys like

14:28

this sort of thing and want to monetize

14:29

your cloud code skills to get your very

14:31

first client, check out Maker School.

14:33

It's my day-by-day roadmap where I'll

14:34

personally guide you through the process

14:36

of getting your very first customer for

14:38

an AI service. An AI service here is

14:40

loosely defined as, you know, systems

14:42

that you build with cloud code, websites

14:44

that you build with codecs, backend

14:45

automations that you build with Naden,

14:47

drag and drop builders, and so on and so

14:49

forth. Um, I actually guarantee your

14:50

first client 90 days or you don't pay me

14:52

a scent. I'll give you all your money

14:53

back. If you guys like this, please

14:55

leave a like and comment down below, and

14:57

I am happy to answer any questions you

14:59

have.

Interactive Summary

The video provides a comprehensive guide on optimizing token usage when using AI models like Fable and Claude. By implementing strategies such as RTK (Rust Token Killer), semantic compression, leveraging SQLite for log searching, blocking large file reads, prioritizing English, enforcing context frugality, and capping adaptive thinking, users can achieve significant cost savings and efficiency without compromising model quality.

Suggested questions

4 ready-made prompts