HomeVideos

Fable Is Here, But Is It Actually Better? | Invest with AI Vibe Check

Now Playing

Fable Is Here, But Is It Actually Better? | Invest with AI Vibe Check

Transcript

1065 segments

0:00

So, I got to admit that I was excited

0:02

when the government had banned Fable out

0:03

of the gates. I think they were all

0:05

things that that like Opus probably

0:07

could have done with a lot of steering.

0:10

But, Fable pretty much one shot at them.

0:13

>> Converting a 120-page

0:15

operating manual on

0:17

>> [music]

0:17

>> building a hedge fund style financial

0:19

model is getting like quite a bit

0:20

easier.

0:27

Welcome to the next episode of Invest

0:30

with AI, where we explore the

0:31

intersection of fundamental investing

0:33

and artificial intelligence.

0:35

>> [music]

0:35

>> We have a short duo episode today where

0:38

we'll do just a little bit around the

0:40

horn vibe check

0:41

on AI in investing. And one of the key

0:45

topics people are discussing these days

0:49

uh is the new launch of Fable. Okay, so

0:52

what uh what are the vibes on on Fable

0:54

by your assessment?

0:56

>> Yeah, great to be here. Uh great to see

0:59

you. So, I got to admit that I was

1:00

excited when the government had banned

1:02

Fable out of the gates. Uh not because I

1:05

believe in any paternalistic

1:07

intervention of any sorts, uh but I was

1:10

just like, "Shoot, I'm not ready."

1:13

Uh

1:13

and I think that you know with all the

1:15

hype on these like big expensive models,

1:18

I I got to be honest with you and I

1:21

wonder how you approach this is I don't

1:22

have a great set of evals.

1:25

And by evals, it is like, do I have a a

1:28

a complex task

1:31

or series of tasks

1:32

with a lot of good context that is hard

1:36

and that prior models have not fully

1:38

succeeded at. And so,

1:41

I don't have like a predetermined set of

1:43

evals, but I do have one thing that I

1:46

run pretty regularly with our whenever

1:49

new model comes out. And that's

1:51

basically I take I have a giant folder

1:54

of like everything my business has ever

1:56

done. So, you're talking like

1:58

uh 200 granola calls,

2:01

50 decks, MCP into my email,

2:05

um

2:07

different like surveys that people have

2:09

filled out, the recordings of all my

2:11

trainings, and I dumped that in and I

2:13

say,

2:15

"Build me a website.

2:18

Create a strategy uh document for me.

2:21

Create a presentation from scratch just

2:24

based on like all of these materials."

2:27

And

2:29

I got to say, the result was pretty

2:31

good. Like, the website was beautiful.

2:34

The deck was pretty good.

2:37

Um but most importantly, the strategy

2:40

doc, cuz that's where, you know, I'm

2:42

intimately connected with my business.

2:44

Like, I know it better than anyone else.

2:46

So, like, if someone's going to give me

2:47

a unique insight, I I can like, that's a

2:49

unique insight.

2:51

And I'll point here, and again, this is

2:52

not the most robust eval, but what I'll

2:55

say is

2:56

I had a call, a prospecting call with

2:58

someone that was not even really a

3:00

prospect, just kind of like half

3:02

networking, half biz dev, but they said

3:05

something to me that I was like, "Ooh,

3:06

that's very

3:09

That's like, I feel seen. Like, you you

3:11

saw something that I didn't see in my

3:13

own business, and it's one of those

3:16

things that's like a little

3:17

uncomfortable where I like where I know

3:20

I should pursue it, but I'm not pursuing

3:22

it. And there's like reasons why I'm not

3:23

pursuing it."

3:24

And it was just one like one minute

3:27

inside probably

3:29

10,000 minutes of phone calls.

3:32

And it jumped right on that one minute

3:34

and it's like, "You should be pushing so

3:36

much harder on this."

3:38

And that was again, I don't want to over

3:42

analyze this, but I had forgotten

3:46

that that was a very salient point when

3:48

they made it. And it was once. It didn't

3:50

happen five times, it happened once. And

3:53

the fact that it read through 10,000

3:56

minutes of calls

3:58

and was like, "This is the thing you

3:59

need to focus on."

4:01

was pretty shocking to me.

4:05

So, that's what I'll say on the like

4:07

formal eval side of things or the

4:09

closest thing I have a formal. Uh I had

4:12

it build a lot of things from scratch.

4:14

So, I had it build like a tool that any

4:18

client of mine can describe themselves

4:20

in a multi-step process. Like, I'm a

4:22

fundamental investor. This is what I

4:24

look at. This is my investment horizon.

4:26

And then it generates like 20 skills

4:30

that they can then copy and paste into

4:32

Claude.

4:33

And I had it use all the skills that I

4:36

had ever created for any of my clients,

4:38

but also comb like YouTube and GitHub

4:42

and like any public repository.

4:45

So, that was pretty cool. I think they

4:47

were all things that that like Opus

4:49

probably could have done with steering.

4:52

But, Fable pretty much one shotted them.

4:55

Now, I can't tell you about the price

4:57

cuz I'm on the subsidy still. So, like

4:59

it's hard to know like was that like

5:01

$500 worth of credits or 30? I I don't

5:04

know.

5:05

Probably closer to 500 than 30 if I had

5:08

to guess. Was it worth $500? Probably

5:10

not. Was it worth 30? Absol- so freaking

5:12

lutely.

5:14

>> Very very interesting.

5:16

Um

5:17

I agree so far that like the eval is

5:20

hard. Like, eval is hard. I have a few

5:22

that I still in the pipeline to run.

5:24

Like, simple things like the level two

5:27

you know, financial modeling tasks that

5:29

are just

5:30

you know, complex and heretofore the the

5:32

models haven't been able to

5:34

accomplish. Uh so, I'll sort of get to

5:36

those.

5:37

Been traveling a little bit.

5:39

Um

5:40

but I was I was running one where I was

5:42

trying to turn

5:44

um uh the process to identify signals

5:47

around good times to short an individual

5:50

stock.

5:51

And Fable did an amazing job of sort of

5:55

going back into the history of good

5:57

times to short that stock and created a

5:59

really coherent, you know, pattern

6:02

recognition engine that historically I

6:05

would have had would have had to steer

6:06

quite a bit. It would have taken me two

6:08

or three iterations to say, "Hey, no, it

6:11

doesn't you know, we're not worried

6:13

about a goodwill impairment or gap

6:15

earnings." These are sort of the things

6:16

that ultimately are triggers to

6:18

to stocks and I may have to feed that

6:20

in, you know, transcripts, etc. or my

6:22

own research and

6:24

it one shotted it, which was which was a

6:26

little bit scary. It sort of catch

6:28

catching that same fundamental dynamic

6:30

that you um, you just mentioned like

6:34

materiality or taste. Like what's the

6:36

what's the white right right right right

6:39

word for it?

6:40

Um,

6:40

>> Yeah, it's like incisive judgment.

6:43

>> Yeah. Yeah.

6:44

>> Like

6:46

>> Yeah, which is it's it's wild. Like I

6:48

guess we would need to bring some

6:49

experts in to explain exactly how

6:53

how they they cook that in. Um, but I am

6:56

start I am launching a new eval, which

6:58

is a real real-life eval. I'm taking 100

7:01

companies

7:02

and I'm running 100 earnings previews as

7:06

earnings season starts next week and I'm

7:09

going to basically give Fable

7:12

two tasks uh,

7:15

forecast revenue EBITDA and EPS

7:17

accurately and accurately forecast where

7:20

the stock's going to trade up or down on

7:22

the earnings print. So, that's a very

7:24

tight sandbox of

7:27

tight sandbox of judgment

7:29

and there's no hot there's no hiding

7:31

from that. I'm going to keep maybe I'll

7:33

maybe we'll talk about that, you know,

7:34

the results once earnings season's over.

7:36

I'll put some of that on Twitter.

7:38

Obviously,

7:39

but I thought that'd would fun I was

7:41

inspired I think the the talk in my

7:43

network

7:44

for the last week was this uh Thinking

7:47

Machines Bridgewater report where

7:51

uh the the researchers took a large

7:54

language model so it which had native

7:57

judgment in financial tasks in the high

7:58

40s, overlaid financial

8:02

sort of

8:03

a judgment layer financial expertise and

8:05

judgment tasks raise up to 75% like the

8:08

mid 70s.

8:09

Um that wasn't the most academically

8:12

rigorous report but I think

8:13

directionally was inspiring to a lot of

8:15

people thinking about this combination

8:17

of

8:18

human judgment with large language model

8:21

judgment that the combination could

8:23

drive better better results.

8:26

Um

8:27

and so I've taken a lot of time to

8:28

codify into my earnings preview process

8:31

the frameworks of judgment that lead to

8:35

my estimation of whether a stock will

8:37

trade up or down.

8:39

Things like setups and expectations and

8:41

buy-side whisper and business momentum

8:43

and guidance trajectory etc.

8:46

Um and so I'm excited to unleash that

8:50

unleash that with the Fable

8:52

uh model in this exercise. Um not I'd be

8:56

shocked if it does really well

8:59

um but I want to just at least kick this

9:00

off cuz then the iteration into version

9:02

2.0 and feeding that back I think is a

9:04

fun data set to start to start uh

9:08

building.

9:09

>> Can I ask so I hadn't seen that paper.

9:11

Um is it are they fine when you say

9:14

they're adding the judgment are they

9:16

fine-tuning the model with that or is it

9:19

in a skill or in a prompt? How is that

9:22

judgment being

9:24

uh incorporated into the

9:25

>> I don't know I don't I don't know the

9:26

specifics. I read it a you know read it

9:29

sort of relatively quickly and maybe we

9:32

need to try and get one of the authors

9:34

in that would be a that would be a fun

9:35

conversation.

9:37

But I think the the conclusion was they

9:39

they took this sort of codified

9:41

analytical judgment of the Bridgewater

9:43

team and overlay that on the large

9:45

language model approach.

9:47

And the com- combination the combination

9:50

was valuable. And sort of in a in a um

9:54

in a in a very few way, I've started to

9:57

see that, right? When I started to see,

9:59

okay, take this take this approach

10:02

almost train it out of think codify some

10:05

of the setup patterns that, you know,

10:08

I've traded in the past. It it's a quick

10:11

study. Like the models have been a quick

10:12

study on those patterns and and

10:14

following on the rails of those

10:17

those that that pattern recognition

10:19

power. There's been a few instances

10:20

where I'm like, "Huh, that was like

10:21

nailed it pretty pretty pretty

10:23

accurately." Not in a scientific way,

10:25

hence I want to start sort of doing this

10:27

in a in a broader subset to see if did I

10:31

just get lucky and heads flip you know,

10:33

head flipped on the coin twice in a row.

10:36

Um so I'm inferring skill, which is why

10:40

um my plan to do this in a more

10:42

systematic way in in the earnings prints

10:45

coming up starting next week.

10:48

>> That's cool. Wow. One thing I'll flag, I

10:50

know we advertise this pod- podcast as

10:53

not having the answers, but us just

10:55

talking through what we're what's on our

10:58

radars and um

10:59

someone

11:01

who knows a lot more about it LLMs than

11:03

I do had mentioned something called I

11:05

think it's called braintrust.dev

11:07

and it's

11:09

it's a platform for running evals.

11:12

And he had actually suggested

11:15

that that I do an exercise kind of like

11:18

the one you're describing. I'm not I'm

11:19

not close enough to fundamental equity

11:21

and investing to to do that. But he had

11:24

suggested that I do that using this

11:26

platform. So just putting it on your

11:28

radar to maybe check it out that that

11:30

there's actually like a structured

11:31

environment to run these evals and we

11:34

can run

11:34

>> Oh, that's interesting.

11:35

That's interesting.

11:37

Um

11:37

yeah, that's one of the one of the areas

11:39

in our work where we're really like I've

11:41

done a I've done evals a lot on vibes,

11:43

right? It's like, you know, you try you

11:45

you shoot you you shoot off a few

11:47

workflows. You're like, "Ah, this looks

11:48

good."

11:49

But the problem is even when you shot

11:51

off five things in Opus 4.8, like you

11:53

kind of like sometimes it works,

11:55

sometimes it doesn't. So, it's very hard

11:57

to scientifically measure vibes.

11:59

>> Totally.

12:00

>> One of the one of the things we're

12:00

trying to do with clients is create a

12:02

more scientific eval set of the various

12:05

MCP inputs, right? So, who are the

12:07

modeling MCP vendors? Let's create an

12:11

eval set so we can more scientifically

12:13

score some of those vendors on on

12:15

accuracy uh accuracy and capability. Uh

12:19

so, it's a I'd say that's like an work

12:21

in progress

12:22

work in work in progress on our on our

12:24

side, but yeah, good call out on the

12:26

brain trust I'll definitely take a take

12:28

a look at that.

12:29

>> Cool.

12:30

On my end, I'm a little late to the

12:32

game, but it's something that I'm having

12:34

a lot of fun with. So,

12:36

we're both Karpathy Andre Karpathy fan

12:39

boys

12:40

and probably, you know, like he drops a

12:42

banger of a tweet every 6 weeks and the

12:44

AI world like it realigns around this

12:47

new idea. So, maybe two or three tweet

12:49

cycles ago, so that might be, you know,

12:51

18 weeks ago,

12:53

he had something he called it kind of

12:55

this like self-updating

12:57

wiki, like a knowledge wiki.

13:00

And and I kind of didn't have time to

13:03

focus on it when it came out and I'll

13:05

circle back to the wiki with a very

13:07

common problem that many of my clients

13:09

are having right now. And that problem

13:12

is they've gotten really good at

13:13

co-work. Some are using code, you know,

13:16

and so now they've got good prompts,

13:19

they've got good skills, all that stuff.

13:22

They've connected their MCPs.

13:24

But let's say they've got an investment

13:26

and in this this like private equity

13:27

fund

13:28

they have an investment and there's 500

13:31

documents in that folder. Let's say

13:33

that's you know, the investment is I

13:35

don't know, Sweetgreen, right? So they

13:37

have this Sweetgreen

13:39

assuming it was private. They point it

13:40

to the Sweetgreen folder, they've got

13:42

500 documents in that folder.

13:45

And it's not even organized and so on.

13:47

And so if they want to ask questions

13:49

about it, they could be very tactical

13:51

questions like

13:52

you know, how has the store same sales

13:55

growth changed over quarter over

13:57

quarter? Or they could be more

13:59

qualitative questions like how should

14:02

this thesis how is this thesis being

14:05

informed by inflation? All right? Our

14:08

our tariffs, right?

14:09

So you point at Cork. Cork's Cork can't

14:13

go through 500 files, right? And so what

14:16

it does is it kind of

14:18

cherry-picks it and what it really does

14:20

is it does like basically like a lot of

14:22

like command F.

14:24

So like let's say you're doing tariffs,

14:25

it's going to do like command F on your

14:27

500 files and be like tariff, command F

14:30

uh inflation, command F and then it's

14:32

going to grab 17 files, pull them into

14:35

context and then try to answer your

14:36

question.

14:38

And it starts to break down more and

14:39

more as you have more and more

14:40

documents, more and more MCP. MCP breaks

14:42

down too because the context windows are

14:45

too small when you're dealing with like

14:46

huge volumes of data.

14:49

So that's the problem that people are

14:50

starting to have and it's even more

14:53

complicated by the fact that there's

14:54

PDFs, there's Excel files, there's you

14:57

know, 200-page, you know,

15:00

uh credit agreements and so on.

15:03

So here's where the Karpathy wiki comes

15:04

in is that the way this this works is it

15:08

says we're not going to we're we're

15:10

going to create a map

15:13

of your information, an an ontology

15:16

of your information. You might have hear

15:18

that word, people throwing that word

15:19

around.

15:20

And so what it does is it basically,

15:22

let's take it takes the 500 files

15:25

and one at a time

15:27

well, first it will convert it into

15:29

markdown.

15:30

And then it will basically read it.

15:33

And as it reads it, it kind of creates a

15:36

summary at the top with tags and

15:39

metadata.

15:40

And then it creates like links of ideas

15:44

inside

15:45

the text that it just read. So, if you

15:48

have something like, I don't know,

15:51

tariffs, right? It will just like find

15:52

all the references to tariffs.

15:55

Then you ingest the next file and it

15:56

does the same thing ex- with one

15:59

exception. As it finds tariffs, it then

16:01

looks in all the other files and says,

16:04

are there any references to tariffs in

16:06

the other files?

16:07

And then they start to link and you kind

16:08

of have this kind of like you could

16:10

think of it as kind of

16:12

like a private worldwide web, right?

16:14

Like a little little wiki a wiki, right?

16:17

And so what's cool about this is that

16:20

the LLM is doing the ingestion.

16:23

So, like taking the file, creating the

16:24

links. And then as you get new

16:26

information

16:27

it just keeps updating it. So, there's

16:29

this like self-improving process.

16:32

Now, I'm going to be honest with you

16:33

like

16:34

again, it's another one of these things

16:35

like hard to know if it's working versus

16:38

the naive approach, which is, you know,

16:39

just pointed at 500 at Fable.

16:43

But that being said

16:45

I I know that some clients have done

16:46

this like, you know, maybe like a

16:48

concentrated like a PE fund that's like

16:50

15 investments. They're like, we're

16:52

going to go through this process for our

16:53

15 investments. We're not going to look

16:55

at the 200 we passed on.

16:57

But we'll do the 15, we understand it

16:59

well, every analyst can kind of own

17:02

their name in the wiki. It's not like

17:04

your style of investing where the turn

17:06

over, you know, you're talking six-year

17:07

investment period, so you're not

17:09

flipping names and there's not data

17:11

you're responding to every second. It's

17:12

a little bit different. But what you

17:15

start to have is this knowledge graph

17:18

that is much easier for an LLM to

17:21

traverse

17:23

without blowing a ton of tokens and by

17:26

increasing the depth of accuracy.

17:29

So, early days early days in my own

17:32

exploration on it. I do have a few

17:34

clients that like I I don't want to say

17:36

like I brought it to them like they kind

17:38

of came up with it we like they came up

17:40

with themselves and we can compared

17:42

notes. You know, a lot of these a lot of

17:44

this is around this idea of like if you

17:46

can get all your files neatly organized

17:49

in markdown with good naming, good

17:51

metadata inside like a nice clean folder

17:54

structure that the LLM can go in and

17:57

read it much better. And so, you're

17:59

starting to see in again

18:02

I encourage everyone or I discourage

18:04

anyone from doing this like oh, we're

18:05

going to like do 20 years of data across

18:08

this like no, no. Pick like five names

18:11

and start. Pick one name and start. And

18:13

so, that's starting to happen and and

18:15

I'll be able to report back on the

18:17

progress. But I have seen some green

18:19

shoots that again, this is more in

18:21

private equity but it works.

18:24

>> It's It's interesting and I've heard

18:26

more about the same concept of

18:29

um

18:30

just making your file system more

18:32

legible to to the agents. How are How

18:37

are clients doing that? Is that a skill

18:39

that co-work or code x can go in and

18:41

organize like that even gives me hives

18:44

when I start to think about organize

18:45

like letting code x rip on my internal

18:48

files to make them

18:49

>> Oh, yeah.

18:50

>> legible or people just pressing the

18:52

button and hoping for the best?

18:54

>> Yeah, I mean I think like any of these

18:56

things you start with a little pilot.

18:59

And so, maybe you've got you know, the

19:01

sweet green folder with 500 files and

19:03

you'll say

19:04

uh find me all the financials and put

19:07

them in the financials folder and then

19:09

give it a descriptive file name.

19:11

And then you go in and you're like "Did

19:13

it do it?" You know, and as the analyst

19:14

you know, you know, sometimes an Excel

19:17

file is a client list. You're like, "I

19:18

hope it didn't take the client list and

19:19

put it in the models folder, right?" So,

19:22

you kind of have to do that testing, but

19:24

it's it's very good.

19:26

And so, uh so I So, then you can kind of

19:30

make this a more

19:31

robust process like

19:34

every time a PDF comes in, immediately

19:36

turn it convert it to markdown.

19:38

Right? That That I know a lot of groups

19:40

are doing that. Um and then add the

19:42

metadata while you're at it. So,

19:45

it's actually quite powerful The models

19:47

are very good at doing that. All right,

19:48

you could even do that with a Sonnet

19:50

model. You don't even need Opus um for

19:52

like the the the just like the the the

19:54

tagging and the cleaning and the moving

19:56

and so on.

19:57

Now, when you want to do it in, you

20:00

know, bigger batches or with more

20:03

complexity to it, more one-shotting, you

20:06

would use a a bigger model, but I would

20:08

say again like kind of start small, but

20:09

you I think you would be very surprised

20:11

that, you know, this is the file And

20:13

again, it all comes back to having good

20:14

skills. Like, this is the

20:16

naming convention that we use for our

20:18

files. Year year year month month day

20:19

day, you know, quarter, you know, if

20:21

it's a financial this.

20:23

Um

20:24

and then another advant- another thing

20:27

you do is like you literally have a

20:28

document that's like, "These are how the

20:30

folders are organized." So, you think

20:32

about the LM comes in, it reads this

20:34

kind of index document, and it's you're

20:36

asking a question that involves models,

20:39

then it's like, "Oh, let me read this."

20:40

And And this file says, "All of the

20:42

models live in the financials folder,

20:44

right?" Uh and so, it just saves the

20:46

model and saves the agent a lot of work

20:49

in like, "Where Where do the models

20:51

live? You know, do I need to open this

20:53

file?"

20:54

>> Yeah. Yeah. That makes uh that makes

20:56

sense. It's interesting we

20:58

I think we both have the same

20:59

fundamental rule. I call it the Karpathy

21:02

rule of wait for Andre to tweet

21:04

something and then trying trying to

21:06

through how it applies to our space. I

21:09

I've used that same Wiki knowledge Wiki

21:12

concept to build skills architectures.

21:15

And a lot of what I've been doing is is

21:17

really like the brain dump of everything

21:19

I know about a certain concept, putting

21:21

that into as deep of a

21:23

sort of knowledge pool as possible. And

21:26

then taking that same knowledge graph

21:28

concept to sort of to make it agent

21:30

legible for skills creation. Uh and

21:32

that's been quite slick uh from a

21:35

process perspective. It makes a lot of

21:37

the hard part actually is like compiling

21:40

compiling the knowledge.

21:42

Um

21:43

but converting uh you know, for example,

21:45

converting a 120-page

21:47

operating manual on building a hedge

21:50

fund style financial model,

21:52

taking that raw data into actual

21:54

operational skills architecture is

21:57

getting like quite a bit easier if you

21:59

take that interim interim step. Um

22:02

so that's been um

22:03

that's that's that's been fun.

22:05

One of the things I've been spending a

22:06

lot of time on is thinking about just

22:08

the the the um

22:10

uh the stack of agents. Like what's the

22:13

right Like we sort of had this concept

22:15

back last year like four to six pages is

22:18

the right

22:19

uh depth for a prompt, right? Because,

22:22

you know, a one-page prompt was probably

22:25

too thin, not descriptive enough, but an

22:27

eight-page prompt was just too much, and

22:29

anything in the middle

22:31

the large language model would lose

22:33

attention.

22:35

You know, that I'm curious your

22:36

perspective on this in skills building,

22:38

too. It's sort of trying to dial in

22:39

myself of like how thin is too thin or

22:42

how thick is too thick for a skill.

22:44

And um one of the concept I've been

22:46

using is a sort of like raw primitives,

22:48

almost like single-purpose skills,

22:50

single-purpose agents

22:52

that are composable into these

22:54

orchestrated pipelines.

22:56

And the sub-agents, like the ability of

22:57

a ramp skill to go pull in a dozen

23:00

different sub-agents,

23:02

individual skills into this workflow

23:05

pipeline.

23:06

I'm getting some really interesting

23:08

stuff out of that. Um that's sort of

23:12

composer composability

23:14

uh process which has been pretty

23:16

exciting to me so far. Exciting and my

23:20

wife would tell me I'm a nerd and then I

23:21

get excited about skill skill

23:23

composition, but it's a little bit of

23:25

like a breakthrough on my side of uh you

23:28

know, consistency and um

23:31

each of each individual piece is sort of

23:33

doing a much better job adhering to that

23:36

individual skill. Whereas, if I load

23:37

everything into one skill, it's like

23:39

things get lost in tran- things get lost

23:41

in translation.

23:43

>> And you know, I haven't spent as much

23:46

time on that. My skills do tend to be

23:48

pretty narrow just by default. But I

23:51

haven't like it's not been a design

23:53

choice, just kind of the way the puck

23:55

landed.

23:56

Um

23:57

but one thing you'll notice, I don't

23:59

know if you notice this on Fable,

24:01

but it does uh it's kind of

24:04

doing a lot of the sub agent work on

24:05

your behalf even without skills.

24:08

And so, for example, if you ask it to,

24:10

you know, create a

24:12

uh I don't know, create a

24:14

uh final memo, you could see it it's

24:17

like

24:18

uh I'm going to create a bunch of

24:19

different sub agents to read these

24:22

documents and then report back.

24:24

And one of the things that I'm just

24:26

starting to realize on sub agents is

24:28

that sub agents use their own context.

24:31

Right? So, they ring-fence their own

24:32

context window so they're not polluting

24:34

the context window of the larger thread.

24:37

And so, if you have these like very

24:38

discrete agents

24:40

driven by skills as you've just

24:42

described, you have this like really

24:44

efficient kind of context gathering

24:47

exercise

24:49

that then extracts the key information

24:51

up to the month, you know, the

24:53

orchestrator or the main thread.

24:55

Uh that then can kind of reassemble it

24:58

and uses like the heavy duty Intel, you

25:00

know, the fable, the open open style

25:02

intelligence. Now,

25:04

again, this is more through observation

25:07

than from things I've like consciously

25:09

architected myself, but I know that

25:12

people, you know, you know, I was a big

25:14

open claw person back in the day. Like

25:16

people who are using open claw were

25:17

thinking through this

25:19

6 months ago of like

25:21

>> Yeah.

25:21

>> when do you need [clears throat] to

25:22

close ring fence contacts, when you need

25:24

to feed contacts into the main chat.

25:26

Like this is why I always love talking

25:29

about open claws cuz they were 6 months

25:31

ahead of the conversation, whether it

25:34

worked or not, right?

25:36

>> Yeah. Yeah, no, it's interesting. It um

25:40

you know, like complicated Excel

25:42

modeling has been one of the areas

25:44

that's been quite disappointing um

25:46

with with AI and I've had to really

25:49

chunk things into individual steps.

25:52

And for whatever reason, AI in the Excel

25:56

wrapper just doesn't listen to my skill.

25:59

It doesn't follow directions well.

26:01

I asked a

26:02

I asked a um a contact about that and

26:06

they're working, we will bring him on.

26:08

Um

26:09

his firm is working on a sub agent

26:10

approach. And that's sort of like that's

26:13

very interesting to me. Like if I could

26:14

take a highly complex LBO model or hedge

26:17

fund style model,

26:19

it's just too much like it's too context

26:22

inefficient, token inefficient today

26:25

to do that in AI Excel system, but if

26:28

you can spawn 30 different sub agents,

26:30

one to go clean up the cash flow

26:32

statement, one to go build out the

26:33

interest schedule,

26:35

and then you bring in a fable on top of

26:36

that to assemble those analyses in,

26:40

um you know, deploying deterministic

26:42

code where it makes sense.

26:44

Um obviously the proof is in the

26:46

pudding, but conceptually that feels

26:48

like a really interesting engineering

26:51

approach

26:52

um

26:53

approach to this this this problem.

26:55

>> And again, I I don't know enough about

26:58

Fable, but it seems like it's trying to

27:00

do with all of that on the fly.

27:04

>> Yeah. Yeah.

27:04

>> Which would be crazy.

27:05

>> Yeah. Yeah, exactly. Exactly.

27:08

>> we're like talking about customizing

27:09

every sub-agent, but

27:11

again, we don't want to get, you know,

27:14

um, get in front of, you know, get in

27:16

front of the story, right? Or mislead.

27:18

But it's there are little breadcrumbs

27:21

that it has

27:22

the capability to do at least some of

27:24

that.

27:25

>> It makes sense. One one question I had,

27:27

um,

27:28

how are you seeing your clients think

27:30

about the tradeoffs between Claude Code

27:33

and Claude Co-work, which is really just

27:35

a different harness? Like where do you

27:37

find Like what do you find are the pros

27:39

and cons of that that that decision to

27:42

be?

27:43

>> It's It's funny you say that. Do my

27:45

clients freaking love Co-work.

27:49

Uh, and Co-work I don't I like Code, but

27:52

I'm kind of nerdy. I like the weirdness

27:54

about it. Um,

27:57

they love Co-work and I'll give you uh,

27:59

so so

28:01

and I teach Code if people want to learn

28:02

it, but now, even after I teach it,

28:06

they're like, "What's the difference

28:07

between Co-work?"

28:08

And there are some differences, but the

28:10

differences are actually like collapsing

28:12

for knowledge work. Like it they'd be

28:13

hard to notice. Like for coders, you

28:16

would know the difference. Like a coder

28:17

would never use Co-work.

28:19

But it's the blur the lines are being

28:21

more blurred. So a few things that are

28:23

pretty cool on on

28:25

on Co-work. One is the There's this

28:28

concept of living artifacts, live

28:30

artifacts,

28:32

where you can run a dashboard off of an

28:35

an Excel spreadsheet.

28:37

Uh, and as you change that Excel

28:39

spreadsheet, the dashboard changes.

28:41

Which is like one of the many one of the

28:43

main reasons why people were using

28:45

Co-work Claude Code anyway was to just

28:48

create more interactive visualizations

28:50

of data.

28:51

Uh

28:52

so, that like knocks out a problem.

28:55

Uh

28:56

another one is that Claude co-work

29:02

is pretty aggressive in solving problems

29:04

with code.

29:06

So, even like if it can't figure

29:08

something out, it might say like, "Can I

29:09

Can I write a Python script to do that?"

29:12

And it it didn't used to do that, or I

29:14

felt more constrained in the past on

29:16

that. So, it's getting a little bit more

29:18

aggressive, co-work.

29:20

Um

29:22

the other reason why people like co-work

29:24

is

29:25

the schedule tasks, right? That's kind

29:27

of how you trigger agentic workflows.

29:30

Um

29:31

they're so clean to run in co-work.

29:34

There's a You can use your mouse, you

29:35

could see the which ones ran, which ones

29:38

didn't want run, which ones need

29:39

approval. When you run schedule tasks in

29:42

code, it just goes into this black hole.

29:44

Like, it's like a it's like a There's a

29:46

markdown file that like, "Did your thing

29:47

run?" Like, you know, and so, the ease

29:51

of the user interface.

29:53

And then, this is something I actually

29:55

just learned the other day cuz you and I

29:57

don't get too deep into the infosec side

29:59

of things. Um we kind of assume that

30:02

people come to us once they've cleared

30:04

the infosec bar.

30:06

Uh information security for those

30:08

unfamiliar with the acronym.

30:10

But, one thing I just recently learned

30:12

is that So, what Claude co-work does

30:15

is it actually creates a virtual

30:17

machine. So, you can think of this like

30:18

a disposable desktop. And so,

30:22

it's actually really hard to blow up

30:23

your computer cuz you're actually using

30:26

a copy of your desktop. And I don't know

30:28

exactly which cloud it lives on.

30:31

But, with co-work with Claude code,

30:33

you're actually giving it root access.

30:36

So, it's not a copy, like this

30:38

cloud-based copy, it's the real shebang.

30:41

It's your computer, right? And so what

30:45

I've heard from IT departments is that

30:48

especially if the gap is converging in

30:50

what they can do, they're like, "I don't

30:52

want to give, you know, Brett root

30:56

access when he's not even a coder.

30:59

I just want him to be able to use the

31:01

cool features of Co-work in this like

31:04

virtual machine environment that's much

31:06

safer." So I didn't actually know this

31:08

distinction until like a couple days

31:10

ago, but

31:11

even since then it's come up twice

31:13

already.

31:14

>> Interesting. Interesting.

31:16

Um

31:17

Yeah, and if you know, if you if you can

31:19

sort of get the same power of Cloud Code

31:21

through the the Co-work wrapper and has

31:23

a few of these bells and whistles that's

31:25

sort of causing people. Do you lose any

31:27

of the flexibility? Like do you have

31:30

more flexibility in Cloud Code to build

31:32

dashboards or is there anything that

31:34

Co-work is constraining constraining of?

31:38

>> Yeah, I for knowledge work per se, no.

31:42

But if you want to

31:45

like,

31:47

you know, some some of the things people

31:48

want to do is like build these like more

31:50

complex data scrapers

31:53

where you need like a heavier build of

31:56

code. I think that I I

31:59

I could be wrong on this, but basically

32:01

in coding there you you can like there's

32:04

these open-source libraries that's like,

32:06

"Oh, if I want to run a scraper, there's

32:07

like 5,000 people have created scraper

32:10

tools on GitHub open-source

32:13

and you could just go grab one. You

32:14

don't have to like rebuild the scraper,

32:16

right? Uh I don't think Co-work has the

32:20

ability to grab those packages

32:23

for the security reason. So I think the

32:25

more you want to reuse like more

32:28

traditional coding components,

32:30

then you'd want to shift to code, but

32:33

again, you'd have to really have more of

32:35

a coding use case than a research

32:39

knowledge work use case.

32:41

>> Yeah, so the develop like it's more more

32:42

developer-minded.

32:45

Person who's going to use Claude Code

32:46

and the more just end-user who's

32:50

working with files and building, you

32:52

know, PDFs and

32:54

uh is

32:56

Excel spreadsheets is more living in

32:58

Claude Co- Claude Co-work. But the

33:00

distinction distinction is is

33:02

compressing anyways from a community.

33:04

>> Definitely compressing.

33:06

>> Yeah, okay.

33:07

>> And I've got some pretty technical

33:08

clients that know how to use code and

33:11

they're just like,

33:12

"I can't be bothered. I just use

33:13

Co-work. It just it does what I need it

33:15

to do." And I to be honest, I'm falling

33:18

more and more in that category. Like,

33:20

let me use my mouse more, show me more

33:22

things on the screen.

33:24

Um and then just today they announced

33:27

that

33:28

they're going to improve the mobile

33:29

capability of Co-work. And so like the

33:32

schedule tasks

33:34

used to only run if your computer was

33:36

on.

33:37

>> Yeah.

33:37

>> And apparently now they're going to run

33:40

if your like laptops off. So they're

33:41

moving some of the those powers to the

33:44

cloud.

33:45

So the the announce it's not out yet,

33:47

but they announced it this morning.

33:50

>> Oh, interesting. Interesting. Okay,

33:52

great.

33:53

Uh any last any last thoughts on on on

33:55

Vibe Check? Like what are how are people

33:58

feeling about AI adoption in in finance

34:01

right now?

34:03

>> I think I I feel like I'm encountering a

34:06

little bit more

34:10

I wouldn't say frustration,

34:12

but I think the skepticism

34:14

level has gone up like

34:17

15% in the past quarter.

34:19

>> Interesting. What

34:21

What are What are the pushbacks? Yeah,

34:22

sorry. What are the pushbacks on that?

34:24

>> I think one is

34:26

the we're spending so much time

34:30

babysitting the AI when we could have

34:32

just done it ourselves. That's one.

34:35

And I feel like after 18 months or so,

34:37

they're like, "Oh, we thought we would

34:38

have like kind of left that period,

34:41

but we're still in it." I don't know if

34:42

you saw Glean had this report called

34:44

bots, they called bot sitting.

34:47

It's the invisible work in just like

34:49

updating your contacts, like checking

34:52

for hallucinations, reprompting, you

34:54

know, I think they It's like, I don't

34:57

know, it's like 30% of your AI time is

34:59

bot sitting, right? So, you're actually

35:01

[snorts] only 70% more effective. So, I

35:04

think there's a little bit of that.

35:06

And and then the cost the cost side of

35:09

things. Just like,

35:11

is this worth the money?

35:14

And I think,

35:15

you know, you're definitely seeing token

35:17

budgets, you know,

35:20

Co-work is great. It uses a lot of

35:22

tokens. I don't think it's I think to

35:25

make it so friendly to non-coders, it

35:27

probably,

35:28

you know, I saw the analogy is probably

35:31

taking a blowtorch to light a cigarette,

35:33

you know, a few times.

35:35

Uh

35:36

so, there's definitely

35:39

a little bit of that. And I think a few

35:40

people have said, "Look, like, they're

35:42

like, "Look, if I'm really honest, I

35:44

didn't notice a difference between Opus

35:46

4.5, 4.6, 4.8. I'm not sure if I'm going

35:50

to notice a difference with Fable.

35:52

I'm not sure if that's like a user

35:55

error, and I put myself in that category

35:57

as well, like I'm not pushing this hard

35:59

enough.

36:00

Or it's the reality that there is some

36:03

kind of plateau for the types of things

36:06

that we're doing, right? Synthesizing

36:07

documents, writing reports, analyzing

36:10

data. That's an open I I

36:12

No one's issuing a verdict on that, but

36:14

there's a little bit more like,

36:16

I'm really honest, I don't know if it's

36:18

really got if AI, like the models,

36:22

have gotten that much better in the past

36:25

3 months, 6 months.

36:26

>> Yeah. Yeah.

36:27

It's interesting. I think um

36:29

you know, the the the the public market

36:32

like

36:33

you know, most of the clients we work

36:34

with are you know, scaled public market

36:39

investors that have a tight coverage

36:41

area. And so chatbots were just like did

36:44

not hit

36:46

you know, effective product market fit

36:48

for that. They're great for generalist

36:50

firms that want to get up to speed and

36:52

do you know, research quickly.

36:54

I think many of my clients are probably

36:56

only in the last three to five months

36:59

like getting really serious about

37:01

implementation.

37:02

So they're probably later on that curve

37:05

than some of your clients who are more

37:06

in the private market space have been at

37:09

this a little bit little bit um

37:11

longer. This concept of a digital twin

37:14

sort of overlaying the entire process

37:16

like a digital analyst uh is something

37:19

that we've been we've been scoping

37:22

um and really was a little bit science

37:25

fiction until

37:27

>> Mhm.

37:27

>> very very very recently.

37:30

Um

37:30

So the vibes, you know, the vibes on

37:32

those conversations I think have gotten

37:34

quite exciting and maybe that hits the

37:37

same wall at some point. Um it's easy to

37:40

have hope and and build prototypes. It's

37:42

another to actually deploy these things

37:45

to to uh the effect of better

37:48

decision-making.

37:50

Um

37:50

>> Yeah.

37:51

>> that's the hill we will we will all work

37:53

to climb and hopefully bring in some

37:56

some more guests to uh help inform that

37:59

inform that climb.

38:01

>> Yeah.

38:02

I'm excited.

38:04

This is a good vibe check.

38:05

>> Good vibe check. Uh

38:07

great great great to connect K and we

38:10

will be back we'll be back next next

38:13

week with another guest and maybe next

38:15

month with another vibe check. We'll

38:17

we'll we'll sort of monitor the vibes

38:19

and bring those vibes back back to you.

38:21

So thanks everyone for us know in the

38:23

comments. Yeah, let us know in the

38:25

comments if what you think of these vibe

38:27

checks, what kind of things you want us

38:28

to cover in future ones.

38:31

>> That sounds great. Sounds great. All

38:33

right. See everyone soon.

Interactive Summary

This episode of Invest with AI explores the practical applications and limitations of the new Fable model in financial contexts. The hosts discuss their personal experiences using AI to build business strategies, conduct earnings previews, and organize data using knowledge wikis. They also compare Claude Code and Claude Co-work as tools for knowledge work, noting that while skepticism about AI's immediate impact has grown, the development of specialized agentic workflows and 'digital twin' concepts remains a promising area for financial decision-making.

Suggested questions

4 ready-made prompts