HomeVideos

Stanford's Method Turns Claude Into a PHD Level Research Team

Now Playing

Stanford's Method Turns Claude Into a PHD Level Research Team

Transcript

418 segments

0:00

So Stanford has a research method called

0:02

storm, which has actually been shown in

0:03

peer-reviewed testing to produce

0:05

articles 25% more organized than the

0:07

next best method. So I put all of those

0:09

storm principles into my own Claude

0:11

skill, which I'm going to give you guys

0:12

for completely free, and you end up with

0:13

the result that looks like this. It is

0:15

an HTML briefing that has been put

0:16

together by five different perspectives

0:18

of agents, and it has been verified.

0:20

Meaning if I scroll down to the bottom,

0:21

you can see that the different

0:22

perspectives are giving analysis on each

0:25

parts of the report. But at the very

0:27

bottom, you can see that we have

0:28

different sources that have been

0:29

confirmed, corrected, or demoted.

0:32

Meaning on the first pass, the briefing

0:34

would have had information in here that

0:36

just wasn't correct. But because our

0:37

skill works in all this verification, on

0:39

V2, we can have a lot more faith in this

0:42

output. So the whole idea of storm is

0:44

that instead of just shooting off one

0:45

prompt and having one angle of research,

0:48

we are utilizing a bunch of different

0:50

angles. Because if you just send off one

0:51

prompt to Claude, there's going to be a

0:52

bunch of blind spots in that research

0:54

plan. So storm utilizes these five

0:56

perspectives. We've got a practitioner,

0:58

an academic, a skeptic, an economist,

1:01

and a historian. And each angle finds a

1:03

hole that the other angles miss. And

1:05

this whole idea of having different

1:06

agents kind of like role-play their own

1:08

personalities and their own, you know,

1:10

backgrounds with different areas of

1:11

expertise, is really, really beneficial.

1:14

If you've seen other videos where I've

1:15

talked about something like the roast

1:16

skill, or how I like to use agent teams

1:18

to basically be a council, it's really,

1:20

really helpful to identify different

1:22

perspectives and, like I said, find

1:24

holes that the other angles are going to

1:26

miss. And so let me just show you a real

1:27

quick example of why that's so

1:28

beneficial. So Claude code natively has

1:31

a feature called deep research, which

1:33

launched with the dynamic workflows. So

1:35

if you come into Claude and you do a

1:37

deep research command like this, you

1:39

will basically be able to enter a

1:40

research topic and it will spin up a

1:41

dynamic workflow, which will kick off

1:43

hundreds of agents in the background. I

1:45

think in this example, there was 103

1:47

different agents running. So this will

1:48

give you a pretty solid deep research

1:49

report. As you can see here at the

1:51

bottom, it didn't actually give me any

1:52

output, it just internalized all that.

1:54

So I said, "Where's the report?" It gave

1:55

me this markdown file, which is decent,

1:57

but it's really not that thorough, and

1:59

there's not as many sources as we'd

2:00

like. There's only two up here, and then

2:02

there's a few more unconfirmed down here

2:04

at the bottom, as well as some open

2:05

questions. And then I took this exact

2:07

prompt that I asked in the deep

2:08

research, and I put it into a Storm

2:10

skill. So, I said, "Hey, Storm research,

2:12

do this." And it said, "Okay, cool.

2:14

Here's the topic. I'm going to run the

2:16

Storm pipeline now. I ran these five

2:18

agents." As you can see, the

2:19

practitioner, the academic, the skeptic,

2:21

the economist, and the historian were

2:23

converging all of that stuff together,

2:24

we're seeing where they disagree, and

2:26

then we're going to run six more agents,

2:27

which are going to verify all those

2:28

facts that you just found.

2:29

Verification's done, and now you have

2:31

this HTML report, which is consistently

2:33

going to look like this every time with

2:35

a 60-second summary key findings. And

2:37

all of these key findings are also

2:38

ranked by reliability. You can see right

2:40

here, reliability high, nine out of 10.

2:43

This one was supported by the academic

2:44

and the skeptic, and it was challenged

2:46

by the practitioner and the economist.

2:47

And it goes like this throughout the

2:49

rest of the entire HTML report here. It

2:51

also calls out the assumption that this

2:53

briefing rests on and the missing six

2:55

lens. All five lenses look at the firm

2:57

from the owner's chair, adoption rates,

2:59

productivity, ROI. None of them sat in

3:01

the seat of the customer or the

3:03

frontline employee. So, that's the

3:04

missing sixth lens here, and I would

3:06

then just say, "Okay, cool. Spin up that

3:08

sixth lens, and run a V3 of this HTML

3:11

report." And then it gives us really

3:12

practical takeaways here. And what's

3:14

cool about this is compared to something

3:16

like the deep research, which is just

3:17

going to basically give you a brain dump

3:19

of a bunch of stats it found, the Storm

3:21

research can really be tailored towards

3:23

you. You can go into the skill and say,

3:25

"Hey, here's what I'm doing. Here's my

3:26

business. Here's what our goals are."

3:27

Every time you run a Storm research

3:29

report, make it tailored towards us, you

3:31

know, what do we actually want to do

3:33

differently now that you've understood

3:34

all of this new data and research. And

3:36

so, in this specific example with the

3:38

deep research and the Storm, I put this

3:40

into Codex, so a completely different AI

3:41

model, and I said, "Hey, which one's

3:43

better?" And it came back and said the

3:45

HTML briefing is better. It's got better

3:46

evidence quality, it's much stronger,

3:48

it's got much stronger source diversity,

3:50

it's got a much stronger thesis, It's

3:52

more actionable. It's got better risk

3:54

control, and it's better for video and

3:57

content. So, in all six of these

3:59

categories here, Codex thought that the

4:01

HTML briefing was better, and I don't

4:04

know the exact metrics here on cost, but

4:06

the storm research was faster to run,

4:09

and it was 100% cheaper because in this

4:12

case we ran about What was this? Maybe

4:13

12 agents total, whereas the deep

4:15

research report this time, this ran like

4:17

over 100 agents. Maybe I should take a

4:19

little easy on the steep research run

4:20

because it did get hit by API rate

4:22

limits, but that's also another point of

4:24

like if you're going to spin up that

4:25

many agents at one time, you might get

4:27

rate limited. Whereas with the storm,

4:28

you know it's always going to be your

4:29

five personas. So, anyways, I think you

4:31

guys now understand the value of this

4:33

report. Let me show you real quick how

4:35

this actually works and how to get the

4:36

skill. So, there's basically four

4:38

prompts. The first one is where we tell

4:40

it to spin up the five different

4:43

angles, right? We've got these five

4:44

which I've talked about. That's prompt

4:45

one. You would just enter in your

4:47

research topic. And then when that comes

4:49

back, you would enter in prompt two,

4:50

which is the contradiction map. So, it's

4:52

saying, "Hey, where do the perspectives

4:54

contradict each other? Which one has

4:56

good evidence? Which one has weak

4:57

evidence?" And basically makes them

4:58

analyze each other's outputs. And so,

5:00

what we're doing here is we're basically

5:01

just chaining together four prompts in a

5:02

row, and then we're getting synthesis,

5:04

and then we're getting the peer review.

5:06

So, what I decided to do was I ran that

5:07

on its own. It worked great. And I said,

5:09

"Cool, package all of that into a skill

5:11

so I can literally just give you a

5:12

prompt, give you a topic, and you do

5:14

that entire thing for me, and you're

5:16

going to give me a consistent template

5:19

so that every time I run this you're

5:20

going to give me an HTML report that

5:21

always looks like this." So, what that

5:23

now looks like is in my dot Claude, I've

5:25

got a bunch of skills as you can see.

5:27

And if I go to my storm research skill

5:29

and I open up the skill.md, this is what

5:31

we've got. So, the storm research, it

5:33

turns one topic into a verified

5:35

multi-perspective HTML briefing. It

5:36

simulates five expert lenses on the

5:38

topic, maps where they contradict each

5:40

other, synthesizes everything into a

5:41

single self-contained HTML report, then

5:43

adversarially peer reviews its own

5:45

outputs, and verifies every citation

5:47

against its primary source before

5:49

delivering. You'll also notice that in

5:51

the skill we have a report template

5:52

HTML, so you guys I will also give you

5:54

guys this for completely free. This is

5:55

referenced in the skill and says, "Hey,

5:57

once you find all the information, just

5:58

put it in HTML and make sure it always

6:00

looks like this." So, that's just for

6:01

consistency on on my end, and I really

6:03

enjoy that. So, I'm going to keep going

6:05

down and explaining how this works, but

6:06

if you guys do want to go ahead and grab

6:07

these two resources, just head over to

6:09

my free school community. The link for

6:10

that is down in the description. All you

6:12

have to do is get in here, click on

6:13

classroom, and click on all YouTube

6:15

resources, and you'll be able to find

6:17

every single YouTube video and all of

6:18

the resources that I've dropped

6:20

associated with them. Once again, that's

6:21

completely free to join. Once you have

6:23

that skill, all you have to do is

6:25

you can give that markdown file and the

6:26

HTML file to Claude and say, "Hey,

6:29

Claude, this is a skill called storm

6:31

research. Put this in the .Claude

6:32

folder, and then

6:34

you're pretty much set up." And if you

6:35

guys don't know what a skill is, it's

6:36

basically just a prompt. This is

6:37

basically just a master prompt that

6:39

every time I say, "Hey, Claude, do storm

6:41

research for me," it's going to invoke

6:43

this skill, it's going to read the whole

6:44

thing, and then just run it for you. So,

6:46

it's very hands-off once you've

6:47

basically installed them. And yes, this

6:49

skill can work with codex or any other

6:51

different type of agent you want. It's

6:52

just that in Claude, it specifically has

6:54

to be in the .Claude folder. But, you

6:56

can see here I've got a folder called

6:57

.codex or .agents, and you can put

6:59

different skills in different types of

7:01

folders based on the coding agent you're

7:03

using. This currently is just a Claude

7:05

code tutorial. So, anyways, from there,

7:07

phase zero is to scope the topic.

7:09

Sometimes, if you don't give it a

7:10

specific enough topic, it will ask a few

7:12

questions before it goes ahead and kicks

7:14

off the storm. Then, it spins up the

7:15

five expert lenses in parallel, and then

7:18

we go into mapping the contradictions,

7:20

synthesizing the report, and then the

7:22

adversarial peer review verification,

7:24

and that's where we get our output. So,

7:26

let me just open up the Claude desktop

7:27

app and start a new session here. And

7:29

I'm just going to say,

7:30

"Hey, Claude, please run a storm

7:32

research for me on voice AI agents."

7:35

And so, what you'll notice here is I

7:37

didn't use a slash command, so it will

7:38

still invoke the skill. And what you'll

7:40

also notice is that

7:42

this isn't very specific, so it might

7:43

ask us some questions. Right here you

7:44

can see it says, "Okay, running the

7:45

skill storm research." So, that's how it

7:47

went ahead and looks through storm

7:48

research. And the argument that it's

7:50

currently aware of is voice AI agents.

7:53

It comes back and says, "Okay, so here's

7:55

the topic, here is the reader." And it

7:57

knows that I am an AI educator and I am

7:59

deciding on potentially whether voice AI

8:01

agents are worth a video or if it's just

8:03

hype. So, the pipeline is now running.

8:04

If I open up this, you can see it is

8:06

going to kick off those five agents, the

8:08

practitioner, the academic, the skeptic,

8:10

and all of the other ones that we need.

8:12

And what's really cool is you can click

8:13

in and see what they're doing. So, if I

8:15

click on the economist, for example,

8:17

this is the prompt that our main session

8:19

kicked off to this subagent. So, now we

8:22

can see the subagent down here is

8:23

browsing the web. It's using a tool.

8:25

It's doing research. We can click on the

8:27

academic. We can see this is the

8:28

academic prompt. And once again, the

8:30

academic subagent is, you know, doing

8:32

all the stuff down here. Now, while this

8:34

is running, let me quickly explain the

8:35

difference between having subagents and

8:39

having agent teams. So, subagents is

8:41

basically where we have one main

8:43

session. So, this session right here,

8:45

this is Claude. This is who we're

8:46

talking to. And all of these subagents

8:48

are working for this main session. So,

8:51

the main session talks to these five,

8:52

but these five cannot talk to each

8:54

other. And that is an important

8:55

distinction because that's what you

8:57

actually have in the agent team world.

8:59

You can spin up teams of agents or

9:00

councils of agents that can not only

9:02

talk to your main session, but they can

9:03

also talk to each other. And that's

9:05

really cool because what I like to do is

9:07

I like to spin up agent teams when I

9:08

need help deciding on certain ideas or

9:10

topics. And I'll have them not only do

9:12

research for me, but then I'll have them

9:14

debate with each other. So, they'll

9:15

literally argue with each other until

9:17

they reach some sort of consensus. Agent

9:19

teams are much more expensive than

9:20

subagents though. So, important

9:22

distinction. I'll bring more videos on

9:23

agent teams later, but if you do want to

9:25

check out a deep dive, check out this

9:26

video that I've tagged right up here.

9:28

And also, if you want to check out

9:30

another video where I've deep dived even

9:31

more on subagents, then you can check

9:32

out this video right up here. Now, you

9:34

can see all of these subagents ran on

9:36

Opus 4.8. If you don't want to do that,

9:38

you don't have to. You can have all of

9:39

these sub-agents run on Haiku or Sonnet

9:41

if you like. But in this case, I liked

9:43

for them to run on Opus. But all five

9:45

lenses are in, so now it's going to go

9:46

ahead and look at the contradictions. So

9:49

you can see it's reading that file,

9:50

which is the report template. It's

9:51

running these agents in the background,

9:53

and now it's going to start verifying

9:54

all of those different citations and

9:56

stats and the different things that our

9:58

initial passive agents had come up with.

10:00

Also, I know in this video I have

10:02

switched between the Claude desktop app

10:04

a little bit with VS Code. If you guys

10:06

have been watching me for a while, you

10:07

know that I typically do like to do most

10:08

of my work in VS Code. Um Claude

10:11

basically works the exact same way in

10:12

both. It's just a difference of UI,

10:14

really. But the reason I chose to show

10:16

some of this video today in the desktop

10:17

app is because I thought it's cool to

10:19

show the actual agents running in here.

10:22

But you can see that our report has come

10:23

back, so let me actually just open this

10:24

up in a browser. You can see this is the

10:27

V2 version, so everything has been

10:28

verified. If I scroll all the way to the

10:29

bottom, you can see the sources were

10:31

either demoted, corrected, or confirmed,

10:33

so that is great. We've got our

10:35

60-second summary. We've got our key

10:36

findings, which obviously once again are

10:38

ranked by reliability. So what I would

10:39

recommend for you guys to do is go grab

10:41

the skill, put it into your own Claude,

10:44

and play with it a little bit. Make it a

10:45

little bit more tailored towards you.

10:46

Maybe you can play with the HTML report

10:48

if you want it to look a certain way.

10:50

But then do it on a topic that you do

10:52

know a lot about and that's important to

10:53

you in your business. And then just read

10:54

through it and see where you need to

10:56

improve it or change it up a little bit.

10:58

And maybe you even want to add a sixth

11:00

lens or a seventh lens. Maybe for me and

11:02

my workflow, it would be helpful to add

11:03

like a beginner in AI because that's a

11:06

lot of people that we're teaching our

11:07

beginners in AI. Or maybe it would be

11:08

good for me to add to the skill a

11:10

content creator or something like that.

11:12

So I guess the reason why I'm saying

11:13

that is because I think what you should

11:14

take away overall from the video is,

11:16

yes, go grab the Storm skill and test it

11:18

out, but it's also less about this

11:21

specific skill and this specific

11:23

Stanford method being the best for

11:24

everybody, but I think the theories that

11:27

you can pull out of it like the idea

11:28

that the more perspectives you have

11:30

doing research and contradicting each

11:32

other, the better and more holistic

11:35

research you're actually going to get.

11:36

Basically, just the whole idea of if you

11:38

don't have subject matter expertise, see

11:39

if you can borrow it in some way. See if

11:41

you can go ahead and kill your own blind

11:43

spots, find the gaps in your knowledge,

11:45

and go use agents to create little

11:47

experts all over so that you've got this

11:48

council of agents that have different

11:51

expertise and different knowledge that

11:52

that have your back, no matter what

11:54

you're doing. So, I know this was a

11:55

quick one today, but hopefully you guys

11:56

enjoyed it or you learned something new.

11:58

If you did, please give it a like. It

11:59

helps me out a ton. And as always, I

12:01

appreciate you guys made it to the end

12:02

of the video, and I'll see you on the

12:03

next one.

12:04

Thanks, guys.

Interactive Summary

This video presents the 'Storm' research method, a Stanford-developed approach that uses multiple diverse expert personas (such as a practitioner, academic, skeptic, economist, and historian) to conduct research, identify blind spots, and verify information. The creator demonstrates how he implemented this as a reusable 'skill' in Claude to generate high-quality, verified HTML briefing reports. By utilizing these multiple perspectives and an adversarial peer-review process, the method creates more organized, actionable, and reliable research compared to standard single-prompt approaches. The video also covers technical implementation, including how to install the skill and how to tailor it with custom lenses to suit specific needs.

Suggested questions

4 ready-made prompts