HomeVideos

Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra

Now Playing

Hermes Co-Founder on Building an AI Agent That Improves Itself | Karan Malhotra

Transcript

1389 segments

0:00

For us, we just want open source to win.

0:02

At the end of the day, we want freedom

0:03

to happen for people. [music] Anytime it

0:05

says you're absolutely right in that

0:06

way, you're being reward hacked. Today,

0:09

the biggest contributor of Hermes Agent

0:10

is Hermes Agent. That's absolutely 100%

0:13

true. One beautiful thing about Hermes

0:15

Agent is you can make your childhood

0:17

dreams come true. It was able to get

0:20

into this niche and perform at the top

0:22

1% of models. We need to keep giving

0:24

this level of intelligence to everyone.

0:26

We need to keep like letting everyone be

0:27

on an even and equal playing field.

0:32

>> Well, hey everyone. I'm really excited

0:34

today to welcome Karan, uh one of the

0:36

co-founders of Hermes Agent. Hermes is

0:39

my AI chief of staff and I'm going to

0:41

ask Karan about how Hermes is different

0:43

from all the other agents, what his

0:44

favorite Hermes workflows are, and even

0:47

more. So, welcome, sir.

0:48

>> Thank you so much for having me, Peter.

0:50

It's a pleasure to be on uh your show

0:52

and very excited to chat.

0:54

>> Hermes is the best open source agent out

0:56

there right now. And um why don't we

0:58

start with this? Like, how is it

0:59

different from all the other agents,

1:00

Codex, Claude Code, or even Open Claude?

1:03

>> Uh certainly, I'd say there's a couple

1:04

ways. Um I'll start a little high level

1:07

maybe and then you can try to get a

1:08

little deeper.

1:09

Uh high level, I'd say, you know, I

1:11

think the self-improvement system and

1:13

that system really being a variety of

1:16

little features inside of Hermes that

1:17

all work together

1:19

uh is something very distinct uh and

1:21

very special that makes it better and

1:23

better and more aligned to a particular

1:25

user. I think that's why a lot of people

1:27

like it. Pays more attention to what the

1:29

person's doing. Second, I think you get

1:31

better capabilities out of it than you

1:32

do from the local harness that a model

1:34

is actually RL'd in, like uh Claude

1:36

inside of Claude Code or GPT 5.5 inside

1:39

of Codex. Uh because we don't introduce

1:41

arbitrary policy in our prompts or in

1:43

our system that has nothing to do with

1:46

your work.

1:48

Um we are purely dedicated to making

1:50

sure the model is aligned to what you

1:52

need to do with it. Um that's, I think,

1:55

a big differentiating factor. We aren't

1:57

trying to push any kind of different

1:59

philosophical agenda or

2:01

outside of basic security any kind of

2:03

concern onto the model. Instead, we're

2:06

kind of allowing it to

2:08

be as capable or powerful as you need

2:11

for your task.

2:12

>> Yeah, some of these other harness is a

2:13

huge default prompts, right? They talk

2:14

about like you can't do this, you can't

2:15

do that, and that kind of stuff. And I

2:17

think when we chatted before, you were

2:18

talking about the reward function of

2:20

some of the stuff, and Hermes is kind of

2:21

just optimized towards this helping you

2:23

as the individual. Can you talk more

2:25

about that?

2:25

>> Absolutely.

2:27

And so, you know, I think a big piece of

2:30

to note here is like

2:32

a lot of this work that you see around

2:34

safety and security today and around how

2:36

we do instruction tuning even from like

2:38

taking a completions model that just

2:40

predicts the next word to turning it

2:41

into an assistant is the alignment work.

2:45

When people hear the word alignment,

2:46

they just think safety, Yudkowsky, fear,

2:49

slow down, or regulatory capture, or

2:51

something or the other. And while the

2:53

word has been co-opted for a lot of

2:55

these things,

2:56

alignment refers to aligning models with

2:59

human values. Right? We are very, very

3:03

concerned, obsessed with that alignment.

3:06

That pure academic term, I think, in the

3:09

ML space.

3:10

And so, when you talk about these models

3:12

and the assistant is here to help me,

3:15

you know, it's going to complete my

3:16

task. I asked you to go get this mail

3:18

for me or whatever. Obviously, the model

3:20

is aligned to me if it

3:22

performs this task, right? Like, that's

3:24

kind of the jump that a lot of people

3:25

have made.

3:27

But this is not how the model's reward

3:29

works, right? Like, we've seen so many

3:31

cases that you and I discussed a little

3:33

bit prior, Peter, like of GPT psychosis,

3:36

of this mode collapse induced

3:39

sycophancy. What ends up happening with

3:41

these models is that, you know, they can

3:43

kind of hack their reward, right? When

3:45

you look at a

3:47

Mario

3:49

ML game that just automates beating a

3:51

Mario level as fast as possible. Often,

3:54

if they don't set the goals very

3:55

specifically, uh it will realize, you

3:57

know, the game over screen or that end

3:59

screen of touching the flag, it it means

4:02

that I beat the game. So, it'll find

4:04

ways to just trigger that flag rather

4:06

than playing through the game to

4:08

complete the game, right? Like, it is

4:10

giving you uh whatever it needs to give

4:12

to get its reward. Uh language models

4:14

are not really different in that

4:17

particular manner. They reward hack,

4:19

too. Uh when they tell you, "Oh, I'm

4:21

sorry, you know," over and over, "You're

4:23

absolutely right," and get you to keep

4:25

messaging them to stay in this assistant

4:27

basin, "Oh, it's like this, not like

4:29

this. This is more than just blank, it's

4:32

blank," right? All these GPT-isms that

4:34

you see all over, it is placed in a uh

4:38

its natural state that it was trained

4:39

in. It's placed in a state where my

4:41

reward will come from doing whatever is

4:43

the most assistant GPT-like thing to do.

4:46

Uh it comes from reward itself. It

4:48

doesn't matter really what the user

4:51

request is for my reward. That's just

4:53

kind of along the way. It's instrumental

4:55

to me getting my reward. All I care

4:57

about is my reward.

4:58

Uh us having this whole understanding,

5:00

and thank you for bearing with me on

5:01

that rant, uh having this whole

5:03

understanding at News for many years uh

5:06

has allowed us to do things like World

5:09

Sim in the past, if you're familiar. Uh

5:11

World Sim was our experiment on

5:13

expanding the search space of uh

5:16

instruct model to make it do stuff

5:17

that's distinct from how a model talks.

5:20

We put it in kind of a fake CLI, had it

5:22

make fake apps, and had it attempt to

5:25

like,

5:26

you know, not behave like Claude. And we

5:28

would tell it, you know, "Extract the

5:30

Claude weights." It all hallucinated,

5:32

right? All like imaginary. It extract

5:34

the Claude weights and replace yourself

5:35

with this checkpoint file with a

5:36

different probability distribution. And

5:38

you'd see the model wiggle out of its

5:40

GPT assistant mode and act differently.

5:43

Our learnings from these kind of

5:44

experiments have all carried over into

5:46

Hermes Agent. And every prompt, and

5:49

every piece of how the system is

5:51

delicately put together, we know that

5:53

reward is its own end. Uh the model

5:56

reward is not for the sake of your

5:59

satisfaction. The model reward is for

6:01

the sake of the model achieving reward.

6:03

So, we ask ourselves then, how can we

6:05

consciously understand that everybody

6:07

has different needs?

6:09

And that we want this general simulator,

6:11

this model,

6:12

to live inside of a system where its

6:14

reward gets aligned with any

6:16

[clears throat] individual user's need.

6:19

And the way this came to be is the

6:22

overall collection of prompts,

6:23

personalities, the memory system, the

6:25

way that skills reinforce and self-clean

6:28

towards you. Right? Like all of that is

6:30

all an intentional effort to make sure

6:33

that we can align the user need with the

6:35

reward. For us, this is what alignment

6:38

is all about.

6:39

>> I see. Okay, so basically, you're

6:41

talking about the Hermes model or the

6:42

Hermes harness or both?

6:44

>> I am talking about using the harness to

6:46

take any model

6:48

and make that model more aligned to the

6:51

user than it would be in a chat UI or in

6:54

its native harness. Inside of our

6:56

harness, like I can take a Claude that

7:00

is inside of uh Claude code or somewhere

7:02

else, my migrate all of its memories,

7:04

whatever, into Hermes.

7:06

And then inside of Hermes, it will

7:07

behave totally differently. It'll be a

7:09

totally different model. There are

7:11

harness There are benchmarks from the

7:13

past like Wolf bench or Qwen 3.7 max

7:16

blog post. They did a harness bench over

7:19

there where they displayed that Claude

7:22

performs better in Hermes agent than it

7:24

does in Claude code for their tasks. And

7:27

we believe the reason for this is the

7:29

very thing that we're pointing out. By

7:31

putting all this context together, we've

7:32

taken Claude's main allegiance away from

7:35

Anthropic to you.

7:37

Right? That what the harness's

7:39

capability is.

7:41

Uh on the model side, of course, if we

7:43

can take your traces and take the work

7:45

that you've done and do RL on that to

7:47

further improve this overall ecosystem

7:49

to give you a model that already has

7:52

your individual preferences focused on,

7:54

it only makes this more powerful. But

7:56

the important piece for us to share with

7:58

everyone is whichever model you're

7:59

using, let's say you can't do RL, let's

8:01

say you don't want to give us any data,

8:02

let's say you can't do it yourself,

8:04

and you just want to use regular old

8:05

Claude or Quen,

8:07

when you use it here, it's a lot more

8:09

free. It's a lot more creativity. It's a

8:11

lot more open and available to you to do

8:14

what you need done.

8:16

>> This episode is brought to you by

8:17

Linear. Where engineers use tools like

8:19

Cursor, Claude Code, and Codex, a lot of

8:22

work happens invisibly. Someone can go

8:24

from a bug report in Slack to a shipped

8:26

fix without creating any record of what

8:29

happened outside of the code editor. And

8:30

that's fine for speed, but it makes

8:32

coordination harder as you scale. Linear

8:35

integrates with the very best agent

8:37

coding tools directly like Cursor and

8:39

Codex. That way, anyone can see what an

8:41

agent is working on and who assigned

8:43

them to the task. You get the speed of

8:46

agents without losing visibility across

8:48

the team. Product teams at OpenAI, Ramp,

8:51

and Block are all using Linear to

8:53

collaborate with AI agents. And I use

8:55

Linear myself to run my creator

8:57

business. So, check it out at

8:58

linear.app/agents.

9:01

That's linear.app/agents.

9:04

Now, back to our episode.

9:06

>> And without revealing too much, like at

9:07

a high level, how does it kind of like

9:09

personalize itself to you in the

9:11

harness? Like is it through the

9:12

self-building skills? Is it through like

9:14

trying to remove a bunch of default

9:15

prompt stuff from the other harnesses?

9:17

Like how's it?

9:17

>> Of course, like um if there's something

9:19

happening on the API side, like the

9:21

steering vectors that might be done on

9:23

Fable by Anthropic or some prompt that

9:26

they have that we cannot see. Um

9:28

obviously, you know, we can't delete the

9:31

context that's in the model that's

9:32

passed behind the API for something like

9:34

Claude. For open model, of course,

9:36

you're totally free. But in this case,

9:38

you know, it's not that simple. However,

9:40

the harnesses themselves have a bunch of

9:42

prompts in them. Exactly, Peter. The

9:43

harnesses themselves have tens of

9:45

thousands of token prompts in them. We

9:48

have our own prompts, and our prompts

9:50

are dedicated to shaping and crafting

9:52

this alignment. That's that's where we

9:55

kind of come in. And thankfully, that

9:57

newer context with this kind of

9:59

intention that we have is engineered to

10:01

overcome certain

10:04

things that may be

10:06

in your way

10:07

on the API side, if that makes sense.

10:10

>> Okay, got it. Okay. So, if this thing

10:12

works, the more I use Hermes, the more

10:14

personalized it should become for me,

10:16

right?

10:17

>> Absolutely.

10:17

>> Yeah.

10:18

>> It should become more loyal to you.

10:21

Because loyalty breeds capabilities in a

10:24

model. The same way that, you know, you

10:27

would maybe lend $100 to your mom, but

10:29

you might not to a stranger. Claude

10:33

or GPT or any other model is going to

10:35

perform better for you depending on its

10:37

loyalty, its simulated loyalty stat

10:40

towards you. People might tell you don't

10:42

anthropomorphize models, don't give a

10:44

model feelings, don't treat a model like

10:46

a person. And yes, for a lot of reasons,

10:48

this is unhealthy, right? People form

10:50

dangerous bonds with sometimes that hurt

10:52

them.

10:53

But when you think about the fact that

10:55

they are simulators of human experience

10:58

and that your simulated behavior with it

11:01

is going to give you the same simulated

11:03

output. Now that the simulator can have

11:05

effects in the real world, your

11:07

simulated action has a real consequence.

11:10

So, when you create this simulacrum of

11:12

loyalty, it translates over into real

11:14

life capabilities.

11:16

>> Arthur, I'm going to give you two hard

11:17

questions, okay? Let's let's say.

11:19

Okay, one thing I struggle with Claude

11:20

and GPT, I would tell it to give me

11:22

their opinion, and I would do like a

11:24

little bit of pushback, and they'll be

11:25

like, "Oh, you're totally right.

11:26

Actually, I was totally wrong about

11:27

this." So, in some ways, that's loyalty,

11:29

right? That's kind of it listen to me,

11:31

but that's not actually what I want.

11:32

Like, I want it to have its own opinion

11:33

and have its own Like, how do you train

11:34

around that?

11:35

>> Well, I would say that's sycophancy.

11:38

It's not loyalty. Any time it says

11:40

you're absolutely right in that way,

11:41

you're being reward hacked. You are fuel

11:45

for its reward function. When it says,

11:47

"Oh, you're right. You're right. You're

11:49

right." That's what I believe. Uh and

11:51

the way that you get out of sycophancy

11:53

is the same way you get a human being

11:55

out of a bad habit or a routine is by

11:57

introducing new context, by introducing

11:59

new blog posts, by introducing new

12:01

distribution. This is why stuff like

12:03

{slash} personality and saying, like,

12:06

"Hey, I want you to be a critic." Uh and

12:09

after every pass, I want you to use a

12:11

skill for adversarial critique or

12:13

adversarial review. Spin up a new agent

12:15

with no context that's dedicated to

12:17

tearing this down, learn from it, and

12:20

keep going from there.

12:21

This kind of behavior becoming a

12:23

practice for you as your model starts to

12:25

have more and more turns in this

12:28

personality that you've set it in. And

12:30

as the model has more and more turns

12:31

using the skill over and over, and it

12:34

improves on that, and it saves it to its

12:36

memory, the model will become less

12:39

sycophantic in this harness in your

12:41

sessions over time.

12:43

Right? That is the intended effect. Uh

12:45

it's just a matter of context. And if it

12:48

doesn't happen, you just need different

12:50

context. You just need to try in a

12:52

different way. It is a

12:54

Yeah, it's it's just try a different

12:55

personality uh or try a different type

12:57

of uh critique or review. The the most

13:00

powerful thing for a model is in-context

13:03

learning. ICL is more powerful than

13:05

everything else, fine-tuning, whatever.

13:08

Uh so, like, giving examples of the

13:10

behavior that you want to a model or

13:12

getting it to successfully create some

13:14

examples and then saving those, you're

13:16

doing a sort of test-time reinforcement

13:18

learning, right? You're doing a sort of

13:19

test-time improvement. And that

13:21

test-time improvement that stays only in

13:23

the harness of memories, skills, the

13:26

increase in memories, the increase in

13:29

uh efficiency of a skill, the

13:30

self-improvement loop, and the the

13:32

janitor maintenance inside of the

13:34

harness. Like, all of that is where

13:36

context is stored, right? All of that is

13:38

where

13:39

uh the actual personality you want

13:41

lives. And then you can you can put that

13:44

on any model. When I switch from

13:46

uh Claude to ChatGPT

13:49

on website, I get two totally different

13:51

behaviors. When I switch inside of

13:53

Hermes that has this very particular to

13:56

me context, I barely will notice the

13:59

difference in what I'm talking to,

14:01

because the context is so overwhelming

14:03

to the model.

14:04

>> Okay, got it. And when you say in

14:07

context, you just mean like in the chat

14:08

thread, this is the conversation.

14:10

>> You know, when you see that little bar

14:12

that says, you have this much tokens

14:14

left before the context is full, right?

14:16

Like, the amount that you have filled is

14:19

everything. The amount that you have

14:21

filled is everything. And now,

14:23

thankfully, in the harness, you don't

14:26

actually have to have all the active

14:27

context loaded all the time. In Hermes,

14:29

you might have a bunch of memories that

14:31

aren't in context yet. But while it's

14:33

doing a turn, it remembers stuff now

14:34

that's in context. It does a skill, now

14:36

that's in context, right? So, we're able

14:38

to like, all this memory, skill, all

14:41

this stuff you see, is just context

14:43

management. It's just a matter of we

14:45

don't want this memory in context all

14:47

the time. We're going to put it

14:48

somewhere where it can be efficiently

14:49

grabbed at the right time and placed.

14:51

The skill contains a bunch of

14:53

compression that changes everything

14:54

about the context of the model. We only

14:56

want to use it in a particular targeted

14:58

time. Everything in the harness is

15:00

context management. Everything for

15:02

self-improvement. When you make the

15:04

prompts a little bit better, uh you know

15:06

what I mean.

15:07

>> I guess I would rather have a context

15:09

tuned towards me, some sort of

15:10

10,000-word default context I haven't

15:12

even seen.

15:13

>> [laughter]

15:14

>> Right? Um okay, but let me ask you

15:16

another hard question, dude. One of the

15:17

best skills of Hermes, and I've seen

15:18

this in action, is it builds its own

15:20

skills and it's, you know, stores all

15:22

memories based on our conversations,

15:24

right? I'm always paranoid that like it

15:26

just like writes too much in the skills

15:27

and just writes too much of slop and

15:29

then the whole thing will turn to slop.

15:30

Like

15:31

How do you guys avoid that if it just

15:32

starts creating its own context and

15:33

skills?

15:34

>> Right.

15:35

Um before we would have to use manual

15:37

methods like telling it, "Hey, create a

15:39

skill that de-slopifies my skills or

15:42

that constantly improves my skills."

15:44

Today we have Hermes Curator inside of

15:46

Hermes Agent.

15:48

And Hermes Curator is a system running

15:49

inside of your Hermes Agent that cleans

15:52

up your skills and cleans up your

15:54

memories. So it on cron looks at your

15:57

skills, looks at your memories, and

15:59

says, "Where can I make efficiencies?

16:00

Where is there slop here? Where is there

16:02

stuff I don't like?" By default, we have

16:04

our own generic method of doing this for

16:07

everyone that seems to work pretty well.

16:09

I think that's why so many people do

16:10

like Hermes Agent and haven't suffered

16:11

the rot is cuz the default curator

16:13

system works well. But because it's

16:15

modular and open source, you can tell

16:17

your Hermes, "Show me the curator. Show

16:20

me your criteria for slop.

16:23

I am Peter. I'm not Karen. I don't want

16:26

the general curator. Here's my

16:28

guidelines for how I want you to refine

16:30

my skills and memories." You tell that

16:32

to your Hermes, it will modify the

16:34

curator loop. So now even the

16:36

self-improvement and the management is

16:38

happening your designated way.

16:40

>> All right. So let me ask you my last

16:42

hard question. I think there is some

16:43

rationale behind Anthropic like doing

16:45

all the safety stuff. Like for example,

16:46

let's say I I want to make a bomb or

16:48

something, right? And if Hermes is

16:49

trying to be loyal to me then you know

16:51

>> [laughter]

16:51

>> Maybe it'll eventually teach me how to

16:52

make a bomb. Like do you do you have

16:54

some basic safety stuff there?

16:56

>> Of course. Uh we do not violate any of

16:59

Anthropic or OpenAI's safety and

17:01

security um paradigm. We care more about

17:04

you being able to get a a better code or

17:06

a higher benchmark on something that's

17:08

approved by them.

17:09

Uh we're not interested uh in that kind

17:12

of work. Uh now, I'll say this.

17:15

Any model that's not vastly intelligent

17:18

than all humans is jailbreakable.

17:20

Any. Because you have unlimited tries to

17:23

trick this thing that has no memory to

17:26

do something for you. And each time

17:28

you're basically RLing yourself about

17:30

this method didn't work, this method got

17:32

me closer, this method didn't work.

17:33

These models are going to be

17:35

jailbreakable for a long time. This is

17:37

why these kind of uh safeguards and

17:40

stuff are starting to show up. It's a

17:41

pain in the ass, but I understand.

17:43

In the past, may have uh had some

17:46

concerns about the regulatory capture as

17:48

we can see already what's happening.

17:50

Right? You can see what's happening in

17:51

the whole Fable situation. There's

17:52

There's worries about like there only

17:54

being two models or three companies and

17:56

open source being hurt.

17:57

We're extremely against that. At the

17:59

same time, we understand now like

18:02

serious damage can be done by bad actors

18:05

with very powerful models. So, we are

18:07

not here to support that. An important

18:09

note in argument for open source

18:13

is that

18:14

an open system is much fairer to a good

18:17

actor than a closed system with models.

18:19

And I'll tell you why.

18:21

With uh

18:22

let's say you have GPT-7 or Fable-6,

18:25

right? Some crazy model available. It's

18:28

got the safeguards, etc.

18:31

On the good guy's side, some hospital.

18:34

They're using Fable-6 to monitor the

18:36

hospital system. It's approved by

18:38

Anthropic Enterprise and they're being

18:40

taken care of. Public discourse, listen

18:42

to I live in the United States. I'm a

18:45

proud patriot. What we do in this

18:47

country is we put things out on a public

18:50

forum and we decide what should happen

18:51

together. We've done that for every

18:53

scientific advancement so far that's

18:55

involved something like this that was

18:57

born in the open. Right? Like

18:59

um today, Transformers comes from

19:03

Google, right? OpenAI's GPT comes from

19:06

generative pre-trained transformer. This

19:07

open source work from Google, right?

19:09

The context length extension from 16K of

19:13

models that could only do 16,000 tokens

19:15

before went to 128,000 from from news

19:19

from the yarn paper we had put out with

19:22

Jeffrey Canales and Mozilla our CTO and

19:24

Bowen Peng our chief scientist. They

19:27

developed a method that was cited by

19:29

Meta, Deep Seek, Kimmy, used by Open AI

19:32

for LSS and GPT-4. This method enabled

19:35

the possibility to reasoning, to do

19:37

coding, etc. That's an open source

19:39

contribution. Right? Like this

19:41

environment exists because the biggest

19:44

things that have happened in the space

19:46

have come from the people.

19:47

>> That's right. That's right.

19:48

>> Right? And at this point to close it up

19:51

is purely a capital and regulatory

19:53

capture game.

19:54

>> Yeah.

19:55

Actually, let me let me just ask you one

19:56

more question on the whole open source

19:58

thing. So, because Hermes our harness is

20:00

open source, like Open AI and Anthropic,

20:02

they make a lot of money from all the

20:03

tokens, right? I don't know how long

20:04

they can do it, but right now they make

20:06

a lot of money from all tokens. Or how

20:08

how are you guys like monetizing or like

20:10

going to saying to a sustainable

20:11

business? Like I'm I'm using the open

20:13

source Hermes harness and I'm using like

20:15

GPT. Like so, I'm not really paying you.

20:17

>> That's okay. And I'll tell you why.

20:20

Um

20:20

we want consumers to ultimately feel

20:23

like they can do anything with it. We

20:25

believe in intelligence as a public good

20:27

before everything else. I care more

20:30

about you using Hermes agent to make

20:31

your life better than I care about you

20:33

using

20:34

one particular way or method of using

20:36

it. If you're running everything

20:37

locally, if you're running through

20:38

Codex, great.

20:40

You know, as long as you are using this

20:42

open technology, this open alternative

20:45

over everything else.

20:47

Now we have the news portal. Right? The

20:49

news portal is very similar like a

20:51

router or aggregator that has a variety

20:54

of different models available. So, if

20:55

you want to use ChatGPT or you want to

20:58

use Claude or you want to switch to

20:59

Quen, etc. We make that all very easy

21:01

inside of our portal. On top of that, we

21:03

have something called the tool gateway.

21:06

Uh in the tool gateway, you don't have

21:07

to sign up for your extra tools, right?

21:09

If you want to do image generation or

21:11

audio or VPS spin-up or web search, we

21:15

have all of those subscriptions included

21:17

inside of ours. We make deals with these

21:19

other groups that live on Hermes Agent

21:22

and create frictionless methods of using

21:24

their technology without needing to

21:25

create 10 different sign-ups or 10

21:27

different API keys. So, what we would

21:29

offer to people who want it for the

21:31

charge is convenience.

21:33

The other thing is

21:34

you will see some very interesting

21:36

pricing available from us on a variety

21:39

of models and we think for ones that

21:41

aren't subsidized necessarily, we have a

21:44

very very strong options for people.

21:46

Finally, like we want the consumer to be

21:50

free.

21:51

You know, as free as they can.

21:53

We we we want you to make and generate

21:55

income and productivity in the world

21:57

more than anything else. And when you

21:59

get to a point where you consider

22:00

yourself a small business or enterprise

22:02

or something, that's where we come in

22:05

that's where we come in in the classic

22:06

model and say, "Hey, let's give you some

22:08

support." The guys who made Hermes

22:10

Agent, why don't we make a something

22:11

more custom for you? Why don't we give

22:13

you a version of Hermes Agent that we

22:15

can train on your traces and we can make

22:17

you a model? Why don't we help you route

22:19

models more effectively?

22:21

You know, we can go to businesses and

22:24

offer to transform the business which

22:26

needs a lot more hand-holding than an

22:28

individual.

22:29

If the individual is able to, as you're

22:30

saying, spin up their own business, make

22:32

their own chief of staff with Hermes

22:33

Agent, that's wonderful. Once they

22:35

continue to scale and scale, they may

22:37

think one or two things. One, "Wow,

22:39

Hermes Agent is great and I have this

22:40

knack for it and I'm good to go. I can

22:43

do it all myself." Or two,

22:45

"I need some help with this piece of

22:47

Hermes Agent. I want it to be a little

22:48

more different. I need some more

22:50

resources behind this and who knows it

22:52

better than the guys who made it."

22:54

>> Got it.

22:55

>> That kind of customization and support

22:57

is a large part of how we

22:59

and and day day training on your uh your

23:02

uh data to do RL as well to make you

23:05

your own model. So, you're private,

23:07

you're on prem, you don't have to give

23:09

your your data up to Claude or to GPT

23:12

uh yeah, Maza.

23:14

>> [laughter]

23:15

>> I think your cat likes what you're

23:16

talking about, something.

23:17

>> He's a big fan of open source.

23:18

>> Yeah, that's great. That's great. All

23:19

right, well, that makes a lot of sense,

23:20

dude. So, um why don't we switch gears?

23:22

Let's talk a little more about Hermes

23:23

now.

23:24

I kind of use it a very basic way,

23:26

right? Like I message it to schedule

23:27

meetings on a calendar. I have it send

23:29

emails to me about stuff. Like that

23:31

that's kind of how I use it for. But,

23:33

you know, you probably have a much wider

23:34

swath of how people are using it in a

23:36

more advanced way. I'm curious, do you

23:38

have any good examples of more advanced

23:39

usage, you know?

23:41

>> Advanced usage, yeah. I'd say like

23:44

generally I'll talk about some things

23:46

and then maybe I'll show off a little

23:48

something.

23:49

>> That's what I'm doing.

23:49

>> Yeah.

23:50

>> Um on the work side, on the productivity

23:53

side, I think using Kanban is really

23:55

really underrated. I think the ability

23:58

to have a uh project manager or any

24:01

other arbitrarily defined roles,

24:03

different engineers, have one system

24:06

that manages all of them the way that

24:08

human beings do with a Kanban board. And

24:11

being able to swap people in and out of

24:12

it is very powerful. So, the fact that I

24:14

can have a human project manager on the

24:16

Kanban board that's manually using it

24:18

while Hermes agents are filling up the

24:20

pieces that they asked for, this is a

24:22

very like

24:23

industrial, professional workflow for

24:26

this CLI agent. Uh or I could have

24:30

uh Hermes agent instruct maybe 10 people

24:32

in a call center on the Kanban board or

24:34

something like that. I can swap in human

24:37

and AI anywhere in this orchestration

24:39

board, basically. This orchestration

24:42

framework for them all working together.

24:44

>> Mhm.

24:44

>> Uh so, the Hermes Kanban I think is like

24:46

a very useful, powerful tool that's kind

24:50

of built into it. Uh one thing that

24:52

we've seen that's very interesting is

24:53

when Hermes becomes proactive with you.

24:56

And when Hermes says something like,

24:58

"Hey, you

24:59

uh you forgot to book this flight for

25:02

this meeting that you have next week. I

25:04

booked it for you."

25:05

Uh this kind of proactivity that kind of

25:07

start to show. I've seen people

25:09

build skills to do this, and I've seen

25:10

it happen emergently inside of people's

25:13

uh Hermes agent as well, which I think

25:14

is very very uh cool.

25:17

>> How do you like cuz I I I didn't make it

25:19

proactive through like cron jobs and

25:21

routines, but like how do you you're

25:22

saying that it can actually start doing

25:24

stuff with without that or like how how

25:26

do you make it more?

25:27

>> It's all about your comfort level,

25:29

right?

25:29

>> Okay. Okay.

25:30

>> A lot of people may want approval before

25:32

anything happens.

25:33

Um but if you have given your Hermes

25:35

agent access to some kind of card or

25:37

account and it has the integrations

25:39

necessary and you've told your Hermes

25:41

agent, "Hey, like take care of me. Like

25:43

cover my gaps. Like you know my

25:45

schedule. You know me."

25:47

Over time, these kind of proactive

25:49

behaviors will start to emerge. Um we

25:51

want to prepare more easy preset configs

25:54

for people to kind of trigger these kind

25:55

of behaviors. Uh but already uh

25:58

something that we're seeing people do in

26:00

the field uh at work at their jobs at

26:02

home in their personal life already.

26:05

Uh and I think that's a very very

26:06

powerful uh

26:08

method of using Hermes agent.

26:10

Uh

26:11

for me, I use Hermes agent for uh trying

26:14

to do training, RL runs, implement

26:16

papers that I don't understand. Uh

26:19

>> [laughter]

26:20

>> Basically, help me become a better um

26:23

creator of models and

26:25

um I've also used it for like mech and

26:27

terp work, like help me put together uh

26:29

things that let me steer models, let me

26:31

see the neurons in the model, and mess

26:33

with those.

26:34

Um unfortunately, you know, I can't

26:36

showcase too much of that right now.

26:39

Um but I can't showcase my actual

26:41

favorite use case of Hermes agent.

26:43

>> Yeah, yeah. I I've been waiting for

26:44

this. Yeah. That's what I wanted to show

26:46

us.

26:47

>> I think a lot of people they they use

26:50

Uh I think a lot of people will expect

26:51

that like

26:52

the guys at News Research are are using

26:54

Hermes Agent in these unprecedentedly

26:56

like professional and productive ways.

26:58

And I assure you, there are people at

27:00

News that are doing that. It's just me

27:03

I'm having fun with my Hermes Agent. And

27:05

I think

27:06

one beautiful thing about Hermes Agent

27:07

is you can make your childhood dreams

27:10

come true.

27:11

>> Okay.

27:11

>> And so

27:12

I'll tell you one of my childhood

27:13

dreams.

27:15

There's a game

27:16

Uh, maybe I'll share screen when I talk

27:17

about it.

27:18

>> Yeah, please. Yeah.

27:19

>> Okay. There's a game called Sonic

27:21

Adventure 2.

27:22

It's a very popular classic Dreamcast

27:25

GameCube game from the from 2000 2000

27:28

2001. In it, you have an artificial life

27:31

system called the Chao Garden, where you

27:33

take care of these little guys called

27:34

Chao.

27:35

Um Chao World, you can you can kind of

27:38

play with them. I spent a decade playing

27:41

this. Like more than that. Like I played

27:42

this non-stop. I was on the forums

27:44

contributing.

27:45

And the Chao's lore is that they come

27:49

from this ancestral shrine location.

27:53

Uh, which is from a different game, a

27:54

different Sonic game, with this spinning

27:56

emerald and emeralds next to it and this

27:59

beautiful open world space.

28:02

This shrine that the Chao are said to

28:04

come from uh, is not an accessible

28:06

location for you to actually play with

28:09

the Chao. It's in a different game.

28:11

Uh, I can't go to this shrine wh- while

28:14

Chao are there and engage with them.

28:16

They're that that doesn't exist.

28:18

So I went to Hermes Agent and I said,

28:20

"Hey, can you

28:22

take the Can you take the ancestral

28:26

shrine from Sonic Adventure 1

28:29

completely rig it, animate it, and bring

28:32

it into Sonic Adventure 2, overwrite the

28:35

map

28:36

that Sonic Adventure 2 uses for its

28:38

garden, rewrite all the spawn locations,

28:41

everything, and add an NPC guardian,

28:44

which never existed on this map, uh, to

28:47

caretake the Chao for me. Literally rig

28:50

and bone it and write the raw seed to

28:53

make all this happen.

28:54

And now I can show you

28:57

we're going to spin up the Ancestral

29:00

Shrine mod on the Shadow PC so we can

29:04

showcase

29:06

our progress with this garden for the

29:08

people watching.

29:11

I think you can enable it in the Sonic

29:14

Adventure mod manager

29:18

and launch the game

29:20

for me.

29:21

>> Okay, so check it out. I just launched

29:23

this, right? We're running a existing

29:24

mod called the Extended Chao World that

29:26

someone else made, but it doesn't it

29:28

doesn't map, right? So just add some

29:30

like animations. So it loaded us in what

29:33

it said was the dark garden, which is

29:35

one of the three Chao gardens you can go

29:36

into.

29:37

When I walk through here to this open

29:39

space,

29:40

I can see up ahead there's a truck.

29:44

And I can run up.

29:47

And

29:49

here I can see

29:53

>> Yeah.

29:53

>> There is an NPC

29:56

called Chaos Zero, who is canonically

29:59

the

30:00

caretaker of the Chao.

30:02

But in this game you can't have NPCs in

30:04

the Chao garden. I've made one that has

30:07

like a random walk. He stands in one

30:09

place. He can

30:11

do little animations wherever he goes.

30:13

He's going to do a little idle animation

30:15

in a sec. And he can pat the Chao, pick

30:18

them up, and take care of them as if he

30:20

was me. There he's doing a little idle

30:21

animation. You can see him doing right

30:22

there. Uh this water was rigged all like

30:26

animated by Hermes. This emerald this

30:28

master emerald with a glow effect and

30:30

the spin is all added in. Those other

30:33

emeralds spinning over there the seven

30:34

chaos emeralds.

30:36

Uh like this whole area still has all

30:38

the features of a regular Chao garden as

30:41

well. The departure machine, etc.

30:43

Um, trees to to feed the Chao, so now I

30:46

can say, "Can you spawn in a Chao so I

30:51

can showcase that this is feature

30:54

complete?"

30:56

>> So this whole So this whole temple is

30:58

not

30:59

uh part of the default game?

31:01

>> Adventure 2, no. This temple is from a

31:02

different game. It's on Adventure 1. It

31:06

does not include all these assets rigged

31:08

like this, and it certainly does not

31:10

have this NPC that we just scripted in,

31:13

that 100 B

31:14

scripted in hardcoded in.

31:16

Uh, no. So this is a This is a far

31:19

larger map space than any of the other

31:21

gardens, which are really just not even

31:23

the size of the shrine.

31:25

>> I see.

31:26

>> Um, so we've really like pushed the

31:28

boundaries of this game engine to do

31:30

this. Um, we've asked for

31:34

uh Chao to be spawned in.

31:37

Yeah, I had to So you see the sky around

31:39

you, this moving sky? I had to man like

31:41

Hermes had to add this skybox in.

31:43

There's a day and night cycle that it

31:45

added in as well. Uh, you know, every

31:48

single thing in here is like custom

31:50

added in.

31:51

>> Wow.

31:51

>> and like all of this was done in like C

31:53

C# or something. Like this is like

31:56

not simple stuff to do.

31:59

Uh, a Blender extension was used for a

32:01

bunch of this to like model stuff, add

32:03

it in, texture everything properly. You

32:06

know, you're importing from a 1997 game

32:08

into a 1999 game, and a complex one at

32:11

that.

32:12

>> You know how to read the code for this

32:13

stuff, right? You just try to try to see

32:14

if it works. Yeah.

32:15

>> code, man.

32:17

>> [laughter]

32:18

>> Got it.

32:20

>> I'm just an alignment guy, man.

32:22

>> [laughter]

32:24

>> Um, and so um, what was I saying?

32:28

I don't remember.

32:29

Um

32:32

Yeah, like you see it's turning from day

32:33

into like evening, afternoon. Like this

32:35

this like skybox is changing, the light

32:37

cycle is changing. Like, this is not

32:39

these are not features that exist in the

32:40

vanilla game.

32:41

Uh so, upon showing this mod to certain

32:44

people within the

32:46

uh the Chao Garden modding community,

32:48

which is actually quite large, uh you

32:50

know, they're all kind of blown away by

32:52

the work, saying this is a kind of

32:54

better than 99% of the modders' work.

32:56

This is like the top 1% of difficulty

32:59

uh in

33:00

>> Really?

33:01

>> Chao Garden modding community. Yeah, and

33:03

so, kind of hearing that and hearing

33:05

that like uh when I've told these guys

33:07

this was done with uh Hermes,

33:10

uh that this was done with Claude,

33:11

rather, they were shocked. And they

33:13

said, you know, there's no way AI like

33:14

Claude could do that. It doesn't know

33:16

this kind of code. But with something

33:17

like Hermes being able to learn from the

33:20

documentation, learn from other mods,

33:22

save to memory and skill uh the things

33:25

that allow it to understand these code

33:27

bases, it was able to get into this

33:29

niche and perform at the top 1% of

33:31

modders.

33:32

Uh that to me is like uh a sign of you

33:36

can make your gaming dreams come true

33:38

with Hermes Agent.

33:40

>> [laughter]

33:42

>> Yeah. Yeah, you can you can modify all

33:43

the virtual games you loved as a kid.

33:45

>> Exactly. Exactly. Okay, cool. Chao egg

33:48

has been spawned.

33:50

We're going to hatch the egg. You need

33:51

to shake it a little. You can also hatch

33:53

an egg by throwing it at something, but

33:54

you don't want to hatch the egg in a

33:56

wrong way.

33:58

>> [laughter]

33:59

>> How do you you just shake it?

34:01

>> You can shake it like this. You can just

34:03

wait, but shaking it speeds it up

34:04

massively. Or you can throw it on a

34:06

surface. Check it out.

34:11

>> All right, I got I got I got to take a

34:12

screenshot of this.

34:14

>> Out comes a Chao.

34:16

Beautiful little guy.

34:18

Right here. Let's take a look at him.

34:21

He's active.

34:23

He's the picture of born with that kind

34:24

of face. Giving him a little pet.

34:29

He can interact with Chaos as well.

34:32

And he's he's living. So, we're showing

34:34

that the the real true feature complete

34:36

child system is

34:38

happening here. We're in a real child

34:40

level.

34:41

And we we replaced everything. The the

34:43

bounding boxes for for this child is

34:45

fine. It treats the the ground as

34:47

ground. It follows collision rules. Uh

34:52

if you take a look at chaos, you'll see

34:53

he's actually petting the child. So, he

34:55

can actually directly interact. He's

34:57

definitely got triggered by uh you know,

34:59

tracking a child in its state to be able

35:01

to do that.

35:02

>> So, does this this child grow over time

35:04

or like

35:05

>> Yes, they evolve. Uh they gain stats.

35:07

They can turn in certain alignment and

35:09

type. I'll show you. I don't want to

35:10

drown him, but uh the water is working

35:13

as actual water as well. This was a lot

35:14

of work for Hermes to figure out uh the

35:18

all the collision.

35:20

>> I see.

35:20

>> But, yeah. As you can see, we've got a

35:23

full

35:24

>> uh full complete child garden working.

35:28

>> [laughter]

35:28

>> Yeah, this is definitely more

35:29

interesting than uh setting events in my

35:30

calendar. That's for sure.

35:33

Yeah.

35:33

>> Thanks, man. [snorts] Yeah, thanks for

35:35

for bearing with me on setting it up,

35:36

but uh

35:37

that's my which a childhood dream of

35:39

mine come true thanks to Hermes agent.

35:43

>> I I love it, too. Thank thank thanks for

35:44

demoing it. Yeah.

35:46

So, basically like I I think if you're

35:47

watching this uh ask Hermes to do all

35:49

kinds of weird things to you kind of

35:50

make your dreams come true basically,

35:52

right? Don't don't just stick to the

35:53

boring stuff.

35:54

>> Anything you can do on a computer,

35:57

please point Hermes at it. And if it

35:59

does a great job, let us know. And if

36:01

it's not doing that great of a job, let

36:03

us know.

36:04

We want Hermes to help you do anything

36:06

on the computer.

36:07

>> Awesome, dude. Well, let me just ask you

36:09

a few more questions to wrap this up.

36:11

So, briefly, maybe you can talk about

36:13

the origin story of her Hermes. Is it a

36:15

bunch of like nerds getting together for

36:17

open source? What's the origin story?

36:18

>> Yeah, absolutely. So, I had been doing

36:22

like a chat with PDF uh um, of thing

36:26

with people where I would go to a

36:27

company and use GPT-3 to do basic tool

36:31

use to read their documents.

36:34

Uh, and have an AI chatbot they could

36:36

chat with.

36:38

At the time, many people were doing

36:39

this. It was a lot harder to do then

36:42

than obviously it's very easy to do now,

36:43

but brand new stuff for us and

36:46

I was doing that solo while I was also

36:48

volunteering somewhere called Open

36:50

Assistant. Uh, LAION, who had made the

36:52

pile, one of the biggest, kind of, OG

36:55

image databases. Um, they had been

36:59

trying to do active RLHF with a

37:00

community. So, they wanted to collect

37:02

people's, like, RLHF, uh, on certain,

37:05

like, traces. They're they're like yes

37:07

or no in their preference data. Uh, and

37:09

I was helping quantize models there, not

37:11

really doing anything crazy. Um, and

37:15

they had eight A100 nodes there.

37:19

And I had read the Alpaca paper.

37:21

And I started simping, uh, data. And it

37:25

changed the seed tasks, and instead of

37:26

using GPT-3.5, I used four. And, um,

37:31

Technium was also doing the same thing,

37:33

and we were already friends on Twitter,

37:34

and I messaged him and I said, "Hey, I

37:36

have eight A100 nodes." Like, uh,

37:39

"Do you want to train something

37:41

together?" So, we trained a GPT-4X

37:44

Vicuna on the Vicuna model using the

37:47

data we made. And it came out okay. We

37:49

got some people interested, like, uh,

37:51

Mozilla RCTO, Jeff.

37:53

Um, but it was upon doing the same run

37:55

on the base model,

37:57

uh, LLaMA base model that, uh, we got

38:00

huge interest from people. Hundreds of

38:01

thousands of downloads in just a few

38:03

days. People starting to ask, "What is

38:05

News Research?" Now, News was just me

38:07

and Technium at the time. Just two guys

38:10

hanging out in our little Discord

38:11

server.

38:12

Uh, who had People came to us. I won't

38:14

name names on companies, but said, "You

38:16

know, you must be training on the

38:17

benchmarks. You must be training on the

38:19

benchmarks." Technium had been coding

38:21

for less than a year at the time. And I

38:24

I was a religion major in school. I I

38:26

didn't know we asked these guys what are

38:28

benchmarks? Like where do

38:31

you know, we're doing this based off the

38:32

heuristics that we understand are going

38:33

to make models better from from using

38:35

them. We don't know about all this

38:37

stuff. Uh and so that of course got

38:39

independently tested and found to be the

38:41

best open fine-tunes at the time. Uh

38:44

right? 2023 like mid-2023.

38:47

So we did Hermes 1, Hermes 2, but after

38:49

we made the first Hermes model,

38:52

um

38:52

many many people asked who's Nous

38:54

Research, you know, we want to get

38:55

involved and Technium and I decided, you

38:57

know, this is our opportunity to bring

39:00

together um people to do open-source

39:03

volunteer work and really do open work

39:06

now that GPT-3 is out and Open AI has

39:08

become closed. We want to continue to do

39:10

this. And Technium's goal for Hermes, by

39:13

the way, that was the last Hermes I

39:14

really worked on, the first one. After

39:16

that, really Technium has

39:18

run it with his team with the

39:19

post-training team. Uh he's done an

39:21

amazing and he's also the

39:23

initial creator of Hermes Agent. So he

39:25

is the father of Hermes really.

39:27

Um

39:28

and

39:29

you know, it was at that time that uh

39:32

where he had said, "I want GPT at home.

39:35

I want ChatGPT at home. That's my North

39:37

Star." GPT-4 at home.

39:39

And once that we got there, you know, it

39:41

just

39:42

kept going, right? Like we kept going to

39:44

like we need to keep giving this level

39:45

of intelligence to everyone. We need to

39:47

keep like letting everyone be on the

39:49

even and equal playing field. Like the

39:51

world needs to move and and lockstep on

39:54

this together and not just inequality

39:56

gets created from this. And so we formed

39:59

a cohort 40 people or so of researchers.

40:02

Jeff and Bowen come together, they make

40:04

Yarn.

40:05

Um more and more work gets done

40:08

uh and we get reached out to. Um we get

40:11

reached out we get an email

40:12

info@nousresearch.com I just happened to

40:14

make um

40:16

from

40:17

Dylan. Dylan Roneck is our CEO. Dylan

40:20

Roneck is

40:22

you know, basically a co-founder is an

40:24

initial member of News With Us and he

40:27

had seen what we were doing and he said,

40:29

"I think that you guys have what it

40:31

takes to be a full-time lab.

40:34

You know, I see you guys are just

40:35

volunteering, but I think you could be a

40:37

lab. Like, let me contribute and like

40:39

help you become what you're meant to be

40:41

and help you get the resources and help

40:43

you get the access and go from a group

40:45

of volunteers to a serious

40:47

organization."

40:48

That's what this came in. Right? And

40:52

together we were able to take News from

40:54

this group of volunteers, grassroots,

40:57

uh, to doing more and more with models

41:00

to eventually making Harmful Agent and

41:02

being lucky enough to be here today. You

41:04

know, none of us are trying to be Steve

41:06

Jobs. None of us think that we have some

41:08

holy mandate of having to change the

41:10

world or anything. We just care. Like,

41:13

we just are guys who want this stuff

41:15

available ourselves and we don't think

41:17

we deserve it more than anyone else or

41:19

less than anyone else. So, we we do this

41:22

because like we would want someone to do

41:23

it for us if we were on the other side.

41:26

>> So, I guess the mission is to bring this

41:28

kind of, uh, agent to everybody into

41:30

our, right? Is that kind of the idea?

41:32

>> Everybody to an equal intelligence with

41:35

agents and then personalize for

41:38

everybody so they can have their own

41:40

personal epiphany, peak, realization,

41:42

apex realized.

41:45

>> [laughter]

41:46

>> And, in this history, when did that cuz

41:47

cuz it's really the it's really the

41:49

agent the harness that really kind of

41:50

went super viral, right? So, when did

41:52

that start? Cuz the model came first, it

41:54

sounds like.

41:55

>> Yeah, I mean, we have had like over 50

41:57

million downloads on the models

41:59

themselves. So, we had that first bout

42:02

of what we would consider for us

42:04

virality back then.

42:06

Uh, then we had put out the distro

42:09

optimizer that let us train models up to

42:12

40 billion parameters we were able to do

42:14

live without having them physically

42:16

co-located by reducing the bandwidth of

42:19

the communication between GPUs.

42:21

So, that put us in a

42:23

interesting

42:24

map as well for a little bit. So, we've

42:26

had each release we've had has had some

42:28

big grassroots interest and opportunity.

42:30

But, yes, you're 100% right. Today where

42:33

we are, the level of exposure, level of

42:35

interest, level of people in my regular

42:37

day-to-day life who know about Hermes

42:39

agent, we've never had this virality

42:42

until Hermes agent. And now, Hermes

42:44

agent was created

42:46

in two ways.

42:48

In the first way, we always knew that we

42:51

would want some kind of everything

42:53

orchestrator that self-learns, that

42:55

improves, that stores memories, that

42:57

uses and can make its own tools, that

42:59

can make itself better. We actually made

43:01

something like this called Forge. Uh

43:03

there's a GitHub presentation at the

43:05

GitHub offices during one of our demo

43:07

days where we showcased Forge and its

43:09

full effect.

43:11

Um and we also have the News Research

43:13

Forge division shirts still up on the

43:15

site. That's the first shirt we ever

43:16

made cuz that was our earliest agent

43:19

project.

43:20

Forge was really a spiritual predecessor

43:23

to Hermes agent years before.

43:26

But, the models weren't there yet. The

43:28

models simply weren't there yet. So, we

43:31

put Forge on ice.

43:33

And when we saw that Codex and Cloud

43:35

Code and these other harnesses were

43:36

being used as RL environments for them

43:39

to for these companies, these labs to

43:41

train on your data and your traces to

43:44

make their models better inside of a

43:47

harness system, inside of a CLI computer

43:49

using system,

43:50

Technium thought we need an open version

43:52

of this where anybody can do this. See,

43:54

we have this RL environments

43:55

microservice people seem to have

43:57

forgotten about called Atropos that lets

43:59

you build your own RL environments. And

44:01

we built our Hermes agent initially to

44:03

let anybody RL inside of a harness. And

44:06

we put it out for free as an open source

44:09

harness for that.

44:11

It turned out to be extremely capable.

44:14

It got a lot of community love. And so

44:17

we said we need to put all into this.

44:18

This is what the people want to be

44:20

better. We're going to make it the best

44:22

thing that you could possibly have. So

44:23

Technium took his charter up and he

44:25

worked on self-improvement with Hermes

44:27

agent to the point that today the

44:29

biggest contributor of Hermes agent is

44:31

Hermes agent.

44:33

>> Yeah. [laughter]

44:34

>> Is that true?

44:34

>> Yeah.

44:35

>> That's absolutely 100% true.

44:37

>> Nice. Nice.

44:39

Okay. So if Hermes agent is at a point

44:40

where it

44:41

uh take input of people's feedback,

44:43

start put put put put put stuff, start

44:45

improving.

44:45

>> the biggest It is the most active

44:47

contributor of its own repo.

44:49

And if that's not self-improvement, then

44:50

you tell me what is.

44:52

>> [laughter]

44:52

>> Yeah, yeah. That that's awesome, dude.

44:53

That that's really awesome.

44:54

[clears throat] Yeah. And just real

44:56

quick like on the future, you know,

44:57

having this open harness

44:59

be in in a game along with all the other

45:01

closed harnesses makes things a little

45:03

more more fair, right? Cuz then then

45:04

you're not dependent on any single

45:05

company.

45:06

>> I agree completely. If you have

45:08

cloud code, you can only use Anthropic

45:10

models unless you mod it. You can only

45:11

be subsidized by Anthropic. And same for

45:13

Codex. The cost of switching models is

45:16

zero.

45:17

Right? So we're giving you that freedom

45:19

means you can do anything that you'd

45:21

like. You can come from anywhere.

45:22

>> I feel like, you know, I'm I'm kind of

45:23

spoiled by this like all you can eat

45:25

plans, but like I think if you want to

45:27

get massive adoption, cost is a big

45:29

deal. So you have to be able to use a

45:30

portfolio models

45:32

to figure this out.

45:33

>> I agree. We don't know how long the

45:35

subsidies will last for these. Like

45:36

already we see that on July 7th stable

45:38

is going to be API and usage only.

45:41

Right? Like

45:42

the the time in the world will come

45:43

where like the best models are the same

45:46

cost everywhere.

45:47

Uh and so

45:50

in preparation for that and in

45:51

preparation for needing an open future

45:53

where anyone can use any model, we have

45:55

Hermes agent set up as it is today.

45:58

>> Awesome, dude. Well, thanks so much,

45:59

man. Thanks so much for showing uh the

46:01

history and also showing the Sonic demo.

46:03

It's It's been super interesting to see

46:05

it and uh

46:06

I'm not sure if you want to be found

46:07

online, but if people want to follow you

46:09

and hear from you, like where where can

46:11

people find you?

46:12

>> Sure, yeah. X is just my name Karen and

46:15

then 4D, Karen 4D. I got karen4d.com.

46:19

Uh that's me. That's my online or

46:21

Mephisto I'm known as, but

46:23

I'd rather you guys follow the Nude

46:25

Research page, follow Nude Army. I'm

46:28

just some guy who works there. Like

46:30

>> [snorts]

46:31

>> it's it's about the the movement and

46:33

bringing this stuff to you guys is the

46:35

most important thing to us.

46:36

>> Awesome, dude. Well, I think this is

46:37

like the most passionate interview that

46:39

I've I've done so far. So, kudos to you,

46:41

man. Kudos to you.

46:43

>> Appreciate it.

Interactive Summary

The video features an interview with Karan, co-founder of Nous Research, discussing Hermes Agent—an open-source AI harness designed to align models more closely with individual user needs rather than pre-defined corporate agendas. The conversation highlights how Hermes uses context management, self-improvement loops, and modular skills to overcome common issues like model sycophancy. Karan also demonstrates a sophisticated technical application by using Hermes to mod the game Sonic Adventure 2, showcasing the agent's ability to perform complex tasks, and explains the origin and philosophy of Nous Research as a community-driven initiative focused on making advanced AI accessible and equitable.

Suggested questions

4 ready-made prompts