HomeVideos

How to Setup Ambient AI (Full Demo)

Now Playing

How to Setup Ambient AI (Full Demo)

Transcript

971 segments

0:00

We own everything except for the

0:01

intelligence. Why can't we own the

0:03

intelligence? So, I got super into local

0:05

AI back in January. And so, I've been

0:08

building this kind of home AI lab. I

0:12

bought three Mac Studios with 512 GB

0:15

[music] each, which now I think are

0:16

rarer than uh yellow diamonds at the

0:19

moment. I built a computer around a 5090

0:21

RTX. So, I have a ton of GPUs, bunch of

0:24

Mac minis, running local intelligence

0:26

and all them. Fable gets banned. You now

0:30

have unlimited 247 intelligence. You can

0:33

create what I'm calling ambient

0:35

intelligence. You can't do these types

0:38

of use cases I'm about to show you with

0:40

Opus [music] with Fable because you're

0:42

going to have a $20,000 bill at the end

0:44

of the month. I built this dashboard. It

0:46

monitors my entire fleet of local

0:48

intelligence. [music] Something that's

0:49

not possible with Frontier. What are the

0:51

few biggest calls that you kind of see

0:54

coming true over the next let's call it

0:56

six to 12 months?

0:57

>> Uh I think six months from now.

0:59

>> I know something that you have been

1:01

preaching for a long time now is the

1:03

idea of sovereign AI and setting up uh

1:06

intelligent models locally like

1:08

literally physically in your office or

1:09

your space. I have so many questions

1:11

about this from why why is it valuable

1:14

to set up models locally versus just

1:17

kind of work with stateofthe-art models

1:19

that kind of they own the intelligence

1:21

um and also it feels very unattainable.

1:24

Um so like what does actually look like

1:26

to to get this all set up in practice?

1:28

>> Yeah, I've uh I've gone hard to this and

1:30

it's it's blown up over the last few

1:33

weeks or so because of some things going

1:34

on with Frontier Intelligence. So, I got

1:37

super into local AI back in January. I

1:40

discover OpenClaw. That was kind of my

1:42

big tipping point moment with my

1:44

content. This is really, really

1:46

interesting having kind of your own

1:48

personal AI agent that knows everything

1:51

about you, that has all these long deep

1:54

memories about you, but it's connected

1:57

to this intelligence that lives in the

1:59

cloud that's controlled by someone else.

2:02

And if someone decides, oh, I'm gonna

2:04

take this model away or I'm gonna change

2:06

how this model works, it kind of

2:08

destroys how your personal agent works,

2:11

right? So, you have this agent that

2:13

lives on your computer, but the one

2:15

thing that runs it all doesn't live on

2:16

your computer and lives somewhere else.

2:19

And so, I got super into local AI and

2:22

being able to kind of own the entire

2:25

stack of your AI agent, right? We own

2:27

everything except for the intelligence.

2:29

Why can't we own the intelligence? And

2:31

so I've been building this kind of home

2:34

AI lab over the last several months. I

2:37

bought three Mac studios with 512

2:39

gigabytes each, which now I think are

2:41

rarer than uh yellow diamonds at the

2:43

moment. Um I bought uh DGX Spark. Uh I

2:48

bought a I built a computer around a

2:51

5090 RTX. So I have a ton of GPUs, bunch

2:54

of Mac minis, running local intelligence

2:56

and all them. And then, you know, I've

2:58

been preaching this for the last several

3:00

months on Twitter and YouTube. You gotta

3:02

get in local AI, getting local AI. I

3:04

think a lot of people made fun of me for

3:05

it. And then there was this big moment

3:08

a few weeks ago that I think started

3:11

shifting the public perception of local

3:14

AI. I think it was two major events. One

3:17

was Fable getting banned, which thank

3:19

god it's back. It came back yesterday.

3:21

Fable gets banned. people kind of

3:24

realize, oh wait, I don't have as much

3:27

control over my intelligence as I

3:28

thought I did, right? These companies

3:30

can just take it away at any given time.

3:33

You've literally no control over it

3:34

whatsoever if it's running through the

3:36

cloud. And then GLM 5.2 comes out, which

3:40

is I think the first kind of open

3:42

weights model that you can run locally

3:44

that basically comes pretty close to

3:48

frontier. A lot of benchmarks have it at

3:51

Opus 48 level or so. I'm running it

3:54

locally right now in a Mac studio. It is

3:56

up there with Opus 48. I can tell you

3:58

that much. And so now the public

4:01

perception is changing. But I still

4:03

think there's a lot of misconceptions. I

4:06

still think expectations aren't really

4:08

set right. And so uh there's still a lot

4:12

to go over. There's still a lot to do.

4:14

I'll show you a lot of what I'm working

4:15

on here. But I think that finally the

4:17

public perception is starting to shift

4:20

on local intelligence and people are

4:21

starting to see okay now I actually do

4:23

need to own my intelligence.

4:25

>> I just want to understand so one are you

4:29

exclusively using local models or are

4:33

you also using frontier models as well?

4:36

>> Yeah this is where people get confused

4:39

and they I tweet like oh Fable is

4:42

incredible and then I get like slam dunk

4:44

in the replies. Oh, but you're telling

4:45

everyone to buy Max Studios. It's like,

4:48

no, that's not it at all. We are not at

4:51

a point right now where you can be 100%

4:54

dependent on local models. There are

4:57

still downsides to local models that

5:00

Frontier models fill in nicely, right?

5:03

Local models slower than Frontier,

5:05

especially if you're running basically

5:07

anything less than an RTX 6000, right?

5:09

If you're running Mac Studios, slower

5:11

local intelligence. Um, they're also not

5:14

as smart, right? Local models, like we

5:17

had this big breakthrough in GLM 52, but

5:21

it's still not quite as smart. It's not

5:22

as smart as Fable. It's getting close to

5:24

Opus 48, but it's, you know, it's still

5:27

behind Frontier Intelligence. Um, and

5:30

then also, you know, it's expensive

5:32

buying more and more computers, but so

5:33

there there is still downsides. So the

5:36

perfect combo is

5:39

using them together using Frontier

5:43

according to its strengths and using

5:45

local according to its strengths. And

5:47

there's also another part to it, right?

5:49

Local is still great. There's still many

5:52

use cases. I'm going to show you many of

5:53

the use cases on this call. At the same

5:57

time, this is also where people really

5:59

struggle. I I don't know why the general

6:01

a lot of people struggle with this. You

6:04

also have to have foresight, right? You

6:06

also have to be able to look a little

6:08

bit in the future. Talk about local mods

6:10

and go, "Oh, but it's slower. It's

6:11

slower. It's slower. It's slower." Yeah,

6:13

but look at the trends, right? Look at

6:15

the way things have been moving the last

6:17

couple years. A couple years ago, local

6:20

models are completely unusable.

6:23

Completely unusable.

6:25

Now, they, you know, several months ago,

6:27

they were six months behind and now

6:29

they're like three months behind. So, if

6:32

you're able to kind of have foresight

6:33

and not be a prisoner of the moment, you

6:36

can see they're getting better, they're

6:38

getting faster, they're getting smarter,

6:40

and if you, you know, project that out

6:43

maybe 6 months, maybe a year,

6:47

>> you can pretty safely say we will have

6:50

fable level intelligence locally at

6:52

pretty decent speeds. And at that point,

6:54

then yeah, we can have a conversation

6:56

around it replacing Frontier. But no, at

6:59

the moment, right, you use them

7:01

together. There's strengths to each and

7:02

we'll go over those strengths in a

7:04

second. But part of also getting into

7:07

local AI is investing in the future.

7:09

Investing in on the bet that local

7:12

intelligence get better, faster,

7:14

smarter, and more efficient where you

7:16

can run it on cheaper computers.

7:19

>> Yep. And and you're doing this

7:21

personally. Do you think this is also

7:22

going to be a trend for businesses?

7:24

>> I I mean, I have zero doubt. Uh again we

7:27

got to look at trends. We got to look at

7:28

the way things are moving. What are the

7:31

strengths of local models? The strengths

7:33

of local models are basically just costs

7:37

the electricity going into your

7:39

computer, right? You're not paying for

7:41

tokens. There's no toll booth like you

7:43

do on the cloud. So you have all these

7:46

companies like Amazon and Meta that are

7:48

now like, "Whoa, whoa, whoa. Stop this.

7:50

Let's burn as many tokens as possible."

7:52

No, that's not the goal anymore. The

7:53

goal is to be efficient, right? So cost

7:56

savings is now becoming a thing for the

7:58

first time ever with AI. Well, local AI,

8:01

it's only the cost of electricity. Two

8:04

is privacy,

8:07

right? Your data when you use the cloud,

8:11

whatever it is, you're building with

8:12

clawed code, you're building with

8:14

codecs, all your code's being exposed to

8:16

anthropic and open AI. And at the same

8:19

time, codes being and listen, I love

8:21

anthropic. I love open AI. They're

8:23

building life-changing technology, but

8:26

at the same time, they're releasing

8:28

competing products with their customers.

8:30

So, you're giving your data to these

8:32

companies who are then turning around,

8:34

and I'm not saying they're looking at

8:36

your code and then, you know, building

8:38

products based on that. I'm not accusing

8:39

them that at all, but what I am saying

8:42

is they just happen to be releasing

8:45

competing products at the exact same

8:47

time.

8:48

>> Yep. And so privacy

8:51

>> is becoming super important. And when

8:53

you run local models, your data does not

8:56

leave your computer, right? Your data

8:58

stays on the computer, doesn't get

9:00

exposed to people releasing competing

9:02

products. And so those two major trends,

9:05

I think it's pretty safe to say almost

9:07

on a guaranteed level within a year or

9:09

two, local LLMs running on edge hardware

9:14

is going to be a primary way companies

9:16

are using AI. Where do you want to start

9:19

with taking us through how to build

9:20

local?

9:20

>> Let me do this. Let me show you first

9:25

the computers I have, the decision

9:27

making behind those computers, what the

9:29

strengths of those computers are for

9:31

kind of,

9:32

>> you know, the people watching. Then I'll

9:34

go into what exactly I'm doing with

9:36

those computers and the use cases and

9:38

strengths around that. Uh, but let's

9:40

first answer kind of some questions that

9:42

people have. What is the computers you

9:44

run this on? What can they do? What

9:46

computer should I buy? I think that's

9:47

like the number one question I get is

9:48

like what computer should I buy to run

9:50

local models? These are the computers I

9:53

run local on. So, as you see, I have

9:56

three Mac Studios with 512 gigabytes

9:59

each. For those not as quite up to speed

10:02

with local models, basically the memory

10:05

is the big component here. The models

10:08

are loaded into local memory. Um, so

10:12

like when you're on a Nvidia computer,

10:15

right, you're running an Nvidia chip,

10:16

it's loaded into the VRAMm. If you're on

10:19

a Mac, it's loaded into the unified

10:21

memory, uh, and it's loaded into the

10:23

memory. And so memory is the big thing

10:25

here. If you have Mac Studios, the

10:27

strengths is it has a lot of memory. So

10:29

I have three of those. And then I also

10:31

have Nvidia computers. The Nvidia, the

10:33

advantage is kind of the the

10:35

architecture around it. Uh, the CUDA

10:37

architecture. Um, and it makes it helps

10:40

run the models extremely fast. Running

10:41

on Nvidia is much faster. So, I have

10:44

different computers for different

10:46

purposes.

10:47

>> And what's the cost of these?

10:49

>> The cost of these? Well, four months ago

10:52

when I bought them, the cost of these

10:54

Mac Studio 512 gigabytes were $10,000

10:57

each. Now, there's been a lot of changes

11:00

in the market. First of all, you can't

11:02

buy them anymore. They don't even sell

11:03

them. Um, they don't sell these

11:04

computers anymore. You can only buy Mac

11:06

Studios, I think, in the 96 gigabyte

11:08

configuration. So, a fifth of the size.

11:10

Um, I told everyone, hey, listen, I told

11:13

everyone in January, you're not going to

11:15

be able to buy these soon. You should go

11:17

out and buy them.

11:18

>> Yeah.

11:18

>> The prediction turned out to be 100%

11:20

correct. Now, if you want to buy them

11:22

resale on like eBay or wherever, they're

11:25

going for four times the price. You're

11:26

paying $40,000 each. I think like

11:29

$30,000 used. So, the prices of these

11:33

have gone up. Still obtainable though

11:36

are uh the DGX Spark. The DGX Spark is a

11:39

fantastic computer and definitely a

11:41

great place to start for a lot of

11:42

people. That has recently been raised to

11:45

I think $4,800 from $4,000.

11:49

You can still run good models on that.

11:51

And then I built this computer myself.

11:54

Um I'm also a little bit of a gamer on

11:56

my uh in my spare time late at night. So

11:59

I wanted to build a nice gaming computer

12:00

with the 5090 in it. And this cost about

12:03

$9,000 to build. The 5090 is a $4,000

12:07

GPU.

12:08

>> Got it.

12:09

>> So you you add on memory and all the

12:10

other components, you get to like 78

12:12

$9,000.

12:13

>> Yeah. Yeah. So we're basically talking

12:14

about like, you know, let's call it

12:16

30ish,000

12:18

of hardware that you built up to kind of

12:20

like have the infrastructure you need to

12:22

run your local models.

12:24

>> Exactly. Y

12:26

>> Exactly.

12:26

>> Cool. And then for those wondering at

12:29

home kind of okay what's the strength of

12:31

each as I was talking about before Mac

12:33

studio high unified memory so you can

12:36

run much larger models right I could run

12:39

GLM52 on a single Mac studio with 512

12:42

gigabytes which again opus level

12:44

technology low bandwidth though and what

12:46

that means is it's kind of an indicator

12:48

of the speed how much uh of the model it

12:51

can run at one time so it's going to be

12:53

slower then you have the AI computers

12:55

which are your plug-andplay DGX X

12:57

Sparks. You plug it in. It's an AI

12:58

workstation. These kind of have medium

13:01

unified memory and decent speed. So, you

13:03

can run decent size models and get

13:05

pretty good speed with it. This is

13:07

probably at this moment in time with the

13:09

price of Mac Studios. This is probably

13:11

the sweet spot for most people are these

13:13

AI computers like DGX Spark. And then

13:16

you have your powerhouse chips. These

13:19

are like your really powerful GPU that

13:21

you had to build computers around. the

13:24

RTX5090s, the 6000 Pros, you get lower

13:27

VRAM, although the 6000 Pro, it's a

13:29

$15,000 chip, so you actually get good

13:31

VRAM and you get very high bandwidth.

13:34

So, you get extremely fast speeds. I

13:36

mean, faster than Cloud Frontier models.

13:39

Um, this is good if you want like all

13:42

right intelligence, but at blazing fast

13:44

speed.

13:45

>> So, you can't run like GLM52 on that.

13:49

>> You cannot. Requires too much space.

13:51

But, you know, again, we're talking

13:53

about trends here. You can run smaller

13:57

Quen 36 models.

13:59

>> Yep.

14:00

>> 6 months ago, a Quen model that's like

14:03

29B, you're not getting good

14:05

intelligence. Today, Quen 36 29B, which

14:09

runs on a 5090,

14:12

you're getting like Sonnet 4 level

14:15

intelligence, which today, yeah, it's

14:17

not the greatest ever. When Sonnet 4

14:20

comes out a year ago, six months ago,

14:22

whatever it was, that was pretty

14:24

amazing.

14:25

>> And now you can run that unlimited

14:27

locally at blazing fast speeds. So, you

14:30

know, again, let's talk about trends

14:32

here. Right now, yeah, the intellig

14:34

before intelligence is bad. Now, it's

14:36

pretty good. You think about a year now,

14:38

a year from now, what you're going to be

14:39

able to run on a consumer GPU you just

14:41

put into a computer.

14:43

>> Again, we're probably talking fable

14:44

level intelligence a year from now

14:46

running locally at blazing fast speeds.

14:49

And I have a fourth category here, which

14:51

is dusty old computers. I get all the

14:53

time, oh, you know, that Mac Mini I've

14:55

had for a few years, can I run local

14:56

intelligence on that? Technically, you

15:00

can. Technically, you can run local

15:02

intelligence. Google's actually been

15:04

doing a spectacular job at this. They've

15:06

been putting a lot of energy and

15:08

research into small, efficient models

15:11

that run on basically any hardware. And

15:14

so you can run like a Gemma 4 model on a

15:17

Mac Mini. Is it Frontier? No, it's

15:20

nowhere close. But it can do small

15:23

vertical tasks that can help you out. It

15:26

can be the memory system for your Open

15:29

Claw, right? It can do embeddings,

15:31

things like that. So there are still

15:33

purposes. And that's why I tell

15:35

everyone, listen, no matter what device

15:37

you have, you need to be getting into

15:39

local AI. Will you be able to run the

15:41

greatest models ever? Maybe not, but

15:44

it's still worth getting into because as

15:46

the trends are going six months from

15:48

now, year from now, no matter what

15:50

device you'll have, you'll be doing

15:51

you'll be able to do something that

15:52

creates economic value for yourself.

15:55

>> Love it.

15:55

>> Now, let's talk about use cases.

15:59

So,

16:01

we talked about frontier intelligence.

16:04

We talked about local models, the

16:05

strengths and weaknesses. All right. So

16:07

if local models

16:10

are slower, yeah,

16:12

>> if local models are not as intelligent,

16:15

>> then why the hell would you use them?

16:18

>> Well, the answer is pretty simple. If

16:22

you have unlimited use of intelligence,

16:25

even if it isn't the smartest

16:27

intelligence ever,

16:29

that unlocks a whole new breath of use

16:33

cases, right? If you now have unlimited

16:37

247 intelligence,

16:39

you can create what I'm calling ambient

16:41

intelligence, which is intelligence that

16:45

is constantly monitoring, constantly

16:48

reacting, and constantly burning tokens,

16:51

right? Even though it's slower and

16:54

dumber, you can't do these types of use

16:57

cases I'm about to show you with Opus,

16:59

with Fable, because you're going to have

17:01

a $20,000 bill at the end of the month.

17:03

So, you have this ambient intelligence.

17:06

What does that mean? Well, first of all,

17:08

I built this dashboard myself. Well, my

17:10

Hermes agent built it. Uh, this is

17:13

basically what I call my fleet

17:14

dashboard. It monitors my entire fleet

17:17

of local intelligence. And these models,

17:21

this local intelligence is doing work

17:24

247, 365 around the clock. Something

17:27

that's not possible with Frontier. What

17:30

are those local models doing? Well, you

17:31

can see kind of my computers here that

17:33

are all running. You can see the local

17:35

models running at the moment. I'm

17:36

constantly switching these out,

17:38

experimenting, running different

17:39

evaluations to see what can do what, but

17:42

they're all doing different things all

17:45

the time around the clock. All of these

17:47

models. So, one example is security. I'm

17:51

building a SAS right now called Henry

17:53

intelligent machines. one of these local

17:55

models every hour or so is running a

18:00

loop where it is picking up a different

18:03

API endpoint in my apps. I also have

18:05

another SAS creator buddy looking at the

18:08

different API endpoints and just doing a

18:10

quick security check. Is there anything

18:12

wrong here? Is there any security issues

18:14

that I should flag? Things like that.

18:16

Just constantly doing security

18:18

surveillance. I have another local model

18:21

that's just looking at different parts

18:23

of the code every 20 minutes. Anything

18:25

we can make more optimized, anything we

18:27

can clean up, anything like that. I have

18:29

another local model that's constantly

18:31

looking at the database for anomalies,

18:35

you know, is there any users using too

18:37

much usage? Anything weird going on? Are

18:40

there churn risks I can address and send

18:42

an email to automatically? constant

18:45

ambient intelligence, just watching all

18:48

parts of my business to find

18:51

opportunities or things to fix. Again,

18:54

not possible with Frontier unless you

18:57

work at one of the Frontier Labs because

18:59

now you're paying thousands and

19:00

thousands of thousands of dollars a

19:01

month um to to use these tokens. When we

19:05

first started talking about this, I

19:07

think people are like, "Okay, Alex, like

19:09

I get it." But when you say like

19:11

sovereign AI is a big thing, I think

19:13

like until people feel the pain of

19:17

intelligence that they don't own being

19:19

taken away from them, that won't

19:21

necessarily land with them. And to the

19:22

point like Fable got taken away, but you

19:25

know, you still had Opus 4.8, you still

19:27

had Sonnet. And so it's like the

19:28

alternatives were still solid. And so I

19:32

think people are still probably like

19:33

okay but does sovereign AI really matter

19:35

especially for individuals. I think for

19:37

companies who are dealing with sensitive

19:38

information totally different argument.

19:42

I think people will feel this one way

19:45

faster which is if you want to actually

19:48

have 247 intelligence and have all these

19:51

different jobs for you

19:53

>> before you either needed a team of a ton

19:55

of people which would have cost a lot of

19:56

money or if you tried doing it right now

19:58

with whether it's Fable 4.8 8 GPT 5.5.

20:03

Either if you're on a pro or max plan,

20:07

you'll end up maxing out your usage very

20:10

quickly and you'll have to wait until

20:12

limits are reset. Or if you're on an

20:14

enterprise plan, it'll end up basically

20:16

costing the salaries of engineers you

20:18

hire just to run these jobs. So you're

20:21

basically saying, how can you do the

20:22

work you want to do but not have it be

20:24

expensive right now,

20:26

>> right? I mean, it's just leverage,

20:28

right? you're creating leverage for

20:30

yourself as an individual and you look

20:33

at the way the economy is going at the

20:35

moment, right? There's there's some job

20:37

loss going on right now, right? You need

20:40

to be able to create value for yourself.

20:43

And if you're a soloreneur, you're

20:45

building something for yourself. The

20:48

more ambient intelligence, the more

20:50

workers you have going for you working

20:51

around the clock, the more leverage you

20:53

have against your competition. And so

20:55

this is just more and more leverage.

20:57

Even if it's really simple tasks, just

20:59

doing a security check every 20 minutes,

21:02

just reviewing small lines of code every

21:05

few minutes. It's just more and more

21:07

leverage, more and more productivity and

21:08

work getting done for yourself that you

21:11

really didn't have unlocked before. And

21:13

it'll only get worse. I mean, look at

21:14

the other trends, right? Big thing here

21:16

is trends. Where's the world moving?

21:18

It's not all about what's going on in

21:19

the world right now. It's about where's

21:20

the world moving. Fable is not going to

21:23

be a part of subscriptions starting July

21:25

7th. Starting July 7th, you're going to

21:28

need to pay for every single token input

21:32

and output for Fable. The

21:34

counterargument of, oh yeah, but I pay

21:37

$20 a month for subscriptions, yet

21:39

you're going to be paying per token

21:41

soon. With Local AI, you're not paying

21:43

per token. You're paying for what? Going

21:45

into your computer, which is a very

21:47

different formula. So that's another

21:49

just huge strength of this is that you

21:52

you you have this unlimited intelligence

21:54

as well.

21:55

>> Do you have a sense of for the amount of

21:57

workloads you're running on all of your

21:59

local um machines right now? What is

22:02

this costing you um a month and then if

22:06

you were to have run all these workloads

22:08

on the state-of-the-art models, how much

22:10

do you think it would be costing you

22:11

roughly? If I was to run this on

22:15

Frontier models, I have no doubt it

22:18

would be thousands of dollars a month,

22:20

right? Uh I I know that the Claude Max

22:23

$200 a month plan, I think the metrics

22:25

came out to get like $3,000 of tokens a

22:28

month. I'm

22:29

>> maxing out that plan with Claude, but

22:32

I'm doing way more work with local. So,

22:35

it has to be at least a few thousand a

22:37

month. So, that's from the cost. How

22:40

much is it costing me from in

22:41

electricity? I'm not a 100% certain to

22:44

tell you. I mean, I think my costs I'm

22:47

in California, so electricity is

22:48

expensive anyway. I think I was paying

22:51

like $300 or $150$160 a month for

22:55

electricity before. Now I'm paying like

22:56

$200 220. So maybe like $60. So I mean

23:00

it's it's very large

23:02

>> totally

23:02

>> cost. And people go, "Oh, but you're

23:04

paying tens of thousands of dollars for

23:06

computers."

23:07

>> You're right. You do pay for some

23:10

upfront, but you know, that's the

23:12

investment into this. And I only believe

23:14

you'll be able to run better and better

23:16

models in the same hardware you own over

23:18

time.

23:18

>> Yeah. And the cost of the hardware is

23:19

going to come down over time as well.

23:21

Like the hardware is the most expensive

23:22

it will ever be for the quality that it

23:25

is now. So you'll maybe have to pay more

23:26

for better quality hardware or you can

23:28

have the same cost hardware but way

23:31

better in the future. So that makes

23:32

sense. And then the other question is

23:34

how are you figuring out with these

23:36

local models and how far you can push

23:39

the work that you're giving your local

23:42

stuff versus having to give to the

23:43

frontier models. It's actually not easy

23:45

right now to know like right you have

23:47

all these things like open router that

23:48

are routing based on the jobs needed.

23:49

But for an individual to figure out

23:51

which workloads they should be using

23:52

with which models is like a ton of trial

23:54

and error and it's very murky. How are

23:56

you figuring that out?

23:58

Yeah, I mean I treat my Hermes agent.

24:02

So, you know, I have OpenClaw and

24:04

Hermes. Hermes is kind of my frontier AI

24:06

agent. That's like my IT guy, my Hermes

24:09

agent. And so, when like new models come

24:11

out, I'll go to my Hermes. I'll say,

24:14

"Hey, go on to all my computers

24:16

areworked through tail scale. I'll say,

24:18

"Hey, go on to my DGX Spark, load up

24:20

these list of like five different

24:22

models, run an eval on each one, and

24:24

then give me a report on what their

24:26

strengths and weaknesses are." I come

24:28

back a couple hours later, and I have a

24:30

full report. It tells me, and because my

24:32

Hermes model knows me so well, it knows

24:33

my businesses, it knows my tasks, it

24:35

knows what I do on a momentto moment

24:37

basis. It goes, "You should be running

24:39

this task on here. You go on your Mac

24:41

Studio, run this task on this model, and

24:43

then go on your 5090 and run this task

24:46

here." And so I'm offloading most of

24:48

that work to my Hermes agent which

24:51

understands on a very deep level my

24:54

entire network, all my hardware, all my

24:56

tasks and can make those decisions for

24:58

me. Uh that's I mean that's another big

25:00

part of this as well I wanted to go over

25:03

which was like what else do you need?

25:04

What do you need to to make this run? If

25:07

you have tail scale and Hermes, for

25:09

those who don't know, tail scale

25:10

basically allows you to create your own

25:12

private network. As long as you have

25:14

tail scale, which connects all your

25:16

devices, and then Hermes or OpenClaw,

25:19

which is basically like your chief IT

25:21

guy who can then move around between all

25:23

your devices. You can figure all that.

25:26

It just figures it out for you. It'll go

25:28

on the devices, figure out which models

25:31

should run on each. Do research on

25:33

Hugging Face, what are the best models

25:35

at the moment, load that in based on

25:37

what it knows about your hardware, you

25:39

know, uh, load it into memory for you so

25:41

you can use it. And then run evals and

25:44

tests so you know which ones are best at

25:46

what. And so if you have these two

25:48

things, tail scale and Hermes, you can

25:50

do any of this and figure out what works

25:52

for you. So, basically the setup is you

25:54

have your uh your three Mac Studios, you

25:58

have two two Mac minis, a DJX Spark, and

26:01

then the Nvidia uh GPU.

26:04

>> Yep.

26:05

>> So, I'm all hooked up to my office. Only

26:07

one of them is connected to a monitor.

26:09

Only my Mac St. Well, my uh Nvidia one's

26:12

connected to this other monitor so I can

26:14

play Cyberpunk 2027 at night. But other

26:17

than that, they have they're all just

26:19

sitting there connected to an outlet.

26:21

None of them are connected to a monitor.

26:22

Y

26:23

>> and but they're all on tail scale. So I

26:25

can go to my Hermes and I can say, "Hey,

26:28

go to my Mac Studio 2, go to my Mac

26:30

Studio 3, go to my DGX Spark, load these

26:32

models up, run all my tasks on them, see

26:35

what does it best, and it just goes over

26:37

and does those things for you on those

26:39

computers. You don't even need to see

26:40

them connected to a monitor ever again

26:42

in your life."

26:43

>> Yeah. And so the whole purpose of tail

26:45

scale like and creating a private

26:47

network just so I understand what it

26:48

kind of means why why it's important in

26:50

practice is

26:51

>> is it so that you can basically

26:54

when you run a job does that basically

26:56

mean based on the workload that's needed

26:59

it can kind of just route to your

27:01

different machines based on the like te

27:03

take me through why it's important to

27:04

even set up a private network versus

27:06

just have these all run individually.

27:07

Why would that be harder? So putting

27:10

them all on the same private network

27:13

basically allows all your computers to

27:15

communicate with each other very easily.

27:17

Yep.

27:18

>> And so what that enables is if you have

27:20

a Hermes agent, I have it running on my

27:22

Mac Studio 1, which just basically my

27:24

only computer connected to a monitor.

27:26

>> I can say go on my DJX Spark, load the

27:29

latest Gwen 36 and move our workload

27:32

over to that. And what it'll do is

27:34

because they're all on the same network

27:36

through tail scale, it can basically

27:39

access the shell of the DGX Spark, so

27:42

basically like root admin access,

27:44

>> communicate with it from the Mac Studio

27:47

1, and from there since everything's

27:49

connected to a CLI now, be able to run

27:52

the CLI on the DGX Spark, download uh

27:56

the model, load it onto the server, and

27:59

start giving it tasks. So it basically

28:01

gives all your computers root access to

28:04

each other. So if you have any agents

28:06

running, they can all communicate with

28:08

each other, run anything they want, and

28:10

do any workloads across any of your

28:12

devices.

28:13

>> How much of this whole setup did you

28:16

have experience in before doing this?

28:18

>> Zero. None. I I had experience in none

28:20

of this. You know, I I think I think the

28:23

most valuable skill right now is the

28:26

ability to be autodidactic.

28:29

>> Yeah. uh the ability to just figure

28:31

things out on your own. And it's never

28:34

been easier to just figure things out on

28:36

your own because all you need to do is

28:39

ask your agent, hey, how the hell do I

28:42

do that? And if you get in, if you can

28:45

train yourself to, you know, most people

28:49

in this world, I'd say 95%, and I don't

28:51

want to kind of underestimate humanity.

28:53

I love humanity, but I'd say 95% of

28:56

people when they run into a challenge,

29:00

they roll over and give up.

29:02

>> I'm trying to build this thing. I need

29:04

it on a database. I don't know how a

29:06

database works. I give up. I quit.

29:08

That's how like 95% of people operate,

29:10

unfortunately.

29:12

>> But if you can train your brain now that

29:14

when you run into issues, I have all

29:16

these computers. How do I connect them?

29:18

I have this computer over here. I have

29:21

but I don't have a monitor connected to

29:22

it. How do I do things with it? If you

29:24

change your mindset from I don't know

29:27

how to do something, I give up to I

29:29

don't know how to do something. I'm

29:30

going to go to AI and just figure out

29:32

how to do it. You can just accomplish so

29:34

much more. I' I've never computers in my

29:37

life.

29:37

>> Yeah.

29:38

>> None of that. I never I still haven't

29:39

loaded an AI model locally in my entire

29:41

life. My Hermes does it all for me. But

29:43

if with AI now, you can really just

29:45

figure out anything you want to do in

29:47

the world.

29:48

>> Love it. Are there any gaps you want to

29:50

fill in on the local side or do you

29:51

think we've covered it?

29:52

>> I think we covered most of it. I mean,

29:54

there's, you know, I talked a lot about

29:56

kind of the technical side of what you

29:58

can do with local models, right? Doing

30:00

security checks, looking at code. You

30:02

know, there's a lot of other really

30:04

interesting, cool things. Even if you're

30:05

not a tech guy, even if you're not vibe

30:07

coding, I have my models every hour

30:11

going scraping Twitter, scraping Reddit,

30:15

uh scraping hacker news, product hunt,

30:17

looking for signals, looking for trends,

30:21

looking for challenges people are having

30:24

and searching for business

30:25

opportunities. Again, this is another

30:27

one of the kind of power of ambient

30:29

intelligence is my intelligence going

30:33

out and just watching the internet all

30:34

day, finding business opportunities for

30:37

me to act on, finding SAS for me to

30:40

build that solves challenges, finding

30:42

content I can write or create based on

30:44

what people are asking about. And so, no

30:47

matter what you're doing, there's always

30:50

advantages to having more intelligence.

30:54

There will never be a lack of demand for

30:56

the more intelligence you have at your

30:58

fingertips the more things you can do.

31:00

And so no matter what you're doing,

31:02

intelligence will only make your job

31:04

better and give you more leverage. And

31:06

so there's just a million different

31:07

things you can do with it. I know uh

31:09

you've generally been pretty right in

31:11

making calls. Like you talked about

31:13

making the call and the the Mac studios

31:15

ended up not being able to buy them

31:16

anymore. or you've made other calls

31:17

around like local which obviously is

31:19

becoming bigger and bigger with Fable

31:21

and now as the open source models are

31:23

just getting better. What uh what are

31:25

the few biggest calls that you kind of

31:28

see coming true over the next let's call

31:30

it 6 to 12 months? Uh I think 6 months

31:34

from now you're having Fable level

31:36

intelligence running on consumer

31:39

hardware, right? Like high-end Mac

31:42

minis, MacBook Pros be able to run that.

31:45

And so you need to be able to come up

31:49

with now

31:50

what use cases you can do within

31:54

intelligence that runs 247 because

31:56

that's not something people really do

31:57

right now. Right? Right now when people

32:00

think AI, it's call and response. I have

32:03

a question, I go to the AI, it responds.

32:06

this concept of AI watching everything

32:09

you're doing and reacting

32:11

proactively for you before you can give

32:13

a prompt or ask a question that doesn't

32:15

exist. So I would say the big thing is

32:18

now I think most people within the next

32:20

year will have intelligence running

32:23

locally on their devices. And so you

32:25

need to think about when that happens

32:27

when you have access to that technology.

32:30

What systems can you set up? If you have

32:32

your own private community, I know you

32:34

have an AI community yourself, right?

32:36

How can I offload a lot of my duties

32:40

with that community to an ambient AI

32:42

that maybe watches every single message

32:44

in that community, then reports back to

32:47

me, and then maybe creates a video or a

32:51

guide or a PDF explanation of what

32:53

everyone's talking about in that

32:54

community, right? That's a use case

32:56

right there you could never do before,

32:57

but unlocks with ambient technology. You

33:00

need to be able to plan right now what

33:03

are those new use cases that are

33:04

unlocked to ambient intelligence that

33:07

you can start planning for so that when

33:08

you do have it running on your devices

33:11

you can just plug it in get it going and

33:13

now you have a huge advantage against

33:15

your competitors.

33:16

>> Love it. So the recap on this and I

33:19

think it would be helpful for people

33:20

because to your point people are going

33:22

to have to think in a way they never

33:23

have because they're going to have

33:24

access to intelligence that can work

33:26

24/7 365 which was never a possibility.

33:31

And so as I think about like what are

33:32

the criteria for good use cases or or

33:36

jobs for ambient intelligence to do it's

33:39

kind of like what is a job that involves

33:42

always on kind of monitoring and

33:45

checking where that would just be kind

33:47

of expensive from a token uh burning

33:50

perspective for just like a traditional

33:52

frontier model. What is something that

33:55

like whether it almost doesn't matter

33:58

the time of day for it to be done. it is

34:00

going to be valuable for you regardless.

34:03

Um, and then what is something that does

34:06

not require frontier level intelligence

34:10

in order to complete the task? Are those

34:11

kind of the the main buckets if

34:13

someone's trying to think about what you

34:15

talked about like security bug fixes or

34:18

signals? Are those the main ingredients

34:20

for kind of finding the use cases for

34:22

ambient intelligence?

34:24

>> I mean, a good way to think about it is

34:26

like this, right? I'll give you a kind

34:29

of an example, then I'll give you an

34:30

exercise you can do. Right? A good

34:33

example of ambient intelligence I think

34:35

people understand kind of intuitively is

34:37

the humanoid robot. That's going to be

34:39

an example of ambient intelligence that

34:42

exists in the next year or so where

34:44

everyone's going to have a robot walking

34:45

around their house, walking around their

34:47

apartment, taking care of tasks, kind of

34:50

low intellect task very easily, doing

34:52

the dishes, folding the laundry, making

34:54

your bed, right? And and so that's going

34:57

to be the same thing but on computers

35:00

with knowledge work. Yeah. But when it

35:01

comes to local models, right? So instead

35:03

of a robot doing the dishes, putting

35:06

away clothes, folding things, it's going

35:08

to be a robot, you know, proactively

35:10

looking at your emails, your calendar.

35:12

One thing you can do is a exercise I

35:15

give a lot of people, tell them to do. I

35:17

write down my all my to-do list are on

35:19

pieces of paper on my desk, right? I got

35:20

all the tasks I do. Spend a day

35:24

writing down on a piece of paper every

35:27

task you do. Right? Everything you do in

35:30

a day, write it down what you're doing,

35:32

right? Oh, I'm I'm managing my

35:33

community. I'm writing posts on this.

35:35

I'm writing tweets. I do this YouTube

35:37

video. I program. I build this. I

35:40

respond to this. You put it all down.

35:42

Then you can go to a model. Go to Fable

35:44

if you want. I find Fable's the best,

35:46

obviously, the best model to talk to

35:47

right now. go, hey, you know, I want to

35:50

prepare for ambient intelligence. I want

35:52

to prepare for, you know, being able to

35:54

run local models. Here's everything I

35:58

did today.

35:59

What can I offload to ambient

36:02

intelligence? What could I offload to

36:04

local models? If I had right now a DGX

36:07

Spark, how would it make my life easier

36:10

with everything I do here? And you

36:11

reverse prompt it and you go, this is

36:14

everything about me, what I do. Now tell

36:16

me what I should be doing and how this

36:18

could help and you will find a lot of

36:20

different things you can be doing with

36:22

it.

36:22

>> Love it. This was great, Alex. Such a

36:25

good breakdown on local AI, why it's

36:28

becoming more important, how to actually

36:29

set it all up from the hardware to tail

36:32

scale to Hermes agent to kind of the

36:35

exercises of finding tasks for ambient

36:38

intelligence to do. So I know I learned

36:40

a ton. I know our viewer will appreciate

36:42

you taking the time and excited to have

36:44

you on again in the future.

36:46

Happy to join anytime, man.

Interactive Summary

The video features a discussion on the growing importance of sovereign, local AI, and why it is a critical investment for individuals and businesses. The speaker explains the transition from reliance on cloud-based 'frontier' models to building personal 'ambient intelligence' labs, highlighting key drivers like privacy, cost-efficiency, and resilience against service bans. The conversation covers the hardware setup (including Macs, NVIDIA GPUs, and DGX Spark), the use of networking tools like Tailscale to manage distributed workloads, and the concept of leveraging AI for 24/7 autonomous monitoring and task execution.

Suggested questions

4 ready-made prompts