How to Setup Ambient AI (Full Demo)
971 segments
We own everything except for the
intelligence. Why can't we own the
intelligence? So, I got super into local
AI back in January. And so, I've been
building this kind of home AI lab. I
bought three Mac Studios with 512 GB
[music] each, which now I think are
rarer than uh yellow diamonds at the
moment. I built a computer around a 5090
RTX. So, I have a ton of GPUs, bunch of
Mac minis, running local intelligence
and all them. Fable gets banned. You now
have unlimited 247 intelligence. You can
create what I'm calling ambient
intelligence. You can't do these types
of use cases I'm about to show you with
Opus [music] with Fable because you're
going to have a $20,000 bill at the end
of the month. I built this dashboard. It
monitors my entire fleet of local
intelligence. [music] Something that's
not possible with Frontier. What are the
few biggest calls that you kind of see
coming true over the next let's call it
six to 12 months?
>> Uh I think six months from now.
>> I know something that you have been
preaching for a long time now is the
idea of sovereign AI and setting up uh
intelligent models locally like
literally physically in your office or
your space. I have so many questions
about this from why why is it valuable
to set up models locally versus just
kind of work with stateofthe-art models
that kind of they own the intelligence
um and also it feels very unattainable.
Um so like what does actually look like
to to get this all set up in practice?
>> Yeah, I've uh I've gone hard to this and
it's it's blown up over the last few
weeks or so because of some things going
on with Frontier Intelligence. So, I got
super into local AI back in January. I
discover OpenClaw. That was kind of my
big tipping point moment with my
content. This is really, really
interesting having kind of your own
personal AI agent that knows everything
about you, that has all these long deep
memories about you, but it's connected
to this intelligence that lives in the
cloud that's controlled by someone else.
And if someone decides, oh, I'm gonna
take this model away or I'm gonna change
how this model works, it kind of
destroys how your personal agent works,
right? So, you have this agent that
lives on your computer, but the one
thing that runs it all doesn't live on
your computer and lives somewhere else.
And so, I got super into local AI and
being able to kind of own the entire
stack of your AI agent, right? We own
everything except for the intelligence.
Why can't we own the intelligence? And
so I've been building this kind of home
AI lab over the last several months. I
bought three Mac studios with 512
gigabytes each, which now I think are
rarer than uh yellow diamonds at the
moment. Um I bought uh DGX Spark. Uh I
bought a I built a computer around a
5090 RTX. So I have a ton of GPUs, bunch
of Mac minis, running local intelligence
and all them. And then, you know, I've
been preaching this for the last several
months on Twitter and YouTube. You gotta
get in local AI, getting local AI. I
think a lot of people made fun of me for
it. And then there was this big moment
a few weeks ago that I think started
shifting the public perception of local
AI. I think it was two major events. One
was Fable getting banned, which thank
god it's back. It came back yesterday.
Fable gets banned. people kind of
realize, oh wait, I don't have as much
control over my intelligence as I
thought I did, right? These companies
can just take it away at any given time.
You've literally no control over it
whatsoever if it's running through the
cloud. And then GLM 5.2 comes out, which
is I think the first kind of open
weights model that you can run locally
that basically comes pretty close to
frontier. A lot of benchmarks have it at
Opus 48 level or so. I'm running it
locally right now in a Mac studio. It is
up there with Opus 48. I can tell you
that much. And so now the public
perception is changing. But I still
think there's a lot of misconceptions. I
still think expectations aren't really
set right. And so uh there's still a lot
to go over. There's still a lot to do.
I'll show you a lot of what I'm working
on here. But I think that finally the
public perception is starting to shift
on local intelligence and people are
starting to see okay now I actually do
need to own my intelligence.
>> I just want to understand so one are you
exclusively using local models or are
you also using frontier models as well?
>> Yeah this is where people get confused
and they I tweet like oh Fable is
incredible and then I get like slam dunk
in the replies. Oh, but you're telling
everyone to buy Max Studios. It's like,
no, that's not it at all. We are not at
a point right now where you can be 100%
dependent on local models. There are
still downsides to local models that
Frontier models fill in nicely, right?
Local models slower than Frontier,
especially if you're running basically
anything less than an RTX 6000, right?
If you're running Mac Studios, slower
local intelligence. Um, they're also not
as smart, right? Local models, like we
had this big breakthrough in GLM 52, but
it's still not quite as smart. It's not
as smart as Fable. It's getting close to
Opus 48, but it's, you know, it's still
behind Frontier Intelligence. Um, and
then also, you know, it's expensive
buying more and more computers, but so
there there is still downsides. So the
perfect combo is
using them together using Frontier
according to its strengths and using
local according to its strengths. And
there's also another part to it, right?
Local is still great. There's still many
use cases. I'm going to show you many of
the use cases on this call. At the same
time, this is also where people really
struggle. I I don't know why the general
a lot of people struggle with this. You
also have to have foresight, right? You
also have to be able to look a little
bit in the future. Talk about local mods
and go, "Oh, but it's slower. It's
slower. It's slower. It's slower." Yeah,
but look at the trends, right? Look at
the way things have been moving the last
couple years. A couple years ago, local
models are completely unusable.
Completely unusable.
Now, they, you know, several months ago,
they were six months behind and now
they're like three months behind. So, if
you're able to kind of have foresight
and not be a prisoner of the moment, you
can see they're getting better, they're
getting faster, they're getting smarter,
and if you, you know, project that out
maybe 6 months, maybe a year,
>> you can pretty safely say we will have
fable level intelligence locally at
pretty decent speeds. And at that point,
then yeah, we can have a conversation
around it replacing Frontier. But no, at
the moment, right, you use them
together. There's strengths to each and
we'll go over those strengths in a
second. But part of also getting into
local AI is investing in the future.
Investing in on the bet that local
intelligence get better, faster,
smarter, and more efficient where you
can run it on cheaper computers.
>> Yep. And and you're doing this
personally. Do you think this is also
going to be a trend for businesses?
>> I I mean, I have zero doubt. Uh again we
got to look at trends. We got to look at
the way things are moving. What are the
strengths of local models? The strengths
of local models are basically just costs
the electricity going into your
computer, right? You're not paying for
tokens. There's no toll booth like you
do on the cloud. So you have all these
companies like Amazon and Meta that are
now like, "Whoa, whoa, whoa. Stop this.
Let's burn as many tokens as possible."
No, that's not the goal anymore. The
goal is to be efficient, right? So cost
savings is now becoming a thing for the
first time ever with AI. Well, local AI,
it's only the cost of electricity. Two
is privacy,
right? Your data when you use the cloud,
whatever it is, you're building with
clawed code, you're building with
codecs, all your code's being exposed to
anthropic and open AI. And at the same
time, codes being and listen, I love
anthropic. I love open AI. They're
building life-changing technology, but
at the same time, they're releasing
competing products with their customers.
So, you're giving your data to these
companies who are then turning around,
and I'm not saying they're looking at
your code and then, you know, building
products based on that. I'm not accusing
them that at all, but what I am saying
is they just happen to be releasing
competing products at the exact same
time.
>> Yep. And so privacy
>> is becoming super important. And when
you run local models, your data does not
leave your computer, right? Your data
stays on the computer, doesn't get
exposed to people releasing competing
products. And so those two major trends,
I think it's pretty safe to say almost
on a guaranteed level within a year or
two, local LLMs running on edge hardware
is going to be a primary way companies
are using AI. Where do you want to start
with taking us through how to build
local?
>> Let me do this. Let me show you first
the computers I have, the decision
making behind those computers, what the
strengths of those computers are for
kind of,
>> you know, the people watching. Then I'll
go into what exactly I'm doing with
those computers and the use cases and
strengths around that. Uh, but let's
first answer kind of some questions that
people have. What is the computers you
run this on? What can they do? What
computer should I buy? I think that's
like the number one question I get is
like what computer should I buy to run
local models? These are the computers I
run local on. So, as you see, I have
three Mac Studios with 512 gigabytes
each. For those not as quite up to speed
with local models, basically the memory
is the big component here. The models
are loaded into local memory. Um, so
like when you're on a Nvidia computer,
right, you're running an Nvidia chip,
it's loaded into the VRAMm. If you're on
a Mac, it's loaded into the unified
memory, uh, and it's loaded into the
memory. And so memory is the big thing
here. If you have Mac Studios, the
strengths is it has a lot of memory. So
I have three of those. And then I also
have Nvidia computers. The Nvidia, the
advantage is kind of the the
architecture around it. Uh, the CUDA
architecture. Um, and it makes it helps
run the models extremely fast. Running
on Nvidia is much faster. So, I have
different computers for different
purposes.
>> And what's the cost of these?
>> The cost of these? Well, four months ago
when I bought them, the cost of these
Mac Studio 512 gigabytes were $10,000
each. Now, there's been a lot of changes
in the market. First of all, you can't
buy them anymore. They don't even sell
them. Um, they don't sell these
computers anymore. You can only buy Mac
Studios, I think, in the 96 gigabyte
configuration. So, a fifth of the size.
Um, I told everyone, hey, listen, I told
everyone in January, you're not going to
be able to buy these soon. You should go
out and buy them.
>> Yeah.
>> The prediction turned out to be 100%
correct. Now, if you want to buy them
resale on like eBay or wherever, they're
going for four times the price. You're
paying $40,000 each. I think like
$30,000 used. So, the prices of these
have gone up. Still obtainable though
are uh the DGX Spark. The DGX Spark is a
fantastic computer and definitely a
great place to start for a lot of
people. That has recently been raised to
I think $4,800 from $4,000.
You can still run good models on that.
And then I built this computer myself.
Um I'm also a little bit of a gamer on
my uh in my spare time late at night. So
I wanted to build a nice gaming computer
with the 5090 in it. And this cost about
$9,000 to build. The 5090 is a $4,000
GPU.
>> Got it.
>> So you you add on memory and all the
other components, you get to like 78
$9,000.
>> Yeah. Yeah. So we're basically talking
about like, you know, let's call it
30ish,000
of hardware that you built up to kind of
like have the infrastructure you need to
run your local models.
>> Exactly. Y
>> Exactly.
>> Cool. And then for those wondering at
home kind of okay what's the strength of
each as I was talking about before Mac
studio high unified memory so you can
run much larger models right I could run
GLM52 on a single Mac studio with 512
gigabytes which again opus level
technology low bandwidth though and what
that means is it's kind of an indicator
of the speed how much uh of the model it
can run at one time so it's going to be
slower then you have the AI computers
which are your plug-andplay DGX X
Sparks. You plug it in. It's an AI
workstation. These kind of have medium
unified memory and decent speed. So, you
can run decent size models and get
pretty good speed with it. This is
probably at this moment in time with the
price of Mac Studios. This is probably
the sweet spot for most people are these
AI computers like DGX Spark. And then
you have your powerhouse chips. These
are like your really powerful GPU that
you had to build computers around. the
RTX5090s, the 6000 Pros, you get lower
VRAM, although the 6000 Pro, it's a
$15,000 chip, so you actually get good
VRAM and you get very high bandwidth.
So, you get extremely fast speeds. I
mean, faster than Cloud Frontier models.
Um, this is good if you want like all
right intelligence, but at blazing fast
speed.
>> So, you can't run like GLM52 on that.
>> You cannot. Requires too much space.
But, you know, again, we're talking
about trends here. You can run smaller
Quen 36 models.
>> Yep.
>> 6 months ago, a Quen model that's like
29B, you're not getting good
intelligence. Today, Quen 36 29B, which
runs on a 5090,
you're getting like Sonnet 4 level
intelligence, which today, yeah, it's
not the greatest ever. When Sonnet 4
comes out a year ago, six months ago,
whatever it was, that was pretty
amazing.
>> And now you can run that unlimited
locally at blazing fast speeds. So, you
know, again, let's talk about trends
here. Right now, yeah, the intellig
before intelligence is bad. Now, it's
pretty good. You think about a year now,
a year from now, what you're going to be
able to run on a consumer GPU you just
put into a computer.
>> Again, we're probably talking fable
level intelligence a year from now
running locally at blazing fast speeds.
And I have a fourth category here, which
is dusty old computers. I get all the
time, oh, you know, that Mac Mini I've
had for a few years, can I run local
intelligence on that? Technically, you
can. Technically, you can run local
intelligence. Google's actually been
doing a spectacular job at this. They've
been putting a lot of energy and
research into small, efficient models
that run on basically any hardware. And
so you can run like a Gemma 4 model on a
Mac Mini. Is it Frontier? No, it's
nowhere close. But it can do small
vertical tasks that can help you out. It
can be the memory system for your Open
Claw, right? It can do embeddings,
things like that. So there are still
purposes. And that's why I tell
everyone, listen, no matter what device
you have, you need to be getting into
local AI. Will you be able to run the
greatest models ever? Maybe not, but
it's still worth getting into because as
the trends are going six months from
now, year from now, no matter what
device you'll have, you'll be doing
you'll be able to do something that
creates economic value for yourself.
>> Love it.
>> Now, let's talk about use cases.
So,
we talked about frontier intelligence.
We talked about local models, the
strengths and weaknesses. All right. So
if local models
are slower, yeah,
>> if local models are not as intelligent,
>> then why the hell would you use them?
>> Well, the answer is pretty simple. If
you have unlimited use of intelligence,
even if it isn't the smartest
intelligence ever,
that unlocks a whole new breath of use
cases, right? If you now have unlimited
247 intelligence,
you can create what I'm calling ambient
intelligence, which is intelligence that
is constantly monitoring, constantly
reacting, and constantly burning tokens,
right? Even though it's slower and
dumber, you can't do these types of use
cases I'm about to show you with Opus,
with Fable, because you're going to have
a $20,000 bill at the end of the month.
So, you have this ambient intelligence.
What does that mean? Well, first of all,
I built this dashboard myself. Well, my
Hermes agent built it. Uh, this is
basically what I call my fleet
dashboard. It monitors my entire fleet
of local intelligence. And these models,
this local intelligence is doing work
247, 365 around the clock. Something
that's not possible with Frontier. What
are those local models doing? Well, you
can see kind of my computers here that
are all running. You can see the local
models running at the moment. I'm
constantly switching these out,
experimenting, running different
evaluations to see what can do what, but
they're all doing different things all
the time around the clock. All of these
models. So, one example is security. I'm
building a SAS right now called Henry
intelligent machines. one of these local
models every hour or so is running a
loop where it is picking up a different
API endpoint in my apps. I also have
another SAS creator buddy looking at the
different API endpoints and just doing a
quick security check. Is there anything
wrong here? Is there any security issues
that I should flag? Things like that.
Just constantly doing security
surveillance. I have another local model
that's just looking at different parts
of the code every 20 minutes. Anything
we can make more optimized, anything we
can clean up, anything like that. I have
another local model that's constantly
looking at the database for anomalies,
you know, is there any users using too
much usage? Anything weird going on? Are
there churn risks I can address and send
an email to automatically? constant
ambient intelligence, just watching all
parts of my business to find
opportunities or things to fix. Again,
not possible with Frontier unless you
work at one of the Frontier Labs because
now you're paying thousands and
thousands of thousands of dollars a
month um to to use these tokens. When we
first started talking about this, I
think people are like, "Okay, Alex, like
I get it." But when you say like
sovereign AI is a big thing, I think
like until people feel the pain of
intelligence that they don't own being
taken away from them, that won't
necessarily land with them. And to the
point like Fable got taken away, but you
know, you still had Opus 4.8, you still
had Sonnet. And so it's like the
alternatives were still solid. And so I
think people are still probably like
okay but does sovereign AI really matter
especially for individuals. I think for
companies who are dealing with sensitive
information totally different argument.
I think people will feel this one way
faster which is if you want to actually
have 247 intelligence and have all these
different jobs for you
>> before you either needed a team of a ton
of people which would have cost a lot of
money or if you tried doing it right now
with whether it's Fable 4.8 8 GPT 5.5.
Either if you're on a pro or max plan,
you'll end up maxing out your usage very
quickly and you'll have to wait until
limits are reset. Or if you're on an
enterprise plan, it'll end up basically
costing the salaries of engineers you
hire just to run these jobs. So you're
basically saying, how can you do the
work you want to do but not have it be
expensive right now,
>> right? I mean, it's just leverage,
right? you're creating leverage for
yourself as an individual and you look
at the way the economy is going at the
moment, right? There's there's some job
loss going on right now, right? You need
to be able to create value for yourself.
And if you're a soloreneur, you're
building something for yourself. The
more ambient intelligence, the more
workers you have going for you working
around the clock, the more leverage you
have against your competition. And so
this is just more and more leverage.
Even if it's really simple tasks, just
doing a security check every 20 minutes,
just reviewing small lines of code every
few minutes. It's just more and more
leverage, more and more productivity and
work getting done for yourself that you
really didn't have unlocked before. And
it'll only get worse. I mean, look at
the other trends, right? Big thing here
is trends. Where's the world moving?
It's not all about what's going on in
the world right now. It's about where's
the world moving. Fable is not going to
be a part of subscriptions starting July
7th. Starting July 7th, you're going to
need to pay for every single token input
and output for Fable. The
counterargument of, oh yeah, but I pay
$20 a month for subscriptions, yet
you're going to be paying per token
soon. With Local AI, you're not paying
per token. You're paying for what? Going
into your computer, which is a very
different formula. So that's another
just huge strength of this is that you
you you have this unlimited intelligence
as well.
>> Do you have a sense of for the amount of
workloads you're running on all of your
local um machines right now? What is
this costing you um a month and then if
you were to have run all these workloads
on the state-of-the-art models, how much
do you think it would be costing you
roughly? If I was to run this on
Frontier models, I have no doubt it
would be thousands of dollars a month,
right? Uh I I know that the Claude Max
$200 a month plan, I think the metrics
came out to get like $3,000 of tokens a
month. I'm
>> maxing out that plan with Claude, but
I'm doing way more work with local. So,
it has to be at least a few thousand a
month. So, that's from the cost. How
much is it costing me from in
electricity? I'm not a 100% certain to
tell you. I mean, I think my costs I'm
in California, so electricity is
expensive anyway. I think I was paying
like $300 or $150$160 a month for
electricity before. Now I'm paying like
$200 220. So maybe like $60. So I mean
it's it's very large
>> totally
>> cost. And people go, "Oh, but you're
paying tens of thousands of dollars for
computers."
>> You're right. You do pay for some
upfront, but you know, that's the
investment into this. And I only believe
you'll be able to run better and better
models in the same hardware you own over
time.
>> Yeah. And the cost of the hardware is
going to come down over time as well.
Like the hardware is the most expensive
it will ever be for the quality that it
is now. So you'll maybe have to pay more
for better quality hardware or you can
have the same cost hardware but way
better in the future. So that makes
sense. And then the other question is
how are you figuring out with these
local models and how far you can push
the work that you're giving your local
stuff versus having to give to the
frontier models. It's actually not easy
right now to know like right you have
all these things like open router that
are routing based on the jobs needed.
But for an individual to figure out
which workloads they should be using
with which models is like a ton of trial
and error and it's very murky. How are
you figuring that out?
Yeah, I mean I treat my Hermes agent.
So, you know, I have OpenClaw and
Hermes. Hermes is kind of my frontier AI
agent. That's like my IT guy, my Hermes
agent. And so, when like new models come
out, I'll go to my Hermes. I'll say,
"Hey, go on to all my computers
areworked through tail scale. I'll say,
"Hey, go on to my DGX Spark, load up
these list of like five different
models, run an eval on each one, and
then give me a report on what their
strengths and weaknesses are." I come
back a couple hours later, and I have a
full report. It tells me, and because my
Hermes model knows me so well, it knows
my businesses, it knows my tasks, it
knows what I do on a momentto moment
basis. It goes, "You should be running
this task on here. You go on your Mac
Studio, run this task on this model, and
then go on your 5090 and run this task
here." And so I'm offloading most of
that work to my Hermes agent which
understands on a very deep level my
entire network, all my hardware, all my
tasks and can make those decisions for
me. Uh that's I mean that's another big
part of this as well I wanted to go over
which was like what else do you need?
What do you need to to make this run? If
you have tail scale and Hermes, for
those who don't know, tail scale
basically allows you to create your own
private network. As long as you have
tail scale, which connects all your
devices, and then Hermes or OpenClaw,
which is basically like your chief IT
guy who can then move around between all
your devices. You can figure all that.
It just figures it out for you. It'll go
on the devices, figure out which models
should run on each. Do research on
Hugging Face, what are the best models
at the moment, load that in based on
what it knows about your hardware, you
know, uh, load it into memory for you so
you can use it. And then run evals and
tests so you know which ones are best at
what. And so if you have these two
things, tail scale and Hermes, you can
do any of this and figure out what works
for you. So, basically the setup is you
have your uh your three Mac Studios, you
have two two Mac minis, a DJX Spark, and
then the Nvidia uh GPU.
>> Yep.
>> So, I'm all hooked up to my office. Only
one of them is connected to a monitor.
Only my Mac St. Well, my uh Nvidia one's
connected to this other monitor so I can
play Cyberpunk 2027 at night. But other
than that, they have they're all just
sitting there connected to an outlet.
None of them are connected to a monitor.
Y
>> and but they're all on tail scale. So I
can go to my Hermes and I can say, "Hey,
go to my Mac Studio 2, go to my Mac
Studio 3, go to my DGX Spark, load these
models up, run all my tasks on them, see
what does it best, and it just goes over
and does those things for you on those
computers. You don't even need to see
them connected to a monitor ever again
in your life."
>> Yeah. And so the whole purpose of tail
scale like and creating a private
network just so I understand what it
kind of means why why it's important in
practice is
>> is it so that you can basically
when you run a job does that basically
mean based on the workload that's needed
it can kind of just route to your
different machines based on the like te
take me through why it's important to
even set up a private network versus
just have these all run individually.
Why would that be harder? So putting
them all on the same private network
basically allows all your computers to
communicate with each other very easily.
Yep.
>> And so what that enables is if you have
a Hermes agent, I have it running on my
Mac Studio 1, which just basically my
only computer connected to a monitor.
>> I can say go on my DJX Spark, load the
latest Gwen 36 and move our workload
over to that. And what it'll do is
because they're all on the same network
through tail scale, it can basically
access the shell of the DGX Spark, so
basically like root admin access,
>> communicate with it from the Mac Studio
1, and from there since everything's
connected to a CLI now, be able to run
the CLI on the DGX Spark, download uh
the model, load it onto the server, and
start giving it tasks. So it basically
gives all your computers root access to
each other. So if you have any agents
running, they can all communicate with
each other, run anything they want, and
do any workloads across any of your
devices.
>> How much of this whole setup did you
have experience in before doing this?
>> Zero. None. I I had experience in none
of this. You know, I I think I think the
most valuable skill right now is the
ability to be autodidactic.
>> Yeah. uh the ability to just figure
things out on your own. And it's never
been easier to just figure things out on
your own because all you need to do is
ask your agent, hey, how the hell do I
do that? And if you get in, if you can
train yourself to, you know, most people
in this world, I'd say 95%, and I don't
want to kind of underestimate humanity.
I love humanity, but I'd say 95% of
people when they run into a challenge,
they roll over and give up.
>> I'm trying to build this thing. I need
it on a database. I don't know how a
database works. I give up. I quit.
That's how like 95% of people operate,
unfortunately.
>> But if you can train your brain now that
when you run into issues, I have all
these computers. How do I connect them?
I have this computer over here. I have
but I don't have a monitor connected to
it. How do I do things with it? If you
change your mindset from I don't know
how to do something, I give up to I
don't know how to do something. I'm
going to go to AI and just figure out
how to do it. You can just accomplish so
much more. I' I've never computers in my
life.
>> Yeah.
>> None of that. I never I still haven't
loaded an AI model locally in my entire
life. My Hermes does it all for me. But
if with AI now, you can really just
figure out anything you want to do in
the world.
>> Love it. Are there any gaps you want to
fill in on the local side or do you
think we've covered it?
>> I think we covered most of it. I mean,
there's, you know, I talked a lot about
kind of the technical side of what you
can do with local models, right? Doing
security checks, looking at code. You
know, there's a lot of other really
interesting, cool things. Even if you're
not a tech guy, even if you're not vibe
coding, I have my models every hour
going scraping Twitter, scraping Reddit,
uh scraping hacker news, product hunt,
looking for signals, looking for trends,
looking for challenges people are having
and searching for business
opportunities. Again, this is another
one of the kind of power of ambient
intelligence is my intelligence going
out and just watching the internet all
day, finding business opportunities for
me to act on, finding SAS for me to
build that solves challenges, finding
content I can write or create based on
what people are asking about. And so, no
matter what you're doing, there's always
advantages to having more intelligence.
There will never be a lack of demand for
the more intelligence you have at your
fingertips the more things you can do.
And so no matter what you're doing,
intelligence will only make your job
better and give you more leverage. And
so there's just a million different
things you can do with it. I know uh
you've generally been pretty right in
making calls. Like you talked about
making the call and the the Mac studios
ended up not being able to buy them
anymore. or you've made other calls
around like local which obviously is
becoming bigger and bigger with Fable
and now as the open source models are
just getting better. What uh what are
the few biggest calls that you kind of
see coming true over the next let's call
it 6 to 12 months? Uh I think 6 months
from now you're having Fable level
intelligence running on consumer
hardware, right? Like high-end Mac
minis, MacBook Pros be able to run that.
And so you need to be able to come up
with now
what use cases you can do within
intelligence that runs 247 because
that's not something people really do
right now. Right? Right now when people
think AI, it's call and response. I have
a question, I go to the AI, it responds.
this concept of AI watching everything
you're doing and reacting
proactively for you before you can give
a prompt or ask a question that doesn't
exist. So I would say the big thing is
now I think most people within the next
year will have intelligence running
locally on their devices. And so you
need to think about when that happens
when you have access to that technology.
What systems can you set up? If you have
your own private community, I know you
have an AI community yourself, right?
How can I offload a lot of my duties
with that community to an ambient AI
that maybe watches every single message
in that community, then reports back to
me, and then maybe creates a video or a
guide or a PDF explanation of what
everyone's talking about in that
community, right? That's a use case
right there you could never do before,
but unlocks with ambient technology. You
need to be able to plan right now what
are those new use cases that are
unlocked to ambient intelligence that
you can start planning for so that when
you do have it running on your devices
you can just plug it in get it going and
now you have a huge advantage against
your competitors.
>> Love it. So the recap on this and I
think it would be helpful for people
because to your point people are going
to have to think in a way they never
have because they're going to have
access to intelligence that can work
24/7 365 which was never a possibility.
And so as I think about like what are
the criteria for good use cases or or
jobs for ambient intelligence to do it's
kind of like what is a job that involves
always on kind of monitoring and
checking where that would just be kind
of expensive from a token uh burning
perspective for just like a traditional
frontier model. What is something that
like whether it almost doesn't matter
the time of day for it to be done. it is
going to be valuable for you regardless.
Um, and then what is something that does
not require frontier level intelligence
in order to complete the task? Are those
kind of the the main buckets if
someone's trying to think about what you
talked about like security bug fixes or
signals? Are those the main ingredients
for kind of finding the use cases for
ambient intelligence?
>> I mean, a good way to think about it is
like this, right? I'll give you a kind
of an example, then I'll give you an
exercise you can do. Right? A good
example of ambient intelligence I think
people understand kind of intuitively is
the humanoid robot. That's going to be
an example of ambient intelligence that
exists in the next year or so where
everyone's going to have a robot walking
around their house, walking around their
apartment, taking care of tasks, kind of
low intellect task very easily, doing
the dishes, folding the laundry, making
your bed, right? And and so that's going
to be the same thing but on computers
with knowledge work. Yeah. But when it
comes to local models, right? So instead
of a robot doing the dishes, putting
away clothes, folding things, it's going
to be a robot, you know, proactively
looking at your emails, your calendar.
One thing you can do is a exercise I
give a lot of people, tell them to do. I
write down my all my to-do list are on
pieces of paper on my desk, right? I got
all the tasks I do. Spend a day
writing down on a piece of paper every
task you do. Right? Everything you do in
a day, write it down what you're doing,
right? Oh, I'm I'm managing my
community. I'm writing posts on this.
I'm writing tweets. I do this YouTube
video. I program. I build this. I
respond to this. You put it all down.
Then you can go to a model. Go to Fable
if you want. I find Fable's the best,
obviously, the best model to talk to
right now. go, hey, you know, I want to
prepare for ambient intelligence. I want
to prepare for, you know, being able to
run local models. Here's everything I
did today.
What can I offload to ambient
intelligence? What could I offload to
local models? If I had right now a DGX
Spark, how would it make my life easier
with everything I do here? And you
reverse prompt it and you go, this is
everything about me, what I do. Now tell
me what I should be doing and how this
could help and you will find a lot of
different things you can be doing with
it.
>> Love it. This was great, Alex. Such a
good breakdown on local AI, why it's
becoming more important, how to actually
set it all up from the hardware to tail
scale to Hermes agent to kind of the
exercises of finding tasks for ambient
intelligence to do. So I know I learned
a ton. I know our viewer will appreciate
you taking the time and excited to have
you on again in the future.
Happy to join anytime, man.
Ask follow-up questions or revisit key timestamps.
The video features a discussion on the growing importance of sovereign, local AI, and why it is a critical investment for individuals and businesses. The speaker explains the transition from reliance on cloud-based 'frontier' models to building personal 'ambient intelligence' labs, highlighting key drivers like privacy, cost-efficiency, and resilience against service bans. The conversation covers the hardware setup (including Macs, NVIDIA GPUs, and DGX Spark), the use of networking tools like Tailscale to manage distributed workloads, and the concept of leveraging AI for 24/7 autonomous monitoring and task execution.
Videos recently processed by our community