Hermes Agent powered by local models on the DGX Spark is basically magic
688 segments
Okay, so this is really sick. I just set
up a Hermes agent on this Nvidia DGX
Spark completely powered by a local
model that's running on it. I now have
an AI agent, a 24/7 AI employee working
for me completely on this local model on
this device, completely private, secure,
and best of all, free. Just the power of
the energy going into this device. And
look how cool this is. Isn't this the
coolest looking computer ever? In this
video, I'm going to show you how to do
the exact same thing. I'm going to show
you how to set up Hermes agent, have it
set up on a local model that will be
running on this DGX Spark. You can put
it on any device you want. For this, I'm
using the DGX Spark. I'll show you how
easy that is. And I'll show you how you
can have your full-time 24/7 AI employee
that's doing work for you at all times.
I'll even show you some really awesome
use cases that can be really helpful.
Whether you've never loaded up a local
model in your life, you're an expert, or
you're just curious in how all of this
stuff works, you're going to learn a ton
in this video. Now, let's lock in and
get into it. So, this is going to be
really fun. I'm going to show you like
the most cutting-edge technology ever.
Quick shout-out. Shout-out Nvidia.
Thanks so much for sponsoring this
video. If you watch this channel, you
know I don't take a lot of sponsorships.
99.9% of my videos are not sponsored. I
try to only work with companies that are
really amazing that I use daily. When
Nvidia reached out, I was super pumped.
One of the greatest companies of all
time. Shout-out Nvidia. Thanks for
sponsoring this. Let's go into the
strengths of local models and why a
local model with Hermes agent is so
sick. even send me a computer. I bought
this one myself months ago. So, that's
why this is really easy for me to do.
But, let's get into this. Let's talk
about local models, what makes them so
amazing, and why it's so powerful with
Hermes agent. Then, we'll quickly jump
into setting it up and having a lot of
fun in setting up experiments and doing
a lot of cool things. If you need to,
chapters down below, jump around if you
need to. Local models, I believe are the
future. I believe very soon everyone
will have their own super intelligence
on their desk. Why is that? Well, local
models are free. You can download them
online, load them onto your computer,
and they're basically free. Just you're
paying the cost of electricity, right?
So, the power powering the computer,
running local models does use up more
electricity. So, you're paying for the
power, you pay for the computer itself,
but unlike AI models running on the
cloud, you're not paying for every token
you use. You're not paying for
subscription plans, nothing like that.
Just the power going into the computer
and the computer itself. So, that's one,
unlimited usage of super intelligence,
pretty incredible. Next up, and this is
a big one, it's completely private. When
you use cloud models, right? So, when
you go to any of the AI websites and you
start chatting with their models, your
prompts and all your chat logs are
stored in the cloud. They're stored on
servers somewhere else. And well, I
don't think any of the employees are
going and reading your chat logs or
anything like that. Having your most
important conversations being private is
really important. So, you get the
privacy. If I unplug the internet from
this DGX Spark right now, I can still
keep using it. I can still chat with the
model and use my Hermes agent cuz it's
all local and private and secure. So,
everything stays local, everything stays
in your computer, no one can snoop on
your chat messages. That's big. It's
customizable. So, the models themselves
are customizable. You can install these
things called LoRAs, which are basically
plugins for the model, which allow you
to customize them really any way you
want. If you want them to really learn
your voice, if you want them to learn of
how to make specific types of images,
you can customize these models any way
you want. I can literally customize
right now so it sounds like me. So, all
of its output sounds like me, which is
amazing.
So, customization, that's big. It's
educational. You're learning about AI,
right? When you install it and start
using it locally, you learn how it
works. And I think this is the most
important technology of all time. I
think it's super important you learn it.
So, when you download it onto your
computer, you start using you learn
about the most important technology I
think humanity has ever invented. That's
huge. It's just straight up fun. It's
fun doing this. It's fun looking at your
desk, seeing a DGX Spark there, any
computer you have, and knowing there's
like superintelligence running on that
device. It's really, really fun. You're
allowed to have fun. You're allowed to
do things for the fun of it. Not
everything needs to be a hyper-optimized
dollar spend. You're allowed to just
have fun. And the last one here before
we get into actually promise we're about
to have a little fun here, use cases.
When you can run a model 24/7 for just
the cost of power, it unlocks so many
use cases. And I'll go into those use
cases in a second, but you really get a
powerful 24/7 AI employee. So, real
quick, why the DGX Spark? Everyone knows
I got a whole bunch of devices. I got
Macs, got everything under the sun here.
Why the DGX Spark though? Why do I like
this device? It is super easy. It makes
this whole process super easy. You buy
it, you plug it in, and it just works.
It turns on and you can immediately
start putting models on it. It's also It
also has the full suite of Nvidia
developer tools on it, so you can do all
the customization that's really only
possible on Nvidia chips. So, if you're
into customization as well, this is like
the best way to do it. So, the Spark is
a great way to run these models and run
Hermes on it. And I'll show you in a
second how easy it is for the agents you
have on your computer to actually manage
this device and install different models
on it. And by the way, that's not even
part of the sponsorship. I just threw
that in cuz I've been using this device
for a long time, and I use it every day,
and I love it. So, let's go into setting
it up. Let's get the device set up. I'll
speed through actually setting up the
computer, and we'll spend most of our
time setting up the agent and the model
itself. Feel free again jump around the
different chapters below if you want to
hear about any specific subject. First
thing you need to do though is actually
plug the computer in. Unlike other
computers out there, you do not need a
monitor for this computer. You can run
this in what is called headless mode.
Headless mode basically means there's no
monitor, you don't have to control it,
you don't have to touch it. It's just
powered on and you're able to use it
basically as a tool for whatever your
main computer is. So, once you have it
plugged into the wall, into the power,
it just turns on and it starts up a
local network. If you look at your
instruction manual with the DGX Spark,
it has actually like network information
on it. You just connect to that network
on your main device. Now your main
device can control that DGX Spark. So,
what you're going to want to do for this
is you're going to want to install
Hermes Agent on your main computer.
Powered by a cloud model at first, we're
going to have the local Hermes Agent
work side-by-side with our cloud agent
so we can get this all set up. Once you
have Hermes Agent set up, you have it
connected to a cloud model, whether it's
a ChatGPT account or using the Anthropic
API, whatever you want, we can now get
to work in setting up our local models
and setting up our new DGX Spark. Now
that we got the Spark plugged in, this
is the prompt I'm going to give to my
Hermes Agent. I'll put this down below
if you're setting this up with me. I
just purchased a new DGX Spark and want
to set it up. I want to run it headless
and I want you to be able to control it.
I'd like you to walk through setting up
the Spark then install Tailscale on it
so you can control it from any device.
So, a whole bunch of things going on
here, let me explain. Number one,
remember as I explained before headless,
that means no monitor or other devices
are going to control it. And number two
is Tailscale. For those not familiar,
Tailscale is like one of the most
incredible free pieces of software out
there. It's basically going to allow all
your devices to create its own private
network. So, your Spark and whatever
other computers you're using will be
able to be on the same kind of virtual
network so they can talk to each other
and control each other. So, what this
prompt will do is have your Hermes agent
walk through setting up your Spark.
It'll have you connect to the network on
the manual. It'll walk through
step-by-step. And then it'll go on there
once you're connected, set up the
computer itself, and then install
Tailscale so that moving forward it's a
lot easier for your Hermes agent to
actually control that device and use it.
It's going to allow your Hermes agent to
go on it, install local models on it,
set them up, run it, and then talk to
the local model anywhere you are in the
world cuz it'll be on that same private
network. So, you hit that, you're going
to be all set up. It's super easy.
That's the beauty of these AI agents.
That's the beauty of using Hermes agent
here is that it makes it super easy to
do whatever you want. Before this, it
would be super intimidating to do
something like this. It'd be super
intimidating to set up a supercomputer
on your desk. People would go, "Oh."
They throw their hands and go, "Oh, I'm
not tech. I'm a non-technical person.
This isn't for me." If you're watching
this video right now and you're
thinking, "Oh, this is not for me. I'm
super non-technical." Throw the word
non-technical in the garbage can. Take
it out of your lexicon. Non-technical
does not exist anymore. Everyone is a
technical person. As long as you got an
AI model, as long as you got a Hermes
agent, it doesn't matter what it is. You
go to your agent, you say, "This is what
I want to accomplish. You accomplish it
for me." And it will figure it out.
Everyone is technical now. I don't want
you to limit yourself by calling
yourself a non-technical person. Every
time someone in my life says, "Oh, I'm a
non-technical person." I I tell them to
kick rocks. Figure it out. Download an
AI model and figure it out. You got
this. I promise you. So, you got the DGX
Spark set up. Next, we're going to
download and install a local model. For
this video, we're going to use Qwen 3.6
27B. Qwen 3.6 is the latest set of
models from Qwen. This is, in my opinion
the strongest local model there is. It
is super fast, it is super efficient,
and at the same time, it is super
intelligent. It is up there with a lot
of the frontier models from the last few
months. So, it is very, very good. So,
here's what we're going to do now that
you have the Spark set up. Say this, "I
want to download and install Qwen 3.6
27B onto the DGX Spark so I can use it
to power AI agents and chat with it."
27B just means it's 27 billion
parameters, which is kind of on the
medium-to-smaller side, but through
tests I've done and through a lot of
tests other people have done, it's still
incredibly smart, but you're going to
get the benefit of really, really good
speeds. So, I'm going to put this down
below as well, so feel free to steal
that. That will go and have your agent
go into your DGX Spark, search the web
on there for that model Qwen 3.6 27B,
find the version that's appropriate for
the Spark, download it, which might take
a little bit. These models are pretty
chunky, but it'll fit onto your new
computer, and then actually load it into
memory. Right, these models, the way
they work is they're loaded onto your
memory, and then you can start using it
and communicating with it. This took
about 20 minutes or so on my computer.
It might take about the same for you
depending on your internet speeds and
all that. So, do that. Feel free to
pause here if you're doing this
alongside of me and come on back. So, if
this is your first time doing this and
you just set it up and it downloaded
Qwen and then got it loaded into the
memory, first of all, congratulations to
you. You just did something that like
0.000001%
of the world has ever done, which is
host your own super intelligence on your
desk. Do yourself a favor, look down at
your computer and go, "Holy crap, I have
super intelligence on my desk. I am a
sovereign individual. Nobody can take
this away from me. No one can cut off my
AI service. No one can take this away.
Even if the internet goes out forever, I
will have super intelligence to get me
through it. So, shout out you, this is
your first time. Congratulations. Now,
let's get into actually using that super
intelligence. So, here's the plan what
we're going to do here. We're going to
set it up inside Hermes agent. I'm going
to show you how to do that. And then I'm
going to show you three really awesome
use cases for Hermes running on local
models. I just like to do one fun thing
real quick before we get into that. And
that is set up a front end chat
interface for our model. This allows you
to just test it real quick and kind of
feel that magical moment of, "Oh my god,
super intelligence is talking to me
locally." So, do this. Please build a
front end for our new local model. Make
it a chat interface and test that it
works. I will put that down below as
well, so you can steal that if need be.
And when you hit enter in that, your
Hermes agent's going to go. It's going
to actually build out a really quick
front end interface, or really simple
code it out, and then it will connect it
to your local model. So, mine is all
built out here. Look at this. This is
awesome. Spark 1 3627B. Let's give it a
test. Hey, are you there? And enter and
now let's see how we do. Is our local
model going to work here? Is it going to
chat with us? Super intelligence, are
you working? Oh, there it is. Yes, I'm
here. How can I help you today? That's
amazing. That came straight from our
computer on our desk. If you're not
geeking out with me, then you're not
alive. That's amazing. So, we've
confirmed it's working. We confirmed our
local model is talking to us. We
confirmed the super intelligence is
doing its thing. Now, let's actually
plug this into Hermes agent so we can
have our local working for us. So, what
makes Hermes agent
really cool is it is built in to be
multi-agent. What that means is it is
just one command away from Hermes being
able to create an entirely new Hermes
agent that can run side by side with
your current one. So, you can do really
cool workflows with multiple agents.
What we're going to do is we're going to
have it set up a second Hermes agent for
us and we're going to say plug it in to
our new Qwen 36 model that's running on
our DGX Spark. So, let's do this. I just
So, I'm going to use this prompt. I'll
put this down below as well if you want
to take it. I just installed Qwen 36 27B
on our DGX Spark. Please set up a new
Hermes profile that is plugged into that
local model and name it Qwen. I like
that name Qwen. So, I'm going to name
our new Hermes agent Qwen. Let's hit
enter on that. Now our Hermes agent's
going to go and it's going to create
this new Hermes profile which is
basically just a second Hermes agent and
it's going to plug it into that model
running on our Spark which is going to
be sick.
All right, here we go. It found the
llama server where the uh local model is
running. It's going to ask us for
permission. Let's give it the okay on
this and it's going to start getting to
work building out that new Hermes agent.
So, we'll have two Hermes agents working
for us, two full-time employees, one
that's going to be completely free run
on the Spark and I'll walk through all
of the different use cases and how do
you do the two models together after
this. All right, looks like it's all
set. All right, so to run this all we
need to do is hermie p Qwen. That's
going to run the Hermes profile for Qwen
which they just set up. So, let's get
this popping. Let's do this. So, I'm
going to open up my terminal here. I'm
going to paste in hermes -p Qwen. I'm
going to hit enter and boom, there we
go. Look at that Hermes agent. Oh,
that's incredible. There it is Qwen 36.
That's the model we got running locally.
We now have an AI employee 24/7 running
locally on our desk on our computer.
Let's say hey here. Let's see how it's
doing. Hey, are you there, Qwen? Come
on, here we go. Hey, I'm here. I'm
actually Hermes agent not Qwen but I'm
ready to help. What do you need? Okay,
so it doesn't know its name but it's
here. It's working. We can talk to it. I
guess when you set up a profile it
doesn't give it its name. So, let's give
it its name now. By the way, your name
is Quen. And that will save it to
memory. So, moving forward, we can refer
to our new AI agent employee by its
proper name, Quen. Got it. Quen it is.
I'll go by that from now on. And then
you can see this is a cool part about
Hermes. It tells you every tool call,
every memory updates. Updating the
memory, the user refers the assistant as
Quen. Really, really cool. Now, let's go
into the three use cases for our local
agent. We have a local 24/7 AI employee.
What do you do with it? It's probably
the number one question I get. It's
like, "Oh, what are the use cases you
do?" Let's go through three. I'm going
to give you a beginner one, a moderate
one, and an advanced one. So, no matter
where you are, you're going to have a
use case that works for you. Let's start
with the beginner one. I strongly
believe everyone should be investing. I
don't think that's a crazy take. I do a
lot of investing myself. My favorite way
to do investing and learn about
investing and get better at it is by
using my AI agents. And so, what we're
going to do is we're going to have Quen
set up a daily report for us that is
going to investigate AI stocks and AI
companies cuz I personally, not
financial advice, think that's a great
place to be investing. So, let's do
this. All right. So, here's the prompt
I'm going to do for this. Please, every
morning at 9:00 a.m. research AI stocks
for me. These are stocks for companies
that stand to benefit from a 10-year AI
build-out. Give me a report on the top
companies that have great moats and
great businesses and tell me why they
have moats. I want to invest in
companies that don't really have
competition or are the very best in
their field. So, that's why I want to
know about these moats. What's great
about Hermes agent is you can have these
scheduled tasks. So, every day at
specific times, it does things for me.
So, I do this. I have this built out.
It's going to now send this to me every
morning at 9:00 a.m. Every morning at
9:00 a.m. it will deliver me this report
so I can read it, see if there's
businesses I want to invest in, make
sure I spend my money wisely, and make
sure I can grow in the future. Now, my
AI employee is going to be a stock
researcher for me. This is really
beginner. Anyone can do this and you're
going to get value out of doing this.
What this is doing in the background is
just scheduling a cron job. For those
that don't know, cron jobs are basically
just scheduled task for computers. In
this case, it is for our agent. And now
every day at 9:00 a.m., the cron job
will fire. It'll tell our agent, which
is running on our local model, to go and
do the research for us and give us that
report. And boom, look at this. Done.
The cron job is scheduled here. The
details, daily AI stock report, every
day at 9:00 a.m. Next run's going to be
tomorrow, 9:00 a.m. And it is all set.
Each morning will search for the current
stock prices and news. That is sick. So,
now we have our daily research reports
up. Let's get into the intermediate use
case for a local AI agent.
And again, this is completely free other
than the cost of the electricity going
into your computer. So, you basically
have your own personal local free AI
researcher, which is amazing. Let's get
into the moderate use case here. This
next one is really cool. It's something
I do often and that is repurpose
content. What I'm going to do is I'm
going to give our local agent a link to
a YouTube video. Even if you're not a
content creator, this will be helpful
for you. What I like to do is I give it
links to my own YouTube videos and say
repurpose this into a newsletter for me.
But what you can do is say, "Hey, check
out this video link, get the transcript,
and tell me what lessons we can learn
from it." You can do it on this video
right now that we have here. So, let's
do this. I'm going to grab one of my
past videos and I'm going to say, "Get
me the transcript of this video, then
repurpose it into a newsletter." And I
paste it in and I'm going to hit enter
and it will be off to the races. It's
going to go And one thing Hermes does
really well is just figure out how to do
things. It'll figure out how to get that
transcript, download it, and then
repurpose it new content for us. Again,
for you, you can have it get the
transcript and then take learnings from
it, right? Say, "Hey, Hermes agent,
check this out and see what we can learn
from this video. Apply new skills. Get
me the top lessons." and it'll do it for
you. So, so many things you guys can do
with this as well. If you want to do as
long as I mean, just take the link to
this video and give it to your agent.
And this is so cool. Look at this. It's
figuring out the skill to get YouTube
content. It's building its own skills.
Got the full transcript. Here's the
newsletter version and boom, it's
writing out the newsletter for me. It
just like so amazing to think about how
this is happening all locally, all just
on the computer on my desk. This is all
just getting figured out. I mean, we
really live in the most amazing time
ever. So incredible. And here's a little
twist on this use case because it's
happening locally, right? Because I'm
not paying for tokens on a cloud, maybe
I schedule this so every hour it
searches YouTube for a new AI video,
downloads it, gets the transcript, gets
the lessons, and now my Hermes agent is
automated improving itself every hour
searching for AI videos, getting lessons
from it, self-improving itself, and
doing that non-stop. That's incredible
to think about. That's something you
can't really do with cloud models cuz
you'd be paying for every single cycle,
for every single token. Because this
just costs the electricity going in
computer, you can just have this go all
day, right? Just constantly downloading
new videos and self-improving itself. It
really is amazing the possibilities to
think about with the Hermes and the
local model. All right, here we go.
Let's get into the advanced use case and
that is vibe coding. We are going to
have our Hermes agent vibe code a to-do
list app for us. So, very simple, just
for demonstration purposes. You can now
have your local models do vibe coding
for you. This is one of the most
compute-intensive activities people do.
Some people spend thousands of dollars a
month on these different vibe coding
tools. You can now get it just for the
cost of electricity. So, let's have Qwen
vibe code for us. Really is amazing
thing about this, you just now have
unlimited vibe coding, but that's one of
the benefits. That's how you get an ROI
on these investments in these computers
is doing things like this. So, let's do
this. Please build me a to-do list app
where I can add new task with priorities
and dates for those tasks. Make the app
beautiful and clean, and we're going to
hit enter. And now our 24/7 AI employee
is going to be building out us out an
entire app. All done locally. Really,
really cool. And what's great is you can
even go and you can plug this local
model in to the many various vibe coding
tools out there. There are ways to plug
it into Claude code, Codex, open code,
whatever you want to use. You can now
leverage this local model in any vibe
coding tool you want, but Hermes is a
great coder. So, using Hermes to build
things out works totally fine. All
right, look at this. It is all complete.
Open in your browser. All right, we're
going to open this up. Here's what it
does. Add task, mark complete, delete
tasks, filter between all active, done.
Wow, it's got a lot in it. This is
amazing. Let's go. And this is all built
locally. How incredible is that? All
right, let's do this. Let's take the
command, and we are going to run it so
we can use our new app. And it opens up.
Let's see what we got. Let's see what
was built 100%
locally. Look at this. This is actually
sick. This looks like it was made by
like a cloud top-of-the-line frontier
model. how good Qwen 3.6 is. Let's do
this.
Edit the DGX Spark Hermes video. Let's
add this task. Boom, there it is. And
you see that nice animation? That
animation was so sick. Let's do I want
to see that again. Make another video.
Add the task. See, oh, that is nice.
That is nice. I love that. Now you have
your free unlimited vibe coding tool
ready to go. You can build whatever you
want. No limitations at all. That is the
beauty of this. This is amazing.
Anything you want to build, you now can
build it. All done locally on your desk.
Really, really well done. If you learned
anything at all, make sure to leave a
like down below. Subscribe. Turn on
notifications. All I do is make amazing
videos about AI. I do live bootcamps on
vibe coding and building things every
single Friday in the vibe coding
academy. Link down below for that. So,
check that out as well. Shout out Nvidia
for partnering up on this. I've had the
DGX Spark forever. I've been running AI
agents on it forever. So, for them to
reach out was amazing. So, thank you so
much for that. Also, another thing, make
sure to hit the link down below for more
information on the Nvidia DGX Spark. An
incredible device to be hosting your own
models on. Appreciate you guys watching.
Let me know down below in the comments
what you want to see next. Do you want
to see more use cases for Hermes agents?
want to see more advanced workflows? I'd
love to hear your opinion. Thank you for
watching. It truly means the world and
I'll see you in the next video.
Ask follow-up questions or revisit key timestamps.
This video demonstrates how to set up an AI-powered 'employee' using a local model on an Nvidia DGX Spark, ensuring a private, secure, and free-to-operate setup. The creator walks through the configuration of an AI agent, the installation of the Qwen 3.6 27B model, and provides examples of practical use cases such as automated stock research, content repurposing, and 'vibe coding' web applications.
Videos recently processed by our community