The BEST AI Video Strategy No One Is Using
539 segments
With today's video models, you can do
whatever you want. Like this. Want to
learn how? Let me show you. So, this
video pipeline is nuts, and I'm really
surprised more people aren't using
something like this. Um this is
extremely powerful and hooks, like you
just saw. It's also really good in ad
creative, as well as organic stuff. Uh I
have a friend who is worth many tens of
millions of dollars that's currently
generating over 2,000 AI videos a week
with something like this. So, it's not
just changing an outfit. I have kind of
this like little USB thing. Um well, I
can do it just like this. Alternatively,
let's say I wanted to, I don't know,
change the time of day. Um re-lighting
and stuff like that is pretty expensive
if you think about typical productions.
Now, I have the ability to change
lighting in a flash. So, I'm going to
show you all that in this video. You're
going to learn everything you need in
order to be able to develop really cool
visuals like this. So, what is the
actual workflow or like pipeline for
this? Well, it's pretty straightforward.
To start, you need some form of source
video. And the source video in our case
is going to be what you guys just saw me
show you earlier. So, if I just head
over to finder here, you can see that I
basically have this folder called AI
video set up. And inside of that folder,
if I just double click, we have that
same video. So, if I just go through
with no audio, you could see that I lean
forward and do the exact same thing that
I did in that intro, and then I even
snap my finger. It's just notice that
when I do the actual snap, nothing
happens. Well, the reason why is because
we're going to use this as the trigger
to change my outfit. Now, that source
footage can be around 10 seconds or so.
I don't think most video models right
now allow you to go over this. You can
also creatively chain together a bunch
of these source videos in a way I'm
going to show you guys with no downside
in a sec. After you have that source
video, you need some sort of hyper
specific prompt. Now, the way that I
define a hyper specific prompt, and I've
just seen a lot of people screw the
pooch on this one, is you just need some
form of trigger, which is typically
time-gated. It's like a thing that that
occurs during the video that results in
a change. And so, in the demo that you
guys saw earlier, for instance, it was
the snap, right? The snap did the
change. You can tie it to specific words
and so on and so forth, too. So, it's
really not that complicated, right? You
have a trigger over here, which is the
moment that the model watches for. And
then you have a change, which is exactly
what is meant to happen next. So, in
this hypothetical example where I'm like
wearing a cloak or something, when the
man reaches towards his shoulders, maybe
like this, have them pull a shimmering
cloak over himself and fade to
invisible, maybe Harry Potter style. So,
if you want to have a good prompt that
works, if you want this to work, you'll
notice that the specificity is very,
very important. And I'll make sure to
show you guys a bunch of examples of
that in a sec. Okay, finally, you
obviously need the AI intelligence. Now,
there are a variety of different models
currently available that you could use.
I'm going to be using one called Omni,
but Kling is also pretty dope. And as
mentioned, people are developing these
things all the time. The key with the
sort of model to use for this is it
needs to be video to video, which is
different from like your typical
run-of-the-mill video generators. Uh,
and let me cover what that means. Right
now, a lot of people here that have
experimented with like AI video
typically get pretty disappointing
results. And the reason why is because
most of the time, you're trying to
generate something from scratch. AKA,
you are just feeding in text, and then
you're expecting a bunch of pixels in
the time dimension to magically appear
and to be totally coherent.
As I'm sure you can imagine, this is
pretty like intellectually complicated.
I mean, you can't I can't do that. Takes
me a hell of a lot longer than like 30
seconds or however long these video
models are. So, rather than generate
from total scratch, okay, and then be
disappointed with the results, we're
working in a medium that we've already
recorded. And so, we we provide some
sort of real footage here. And then all
we do is we actually just modify that
footage ever so slightly with a prompt.
Um, and so that's sort of what we're
going to do over here with Omni. We're
actually going to like feed in that the
pixels, the lighting, everything's
already going to be there. All we need
to do is just basically insert something
into this world that we've already
created. Obviously, inserting something
into a world that we've already created
is a lot easier than creating an
entirely new world. Makes sense. Now,
finally, when all is said and done, um,
we're going to output some sort of 720p
render. For those of you guys that are
unfamiliar, most video nowadays,
especially video on YouTube and other
social platforms, is higher than 720p.
And the reason why I bring this up is
because the people that build AI video
into their workflows that don't account
for the fact that the quality is
significantly degraded, at least as of
the time of this recording, I'm sure
it'll get better eventually. The people
that don't learn how to work around this
fact are going to get significantly
crappier results and their pipeline is
just going to be terrible. What you need
to do in order to make this work for any
ad, any organic video, whatever, is you
need to hide the 720p seam with an
intelligent cut to a different scene.
Now, what do I mean? I mean, if you guys
have a crappy AI shot at 720p or, you
know, like a more compressed version and
then you immediately cut back to a real
shot, one that has not been AI edited,
it'll be naturally disjointed. Um, a
really good example of that is like
this. I mean, this just looks really
weird all of the sudden because
obviously, you know, it gotten a lot
more compressed. So, what that means is
instead, you need to change the actual
nature of the shot. Rather than going
from like full screen talking face to
full screen talking face, you need to go
from like full screen talking face to
just another scene entirely. Um, the way
that I do it and the way that I've done
it many times in this video, if you
haven't seen, is I go from the 720p AI
shot to a screen share. Can have some
super crazy happen in the
background, but then cut back to the
video, uh, you know, you have no idea
that that occurred. You can't really
tell the pixels are different. Nobody
can really say, "Hey, you know, what
happened with that weird scene?" And
that's it. Now that you understand
everything that goes into this, let me
show you how to actually do it
yourselves. Now, you need that video
model like I talked about. The first one
I'm going to be using is Gemini Omni.
And if you head over to
deepmind.google/models/gemini-omni,
um, you can go to this page right here.
You can try it directly in Gemini, also
Google Flow, and then they have a little
build with Gemini Omni feature down
below. And you know, you guys can see
the scope of changes you're capable of
making. You could turn your whole room
into bubbles if you wanted. Really, it's
more about the ideas than it is anything
else. Um, you just need to like know
what to tell the model in order to make
it work the way that you want. This guy
turned himself into like Neo from The
Matrix, basically.
Um, now, I'll be honest, it's not free,
obviously. We're going to be doing
something that's pretty computationally
intense. Because we're going to be doing
something that's pretty computationally
intense, you guys can expect it to cost
a little bit of money. Um we're working
with pixels here and not just static
pixels like with images, we're working
with a time dimension. Keep in mind that
on average a video is 24 to 30 frames a
second. So you're technically generating
24 to 30 images a second. We had 8 to 10
second video clips which I'm going to
show you guys in a second. You know, one
of these things might realistically cost
you 50 cents in order to generate. And
what you'll see is you can't just
generate one, you typically have to
generate multiple simultaneously to take
advantage of parallelism and then
exploring a vast solution space. More on
that in a sec. Now because of this,
because of the fact that it's expensive
and then I got to come up with all these
prompts and stuff like that on my own
and to be honest, I'm not a very
creative person. I don't actually use um
Gemini Omni directly in their builder. I
use it through like a third-party
platform, in my case Higgsfield. There
are a couple of these but essentially
what they are are model aggregators that
just slam together all available current
video models, state-of-the-art systems
and stuff like that, just into one place
so that you can build a pipeline where
you pass it through, let's say model A
first and then go to like model B after
and then do model C after that. And you
know, with so many different models out
there, this is really their value prop.
It's like, oh you know, just come to us
and we'll deal with all of it. So I like
Higgsfield. I'm going to use it just to
show you guys more or less how this
stuff works. But I want you to know you
don't have to be tied to this. You can
also just use it in the Google Omni sort
of dashboard. When you sign up to this
and then you click on video in the top
left-hand corner and that's what I would
do. Don't click on a specific one, just
go video. You'll be taken to a page that
looks something like this. If you head
to the top left-hand corner, click the
button, you'll get something that looks
like this which is just like a big list
of prompts that you can use to one-shot.
I'm not going to do any of that right
now but it's pretty neat. You can just
like select one, um have a really
engaging intro hook for your ad or
whatever and then immediately arbitrage
tokens and then like the increased CPCs.
It's actually kind of nuts how you can
do that nowadays.
In terms of the actual workflow, it's
pretty straightforward. So I'm just
going to exit out of this, pretend I
haven't actually uploaded my media yet.
Um I'll head over to upload media and
you'll see here that I have like a video
combined 2026 one that I was talking
about um, that I'm going to have
uploaded. After it's done this little
uploading dialogue, all you need to do
is just click on it, and then that'll uh
put it right over here in the left-hand
side, added to prompt box, and then just
enter whatever it is that you want. And
in my case, what I'm going to do with my
hyper-specific prompt is I'll say,
"Right after the man says this at
exactly 2.9 seconds, change his outfit
to a cool-looking hoodie with a chain."
Then I'm going to click generate.
Now, one thing I want to stop and
disvalue of right now if I'm thinking is
that this is going to work 100% of the
time. This is going to work like 20% of
the time, which means on average you're
probably going to need to rerun this
thing five times or so if you really
want to get the smoothest effect
possible. Nobody's telling you this
right now because they want you to think
that this is way cheaper than it
actually is, um, but you're going to
have to run this multiple times. So,
just knowing right off the bat that I'm
probably going to have to run it
multiple times, rather than just like
make a prompt, click generate, wait,
wait for the results, see the results,
be dissatisfied with the results, and do
it all over again, rather than spend
like 30 seconds on that, I'm actually
just going to generate a bunch of these
simultaneously.
And I'm going to do five or so, and I'm
just going to see which one is the best.
Um, this also takes a fair amount of
time cuz as mentioned, we're not just
doing, you know, like one image, even
images take a lot of time. We're doing
like 24 to 30 images a second for
however long the length of the video is.
And you'll [snorts] also notice that
this isn't free. I mean, this costs 15
tokens. I don't know exactly what the
token to dollar conversion is here, but
you know, if you do this four or five
times, that's like 60, 70 tokens. That
can actually add up a fair amount. You
should expect to spend maybe 50 cents to
a dollar for this. Obviously, as time
goes on, this is going to go way
cheaper, but, um, some ways that you can
significantly reduce the cost if you
guys are cost-bottlenecked, is you can
upload a shorter video. Uh, if your
video is like 10 seconds, you're going
to consume more than 15. If it's like 5
seconds, you're going to consume less
than 15. Um, I'm doing 16 by 9 here
because that's wide screen, uh, but
obviously, you know, you guys can select
the smallest aspect ratio for whatever
the specific model is that you want. And
then, yeah, I mean, like the the shorter
the video, uh experiment with different
types of models if, you know, cost the
major bottleneck. And then you can
eventually get to the point where I'm
going to show you here. After all said
and done, you get something like this.
This just generated. I'm just going to
turn on my audio. Hopefully, you guys
can hear this.
With today's video models, you can do
whatever you want. Like this. Want to
learn how? Let me show you. And you'll
see that, you know, it hallucinate from
time to time. Like what happened there?
I didn't even snap my fingers and then
it changed my outfit. And then, not only
did it change my outfit once, it changed
my outfit again. And then it changed my
outfit again to like a half blue suit,
right? So, that's not actually what I
want. I just want like the the super
finesse chain switch. So, I'm going to
do the same thing here.
With today's video models, you can do
whatever you want. Like this. Want to
learn how? Let me show you.
With today's video models, you can do
This did two things. I mean, like that I
didn't really like. The first thing is
it sort of like slowly
like spanned the hoodie over me, which
is actually kind of dope if you go frame
by frame. But, um then it like put the
mic in my hands, which I don't really
like. I just wanted it to look exactly
like you guys expected.
With today's video models, you can do
whatever you want. Like this. Want to
learn how? Let me show you. This one was
today's video
This one was pretty good. I think this
is probably like 80 90% of the way
there. I could probably select this and
move on. I'm just going to poke around
and see if the other couple of gens that
I made are okay and then I'll loop back.
Okay, and then the end output, the thing
that I liked actually ended up looking
like this. With today's video models,
you can do whatever you want. Like this.
Want to learn how? Let me show you. So,
as you guys have already seen this cuz I
put this in the intro. Um that snap, it
happened immediately, that millisecond
that I did it. We have to do is we have
to weave this into whatever other
footage that we're doing with the scene
change. So, you know, in this case, I'm
going to use Premiere Pro. Uh you guys
don't have to be video editor pros or
anything like that. Uh Premiere Pro also
costs money like, you know, Higgs Field
like Omni and stuff like that. Um so, to
be clear, if you want to work with AI
video, you know, it's not going to be
cheap. It's going to cost you quite a
pretty penny.
Uh I just like contrast that to like
token-based pricing, right? But anyway,
I'm just going to open up Premiere Pro
and then I'm going to feed in my
content. So, I'm going to go shorter
over here. It's just like the name of my
project template for whatever reason.
I'm going to open that puppy up.
And then what I'll do, and keep in mind,
you can do this again in whatever thing
you want. Okay, I'm going to feed in the
original footage.
And we're just going to keep existing
settings. Then I'm also going to feed in
this Higgs Field edited footage. So,
just so you guys could see the actual
difference.
Okay, so this is the high-quality intro.
And let me just make sure that it's on
full. Do you guys see how like the
pixels are really defined and stuff like
that?
This is the Higgs Field 720p.
And so, I mean like if we kind of, I
don't know, hypothetically were to
overlay this one on top of each other,
like this. And then if I were to go back
here, okay, and then play, you'll also
notice that the length of the videos are
just a little bit different. Do you see
how this one here goes to here, whereas
this one here is sort of different?
That's because the AI is actually
pretty, like it's kind of changed the
very nature of the video itself,
including like the audio waveform and
when things are in the in the audio
file. So, you know, what I'm going to do
is I'm just going to very try very
carefully try and drag this over so it's
like one-to-one. And then at any point
in the video,
I'm just going to go back and forth so
you guys could see. So, this is the Omni
one. This is the real one. Omni, real.
Omni, real. You can see there are very
slight differences, but we're nowhere
near like uncanny valley territory like
we are when, you know, we realistically
try and AI generate a totally new video
clip. Um, so now that you have that,
what you need to do is you need to weave
that into like a piece of footage where,
uh, you know, we're not actually
transitioning right back to another
talking head screen cuz if we do that,
it'll kind of be broken. Okay, now I'm
just going to find a really simple video
I could use as an example. Why don't we
use this one here?
This one looks like it's kind of cutting
into a bunch of stuff, so that looks
nice.
And I'll go back here. All right. Okay,
so now what we have, I just need this
track cuz we have this sort of magical,
you know, snap my finger thing. And then
notice how I'm transitioning to a new
scene.
And so, you can't actually tell that we
just went from like AI to something
else. And obviously ideally whatever the
footage is that you have if you're doing
like a screen share talking head thing,
that would be pretty similar to.
Okay, so that's more or less it in a
nutshell. The cool thing is you can
actually proceduralize this with Claude
if you guys are using Claude code or a
codex if you guys are using codex. Um
the way that you do that is at least
Text Field anyway has like an MCP which
is just a series of eight API connectors
essentially that you can call that can
do all of this stuff for you. So you
could actually say like yo, I got this
thing on my computer, you know, I want
to edit it 10 times and then I'm just
going to select the best one. And that's
pretty much it. If you guys like this
sort of video, let me know down below.
As mentioned, the key here is you need
to be pretty specific with your prompt.
So ideally you'd either denote the exact
moment that you want something to happen
or the exact trigger condition with the
change. You also need to retry a bunch
of times. Like this isn't going to
happen immediately for you. That's just
the state of AI video.
That said, if you guys do get to a point
where the pipeline works really well for
you, whatever the effect that you're
adding and and so on and so forth is,
this is extraordinarily effective right
now. Like you can arbitrage the credits
that you spent on this to like massively
improve, you know, your ad CPCs, you
know, the the watch time and retention
of your content in specific places and
so on and so forth. And basically
nobody's doing it right now which is
wild. So yeah, let me know down below if
you guys like that sort of stuff. I'll
make tons more AI video content for you
if so. Also if you guys want all the
resources from this video, I've actually
uploaded them in the classroom of Maker
Zero. It's my free school community
where you essentially just upload
everything to make it really easy and
centralized. At the same time there are
a bunch of additional benefits including
comprehensive courses, good prompts and
and workflows and loops and stuff like
that. So if you want that, just head
over to Maker Zero, click classroom, and
then head over to recently uploaded for
everything you need from today's
session. Thank you very much for
watching and have a lovely rest of the
day.
Ask follow-up questions or revisit key timestamps.
This video details a professional workflow for creating high-quality AI-augmented video content. The creator explains how to use 'video-to-video' generation models to modify existing footage rather than creating from scratch. Key aspects of the pipeline include defining a specific trigger for changes (like a snap), understanding the limitations of current AI video quality (e.g., 720p output), and how to use editing techniques, such as scene transitions, to hide quality discrepancies. The creator also discusses the costs, the importance of trial and error (multiple generations), and how to potentially automate the process using API-based tools.
Videos recently processed by our community