HomeVideos

The BEST AI Video Strategy No One Is Using

Now Playing

The BEST AI Video Strategy No One Is Using

Transcript

539 segments

0:00

With today's video models, you can do

0:01

whatever you want. Like this. Want to

0:03

learn how? Let me show you. So, this

0:04

video pipeline is nuts, and I'm really

0:06

surprised more people aren't using

0:07

something like this. Um this is

0:09

extremely powerful and hooks, like you

0:10

just saw. It's also really good in ad

0:12

creative, as well as organic stuff. Uh I

0:14

have a friend who is worth many tens of

0:16

millions of dollars that's currently

0:17

generating over 2,000 AI videos a week

0:19

with something like this. So, it's not

0:21

just changing an outfit. I have kind of

0:23

this like little USB thing. Um well, I

0:25

can do it just like this. Alternatively,

0:27

let's say I wanted to, I don't know,

0:28

change the time of day. Um re-lighting

0:30

and stuff like that is pretty expensive

0:32

if you think about typical productions.

0:34

Now, I have the ability to change

0:35

lighting in a flash. So, I'm going to

0:37

show you all that in this video. You're

0:38

going to learn everything you need in

0:39

order to be able to develop really cool

0:40

visuals like this. So, what is the

0:42

actual workflow or like pipeline for

0:44

this? Well, it's pretty straightforward.

0:46

To start, you need some form of source

0:48

video. And the source video in our case

0:50

is going to be what you guys just saw me

0:52

show you earlier. So, if I just head

0:54

over to finder here, you can see that I

0:56

basically have this folder called AI

0:58

video set up. And inside of that folder,

1:00

if I just double click, we have that

1:02

same video. So, if I just go through

1:04

with no audio, you could see that I lean

1:06

forward and do the exact same thing that

1:07

I did in that intro, and then I even

1:08

snap my finger. It's just notice that

1:10

when I do the actual snap, nothing

1:12

happens. Well, the reason why is because

1:14

we're going to use this as the trigger

1:15

to change my outfit. Now, that source

1:17

footage can be around 10 seconds or so.

1:19

I don't think most video models right

1:21

now allow you to go over this. You can

1:22

also creatively chain together a bunch

1:24

of these source videos in a way I'm

1:25

going to show you guys with no downside

1:27

in a sec. After you have that source

1:28

video, you need some sort of hyper

1:29

specific prompt. Now, the way that I

1:31

define a hyper specific prompt, and I've

1:33

just seen a lot of people screw the

1:34

pooch on this one, is you just need some

1:35

form of trigger, which is typically

1:37

time-gated. It's like a thing that that

1:39

occurs during the video that results in

1:41

a change. And so, in the demo that you

1:43

guys saw earlier, for instance, it was

1:44

the snap, right? The snap did the

1:46

change. You can tie it to specific words

1:48

and so on and so forth, too. So, it's

1:49

really not that complicated, right? You

1:51

have a trigger over here, which is the

1:52

moment that the model watches for. And

1:54

then you have a change, which is exactly

1:56

what is meant to happen next. So, in

1:58

this hypothetical example where I'm like

1:59

wearing a cloak or something, when the

2:01

man reaches towards his shoulders, maybe

2:03

like this, have them pull a shimmering

2:05

cloak over himself and fade to

2:07

invisible, maybe Harry Potter style. So,

2:09

if you want to have a good prompt that

2:10

works, if you want this to work, you'll

2:12

notice that the specificity is very,

2:14

very important. And I'll make sure to

2:15

show you guys a bunch of examples of

2:16

that in a sec. Okay, finally, you

2:17

obviously need the AI intelligence. Now,

2:19

there are a variety of different models

2:21

currently available that you could use.

2:22

I'm going to be using one called Omni,

2:24

but Kling is also pretty dope. And as

2:26

mentioned, people are developing these

2:27

things all the time. The key with the

2:29

sort of model to use for this is it

2:30

needs to be video to video, which is

2:33

different from like your typical

2:34

run-of-the-mill video generators. Uh,

2:36

and let me cover what that means. Right

2:37

now, a lot of people here that have

2:38

experimented with like AI video

2:40

typically get pretty disappointing

2:41

results. And the reason why is because

2:43

most of the time, you're trying to

2:44

generate something from scratch. AKA,

2:46

you are just feeding in text, and then

2:47

you're expecting a bunch of pixels in

2:49

the time dimension to magically appear

2:51

and to be totally coherent.

2:53

As I'm sure you can imagine, this is

2:54

pretty like intellectually complicated.

2:55

I mean, you can't I can't do that. Takes

2:57

me a hell of a lot longer than like 30

2:58

seconds or however long these video

2:59

models are. So, rather than generate

3:02

from total scratch, okay, and then be

3:03

disappointed with the results, we're

3:05

working in a medium that we've already

3:06

recorded. And so, we we provide some

3:08

sort of real footage here. And then all

3:10

we do is we actually just modify that

3:11

footage ever so slightly with a prompt.

3:14

Um, and so that's sort of what we're

3:15

going to do over here with Omni. We're

3:16

actually going to like feed in that the

3:17

pixels, the lighting, everything's

3:18

already going to be there. All we need

3:20

to do is just basically insert something

3:22

into this world that we've already

3:23

created. Obviously, inserting something

3:25

into a world that we've already created

3:26

is a lot easier than creating an

3:28

entirely new world. Makes sense. Now,

3:29

finally, when all is said and done, um,

3:31

we're going to output some sort of 720p

3:34

render. For those of you guys that are

3:35

unfamiliar, most video nowadays,

3:37

especially video on YouTube and other

3:38

social platforms, is higher than 720p.

3:41

And the reason why I bring this up is

3:42

because the people that build AI video

3:45

into their workflows that don't account

3:47

for the fact that the quality is

3:48

significantly degraded, at least as of

3:50

the time of this recording, I'm sure

3:51

it'll get better eventually. The people

3:53

that don't learn how to work around this

3:54

fact are going to get significantly

3:56

crappier results and their pipeline is

3:57

just going to be terrible. What you need

3:58

to do in order to make this work for any

4:00

ad, any organic video, whatever, is you

4:02

need to hide the 720p seam with an

4:05

intelligent cut to a different scene.

4:08

Now, what do I mean? I mean, if you guys

4:11

have a crappy AI shot at 720p or, you

4:14

know, like a more compressed version and

4:16

then you immediately cut back to a real

4:19

shot, one that has not been AI edited,

4:21

it'll be naturally disjointed. Um, a

4:23

really good example of that is like

4:24

this. I mean, this just looks really

4:26

weird all of the sudden because

4:27

obviously, you know, it gotten a lot

4:29

more compressed. So, what that means is

4:30

instead, you need to change the actual

4:33

nature of the shot. Rather than going

4:35

from like full screen talking face to

4:37

full screen talking face, you need to go

4:39

from like full screen talking face to

4:41

just another scene entirely. Um, the way

4:43

that I do it and the way that I've done

4:44

it many times in this video, if you

4:45

haven't seen, is I go from the 720p AI

4:47

shot to a screen share. Can have some

4:49

super crazy happen in the

4:50

background, but then cut back to the

4:52

video, uh, you know, you have no idea

4:54

that that occurred. You can't really

4:56

tell the pixels are different. Nobody

4:57

can really say, "Hey, you know, what

4:59

happened with that weird scene?" And

5:00

that's it. Now that you understand

5:01

everything that goes into this, let me

5:03

show you how to actually do it

5:04

yourselves. Now, you need that video

5:05

model like I talked about. The first one

5:07

I'm going to be using is Gemini Omni.

5:09

And if you head over to

5:10

deepmind.google/models/gemini-omni,

5:14

um, you can go to this page right here.

5:16

You can try it directly in Gemini, also

5:18

Google Flow, and then they have a little

5:20

build with Gemini Omni feature down

5:22

below. And you know, you guys can see

5:24

the scope of changes you're capable of

5:25

making. You could turn your whole room

5:27

into bubbles if you wanted. Really, it's

5:28

more about the ideas than it is anything

5:30

else. Um, you just need to like know

5:32

what to tell the model in order to make

5:33

it work the way that you want. This guy

5:35

turned himself into like Neo from The

5:36

Matrix, basically.

5:38

Um, now, I'll be honest, it's not free,

5:41

obviously. We're going to be doing

5:43

something that's pretty computationally

5:44

intense. Because we're going to be doing

5:46

something that's pretty computationally

5:47

intense, you guys can expect it to cost

5:49

a little bit of money. Um we're working

5:51

with pixels here and not just static

5:52

pixels like with images, we're working

5:53

with a time dimension. Keep in mind that

5:55

on average a video is 24 to 30 frames a

5:58

second. So you're technically generating

5:59

24 to 30 images a second. We had 8 to 10

6:03

second video clips which I'm going to

6:04

show you guys in a second. You know, one

6:06

of these things might realistically cost

6:08

you 50 cents in order to generate. And

6:10

what you'll see is you can't just

6:11

generate one, you typically have to

6:13

generate multiple simultaneously to take

6:15

advantage of parallelism and then

6:17

exploring a vast solution space. More on

6:19

that in a sec. Now because of this,

6:20

because of the fact that it's expensive

6:22

and then I got to come up with all these

6:23

prompts and stuff like that on my own

6:25

and to be honest, I'm not a very

6:25

creative person. I don't actually use um

6:28

Gemini Omni directly in their builder. I

6:30

use it through like a third-party

6:31

platform, in my case Higgsfield. There

6:34

are a couple of these but essentially

6:35

what they are are model aggregators that

6:37

just slam together all available current

6:40

video models, state-of-the-art systems

6:42

and stuff like that, just into one place

6:44

so that you can build a pipeline where

6:46

you pass it through, let's say model A

6:48

first and then go to like model B after

6:50

and then do model C after that. And you

6:53

know, with so many different models out

6:54

there, this is really their value prop.

6:56

It's like, oh you know, just come to us

6:57

and we'll deal with all of it. So I like

6:59

Higgsfield. I'm going to use it just to

7:01

show you guys more or less how this

7:02

stuff works. But I want you to know you

7:04

don't have to be tied to this. You can

7:05

also just use it in the Google Omni sort

7:07

of dashboard. When you sign up to this

7:09

and then you click on video in the top

7:10

left-hand corner and that's what I would

7:11

do. Don't click on a specific one, just

7:13

go video. You'll be taken to a page that

7:15

looks something like this. If you head

7:16

to the top left-hand corner, click the

7:17

button, you'll get something that looks

7:18

like this which is just like a big list

7:20

of prompts that you can use to one-shot.

7:22

I'm not going to do any of that right

7:23

now but it's pretty neat. You can just

7:25

like select one, um have a really

7:27

engaging intro hook for your ad or

7:28

whatever and then immediately arbitrage

7:30

tokens and then like the increased CPCs.

7:32

It's actually kind of nuts how you can

7:33

do that nowadays.

7:35

In terms of the actual workflow, it's

7:36

pretty straightforward. So I'm just

7:36

going to exit out of this, pretend I

7:38

haven't actually uploaded my media yet.

7:40

Um I'll head over to upload media and

7:41

you'll see here that I have like a video

7:44

combined 2026 one that I was talking

7:45

about um, that I'm going to have

7:47

uploaded. After it's done this little

7:49

uploading dialogue, all you need to do

7:51

is just click on it, and then that'll uh

7:54

put it right over here in the left-hand

7:55

side, added to prompt box, and then just

7:57

enter whatever it is that you want. And

7:59

in my case, what I'm going to do with my

8:01

hyper-specific prompt is I'll say,

8:02

"Right after the man says this at

8:04

exactly 2.9 seconds, change his outfit

8:07

to a cool-looking hoodie with a chain."

8:09

Then I'm going to click generate.

8:11

Now, one thing I want to stop and

8:13

disvalue of right now if I'm thinking is

8:15

that this is going to work 100% of the

8:17

time. This is going to work like 20% of

8:19

the time, which means on average you're

8:20

probably going to need to rerun this

8:22

thing five times or so if you really

8:24

want to get the smoothest effect

8:25

possible. Nobody's telling you this

8:26

right now because they want you to think

8:27

that this is way cheaper than it

8:28

actually is, um, but you're going to

8:30

have to run this multiple times. So,

8:32

just knowing right off the bat that I'm

8:33

probably going to have to run it

8:34

multiple times, rather than just like

8:36

make a prompt, click generate, wait,

8:39

wait for the results, see the results,

8:41

be dissatisfied with the results, and do

8:42

it all over again, rather than spend

8:44

like 30 seconds on that, I'm actually

8:46

just going to generate a bunch of these

8:48

simultaneously.

8:49

And I'm going to do five or so, and I'm

8:50

just going to see which one is the best.

8:52

Um, this also takes a fair amount of

8:53

time cuz as mentioned, we're not just

8:55

doing, you know, like one image, even

8:57

images take a lot of time. We're doing

8:58

like 24 to 30 images a second for

9:01

however long the length of the video is.

9:03

And you'll [snorts] also notice that

9:04

this isn't free. I mean, this costs 15

9:06

tokens. I don't know exactly what the

9:07

token to dollar conversion is here, but

9:09

you know, if you do this four or five

9:10

times, that's like 60, 70 tokens. That

9:12

can actually add up a fair amount. You

9:13

should expect to spend maybe 50 cents to

9:15

a dollar for this. Obviously, as time

9:17

goes on, this is going to go way

9:18

cheaper, but, um, some ways that you can

9:19

significantly reduce the cost if you

9:21

guys are cost-bottlenecked, is you can

9:23

upload a shorter video. Uh, if your

9:25

video is like 10 seconds, you're going

9:27

to consume more than 15. If it's like 5

9:30

seconds, you're going to consume less

9:31

than 15. Um, I'm doing 16 by 9 here

9:34

because that's wide screen, uh, but

9:35

obviously, you know, you guys can select

9:37

the smallest aspect ratio for whatever

9:38

the specific model is that you want. And

9:41

then, yeah, I mean, like the the shorter

9:42

the video, uh experiment with different

9:44

types of models if, you know, cost the

9:45

major bottleneck. And then you can

9:46

eventually get to the point where I'm

9:48

going to show you here. After all said

9:49

and done, you get something like this.

9:50

This just generated. I'm just going to

9:52

turn on my audio. Hopefully, you guys

9:53

can hear this.

9:55

With today's video models, you can do

9:56

whatever you want. Like this. Want to

9:58

learn how? Let me show you. And you'll

10:00

see that, you know, it hallucinate from

10:01

time to time. Like what happened there?

10:03

I didn't even snap my fingers and then

10:05

it changed my outfit. And then, not only

10:07

did it change my outfit once, it changed

10:09

my outfit again. And then it changed my

10:10

outfit again to like a half blue suit,

10:12

right? So, that's not actually what I

10:13

want. I just want like the the super

10:15

finesse chain switch. So, I'm going to

10:17

do the same thing here.

10:18

With today's video models, you can do

10:20

whatever you want. Like this. Want to

10:22

learn how? Let me show you.

10:23

With today's video models, you can do

10:25

This did two things. I mean, like that I

10:27

didn't really like. The first thing is

10:28

it sort of like slowly

10:31

like spanned the hoodie over me, which

10:33

is actually kind of dope if you go frame

10:34

by frame. But, um then it like put the

10:37

mic in my hands, which I don't really

10:38

like. I just wanted it to look exactly

10:39

like you guys expected.

10:41

With today's video models, you can do

10:43

whatever you want. Like this. Want to

10:45

learn how? Let me show you. This one was

10:47

today's video

10:48

This one was pretty good. I think this

10:49

is probably like 80 90% of the way

10:51

there. I could probably select this and

10:52

move on. I'm just going to poke around

10:54

and see if the other couple of gens that

10:55

I made are okay and then I'll loop back.

10:57

Okay, and then the end output, the thing

10:59

that I liked actually ended up looking

11:00

like this. With today's video models,

11:02

you can do whatever you want. Like this.

11:05

Want to learn how? Let me show you. So,

11:07

as you guys have already seen this cuz I

11:08

put this in the intro. Um that snap, it

11:10

happened immediately, that millisecond

11:12

that I did it. We have to do is we have

11:13

to weave this into whatever other

11:15

footage that we're doing with the scene

11:16

change. So, you know, in this case, I'm

11:18

going to use Premiere Pro. Uh you guys

11:20

don't have to be video editor pros or

11:22

anything like that. Uh Premiere Pro also

11:23

costs money like, you know, Higgs Field

11:25

like Omni and stuff like that. Um so, to

11:26

be clear, if you want to work with AI

11:27

video, you know, it's not going to be

11:29

cheap. It's going to cost you quite a

11:30

pretty penny.

11:31

Uh I just like contrast that to like

11:32

token-based pricing, right? But anyway,

11:34

I'm just going to open up Premiere Pro

11:35

and then I'm going to feed in my

11:36

content. So, I'm going to go shorter

11:37

over here. It's just like the name of my

11:39

project template for whatever reason.

11:41

I'm going to open that puppy up.

11:43

And then what I'll do, and keep in mind,

11:44

you can do this again in whatever thing

11:46

you want. Okay, I'm going to feed in the

11:48

original footage.

11:50

And we're just going to keep existing

11:52

settings. Then I'm also going to feed in

11:54

this Higgs Field edited footage. So,

11:56

just so you guys could see the actual

11:57

difference.

11:59

Okay, so this is the high-quality intro.

12:01

And let me just make sure that it's on

12:02

full. Do you guys see how like the

12:03

pixels are really defined and stuff like

12:05

that?

12:05

This is the Higgs Field 720p.

12:08

And so, I mean like if we kind of, I

12:10

don't know, hypothetically were to

12:11

overlay this one on top of each other,

12:13

like this. And then if I were to go back

12:15

here, okay, and then play, you'll also

12:18

notice that the length of the videos are

12:20

just a little bit different. Do you see

12:22

how this one here goes to here, whereas

12:23

this one here is sort of different?

12:25

That's because the AI is actually

12:26

pretty, like it's kind of changed the

12:28

very nature of the video itself,

12:30

including like the audio waveform and

12:32

when things are in the in the audio

12:34

file. So, you know, what I'm going to do

12:36

is I'm just going to very try very

12:37

carefully try and drag this over so it's

12:39

like one-to-one. And then at any point

12:40

in the video,

12:42

I'm just going to go back and forth so

12:43

you guys could see. So, this is the Omni

12:44

one. This is the real one. Omni, real.

12:47

Omni, real. You can see there are very

12:48

slight differences, but we're nowhere

12:50

near like uncanny valley territory like

12:52

we are when, you know, we realistically

12:54

try and AI generate a totally new video

12:56

clip. Um, so now that you have that,

12:58

what you need to do is you need to weave

12:59

that into like a piece of footage where,

13:01

uh, you know, we're not actually

13:03

transitioning right back to another

13:05

talking head screen cuz if we do that,

13:06

it'll kind of be broken. Okay, now I'm

13:08

just going to find a really simple video

13:10

I could use as an example. Why don't we

13:11

use this one here?

13:13

This one looks like it's kind of cutting

13:15

into a bunch of stuff, so that looks

13:16

nice.

13:18

And I'll go back here. All right. Okay,

13:19

so now what we have, I just need this

13:20

track cuz we have this sort of magical,

13:22

you know, snap my finger thing. And then

13:24

notice how I'm transitioning to a new

13:26

scene.

13:27

And so, you can't actually tell that we

13:29

just went from like AI to something

13:31

else. And obviously ideally whatever the

13:32

footage is that you have if you're doing

13:34

like a screen share talking head thing,

13:35

that would be pretty similar to.

13:37

Okay, so that's more or less it in a

13:39

nutshell. The cool thing is you can

13:41

actually proceduralize this with Claude

13:44

if you guys are using Claude code or a

13:46

codex if you guys are using codex. Um

13:48

the way that you do that is at least

13:49

Text Field anyway has like an MCP which

13:51

is just a series of eight API connectors

13:54

essentially that you can call that can

13:55

do all of this stuff for you. So you

13:56

could actually say like yo, I got this

13:57

thing on my computer, you know, I want

13:59

to edit it 10 times and then I'm just

14:00

going to select the best one. And that's

14:01

pretty much it. If you guys like this

14:03

sort of video, let me know down below.

14:05

As mentioned, the key here is you need

14:07

to be pretty specific with your prompt.

14:08

So ideally you'd either denote the exact

14:10

moment that you want something to happen

14:12

or the exact trigger condition with the

14:14

change. You also need to retry a bunch

14:16

of times. Like this isn't going to

14:17

happen immediately for you. That's just

14:18

the state of AI video.

14:20

That said, if you guys do get to a point

14:22

where the pipeline works really well for

14:23

you, whatever the effect that you're

14:24

adding and and so on and so forth is,

14:27

this is extraordinarily effective right

14:29

now. Like you can arbitrage the credits

14:31

that you spent on this to like massively

14:33

improve, you know, your ad CPCs, you

14:36

know, the the watch time and retention

14:38

of your content in specific places and

14:39

so on and so forth. And basically

14:41

nobody's doing it right now which is

14:42

wild. So yeah, let me know down below if

14:44

you guys like that sort of stuff. I'll

14:46

make tons more AI video content for you

14:47

if so. Also if you guys want all the

14:49

resources from this video, I've actually

14:51

uploaded them in the classroom of Maker

14:53

Zero. It's my free school community

14:54

where you essentially just upload

14:55

everything to make it really easy and

14:57

centralized. At the same time there are

14:59

a bunch of additional benefits including

15:00

comprehensive courses, good prompts and

15:03

and workflows and loops and stuff like

15:04

that. So if you want that, just head

15:05

over to Maker Zero, click classroom, and

15:07

then head over to recently uploaded for

15:09

everything you need from today's

15:10

session. Thank you very much for

15:11

watching and have a lovely rest of the

15:13

day.

Interactive Summary

This video details a professional workflow for creating high-quality AI-augmented video content. The creator explains how to use 'video-to-video' generation models to modify existing footage rather than creating from scratch. Key aspects of the pipeline include defining a specific trigger for changes (like a snap), understanding the limitations of current AI video quality (e.g., 720p output), and how to use editing techniques, such as scene transitions, to hide quality discrepancies. The creator also discusses the costs, the importance of trial and error (multiple generations), and how to potentially automate the process using API-based tools.

Suggested questions

3 ready-made prompts