HomeVideos

YOLO explained step by step - Object detection

Now Playing

YOLO explained step by step - Object detection

Transcript

773 segments

0:00

In this video, we'll talk about the Yoli

0:02

model that is used for object detection.

0:06

I will mainly focus on the structure of

0:08

the model and how it can be trained

0:10

based on its loss function and then show

0:13

some simple Python code on how you can

0:15

train a simple YOLO model. I will focus

0:19

on the original Yoli model from the

0:21

paper you only look once. Note that

0:24

newer versions have been released since

0:27

2015 with a lot of improvements. I will

0:31

here assume that you are familiar with

0:33

CNN's and image classification. If not,

0:37

have a look at my video about transfer

0:39

learning. I also recommend that you

0:42

watch my video about object detection

0:44

and bounding box regression which

0:46

explains the basics. To understand YOLO,

0:50

we here use the following example image

0:53

with 224 pixels in width and 224 pixels

0:57

in height. This is the pixel coordinate

1:00

in the top left corner. And this is the

1:03

pixel coordinate in the bottom right

1:05

corner. YOLO adds a 7 * 7 grid on the

1:10

image like this, which means that the

1:12

image is divided into 49 grid cells.

1:16

This is an example of one grid cell. To

1:20

train a YOLA model, we need to manually

1:23

draw a bounding box around each object

1:26

in the image. This handdrawn bounding

1:29

box is referred to as the ground truth

1:32

box or the target box.

1:35

Here are a few examples of tools that

1:37

you can use to draw boxes to annotate

1:40

your images.

1:42

For each box including an object, we

1:45

also need to define the class of this

1:48

object. So we here say that the object

1:51

is a cat.

1:53

To train a YOLO model, we should provide

1:56

the center X and Y pixel coordinates of

1:59

the target box and the width and height

2:02

of the box

2:04

and the class. In this example, we will

2:08

have cats and dogs in the images. So the

2:11

number of classes is here too. So the

2:15

training target for this image should

2:17

look like this where these numbers are

2:20

the center coordinates of the bounding

2:23

box and these numbers correspond to the

2:26

width and height of the box.

2:29

Note that these values are normalized by

2:31

dividing by the width and height of the

2:34

image.

2:35

Finally, we say that the box contains a

2:38

cat.

2:40

If you would have a second object in the

2:42

image,

2:44

we need to draw a bounding box around

2:46

this object as well and determine the

2:50

center coordinates of this box and its

2:53

width and height

2:56

and its class.

2:58

So for this image, we need the following

3:01

inputs for the model in order to train

3:04

it so that the model knows where it

3:07

should put the predicted boxes and how

3:09

to predict the class.

3:12

This is the YOLO network from the

3:14

original paper. You only look once.

3:18

The input is a color image that is

3:21

resized to 448* 448 pixels. Remember

3:26

that a color image has three channels,

3:30

one for red, one for green, and one for

3:32

blue. A filter of size 7 by seven with a

3:37

depth of three is sliding over the image

3:40

with a stride of two. And after max

3:44

pooling, we get the feature map of size

3:47

112*

3:48

112. After the convolutions, the last

3:52

feature maps are flattened into a vector

3:55

which is connected to a fully connected

3:58

layer with 4,96

4:00

nodes. The 4,96

4:03

nodes are then connected to the

4:05

following output layer. So what we have

4:09

here is basically just a simple

4:11

convolutional neural network which you

4:14

should be familiar with if you have seen

4:15

my video about transfer learning. The

4:18

new thing here is this strange output

4:21

layer. If we were only to use this

4:24

network for image classification to

4:27

predict if the image contains a cat or a

4:30

dog, we would only need two output nodes

4:33

or in this case it would be enough with

4:35

just one since it is a simple binary

4:38

classification. If you also like to

4:41

learn the network to predict bounding

4:43

boxes around the cats and dogs, we need

4:47

to add output nodes that predict the

4:49

center x and y coordinates of the box as

4:53

well as its width and height.

4:56

This type of network only works if there

4:59

is one object per image.

5:02

If we have two objects in the image, we

5:05

need to have a unique output for each

5:08

box.

5:09

Let's rotate the output layer like this.

5:13

Now in the original YOLO model, they had

5:16

two boxes in each cell. Since there are

5:19

49 cells in the image, 96 boxes exist

5:24

for each image.

5:26

In contrast, the classification was

5:29

based on a grid cell, which means that

5:31

the original YOLO model could only

5:34

identify a maximum of 49 objects in an

5:37

image.

5:38

The reason for having two boxes per cell

5:42

was to let these compete to become the

5:44

responsible box which best fits the

5:47

target box. In later YOLO versions,

5:51

anchor boxes are used so that each grid

5:53

cell can predict multiple bounding boxes

5:56

allowing multiple objects and class

5:59

predictions within the same cell. So

6:02

every cell in the original Jolie model

6:05

has the following output nodes.

6:08

We can denote the output layer like this

6:12

because the image consists of 49 cells

6:15

where each cell is connected to 10

6:18

output nodes. Also, YOLO has one

6:22

additional output that predicts the

6:24

confidence which I will explain later

6:26

on.

6:28

So in our case where we only have two

6:31

classes cats and dogs, we'll have the

6:34

following output layer. This output

6:36

layer has 12 output nodes for each of

6:40

the 49 grid cells. The original YOLO

6:43

network was trained on the Pascal data

6:46

set which consists of 20 different

6:48

classes such as persons, dogs, cars,

6:52

airplanes, and trains. This now explains

6:56

the output layer shown in the original

6:58

YOLO paper because every cell in the

7:01

image has 30 output nodes, two for each

7:05

box that predict the center of the

7:07

bounding box, the width and the height

7:10

of the two boxes and the confidence of

7:13

the two boxes. Since the data set had 20

7:17

different classes,

7:19

30 output nodes are used for each cell

7:22

in the image.

7:24

Now let's try to understand how the

7:26

confidence is computed. But to

7:28

understand the confidence, we first need

7:30

to understand the measure intersection

7:32

over a union. This is our ground truth

7:36

or target bounding box. If you multiply

7:39

the width and the height of the box, we

7:42

see that the box covers 9,000 pixels.

7:46

When the network processes this image to

7:48

compute the two boxes in a given cell,

7:51

it may predict the following box given

7:54

the current weights of the network. The

7:57

predicted box has an area of 4,96 pixels

8:02

and the intersection covers 2,560

8:06

pixels. We can now calculate the

8:09

intersection of a union where we plug in

8:11

the area of the intersection here and

8:15

the areas of the target box and the

8:17

predicted box. This gives us an

8:20

intersection of a union score of about

8:23

0.24.

8:25

If the predicted box were over here

8:28

where it does not intersect the target

8:30

box, the intersection of a union score

8:33

would be zero.

8:35

And if the predicted box would have the

8:37

same size and location as the target

8:39

box, the intersection of a union would

8:42

be one. We therefore know that the

8:45

intersection of a union spans between

8:47

zero and one. And that the closer to one

8:50

we are, the better the predicted box

8:52

fits the target box. Now remember that

8:56

each cell has two boxes. But only the

9:00

two boxes that are within the cell that

9:03

includes the center coordinates of the

9:05

target box are important here. Although

9:08

these two boxes are within the target

9:11

box, they are less important because

9:13

they belong to a grid cell that does not

9:16

include the center of the target box.

9:19

The two boxes are equally useless for

9:22

this image as these two. The area of box

9:26

one is 1,000

9:28

and the area of box two is 225.

9:32

We previously saw that the area of the

9:34

target box was 9,000.

9:37

The intersection of a union of box one

9:40

is about 0.11.

9:43

Whereas the intersection of a union of

9:45

box two is about 0.025.

9:49

Since box one has a larger intersection

9:52

of a union compared to box two, it is

9:55

here considered as the responsible box.

9:59

Now suppose that the network processes

10:02

the image and currently predicts the

10:04

following confidence scores for box one

10:07

and two. Then YOLO constructs a target

10:11

confidence for all the boxes.

10:14

P object is set to one if the box is

10:17

responsible

10:19

and to zero if the boxes are not

10:21

responsible.

10:24

So for boxes that are not associated

10:27

with the cell where the center of the

10:29

target boxes, the target P object is set

10:33

to zero.

10:35

The only box that gets a target P object

10:38

of one is this box because it is a box

10:42

within the cell that includes the center

10:44

of the target box. And out of the two

10:47

boxes in this cell, it has the largest

10:50

intersection of union.

10:53

So the target confidence of box one is

10:56

0.111.

10:58

Whereas the target confidence for box

11:00

two is zero which is true also for the

11:04

other boxes in all the other grid cells.

11:08

The total confidence loss of this cell

11:11

is the square difference between the

11:13

target confidence and the predicted

11:15

confidence for the responsible box and

11:18

for the box that is not responsible.

11:21

For no object boxes or boxes that are

11:25

not responsible, the square difference

11:27

is multiplied by a weight which they set

11:31

to 0.5 in the paper.

11:34

Since the network is trained to reduce

11:36

the loss by changing the values of its

11:39

weights, the weights will be changed so

11:42

that the predicted confidences get

11:44

closer to the target confidence scores.

11:48

This is the full loss function that was

11:51

described in the original Jolu paper.

11:54

Before any loss is computed, the image

11:57

coordinates are normalized so that the

12:00

coordinates span between zero and one.

12:04

Lambda chord is a weight that puts more

12:07

weight on fitting the predicted box

12:10

closer to the target box. The value of

12:13

this weight was set to five in the

12:16

original paper.

12:18

These summations denote that we sum over

12:21

all the 49 grid cells. And these

12:24

summations denote that we sum over all

12:27

boxes within each cell. Since there are

12:30

only two boxes per cell, this means that

12:33

we sum the loss of the two boxes per

12:36

cell. This denotes that if an object

12:40

appears in cell I, it is set to one.

12:44

Else it is set to zero. So this means

12:47

that the class loss will only be

12:49

computed within a cell that includes the

12:52

center of the target box because this

12:55

notation will be set to zero for all

12:58

other cells.

13:01

This notation denotes that if box J in

13:05

cell I is responsible

13:08

it will be set to one else to zero.

13:12

This means that this notation will be

13:14

set to one only for this box. Which

13:18

means that mainly this box will update

13:21

its location and size during the current

13:24

training step. The other 97 boxes will

13:28

focus mainly on updating this confidence

13:31

loss term where its target value is zero

13:35

which means that the current confidence

13:37

is forced down to zero. So for no

13:41

responsible boxes, this is set to one.

13:45

This means that the total loss is mainly

13:48

dependent on box one which is the

13:51

responsible box as well as the cell that

13:54

includes the center of the target box

13:58

and very little on nonresponsible boxes.

14:02

Suppose that we will compute the class

14:04

loss of this cell which we know contains

14:07

the center of the object with class cat.

14:11

The target value for this output node is

14:13

therefore one and the target value for

14:16

this node is zero because the given grid

14:19

cell does not include a dog.

14:22

When the network processes the image, it

14:25

might compute the predicted probability

14:28

of 0.8. Since the network tries to

14:31

minimize the loss during training, it

14:34

will update its weights so that the

14:37

predicted probability for a cat gets

14:40

closer to one and the predicted

14:42

probability for a dog gets closer to

14:45

zero. At the same time, it will update

14:49

the location and size of the responsible

14:52

box so that it gets closer to the target

14:55

box

14:56

so that the square difference between

14:58

the center of the responsible box and

15:01

the target box as well as the difference

15:04

between a width and a height are as

15:06

small as possible. The reason why they

15:09

based the loss function on the square

15:12

root of the width and height was that

15:14

small deviations in large boxes should

15:17

matter less than in small boxes. The

15:21

idea with the training is that the

15:23

weights are updated so that the

15:25

responsible box fits as good as possible

15:28

with the target box and that the cell

15:31

correctly predicts the class cat. Note

15:34

that the network does not see the target

15:37

box during training. It only sees the

15:40

image. The information of the target box

15:43

is only used as target values in the

15:46

output nodes. The network therefore

15:49

learns to identify what is presented in

15:52

the image in relation to the location of

15:55

the target box. We'll now discuss some

15:58

of the performance metrics that are used

16:01

in object detection. In this image, we

16:04

have three cats with ground truth boxes

16:07

in green. The blue box is the predicted

16:11

box by the network and the label is the

16:14

predicted class. This is an example of a

16:18

true positive result because the

16:20

predicted box fits nicely with the

16:22

ground truth box and the predicted class

16:24

cat is correct. This is an example of a

16:29

false positive result because the

16:31

predicted class dog is incorrect.

16:35

Whereas this is a false negative result

16:38

because there is no predicted box

16:40

associated with a ground truth box. This

16:44

is an example of a true negative result

16:47

which does not make sense because a

16:50

ground truth box without any object

16:52

should not exist.

16:54

We can now calculate the precision which

16:57

measures the proportion of predicted

16:59

bounding boxes that are correct which is

17:03

the true positives among all predicted

17:06

detections.

17:07

Note that this metric evaluates only the

17:10

predictions made by the model. In

17:14

comparison, the recall measures the

17:16

proportion of ground trth boxes that are

17:20

correctly detected by the model. It

17:23

indicates how many of the actual objects

17:25

in the data set that are successfully

17:28

predicted.

17:29

Only one ground truth box was correctly

17:32

detected out of three.

17:36

Note that this results in both one false

17:38

positive and one false negative result

17:43

because the predicted box has incorrect

17:46

class

17:47

and the ground truth box was incorrectly

17:50

detected.

17:52

But as we have seen before, there should

17:54

be 96 predicted boxes in this image. So

17:58

where are those? Well, to count or

18:01

display a predicted box, we use a

18:04

confidence threshold that is here set to

18:06

0.7. This means that we only display

18:10

predicted boxes with a confidence score

18:12

greater than 0.7.

18:15

If you increase this threshold to 0.9,

18:18

the following predicted box will

18:20

disappear.

18:21

and the result will change to only a

18:24

false negative result and the precision

18:27

will therefore change value. If you use

18:29

a higher confidence threshold, we expect

18:32

higher precision because fewer predicted

18:34

bounding boxes will be considered and

18:37

those that remain are more likely to be

18:40

correct. However, the recall is expected

18:43

to decrease because fewer ground truth

18:46

boxes will be successfully detected. If

18:49

instead reduce the confidence threshold

18:51

to 0.3,

18:53

we will see more boxes.

18:56

The problem now is whether we should

18:58

count this predicted box as a false

19:01

positive or true positive because the

19:03

class prediction is correct but the

19:06

predicted box does not fit well with the

19:08

ground truth box. Usually one uses an

19:12

additional cutoff based on the

19:14

intersection of a union. This threshold

19:17

is commonly set to 0.5.

19:20

So if the intersection of a union in

19:23

this case is 0.4 four which is below the

19:26

threshold. This will be counted both as

19:30

a false positive result and a false

19:33

negative result because the ground truth

19:35

box is not correctly detected and the

19:38

predicted box does not properly identify

19:41

the object.

19:43

The precision and recall will therefore

19:45

change values like this.

19:48

If you reduce the intersection of a

19:50

union threshold to 0.3,

19:53

this will now be counted as a true

19:55

positive result because the predicted

19:58

box correctly identifies the object and

20:01

the precision and recall will now show

20:04

these values.

20:06

So we now know that the position and

20:09

recall depend on these cutoff values.

20:12

To get an overall measure of the

20:14

performance, we can change the cutoff

20:17

and compute the precision and recall for

20:20

a number of different threshold values.

20:23

We compute these measures for all

20:26

images, not just this one.

20:29

Then we can plot how the precision and

20:32

recall change when we change the

20:34

confidence threshold. This curve is

20:36

called the precision recall curve or the

20:39

PR curve. For example, this point on the

20:43

curve may correspond to a confidence

20:45

threshold of 0.99.

20:48

At such high threshold, precision is

20:51

expected to be close to one since only

20:54

very confidence predictions are kept

20:57

while the recall is expected to be close

21:00

to zero because many ground truth

21:02

objects will not be detected.

21:05

For a lower threshold, the precision is

21:08

expected to reduce and the recall to

21:11

increase.

21:12

Our final measure of performance is the

21:15

average precision, which is the area

21:17

below this curve. The area under this

21:20

curve spans between zero and one. Since

21:24

this curve is based on the intersection

21:26

of union threshold of 0.5,

21:30

it is called AP0.5 or AP50.

21:35

If you instead used a threshold of 0.75,

21:38

the curve will change and the measure

21:41

would then be called AP0.75

21:44

or AP75.

21:46

One can even compute the curve for

21:48

thresholds going from 0.5 to 0.95

21:53

with 0.05 increments and compute the

21:56

average of these.

21:59

Then the measure will be denoted like

22:01

this.

22:03

Suppose that our AP50 score is 0.8. In

22:07

this case,

22:09

what we want is a score close to one

22:12

because at the certain confidence

22:14

threshold, we will have a precision and

22:17

recall that are close to one.

22:20

Note that we had so far only computed

22:23

the AP50 for the cats.

22:26

We also need to compute the same for the

22:29

dogs for all our images which might

22:32

result in the AP50 score of for example

22:35

0.85.

22:37

So our final measure of the performance

22:40

of the object detection is the average

22:42

AP50 value of all the classes which is

22:46

called the mean AP50.

22:49

We'll now have a look at some basic

22:51

Python code so that you can train your

22:54

first Jolo model.

22:56

Before we discuss the training, let's

22:58

first see how we can use an existing

23:00

YOLO model to do object detection. We

23:04

first begin to install the ultral litic

23:07

package with pip install. Then we import

23:10

the yolo class

23:14

and load the pre-trained model. We'll

23:17

here use the model called YOLO 26N which

23:21

was released in September 2025.

23:24

This model is the smallest model in it

23:27

26 family since it is the smallest

23:30

network. It is the fastest for training

23:33

and for inference. But the smaller size

23:36

comes with the cost because it performs

23:39

worse than larger models. Anyway, it is

23:43

a good model to start with.

23:46

To use the existing model to run

23:48

inference on an image, we simply provide

23:51

the path to the image. You may change

23:54

this path to an image that you would

23:56

like to try on. This is the confidence

23:59

threshold for displaying the predicted

24:01

bounding boxes. Then we simply display

24:04

the result with this line of code.

24:08

A common mistake is to interpret this

24:10

value as the probability that it is a

24:13

cat.

24:15

This value represents the detection

24:17

confidence which is the probability that

24:20

some object exists inside the box times

24:24

the probability that the object belongs

24:26

to a specific class given that the

24:29

object exists.

24:31

For example, suppose that the model

24:33

computes a probability of 0.98 that an

24:36

object exists in this predicted box and

24:40

that the probability is 0.96 that the

24:43

object is a cat given that there is an

24:46

object inside the box. This will result

24:49

in the following detection confidence.

24:52

The default confidence value for

24:54

displaying a box is 0.55.

24:58

If you change this value to 0.94,

25:01

no box will be shown around the cat. And

25:05

if you set this cutoff to zero, a lot of

25:08

irrelevant boxes will be shown. We will

25:11

now see how to train a Yoli model for a

25:14

certain task. To train a YOLO model, we

25:17

need images that include ground truth

25:19

bounding boxes around the objects. We

25:23

may manually draw such bounding boxes

25:25

around the objects and define the class

25:28

of such objects with tools like these.

25:33

But we will here use the Oxford pet data

25:35

set that includes bounding boxes around

25:38

the heads of cats and dogs that you can

25:41

download from the following website. You

25:44

may download the images of cats and dogs

25:47

here.

25:48

This folder contains information about

25:51

the bounding box and segmentation.

25:53

However, we will not use that folder

25:55

because these XML files do not match the

25:58

files required for YOLO.

26:01

Instead, I prepared a CSV file based on

26:05

those XML files that you can download

26:07

from my homepage. Here you will also

26:10

find the code that I will show later on.

26:15

The CSV file includes the file names of

26:18

the images in the Oxford PET data set,

26:21

the width and the height of the images.

26:26

It also includes the xcoordinate of the

26:28

top left corner of the bounding box

26:30

around the head and the corresponding

26:33

y-coordinate as well as the width and

26:36

height.

26:38

This column includes the breed which we

26:41

will not use. Instead, we will use this

26:44

column which tells if the image contains

26:46

a cat or a dog. Note that we will later

26:50

change these coordinates to center

26:52

coordinates because this is what Yol

26:54

expects.

26:56

So I first suggest that you create a

26:58

folder that you call Oxford pet data set

27:02

somewhere on your computer with the

27:04

exact same name as shown here because

27:08

the Python code that I will show later

27:10

on uses the name of this folder.

27:13

Note that you need to change this path

27:15

to where you have created the folder

27:17

Oxford pet data set on your computer.

27:21

After you have downloaded the zipped

27:23

file with the images, unzip it and place

27:27

it in the Oxford pet data set folder

27:30

with the name images.

27:33

Make sure that this folder contains the

27:35

image files with cats and dogs. You

27:38

should also put the CSV file that you

27:41

downloaded from my homepage in this

27:43

folder. This is the file that includes

27:46

the information about the bounding boxes

27:48

and labels.

27:51

YOLO expects a certain structure of the

27:54

data. We will use a Python code to

27:57

construct the folders and files

27:59

automatically.

28:01

Inside this folder, the code will create

28:04

a JML file which will contain the

28:07

following information.

28:09

But this line shows the path to the YOLO

28:11

data set folder and a path to the

28:15

training and validation images inside

28:17

the YOLO data set folder.

28:20

and all the classes you have in your

28:22

data set. Cats are herecoded as zeros

28:26

and dogs as ones.

28:29

You may also include a test data set,

28:32

but this is optional. We will not use

28:34

the test data set in this example. I

28:37

have created a Python code that sets up

28:40

this structure, so you do not need to do

28:42

this manually.

28:45

Inside the training folder, the code

28:47

will create the following two folders

28:51

where the image folder will contain the

28:53

training images, whereas the labels

28:56

folder will include the txt files that

28:59

include information about the ground

29:01

truth boxes from the CSV file.

29:05

Note that the file name of the txt file

29:08

should be the same as the corresponding

29:10

JPEG file.

29:12

For example, the first txt file will

29:15

include a zero or a one if it is a cat

29:19

or a dog. Remember that cats were coded

29:23

as zeros. So this value tells us that it

29:26

is a cat in the image. There will be one

29:30

row for each object in the image. Since

29:33

the Oxford pet data set only has one

29:36

object per image, there will only be one

29:39

row in each txt file.

29:42

These are the normalized centered

29:44

coordinates of the bounding box. And

29:47

these are the normalized width and

29:49

height of the bounding box. I will now

29:52

show the code that sets up everything we

29:55

need to train the YOLO model.

29:58

You can find this code on my homepage

30:01

where you can also download the

30:03

necessary CSV file.

30:06

You should paste this CSV file in the

30:08

Oxford PET data set folder.

30:11

This line saves the path to the CSV

30:14

file. Note that you might need to change

30:16

this path to where you have saved this

30:19

file on your computer. This line saves

30:22

the path to the folder where we have the

30:24

images of cats and dogs.

30:28

This line determines the path to the new

30:30

folder YOLO data set that will be

30:33

created later on. After you have run

30:36

these lines, you should see the

30:39

following new folder which contains the

30:42

empty folders for the training and

30:44

validation.

30:46

We will now read in this CSV file with

30:48

the labels and information about the

30:51

bounding boxes.

30:54

The following line reads in the CSV file

30:57

and saves it as a pandas data frame.

31:00

Then we save the file names from the

31:02

data frame and randomly shuffle the

31:05

order of the file names so that the

31:08

generated training and validation data

31:11

do not depend on the order of the files.

31:14

Note that you can also use stratified

31:16

sampling to get the same proportion of

31:19

classes in the two sets.

31:23

With this code, we say that 80% of the

31:25

images should be used for training data

31:28

and 20% for the validation data. Then we

31:32

loop through all the images and place

31:35

them in the correct folders

31:38

and create the txt files.

31:41

Note that we convert the original pixel

31:43

coordinates in the CSV file from the top

31:46

left corner to coordinates that instead

31:49

represent the center of the bounding

31:51

box. We also normalize the coordinates

31:54

based on the width and height of the

31:56

image.

31:59

We also normalize the width and height

32:01

of the bounding box and code cats as

32:04

zeros and dogs as ones.

32:07

Finally, we write this information to

32:10

the txt file with the same name as the

32:13

image file.

32:15

After you run this code, you should see

32:17

the image files and the text files in

32:20

the correct folder.

32:22

Finally, we create the JML file.

32:26

Note that you here need to manually type

32:28

the names of the classes if you would

32:30

like to apply the code on another data

32:33

set.

32:35

After you run the code, you should see

32:38

the following folders and the JML file

32:41

in the YOLO data set folder. Make sure

32:44

that the files look okay and that the

32:46

folders contains the images and the txt

32:49

files. You may also open this file in

32:53

the text editor. Make sure it looks

32:55

exactly as I showed before. It is

32:58

important that the file extension is jl.

33:03

It is now time to train the yol model.

33:06

We first import the yol class from the

33:09

ultralitics package.

33:12

Then we load a pre-trained YOLO 26N

33:15

model. This model has been trained on

33:18

the Cocoa data set. This model therefore

33:22

already knows how to detect cats. But we

33:25

will here fine-tune this model so that

33:27

it instead place a box around the heads

33:31

and focus the training only on cats and

33:34

dogs.

33:35

When you run this code, the training

33:38

should start.

33:39

This is the path to the JML file. We

33:43

here run 20 epochs to save time, but you

33:47

can run as long as the performance on

33:49

the validation data set continues to

33:51

improve. To save time for training, we

33:55

here use a smaller input image size than

33:58

a default 640. This may reduce the

34:01

performance, but the training goes a lot

34:04

faster.

34:07

During the training, you should observe

34:09

that the loss for the bounding boxes and

34:12

the class prediction decreases towards

34:14

zero.

34:16

The data set contains a small number of

34:19

corrupted images, but it is still

34:22

possible to train the model.

34:25

After each epoch, you should see the

34:28

performance based on the 738 images in

34:31

the validation data set. This is the

34:34

precision and recall for a certain

34:36

threshold and this is the MAP50 value

34:41

that we discussed previously. After 20

34:44

epochs, this value is close to one,

34:47

which is exactly what we aim for. After

34:50

the training is done, Yoli saves a

34:54

number of files in your active Python

34:56

working directory. For example, you can

34:59

find the performance metrics during the

35:01

training,

35:03

a confusion matrix and the precision

35:06

recall curve and much more.

35:10

In the fuller weights, you will find

35:12

your train Yolo model.

35:16

This model is based on the weights that

35:18

have been optimized up to the last

35:21

epoch. Whereas this file includes the

35:24

model with the weights that resulted in

35:26

the best performance based on the

35:28

validation data for a given epoch.

35:31

You may copy this file if you like to

35:33

use your model for future purposes.

35:37

If we for example place the best pt file

35:40

in the Oxford pet data set folder, we

35:43

can load it like this

35:46

and try it on any image that we like to

35:49

use.

35:52

This line computes the inference on all

35:55

the 738 images in the validation data

35:58

set based on our trained model.

36:02

Then we can show the result of the first

36:04

image from the validation data set. It

36:07

seems to work fine because the model

36:09

predicts that it is a cat and the box is

36:13

placed around the head.

36:16

Here are some other examples. Although

36:19

it looks very nice,

36:22

it makes mistakes because it should not

36:25

put a box around the tail here, but we

36:28

may change the confidence threshold from

36:30

the default value of 0.25 to 0.5 to

36:34

avoid bad predicted boxes. But using a

36:38

higher threshold may then result in some

36:41

images not showing any bounding box.

36:45

This was the end of this video about

36:47

object detection with YOLO. Thanks for

36:50

watching.

Interactive Summary

This video provides a comprehensive introduction to the YOLO (You Only Look Once) model for object detection. It covers the foundational architecture of the original YOLO paper, including how the image is divided into a grid, the definition of bounding box targets, and how confidence and loss functions are calculated. The video further explores performance metrics like precision, recall, and Average Precision (AP). Finally, it includes a practical demonstration using Python and the ultralytics library to fine-tune a pre-trained YOLO model on the Oxford Pet dataset, detailing the required data preparation and training process.

Suggested questions

4 ready-made prompts