YOLO explained step by step - Object detection
773 segments
In this video, we'll talk about the Yoli
model that is used for object detection.
I will mainly focus on the structure of
the model and how it can be trained
based on its loss function and then show
some simple Python code on how you can
train a simple YOLO model. I will focus
on the original Yoli model from the
paper you only look once. Note that
newer versions have been released since
2015 with a lot of improvements. I will
here assume that you are familiar with
CNN's and image classification. If not,
have a look at my video about transfer
learning. I also recommend that you
watch my video about object detection
and bounding box regression which
explains the basics. To understand YOLO,
we here use the following example image
with 224 pixels in width and 224 pixels
in height. This is the pixel coordinate
in the top left corner. And this is the
pixel coordinate in the bottom right
corner. YOLO adds a 7 * 7 grid on the
image like this, which means that the
image is divided into 49 grid cells.
This is an example of one grid cell. To
train a YOLA model, we need to manually
draw a bounding box around each object
in the image. This handdrawn bounding
box is referred to as the ground truth
box or the target box.
Here are a few examples of tools that
you can use to draw boxes to annotate
your images.
For each box including an object, we
also need to define the class of this
object. So we here say that the object
is a cat.
To train a YOLO model, we should provide
the center X and Y pixel coordinates of
the target box and the width and height
of the box
and the class. In this example, we will
have cats and dogs in the images. So the
number of classes is here too. So the
training target for this image should
look like this where these numbers are
the center coordinates of the bounding
box and these numbers correspond to the
width and height of the box.
Note that these values are normalized by
dividing by the width and height of the
image.
Finally, we say that the box contains a
cat.
If you would have a second object in the
image,
we need to draw a bounding box around
this object as well and determine the
center coordinates of this box and its
width and height
and its class.
So for this image, we need the following
inputs for the model in order to train
it so that the model knows where it
should put the predicted boxes and how
to predict the class.
This is the YOLO network from the
original paper. You only look once.
The input is a color image that is
resized to 448* 448 pixels. Remember
that a color image has three channels,
one for red, one for green, and one for
blue. A filter of size 7 by seven with a
depth of three is sliding over the image
with a stride of two. And after max
pooling, we get the feature map of size
112*
112. After the convolutions, the last
feature maps are flattened into a vector
which is connected to a fully connected
layer with 4,96
nodes. The 4,96
nodes are then connected to the
following output layer. So what we have
here is basically just a simple
convolutional neural network which you
should be familiar with if you have seen
my video about transfer learning. The
new thing here is this strange output
layer. If we were only to use this
network for image classification to
predict if the image contains a cat or a
dog, we would only need two output nodes
or in this case it would be enough with
just one since it is a simple binary
classification. If you also like to
learn the network to predict bounding
boxes around the cats and dogs, we need
to add output nodes that predict the
center x and y coordinates of the box as
well as its width and height.
This type of network only works if there
is one object per image.
If we have two objects in the image, we
need to have a unique output for each
box.
Let's rotate the output layer like this.
Now in the original YOLO model, they had
two boxes in each cell. Since there are
49 cells in the image, 96 boxes exist
for each image.
In contrast, the classification was
based on a grid cell, which means that
the original YOLO model could only
identify a maximum of 49 objects in an
image.
The reason for having two boxes per cell
was to let these compete to become the
responsible box which best fits the
target box. In later YOLO versions,
anchor boxes are used so that each grid
cell can predict multiple bounding boxes
allowing multiple objects and class
predictions within the same cell. So
every cell in the original Jolie model
has the following output nodes.
We can denote the output layer like this
because the image consists of 49 cells
where each cell is connected to 10
output nodes. Also, YOLO has one
additional output that predicts the
confidence which I will explain later
on.
So in our case where we only have two
classes cats and dogs, we'll have the
following output layer. This output
layer has 12 output nodes for each of
the 49 grid cells. The original YOLO
network was trained on the Pascal data
set which consists of 20 different
classes such as persons, dogs, cars,
airplanes, and trains. This now explains
the output layer shown in the original
YOLO paper because every cell in the
image has 30 output nodes, two for each
box that predict the center of the
bounding box, the width and the height
of the two boxes and the confidence of
the two boxes. Since the data set had 20
different classes,
30 output nodes are used for each cell
in the image.
Now let's try to understand how the
confidence is computed. But to
understand the confidence, we first need
to understand the measure intersection
over a union. This is our ground truth
or target bounding box. If you multiply
the width and the height of the box, we
see that the box covers 9,000 pixels.
When the network processes this image to
compute the two boxes in a given cell,
it may predict the following box given
the current weights of the network. The
predicted box has an area of 4,96 pixels
and the intersection covers 2,560
pixels. We can now calculate the
intersection of a union where we plug in
the area of the intersection here and
the areas of the target box and the
predicted box. This gives us an
intersection of a union score of about
0.24.
If the predicted box were over here
where it does not intersect the target
box, the intersection of a union score
would be zero.
And if the predicted box would have the
same size and location as the target
box, the intersection of a union would
be one. We therefore know that the
intersection of a union spans between
zero and one. And that the closer to one
we are, the better the predicted box
fits the target box. Now remember that
each cell has two boxes. But only the
two boxes that are within the cell that
includes the center coordinates of the
target box are important here. Although
these two boxes are within the target
box, they are less important because
they belong to a grid cell that does not
include the center of the target box.
The two boxes are equally useless for
this image as these two. The area of box
one is 1,000
and the area of box two is 225.
We previously saw that the area of the
target box was 9,000.
The intersection of a union of box one
is about 0.11.
Whereas the intersection of a union of
box two is about 0.025.
Since box one has a larger intersection
of a union compared to box two, it is
here considered as the responsible box.
Now suppose that the network processes
the image and currently predicts the
following confidence scores for box one
and two. Then YOLO constructs a target
confidence for all the boxes.
P object is set to one if the box is
responsible
and to zero if the boxes are not
responsible.
So for boxes that are not associated
with the cell where the center of the
target boxes, the target P object is set
to zero.
The only box that gets a target P object
of one is this box because it is a box
within the cell that includes the center
of the target box. And out of the two
boxes in this cell, it has the largest
intersection of union.
So the target confidence of box one is
0.111.
Whereas the target confidence for box
two is zero which is true also for the
other boxes in all the other grid cells.
The total confidence loss of this cell
is the square difference between the
target confidence and the predicted
confidence for the responsible box and
for the box that is not responsible.
For no object boxes or boxes that are
not responsible, the square difference
is multiplied by a weight which they set
to 0.5 in the paper.
Since the network is trained to reduce
the loss by changing the values of its
weights, the weights will be changed so
that the predicted confidences get
closer to the target confidence scores.
This is the full loss function that was
described in the original Jolu paper.
Before any loss is computed, the image
coordinates are normalized so that the
coordinates span between zero and one.
Lambda chord is a weight that puts more
weight on fitting the predicted box
closer to the target box. The value of
this weight was set to five in the
original paper.
These summations denote that we sum over
all the 49 grid cells. And these
summations denote that we sum over all
boxes within each cell. Since there are
only two boxes per cell, this means that
we sum the loss of the two boxes per
cell. This denotes that if an object
appears in cell I, it is set to one.
Else it is set to zero. So this means
that the class loss will only be
computed within a cell that includes the
center of the target box because this
notation will be set to zero for all
other cells.
This notation denotes that if box J in
cell I is responsible
it will be set to one else to zero.
This means that this notation will be
set to one only for this box. Which
means that mainly this box will update
its location and size during the current
training step. The other 97 boxes will
focus mainly on updating this confidence
loss term where its target value is zero
which means that the current confidence
is forced down to zero. So for no
responsible boxes, this is set to one.
This means that the total loss is mainly
dependent on box one which is the
responsible box as well as the cell that
includes the center of the target box
and very little on nonresponsible boxes.
Suppose that we will compute the class
loss of this cell which we know contains
the center of the object with class cat.
The target value for this output node is
therefore one and the target value for
this node is zero because the given grid
cell does not include a dog.
When the network processes the image, it
might compute the predicted probability
of 0.8. Since the network tries to
minimize the loss during training, it
will update its weights so that the
predicted probability for a cat gets
closer to one and the predicted
probability for a dog gets closer to
zero. At the same time, it will update
the location and size of the responsible
box so that it gets closer to the target
box
so that the square difference between
the center of the responsible box and
the target box as well as the difference
between a width and a height are as
small as possible. The reason why they
based the loss function on the square
root of the width and height was that
small deviations in large boxes should
matter less than in small boxes. The
idea with the training is that the
weights are updated so that the
responsible box fits as good as possible
with the target box and that the cell
correctly predicts the class cat. Note
that the network does not see the target
box during training. It only sees the
image. The information of the target box
is only used as target values in the
output nodes. The network therefore
learns to identify what is presented in
the image in relation to the location of
the target box. We'll now discuss some
of the performance metrics that are used
in object detection. In this image, we
have three cats with ground truth boxes
in green. The blue box is the predicted
box by the network and the label is the
predicted class. This is an example of a
true positive result because the
predicted box fits nicely with the
ground truth box and the predicted class
cat is correct. This is an example of a
false positive result because the
predicted class dog is incorrect.
Whereas this is a false negative result
because there is no predicted box
associated with a ground truth box. This
is an example of a true negative result
which does not make sense because a
ground truth box without any object
should not exist.
We can now calculate the precision which
measures the proportion of predicted
bounding boxes that are correct which is
the true positives among all predicted
detections.
Note that this metric evaluates only the
predictions made by the model. In
comparison, the recall measures the
proportion of ground trth boxes that are
correctly detected by the model. It
indicates how many of the actual objects
in the data set that are successfully
predicted.
Only one ground truth box was correctly
detected out of three.
Note that this results in both one false
positive and one false negative result
because the predicted box has incorrect
class
and the ground truth box was incorrectly
detected.
But as we have seen before, there should
be 96 predicted boxes in this image. So
where are those? Well, to count or
display a predicted box, we use a
confidence threshold that is here set to
0.7. This means that we only display
predicted boxes with a confidence score
greater than 0.7.
If you increase this threshold to 0.9,
the following predicted box will
disappear.
and the result will change to only a
false negative result and the precision
will therefore change value. If you use
a higher confidence threshold, we expect
higher precision because fewer predicted
bounding boxes will be considered and
those that remain are more likely to be
correct. However, the recall is expected
to decrease because fewer ground truth
boxes will be successfully detected. If
instead reduce the confidence threshold
to 0.3,
we will see more boxes.
The problem now is whether we should
count this predicted box as a false
positive or true positive because the
class prediction is correct but the
predicted box does not fit well with the
ground truth box. Usually one uses an
additional cutoff based on the
intersection of a union. This threshold
is commonly set to 0.5.
So if the intersection of a union in
this case is 0.4 four which is below the
threshold. This will be counted both as
a false positive result and a false
negative result because the ground truth
box is not correctly detected and the
predicted box does not properly identify
the object.
The precision and recall will therefore
change values like this.
If you reduce the intersection of a
union threshold to 0.3,
this will now be counted as a true
positive result because the predicted
box correctly identifies the object and
the precision and recall will now show
these values.
So we now know that the position and
recall depend on these cutoff values.
To get an overall measure of the
performance, we can change the cutoff
and compute the precision and recall for
a number of different threshold values.
We compute these measures for all
images, not just this one.
Then we can plot how the precision and
recall change when we change the
confidence threshold. This curve is
called the precision recall curve or the
PR curve. For example, this point on the
curve may correspond to a confidence
threshold of 0.99.
At such high threshold, precision is
expected to be close to one since only
very confidence predictions are kept
while the recall is expected to be close
to zero because many ground truth
objects will not be detected.
For a lower threshold, the precision is
expected to reduce and the recall to
increase.
Our final measure of performance is the
average precision, which is the area
below this curve. The area under this
curve spans between zero and one. Since
this curve is based on the intersection
of union threshold of 0.5,
it is called AP0.5 or AP50.
If you instead used a threshold of 0.75,
the curve will change and the measure
would then be called AP0.75
or AP75.
One can even compute the curve for
thresholds going from 0.5 to 0.95
with 0.05 increments and compute the
average of these.
Then the measure will be denoted like
this.
Suppose that our AP50 score is 0.8. In
this case,
what we want is a score close to one
because at the certain confidence
threshold, we will have a precision and
recall that are close to one.
Note that we had so far only computed
the AP50 for the cats.
We also need to compute the same for the
dogs for all our images which might
result in the AP50 score of for example
0.85.
So our final measure of the performance
of the object detection is the average
AP50 value of all the classes which is
called the mean AP50.
We'll now have a look at some basic
Python code so that you can train your
first Jolo model.
Before we discuss the training, let's
first see how we can use an existing
YOLO model to do object detection. We
first begin to install the ultral litic
package with pip install. Then we import
the yolo class
and load the pre-trained model. We'll
here use the model called YOLO 26N which
was released in September 2025.
This model is the smallest model in it
26 family since it is the smallest
network. It is the fastest for training
and for inference. But the smaller size
comes with the cost because it performs
worse than larger models. Anyway, it is
a good model to start with.
To use the existing model to run
inference on an image, we simply provide
the path to the image. You may change
this path to an image that you would
like to try on. This is the confidence
threshold for displaying the predicted
bounding boxes. Then we simply display
the result with this line of code.
A common mistake is to interpret this
value as the probability that it is a
cat.
This value represents the detection
confidence which is the probability that
some object exists inside the box times
the probability that the object belongs
to a specific class given that the
object exists.
For example, suppose that the model
computes a probability of 0.98 that an
object exists in this predicted box and
that the probability is 0.96 that the
object is a cat given that there is an
object inside the box. This will result
in the following detection confidence.
The default confidence value for
displaying a box is 0.55.
If you change this value to 0.94,
no box will be shown around the cat. And
if you set this cutoff to zero, a lot of
irrelevant boxes will be shown. We will
now see how to train a Yoli model for a
certain task. To train a YOLO model, we
need images that include ground truth
bounding boxes around the objects. We
may manually draw such bounding boxes
around the objects and define the class
of such objects with tools like these.
But we will here use the Oxford pet data
set that includes bounding boxes around
the heads of cats and dogs that you can
download from the following website. You
may download the images of cats and dogs
here.
This folder contains information about
the bounding box and segmentation.
However, we will not use that folder
because these XML files do not match the
files required for YOLO.
Instead, I prepared a CSV file based on
those XML files that you can download
from my homepage. Here you will also
find the code that I will show later on.
The CSV file includes the file names of
the images in the Oxford PET data set,
the width and the height of the images.
It also includes the xcoordinate of the
top left corner of the bounding box
around the head and the corresponding
y-coordinate as well as the width and
height.
This column includes the breed which we
will not use. Instead, we will use this
column which tells if the image contains
a cat or a dog. Note that we will later
change these coordinates to center
coordinates because this is what Yol
expects.
So I first suggest that you create a
folder that you call Oxford pet data set
somewhere on your computer with the
exact same name as shown here because
the Python code that I will show later
on uses the name of this folder.
Note that you need to change this path
to where you have created the folder
Oxford pet data set on your computer.
After you have downloaded the zipped
file with the images, unzip it and place
it in the Oxford pet data set folder
with the name images.
Make sure that this folder contains the
image files with cats and dogs. You
should also put the CSV file that you
downloaded from my homepage in this
folder. This is the file that includes
the information about the bounding boxes
and labels.
YOLO expects a certain structure of the
data. We will use a Python code to
construct the folders and files
automatically.
Inside this folder, the code will create
a JML file which will contain the
following information.
But this line shows the path to the YOLO
data set folder and a path to the
training and validation images inside
the YOLO data set folder.
and all the classes you have in your
data set. Cats are herecoded as zeros
and dogs as ones.
You may also include a test data set,
but this is optional. We will not use
the test data set in this example. I
have created a Python code that sets up
this structure, so you do not need to do
this manually.
Inside the training folder, the code
will create the following two folders
where the image folder will contain the
training images, whereas the labels
folder will include the txt files that
include information about the ground
truth boxes from the CSV file.
Note that the file name of the txt file
should be the same as the corresponding
JPEG file.
For example, the first txt file will
include a zero or a one if it is a cat
or a dog. Remember that cats were coded
as zeros. So this value tells us that it
is a cat in the image. There will be one
row for each object in the image. Since
the Oxford pet data set only has one
object per image, there will only be one
row in each txt file.
These are the normalized centered
coordinates of the bounding box. And
these are the normalized width and
height of the bounding box. I will now
show the code that sets up everything we
need to train the YOLO model.
You can find this code on my homepage
where you can also download the
necessary CSV file.
You should paste this CSV file in the
Oxford PET data set folder.
This line saves the path to the CSV
file. Note that you might need to change
this path to where you have saved this
file on your computer. This line saves
the path to the folder where we have the
images of cats and dogs.
This line determines the path to the new
folder YOLO data set that will be
created later on. After you have run
these lines, you should see the
following new folder which contains the
empty folders for the training and
validation.
We will now read in this CSV file with
the labels and information about the
bounding boxes.
The following line reads in the CSV file
and saves it as a pandas data frame.
Then we save the file names from the
data frame and randomly shuffle the
order of the file names so that the
generated training and validation data
do not depend on the order of the files.
Note that you can also use stratified
sampling to get the same proportion of
classes in the two sets.
With this code, we say that 80% of the
images should be used for training data
and 20% for the validation data. Then we
loop through all the images and place
them in the correct folders
and create the txt files.
Note that we convert the original pixel
coordinates in the CSV file from the top
left corner to coordinates that instead
represent the center of the bounding
box. We also normalize the coordinates
based on the width and height of the
image.
We also normalize the width and height
of the bounding box and code cats as
zeros and dogs as ones.
Finally, we write this information to
the txt file with the same name as the
image file.
After you run this code, you should see
the image files and the text files in
the correct folder.
Finally, we create the JML file.
Note that you here need to manually type
the names of the classes if you would
like to apply the code on another data
set.
After you run the code, you should see
the following folders and the JML file
in the YOLO data set folder. Make sure
that the files look okay and that the
folders contains the images and the txt
files. You may also open this file in
the text editor. Make sure it looks
exactly as I showed before. It is
important that the file extension is jl.
It is now time to train the yol model.
We first import the yol class from the
ultralitics package.
Then we load a pre-trained YOLO 26N
model. This model has been trained on
the Cocoa data set. This model therefore
already knows how to detect cats. But we
will here fine-tune this model so that
it instead place a box around the heads
and focus the training only on cats and
dogs.
When you run this code, the training
should start.
This is the path to the JML file. We
here run 20 epochs to save time, but you
can run as long as the performance on
the validation data set continues to
improve. To save time for training, we
here use a smaller input image size than
a default 640. This may reduce the
performance, but the training goes a lot
faster.
During the training, you should observe
that the loss for the bounding boxes and
the class prediction decreases towards
zero.
The data set contains a small number of
corrupted images, but it is still
possible to train the model.
After each epoch, you should see the
performance based on the 738 images in
the validation data set. This is the
precision and recall for a certain
threshold and this is the MAP50 value
that we discussed previously. After 20
epochs, this value is close to one,
which is exactly what we aim for. After
the training is done, Yoli saves a
number of files in your active Python
working directory. For example, you can
find the performance metrics during the
training,
a confusion matrix and the precision
recall curve and much more.
In the fuller weights, you will find
your train Yolo model.
This model is based on the weights that
have been optimized up to the last
epoch. Whereas this file includes the
model with the weights that resulted in
the best performance based on the
validation data for a given epoch.
You may copy this file if you like to
use your model for future purposes.
If we for example place the best pt file
in the Oxford pet data set folder, we
can load it like this
and try it on any image that we like to
use.
This line computes the inference on all
the 738 images in the validation data
set based on our trained model.
Then we can show the result of the first
image from the validation data set. It
seems to work fine because the model
predicts that it is a cat and the box is
placed around the head.
Here are some other examples. Although
it looks very nice,
it makes mistakes because it should not
put a box around the tail here, but we
may change the confidence threshold from
the default value of 0.25 to 0.5 to
avoid bad predicted boxes. But using a
higher threshold may then result in some
images not showing any bounding box.
This was the end of this video about
object detection with YOLO. Thanks for
watching.
Ask follow-up questions or revisit key timestamps.
This video provides a comprehensive introduction to the YOLO (You Only Look Once) model for object detection. It covers the foundational architecture of the original YOLO paper, including how the image is divided into a grid, the definition of bounding box targets, and how confidence and loss functions are calculated. The video further explores performance metrics like precision, recall, and Average Precision (AP). Finally, it includes a practical demonstration using Python and the ultralytics library to fine-tune a pre-trained YOLO model on the Oxford Pet dataset, detailing the required data preparation and training process.
Videos recently processed by our community