Key Concepts
- Object Detection: Identifying and locating objects within an image or video.
- Sign Language Detection: Specifically, detecting and classifying different sign language signs.
- DETR (DEtection TRansformer): A deep learning architecture for object detection that uses transformers.
- DR Architecture: A simplified training algorithm used with DETR.
- Real-time Detection: Processing and displaying object detections in real-time.
- Pre-trained Weights: Pre-existing model weights from a previous training run, used to accelerate and improve new training.
- Fine-tuning: Adjusting a pre-trained model with new data to improve performance on a specific task.
- YOLO (You Only Look Once): A popular object detection algorithm and a specific data format for bounding box annotations.
- Label Studio: A tool for labeling images and creating training data for machine learning models.
- Hungarian Matcher (Linear Sum Assignment): An algorithm used in DETR to match predicted object queries to ground truth objects.
- Loss Function: A function that quantifies the difference between predicted and actual values, used to guide model training.
- Backbone Layer: The foundational convolutional neural network (CNN) used for feature extraction in the DETR model (e.g., ResNet50).
- Transformer: A neural network architecture that uses self-attention mechanisms to process sequential data.
- Inference Time: The time it takes for a model to make a prediction on a single input.
- Frames Per Second (FPS): The number of frames processed per second, a measure of real-time performance.
- UV: A python package and environment manager.
- Naughty: The name of the custom library created for this project.
1. Setting up Real-time Detection with Pre-trained Weights
- Cloning the GitHub Repository: The first step involves cloning the provided GitHub repository containing the code and pre-trained weights using the command
git clone <repository_link>. - Navigating to the Directory: After cloning, the user navigates into the cloned directory using
cd sign-detr. - Opening in VS Code: The project is then opened in VS Code using the command
code .. - File Structure Overview: The repository contains:
data/pre-train: Images and labeled data.- Checkpoints: Pre-trained model weights (trained for 4426 epochs in approximately 24 hours on a MacBook).
source: Contains the core code, includingrealtime.py,data.py,model.py,train.py,test.py, andutils.
- Running Real-time Detection: The
realtime.pyfile is executed usinguv run source realtime.py. This requires UV to be installed (pip install uv). - Camera Device Selection: The
realtime.pyscript allows specifying the camera device to use (default is device 0). - Logging Information: The script provides logging information, including:
- Model name.
- Number of classes.
- Hidden dimensions.
- Inference time.
- Frames per second (FPS).
- Demonstration: The video demonstrates real-time detection of three sign language signs: "Hello," "I Love You," and "Thank You."
2. Data Collection and Preparation
- Motivation for Custom Data: The pre-trained model may not perform optimally under different lighting conditions or for different individuals, necessitating fine-tuning with custom data.
- Data Folder Structure: The
datafolder containstrainandtestsubfolders, each withimagesandlabelssubfolders. - YOLO Label Format: The labels are stored in YOLO format, with each line representing an object and containing the class ID, center coordinates (x, y), width, and height.
- Collecting Images: A utility script (
utils/collect_images.py) is used to capture images from the camera.- The script allows specifying the camera device and the number of images to capture per class.
- The captured images are saved into the
trainortestfolders.
- Deleting Existing Data: The existing images and labels in the
trainandtestfolders are deleted before collecting new data. - Labeling Images with Label Studio:
- Label Studio is used to annotate the images with bounding boxes.
- Label Studio is started using
uv run label-studio. - A new project is created in Label Studio, and the images are imported.
- The "Object Detection with Bounding Boxes" template is selected.
- Custom labels are added, matching the order in the
config.pyfile: "Hello," "I Love You," and "Thank You." - The images are then manually labeled by drawing bounding boxes around the signs and assigning the appropriate class.
- Exporting Labeled Data: The labeled data is exported from Label Studio in YOLO format with images.
- Replacing Existing Data with Labeled Data: The exported images and labels are copied into the
trainandtestfolders, replacing the original files. - Data Validation: The
data.pyscript can be run (uv run source data.py) to visualize the data and transformations applied during training. Thetrain=Falseargument can be used to disable training augmentations and view the raw data.
3. Model Training
- Model Architecture (model.py):
- The model uses a DETR architecture with a ResNet50 backbone for feature extraction.
- Positional embeddings are added to the extracted features.
- The features are then passed through a transformer encoder and decoder.
- Two prediction heads are used: one for classification and one for bounding box regression.
- The model architecture can be visualized using
uv run source model.py, which uses thetorchinfolibrary to display the model's layers and parameters (approximately 26.9 million trainable parameters).
- Loss Function (Hungarian Matcher):
- The loss function uses the Hungarian matcher (linear sum assignment) to match predicted object queries to ground truth objects.
- The loss is composed of three components:
- Classification loss (cross-entropy loss).
- Bounding box loss (L1 loss for coordinates).
- Generalized Intersection over Union (GIoU) loss (area overlap).
- Training Script (train.py):
- The
train.pyscript contains the training loop and configuration options. - Pre-trained weights can be loaded to accelerate training.
- The learning rate and loss component weights can be adjusted.
- The number of training epochs can be set.
- The script saves checkpoints every 10 epochs.
- The
- Starting Training: The training process is started using the command
uv run source train.py. - Logging: The training script provides detailed logging information, including:
- Data set statistics.
- Data transforms.
- Model architecture.
- Training progress (epoch, train loss, test loss, elapsed time).
- Checkpoints: Model checkpoints are saved in the
checkpointsfolder.
4. Model Testing and Evaluation
- Testing Script (test.py):
- The
test.pyscript is used to evaluate the trained model on the test data set. - The path to the model weights must be specified in the script.
- The script loads an image from the test partition, makes a prediction, and displays the results.
- The script prints the predicted class, confidence score, and bounding box coordinates.
- The
- Real-time Testing (realtime.py):
- The
realtime.pyscript is used to perform real-time object detection with the trained model. - The path to the model weights must be specified in the script.
- The script captures video from the camera, makes predictions on each frame, and displays the results.
- A confidence threshold can be set to filter out low-confidence predictions.
- The
- Running Testing: The test script is executed using
uv run source test.pyand the real-time script is executed usinguv run source realtime.py. - Confidence Threshold: The confidence threshold for displaying predictions can be adjusted in the
realtime.pyscript.
5. Conclusion
The video provides a comprehensive guide to training a sign language object detection model from scratch using the DETR architecture. It covers data collection, labeling, model training, and testing, with a focus on practical implementation and ease of use. The use of pre-trained weights, custom data, and detailed logging makes it possible to train a reasonably accurate model in a relatively short amount of time, even without a dedicated GPU. The "Naughty" library and the detailed explanations aim to simplify the process and make it accessible to a wider audience.
AI summaries can miss context or contain errors. Check important details against the original video.





