I trained a Sign Language Detection Transformer (here's how you can do it too!)

Nicholas RenotteAbout 6 min readAug 18, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Object Detection: Identifying and locating objects within an image or video.
  • Sign Language Detection: Specifically, detecting and classifying different sign language signs.
  • DETR (DEtection TRansformer): A deep learning architecture for object detection that uses transformers.
  • DR Architecture: A simplified training algorithm used with DETR.
  • Real-time Detection: Processing and displaying object detections in real-time.
  • Pre-trained Weights: Pre-existing model weights from a previous training run, used to accelerate and improve new training.
  • Fine-tuning: Adjusting a pre-trained model with new data to improve performance on a specific task.
  • YOLO (You Only Look Once): A popular object detection algorithm and a specific data format for bounding box annotations.
  • Label Studio: A tool for labeling images and creating training data for machine learning models.
  • Hungarian Matcher (Linear Sum Assignment): An algorithm used in DETR to match predicted object queries to ground truth objects.
  • Loss Function: A function that quantifies the difference between predicted and actual values, used to guide model training.
  • Backbone Layer: The foundational convolutional neural network (CNN) used for feature extraction in the DETR model (e.g., ResNet50).
  • Transformer: A neural network architecture that uses self-attention mechanisms to process sequential data.
  • Inference Time: The time it takes for a model to make a prediction on a single input.
  • Frames Per Second (FPS): The number of frames processed per second, a measure of real-time performance.
  • UV: A python package and environment manager.
  • Naughty: The name of the custom library created for this project.

1. Setting up Real-time Detection with Pre-trained Weights

  • Cloning the GitHub Repository: The first step involves cloning the provided GitHub repository containing the code and pre-trained weights using the command git clone <repository_link>.
  • Navigating to the Directory: After cloning, the user navigates into the cloned directory using cd sign-detr.
  • Opening in VS Code: The project is then opened in VS Code using the command code ..
  • File Structure Overview: The repository contains:
    • data/pre-train: Images and labeled data.
    • Checkpoints: Pre-trained model weights (trained for 4426 epochs in approximately 24 hours on a MacBook).
    • source: Contains the core code, including realtime.py, data.py, model.py, train.py, test.py, and utils.
  • Running Real-time Detection: The realtime.py file is executed using uv run source realtime.py. This requires UV to be installed (pip install uv).
  • Camera Device Selection: The realtime.py script allows specifying the camera device to use (default is device 0).
  • Logging Information: The script provides logging information, including:
    • Model name.
    • Number of classes.
    • Hidden dimensions.
    • Inference time.
    • Frames per second (FPS).
  • Demonstration: The video demonstrates real-time detection of three sign language signs: "Hello," "I Love You," and "Thank You."

2. Data Collection and Preparation

  • Motivation for Custom Data: The pre-trained model may not perform optimally under different lighting conditions or for different individuals, necessitating fine-tuning with custom data.
  • Data Folder Structure: The data folder contains train and test subfolders, each with images and labels subfolders.
  • YOLO Label Format: The labels are stored in YOLO format, with each line representing an object and containing the class ID, center coordinates (x, y), width, and height.
  • Collecting Images: A utility script (utils/collect_images.py) is used to capture images from the camera.
    • The script allows specifying the camera device and the number of images to capture per class.
    • The captured images are saved into the train or test folders.
  • Deleting Existing Data: The existing images and labels in the train and test folders are deleted before collecting new data.
  • Labeling Images with Label Studio:
    • Label Studio is used to annotate the images with bounding boxes.
    • Label Studio is started using uv run label-studio.
    • A new project is created in Label Studio, and the images are imported.
    • The "Object Detection with Bounding Boxes" template is selected.
    • Custom labels are added, matching the order in the config.py file: "Hello," "I Love You," and "Thank You."
    • The images are then manually labeled by drawing bounding boxes around the signs and assigning the appropriate class.
  • Exporting Labeled Data: The labeled data is exported from Label Studio in YOLO format with images.
  • Replacing Existing Data with Labeled Data: The exported images and labels are copied into the train and test folders, replacing the original files.
  • Data Validation: The data.py script can be run (uv run source data.py) to visualize the data and transformations applied during training. The train=False argument can be used to disable training augmentations and view the raw data.

3. Model Training

  • Model Architecture (model.py):
    • The model uses a DETR architecture with a ResNet50 backbone for feature extraction.
    • Positional embeddings are added to the extracted features.
    • The features are then passed through a transformer encoder and decoder.
    • Two prediction heads are used: one for classification and one for bounding box regression.
    • The model architecture can be visualized using uv run source model.py, which uses the torchinfo library to display the model's layers and parameters (approximately 26.9 million trainable parameters).
  • Loss Function (Hungarian Matcher):
    • The loss function uses the Hungarian matcher (linear sum assignment) to match predicted object queries to ground truth objects.
    • The loss is composed of three components:
      • Classification loss (cross-entropy loss).
      • Bounding box loss (L1 loss for coordinates).
      • Generalized Intersection over Union (GIoU) loss (area overlap).
  • Training Script (train.py):
    • The train.py script contains the training loop and configuration options.
    • Pre-trained weights can be loaded to accelerate training.
    • The learning rate and loss component weights can be adjusted.
    • The number of training epochs can be set.
    • The script saves checkpoints every 10 epochs.
  • Starting Training: The training process is started using the command uv run source train.py.
  • Logging: The training script provides detailed logging information, including:
    • Data set statistics.
    • Data transforms.
    • Model architecture.
    • Training progress (epoch, train loss, test loss, elapsed time).
  • Checkpoints: Model checkpoints are saved in the checkpoints folder.

4. Model Testing and Evaluation

  • Testing Script (test.py):
    • The test.py script is used to evaluate the trained model on the test data set.
    • The path to the model weights must be specified in the script.
    • The script loads an image from the test partition, makes a prediction, and displays the results.
    • The script prints the predicted class, confidence score, and bounding box coordinates.
  • Real-time Testing (realtime.py):
    • The realtime.py script is used to perform real-time object detection with the trained model.
    • The path to the model weights must be specified in the script.
    • The script captures video from the camera, makes predictions on each frame, and displays the results.
    • A confidence threshold can be set to filter out low-confidence predictions.
  • Running Testing: The test script is executed using uv run source test.py and the real-time script is executed using uv run source realtime.py.
  • Confidence Threshold: The confidence threshold for displaying predictions can be adjusted in the realtime.py script.

5. Conclusion

The video provides a comprehensive guide to training a sign language object detection model from scratch using the DETR architecture. It covers data collection, labeling, model training, and testing, with a focus on practical implementation and ease of use. The use of pre-trained weights, custom data, and detailed logging makes it possible to train a reasonably accurate model in a relatively short amount of time, even without a dedicated GPU. The "Naughty" library and the detailed explanations aim to simplify the process and make it accessible to a wider audience.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.