NVIDIA NeMo Microservices: ULTIMATE Guide for Model Fine-Tuning!

Mervin PraisonAbout 4 min readMay 17, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

NVIDIA Nemo Microservices, Data Flywheel, Nemo Curator (data processing), Nemo Customizer (model customization/fine-tuning), Nemo Evaluator (model evaluation), Nemo Guardrails, Nemo Retriever (RAG pipeline), Llama 3 2.1 billion instruct model, Function calling/Tool calling, XLAM Salesforce dataset, NGC API key, MiniKube, Helm Chart, Data Store, Entity Store, LoRA fine-tuning, SFT (Supervised Fine-Tuning).

1. Introduction to NVIDIA Nemo Microservices

The video introduces NVIDIA Nemo Microservices as a solution to simplify the data flywheel for running models. The data flywheel consists of several components: Nemo Curator for data processing, Nemo Customizer for model customization, Nemo Evaluator for model evaluation, Nemo Guardrails, and Nemo Retriever for RAG pipelines. The video focuses on Nemo Curator and Nemo Customizer, demonstrating data processing and model fine-tuning.

2. Performance Stats of Nemo Microservices

NVIDIA Nemo microservices offer significant performance improvements:

  • Nemo Customizer: 1.8x faster post-training.
  • Nemo Evaluator: 3x reduction in APIs.
  • Nemo Guardrails: 1.4x higher safety compliance with minimal latency.

3. Setting up the Environment

3.1. Prerequisites and API Key

The first step involves obtaining an NGC API key from the provided link. This key is essential for downloading the Llama 3 2.1 billion instruct model. The video mentions using an NVIDIA H100 configuration (with two GPUs). A beginner's tutorial link is provided for detailed hardware and software requirements, including GPU specifications, disk space, and necessary components.

3.2. Installation Process

To simplify the installation, a script is provided to automatically download and install the Helm chart. Alternatively, manual installation using MiniKube is possible, but requires exporting the NGC API key in the terminal. After installation, the running pods can be viewed, and the IP address for MiniKube is identified.

4. Data Preparation with Nemo Curator

4.1. Downloading and Configuring the Notebook

A Python notebook for data preparation is downloaded and opened. The requirements.txt file is used to install necessary dependencies, which are then upgraded to the latest versions. The configuration file (config.py) needs to be updated with the correct NDS URL (Data Store Nemo URL), Nemo URL, and NIM URL. These URLs correspond to the MiniKube IP address and mapped ports, which can be viewed using kubectl get svc.

4.2. Running the Jupyter Notebook

The Jupyter Lab notebook is started using the specified command. The downloaded data preparation notebook is opened, and the process involves downloading the Salesforce tool calling dataset, preparing the data for customization, and preparing data for evaluation. A Hugging Face token and endpoint are required to access the dataset.

4.3. Data Processing Steps

The data preparation notebook performs the following steps:

  1. Downloading the Dataset: The Salesforce tool calling dataset is downloaded after granting access on Hugging Face.
  2. Data Conversion: The dataset is converted to the OpenAI specification format.
  3. Data Saving: The processed data is saved into training.jsonl (for training) and validation.jsonl (for validation). A portion of the data is also prepared for evaluation and saved in a separate file.

The data format in training.jsonl includes "role," "user content," "content," and "tool call information" in JSON format.

5. Model Fine-tuning with Nemo Customizer

5.1. Setting up the Fine-tuning Notebook

A separate notebook for fine-tuning and inference using Nemo Customizer is downloaded and opened in Jupyter. The NDS URL, Nemo URL, and NIM URL are defined, consistent with the config.py file.

5.2. Uploading Data to Nemo Data Store

The prepared data (training, validation, and XLAM) is uploaded to the Nemo Data Store. A repository is created, and the files are uploaded to this repository. The data is then registered with the Nemo Entity Store.

5.3. LoRA Fine-tuning Process

The fine-tuning process involves creating a job using Nemo Customizer. The Llama 3 2.1 billion instruct model is downloaded locally to the Nemo URL. The training is triggered using the downloaded model and the uploaded dataset. The configuration includes parameters such as:

  • sft_training_type: LoRA
  • finetuning_type: LoRA
  • number_of_epochs
  • batch_size
  • learning_rates

These parameters can be adjusted based on specific requirements.

5.4. Monitoring the Job Status

The job status can be monitored, transitioning from "created" to "running" and finally "completed." Logs are available for review. After completion, the customized model can be verified to ensure it is available.

6. Conclusion

The video demonstrates how to fine-tune the Llama 3 2.1 billion parameter model using the XLAM function calling dataset, enabling it with function calling capabilities. While the initial setup may seem complex, the process becomes easier with familiarity. Fine-tuning with different datasets involves simply changing the dataset name during data preparation. The presenter encourages viewers to share their thoughts and explore another video on vulnerability scanning using Nemo.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.