Stanford CS230 | Autumn 2025 | Lecture 3: Full Cycle of a DL project
By Unknown Author
Key Concepts
- Full Cycle of a Deep Learning Project: The entire process from problem definition to deployment and maintenance.
- AI Projects vs. Traditional Software Engineering: AI projects involve both code and data, with data's richness and unpredictability making them iterative.
- Iterative Development: The core methodology for AI projects, involving building, observing, and refining.
- Data Richness and Unpredictability: The inherent complexity and unknown factors within data that influence AI system behavior.
- Empirical/Experimental Process: AI development relies heavily on doing, observing, and learning from outcomes.
- Siamese Network: A neural network architecture commonly used for face recognition, which compares two images to determine if they belong to the same person.
- Data Collection Speed: A critical factor for startup success, prioritizing rapid data acquisition for quick iteration.
- Error Analysis: The process of identifying where a model performs poorly to guide improvements.
- Data-Centric AI: A discipline focused on systematically engineering data to build successful AI systems.
- Deployment Optimization: Techniques to make AI systems computationally feasible for real-world applications.
- Visual Activity Detection (VAD): A low-cost, low-power method to detect human presence before engaging more computationally intensive recognition systems.
- Data Drift/Concept Drift: Changes in the real-world data distribution or input-output mapping that can degrade model performance over time.
- Monitoring and Maintenance: Essential post-deployment activities to ensure continued system effectiveness.
- Human-Level Performance: A common benchmark for evaluating AI model accuracy.
The Full Cycle of a Deep Learning Project
This discussion outlines the comprehensive lifecycle of a deep learning project, emphasizing its iterative and empirical nature, which distinguishes it from traditional software engineering. The core argument is that successful AI development requires a holistic approach beyond just model training, encompassing data, deployment, and ongoing maintenance.
1. The Nature of AI Projects: Code vs. Data
A fundamental difference between AI projects and traditional software engineering lies in the role of data. While traditional projects offer direct control over code, AI projects involve both code and data. The inherent richness and unpredictability of data mean that developers often cannot foresee all the nuances and challenges their algorithms will encounter.
- Example: In a face recognition system, it's difficult to predict in advance how variations in lighting, hairstyles, facial expressions, or accessories like glasses will affect performance.
- Key Point: This unpredictability necessitates an iterative development process where systems are built, observed, and refined based on real-world performance. This applies equally to deep learning models and large language models (LLMs), which are trained on vast, unexamined datasets.
2. The AI Project Workflow: Beyond Model Training
While academic focus often centers on model building and evaluation, a complete AI system development involves several crucial stages:
- Problem Specification: Clearly defining the task and its objectives.
- Data Collection: Acquiring relevant data for training.
- Model Design: Choosing or creating an appropriate model architecture.
- Model Training: Using the data to train the model.
- Iteration: Repeating the design, training, and analysis steps until satisfactory performance is achieved.
- Deployment: Integrating the trained model into a functional system.
- Maintenance: Monitoring and updating the system over time.
The speaker emphasizes moving beyond the "small box of models" to understand the broader development landscape.
3. Illustrative Example: Face Recognition for Door Access
The running example used throughout the discussion is building a face recognition system for unlocking doors. This system would take a picture of an approaching person and decide whether to grant access. A common commercial application discussed is using face recognition to verify that the person swiping a key card is indeed the authorized individual, enhancing security.
4. The Iterative Development Loop
The core of machine learning development is a rapid, iterative loop:
- Design Model/Data: Define the model architecture and plan data collection/preparation.
- Train Model: Train the model using the collected data.
- Analyze Results: Evaluate the model's performance and identify areas for improvement.
- Update Model/Data: Modify the model architecture, data collection strategy, or data itself based on the analysis.
- Repeat: Continue this loop until the model meets performance requirements.
5. Face Recognition Architecture: Siamese Networks
For face recognition, a common neural network architecture is the Siamese network.
- Functionality: It takes two images as input and determines if they depict the same person.
- Advantage: This approach avoids retraining the network for each individual. Instead, registration pictures of authorized users are stored, and new incoming images are compared against these.
- Application: In the door access scenario, a new image is compared against registered images of authorized individuals. For the key card example, a picture taken upon card swipe is compared against the registered photo of the cardholder.
6. Data Collection Strategies: Prioritizing Speed
A key challenge in AI projects is acquiring sufficient and representative data. The speaker poses a question to the audience: "How would you go about getting data to train the system if you can't download from the internet?"
- Audience Suggestions: Video streaming services (like Zoom), placing cameras and asking for opt-in, using personal networks (friends, colleagues), and leveraging university resources (like Stanford).
- Guiding Principle: Speed of Execution: The speaker advocates for prioritizing speed in data collection, especially for startups. The rationale is that rapid data acquisition allows for quicker iteration, faster discovery of data issues, and more efficient problem-solving.
- Timeframe: Aiming for data collection within one to two days, even if the dataset is smaller or of lower quality initially.
- Rationale: The value of data is hard to predict in advance. Quick iteration helps uncover what's truly important in the data. Spending excessive time on data collection can create a bottleneck, especially when model training can be relatively fast.
- Real-world Example: A CEO spent over $100 million buying a company for its data, only to later question its monetization potential, highlighting the difficulty of pre-assessing data value.
- Practical Approach: For campus environments, leveraging high foot-traffic areas like cafeterias and politely asking individuals for permission to take pictures or record voice samples is suggested. The emphasis is on informed consent and respecting privacy.
- Time Allocation: A common practice is to set short, fixed timeframes (e.g., 48 hours) for data collection, forcing creative and efficient solutions. This contrasts with asking "how long will it take?" which often leads to slower progress.
7. The Empirical Nature of AI and LLMs
The speaker reiterates that machine learning is an empirical (experimental) process. This means building, observing, and learning from outcomes.
- LLMs: Even with LLMs, where prompts are used, the process is experimental because the training data is vast and unknown. Prompt engineering involves trying different prompts and observing the results to refine them.
- Responsible AI: Building and testing AI systems in a sandbox environment before full deployment is crucial for identifying potential safety issues, inappropriate responses, or unforeseen consequences.
8. Error Analysis and Data Improvement
When a model performs poorly, error analysis is key to understanding why and how to improve it.
- Strategies:
- Change Model Architecture: Explore different neural network designs.
- Change Data: This is often the most impactful.
- Data-Centric AI: This discipline focuses on systematically engineering data. If a face recognition system struggles with people wearing hats, the solution might be to collect more data of people wearing hats.
- Targeted Data Collection: Blindly collecting more data is inefficient. Error analysis helps identify specific subsets of data where the model struggles (e.g., people with long hair, wearing glasses, or in specific poses) and guides targeted data acquisition.
- LLM Data Quality: For LLMs, high-quality, edited content (like books) is more valuable than random internet chat. Similarly, for face recognition, sharp, in-focus images are better than blurry ones.
9. Data Distribution and Model Performance
The similarity between training data distribution and real-world data distribution is important, but perhaps less critical than commonly believed, especially for large neural networks.
- Large Neural Networks: These models have a high capacity to absorb diverse data, including slightly irrelevant examples, without significant performance degradation. They can learn core tasks even with some "distracting" data.
- Historical View: In the past, with smaller models and limited computational resources, strict adherence to matching training and testing distributions was more critical.
- Modern Approach: With larger, more capable neural networks, it's more acceptable to include a broader range of data, as long as it's not outright incorrect. The analogy is made to the human brain, which can learn multiple skills without one hindering the other.
10. Deployment: Optimizing for Practicality
Deploying AI models often involves significant software engineering work. For practical applications like face recognition, computational cost is a major consideration.
- Challenge: Streaming high-resolution video 24/7 to the cloud for classification at 30 frames per second can be prohibitively expensive and slow.
- Optimization: Visual Activity Detection (VAD): A common strategy is to use a low-cost, low-power VAD system to quickly detect potential human presence. Only when VAD triggers is the more computationally intensive face recognition model engaged.
- VAD Options:
- Non-ML Based: Detecting changes in pixel values above a certain threshold. This is simple and fast but can be triggered by non-human activity (e.g., swaying trees, passing cars).
- Small ML Model: Training a lightweight neural network to detect human presence. This is more sophisticated but requires training.
- Decision-Making: The choice between VAD options depends on factors like the environment (e.g., busy street vs. quiet home) and cost. The speaker advocates for starting with the quickest implementation (Option 1) to gather insights and then iterating.
- Discovery through Implementation: Implementing a system, even a simple one, reveals practical challenges. For face recognition, it was discovered that selecting high-resolution, in-focus frames significantly boosts accuracy, a detail not easily predicted beforehand. This led to the development of models that not only detect faces but also assess focus.
11. Performance Benchmarking
- Human-Level Performance: A common benchmark for AI systems, especially in tasks where humans excel.
- Face Recognition: AI systems can surpass human performance in controlled environments for distinguishing between two images of the same person.
- Challenging Tasks: For tasks where humans are not inherently good (e.g., recommending movies), establishing a baseline is harder, but AI can potentially outperform humans.
12. Monitoring and Maintenance: Addressing Data/Concept Drift
Once a model is deployed, continuous monitoring and maintenance are crucial due to data drift and concept drift.
- Data Drift: The distribution of incoming data changes over time (e.g., seasonal changes affecting clothing, new trends, deployment in different geographical locations).
- Concept Drift: The underlying relationship between input and output changes (e.g., traffic light designs varying by region).
- Consequences: These drifts can degrade model performance, even if the model performed well on the original test set.
- AI Engineer's Role: The job is not just to perform well on a test set but to build a system that works in the real world, which includes adapting to changing conditions.
- Examples of Drift:
- Face Recognition: Seasonal changes (scarves, sunglasses), different cultural attire.
- Web Search: New events, viral content, or popular culture phenomena.
- Factory Inspection: Changes in materials or new machinery introducing new defect types.
- Self-Driving Cars: Differences in traffic light designs between California and Texas.
- Robustness: Simple, non-ML models (like the pixel-change threshold) can be more robust to drift due to their simplicity. Complex ML models, especially those prone to overfitting, may require more frequent updates.
- LLM Data Feedback: For LLMs, obtaining user permission to stream back anonymized data (while respecting privacy) can help monitor performance and identify issues.
- Proactive Monitoring:
- Brainstorming Potential Failures: Teams should brainstorm all possible ways the system could fail or data could change.
- Rich Dashboards: Create comprehensive dashboards tracking various metrics (latency, acceptance/rejection rates, re-authentication frequency) over time.
- Pruning Metrics: Initially plot many metrics and then prune to the most informative ones as patterns emerge.
- Alarms: Set up upper and lower bounds for key metrics to trigger alerts when performance deviates significantly.
Conclusion
The full cycle of a deep learning project is a dynamic, iterative, and empirical journey. Success hinges on a deep understanding of data's unpredictability, a commitment to rapid iteration, strategic data collection, robust deployment optimization, and continuous monitoring and maintenance to adapt to the ever-changing real world. The emphasis on speed, error analysis, and proactive problem-solving is paramount for building effective and reliable AI systems.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development