Stanford CS231N Deep Learning for Computer Vision | Spring 2025 | Lecture 15: 3D Vision

Unknown AuthorAbout 5 min readSep 3, 2025Watch original
THE SUMMARYAI-generated

3D Vision Lecture Summary

Key Concepts:

  • 3D Representations: Explicit (Point Clouds, Polygon Meshes, Parametric Representations) vs. Implicit (Level Sets, Algebraic Surfaces, Distance Functions)
  • Point Clouds: Collection of 3D points, often raw output from 3D sensors.
  • Polygon Meshes: Collection of vertices, edges, and faces, widely used in graphics engines.
  • Parametric Representations: Representing shapes as functions with underlying degrees of freedom (e.g., Bézier curves, Bézier surfaces).
  • Implicit Representations: Representing 3D objects as functions where points on the surface satisfy a certain relationship (f(x, y, z) = 0).
  • Voxels: 3D matrix representing space, where each element indicates whether a point is inside or outside an object.
  • ShapeNet: Large-scale 3D model dataset.
  • Neural Rendering: Techniques to render 3D scenes using neural networks, often differentiable.
  • Deep Implicit Functions: Using deep neural networks to represent implicit functions for 3D geometry.
  • NeRF (Neural Radiance Fields): Using neural networks to represent radiance and density of a scene, enabling photorealistic rendering from 2D images.
  • Gaussian Splats: Representing scenes with 3D Gaussian blobs, enabling efficient rendering.
  • Symmetric Functions: Functions invariant to the order of their inputs (e.g., max, sum).
  • Chamfer Distance & Earth Mover Distance: Metrics for comparing point clouds.

3D Representations: Explicit vs. Implicit

The lecture begins by addressing the fundamental question of how to represent 3D objects, contrasting it with the straightforward pixel-based representation in 2D. 3D objects are diverse in scale and complexity, requiring different representations to capture fine details.

  • Explicit Representations: Directly represent parts of the object.
    • Point Clouds: Simplest representation, a collection of 3D points (x, y, z coordinates). Can include surface normals (surfels). Often results from 3D scanners and can be noisy. Flexible but lacks topological information.
    • Polygon Meshes: Collection of vertices, edges, and faces. Most widely used in graphics engines and computer games. Supports operations like subdivision and simplification. Can be very complex (millions or trillions of triangles).
    • Parametric Representations: Represent shapes as functions. Useful for objects with straight lines or regular shapes. Examples include Bézier curves and surfaces. Capture underlying lower dimensionalities of surfaces.
  • Implicit Representations: Represent 3D objects as functions.
    • Level Sets, Algebraic Surfaces, Distance Functions: Represent objects as functions where points on the surface satisfy a certain relationship (f(x, y, z) = 0).
    • Explicit vs. Implicit Trade-offs: Explicit representations are easy to sample points from but hard to test if a point is inside or outside the object. Implicit representations are the opposite.

Choosing a Representation

The choice of representation depends on factors like storage efficiency, support for creating new shapes, editing capabilities (simplification, smoothing, filtering), rendering, animation, and integration with deep learning methods.

Point Clouds in Detail

  • Represented as a 3 x N matrix, where N is the number of points.
  • Surface normals provide orientation information.
  • Benefit: Raw format from 3D sensors.
  • Limitations: No connectivity, doesn't directly support smoothing or subdivision, no topological information.

Polygon Meshes in Detail

  • Represents objects as a collection of points and their connections (faces).
  • Widely used in graphics engines and computer games.
  • Supports subdivision and simplification.
  • Challenge: Integrating irregular mesh data with neural networks.

Parametric Representations in Detail

  • Represent shapes as functions.
  • Useful for objects with straight lines or regular shapes.
  • Examples: Bézier curves and surfaces.
  • Capture underlying lower dimensionalities of surfaces.

Implicit Representations in Detail

  • Represent objects as functions where points on the surface satisfy a certain relationship (f(x, y, z) = 0).
  • Easy to test if a point is inside or outside the object.
  • Hard to sample points on the surface.
  • Easy to compose using logical operations (unions, intersections, differences).
  • Can be used to smoothly blend shapes using distance functions.

Level Set Methods and Voxels

  • Level Set Methods: Pre-querying a 3D space and storing distance values in a matrix. Boundaries are where adjacent values change sign.
  • Voxels: Binarized representation where each element indicates whether a point is inside (1) or outside (0) an object. Analogous to pixels in 2D images.

Data Sets for 3D Vision

  • Princeton Shape Benchmark: Early dataset with 1,800 models in 180 categories.
  • ShapeNet: Large-scale dataset with 3 million models (ShapeNet Core: 50,000 models in 55 categories).
  • Objaverse & Objaverse Extra Large: Datasets with 1 million and 10 million models, respectively.
  • Real-world datasets from 3D scans (e.g., Redwood dataset, Meta/Oxford dataset).
  • PartNet: Dataset with annotated object parts and their mobility.
  • ScanNet: Dataset with 3D scans of indoor scenes.

Tasks in 3D Vision

  • Generative Modeling: Generating 3D shapes and scenes, conditioned on language or images.
  • Discriminative Models: Classifying 3D shapes (e.g., chair vs. table).
  • Joint Modeling of 2D and 3D Data: Leveraging priors from 2D foundation models for 3D reconstruction.
  • Multimodal Fusion: Combining visual data with text, tactile, LiDAR, or depth data.

Deep Learning on 3D Data: From Voxels to NeRF

  • Multi-View Approach: Rendering 3D objects from different views and applying 2D convolutional neural networks.
  • Volumetric Convolutional Neural Networks: Applying 3D convolutional filters to voxel data.
  • PointNet: Deep network that directly works with 3D point clouds, using symmetric functions (max, sum) to achieve permutation invariance.
  • AtlasNet: Learning transformations from a low-dimensional space to a higher-dimensional space, representing object surfaces.
  • Deep Implicit Functions: Using deep neural networks to represent implicit functions for 3D geometry.
  • NeRF (Neural Radiance Fields): Using neural networks to represent radiance and density of a scene, enabling photorealistic rendering from 2D images.
    • Queries the network for x, y, z coordinates and viewing directions.
    • Outputs density and color values.
    • Trained on 2D images using differentiable volume rendering.
  • Gaussian Splats: Representing scenes with 3D Gaussian blobs, enabling efficient rendering.

Object Geometry Structures

  • Representing objects as a collection of simple geometric parts.
  • Modeling relationships between parts.
  • Hierarchical graph representations.
  • Using programs to synthesize object shapes.
  • Using large language models to output programs for generating 3D shapes.

Conclusion

The lecture provides a comprehensive overview of 3D vision, covering various 3D representations, their trade-offs, and the evolution of deep learning methods for 3D data. It highlights the shift from explicit representations like voxels and point clouds to implicit representations and neural radiance fields, emphasizing the importance of data sets and the integration of 2D and 3D information. The lecture concludes by discussing emerging trends in using large language models for 3D shape generation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.