But what is a neural network? | Deep learning chapter 1
By 3Blue1Brown
Key Concepts
- Neural Networks: A computational model inspired by the structure and function of biological neural networks.
- Neurons: Basic units of a neural network that hold a number (activation) between 0 and 1.
- Layers: Organized groups of neurons (input, hidden, output).
- Activations: The numerical value held by a neuron, representing its state.
- Weights: Numerical values assigned to connections between neurons, indicating the strength of the connection.
- Bias: A numerical value added to the weighted sum of inputs to a neuron, influencing its activation threshold.
- Sigmoid Function: A mathematical function that squishes real numbers into the range between 0 and 1, used to determine neuron activation.
- ReLU (Rectified Linear Unit): An activation function that outputs the input if it is positive, otherwise outputs zero.
- Training: The process of adjusting the weights and biases of a neural network to improve its performance.
- Matrix Multiplication: A linear algebra operation used to efficiently compute the weighted sum of activations in a neural network.
Neural Network Structure and Function
Introduction to Neural Networks
The video introduces neural networks as a means of enabling computers to recognize patterns, specifically handwritten digits. The core idea is to create a system that can take a grid of pixels as input and output a number representing the digit it identifies. The video aims to demystify neural networks by explaining their structure and how they "learn."
The Structure of a Simple Neural Network
The video focuses on a simple, "plain vanilla" neural network. This network consists of:
- Input Layer: 784 neurons, corresponding to the 28x28 pixels of an input image. Each neuron's activation represents the grayscale value of the corresponding pixel (0 for black, 1 for white).
- Hidden Layers: Two layers, each with 16 neurons. The number of layers and neurons per layer is somewhat arbitrary and can be experimented with.
- Output Layer: 10 neurons, each representing a digit (0-9). The activation of each neuron represents the network's confidence that the input image corresponds to that digit.
How Activations Propagate Through the Network
Activations in one layer determine the activations in the next layer. The network is "trained" to recognize digits, meaning that when an image is fed into the input layer, the pattern of activations propagates through the hidden layers to produce a specific pattern in the output layer. The neuron with the highest activation in the output layer represents the network's "choice" for the digit.
The Hope for Hidden Layers: Feature Detection
The video presents a conceptual model of what the hidden layers might be doing:
- Second Layer (Edges): Neurons in the second layer might correspond to small edges.
- Third Layer (Patterns): Neurons in the third layer might correspond to combinations of edges, forming patterns like loops or lines.
- Output Layer (Digits): The output layer then learns which combinations of these patterns correspond to which digits.
This layered approach is analogous to how humans recognize digits by piecing together components. The ability to detect edges and patterns could be useful for other image recognition tasks and even tasks like speech parsing.
How One Layer Influences the Next
Weights and Biases
The key to how one layer influences the next lies in weights and biases:
- Weights: Each connection between a neuron in one layer and a neuron in the next layer has a weight associated with it. The weight determines the strength of the connection. Positive weights (represented by green pixels) indicate excitatory connections, while negative weights (represented by red pixels) indicate inhibitory connections.
- Weighted Sum: The activation of a neuron in the next layer is determined by the weighted sum of the activations of the neurons in the previous layer.
- Bias: A bias is added to the weighted sum. The bias determines the activation threshold for the neuron.
Sigmoid Function
The weighted sum plus the bias is then passed through a sigmoid function (also known as a logistic curve). The sigmoid function squishes the real number line into the range between 0 and 1, ensuring that the neuron's activation is between 0 and 1.
Number of Weights and Biases
The network has a large number of weights and biases. For example, with a hidden layer of 16 neurons, there are 784 (input neurons) * 16 (hidden neurons) = 12,544 weights, plus 16 biases. In total, the network has almost 13,000 weights and biases.
Compact Notation Using Linear Algebra
The video introduces a more compact notation for representing the connections between layers using linear algebra:
- Activations from one layer are organized into a column vector.
- Weights are organized into a matrix, where each row corresponds to the connections between one layer and a particular neuron in the next layer.
- The weighted sum of activations is computed using matrix-vector multiplication.
- Biases are organized into a vector and added to the result of the matrix-vector multiplication.
- The sigmoid function is applied to each component of the resulting vector.
This compact notation makes the code simpler and faster, as matrix multiplication is highly optimized in many libraries.
The Network as a Function
The entire network can be viewed as a function that takes 784 numbers as input and outputs 10 numbers. This function is complex, involving 13,000 parameters (weights and biases) and iterating many matrix-vector products and the sigmoid function.
Modern Activation Functions: ReLU
Lisha Li, a deep learning expert, discusses the sigmoid function and its limitations. While early networks used the sigmoid function, modern networks often use ReLU (Rectified Linear Unit). ReLU is defined as max(0, a), where 'a' is the weighted sum of inputs. ReLU is easier to train than sigmoid, especially for deep neural networks. ReLU is motivated by the biological analogy of neurons being either activated or not.
Conclusion
The video provides a detailed explanation of the structure and function of a simple neural network. It explains how activations propagate through the network, how weights and biases influence neuron activation, and how the network can be viewed as a complex function. The video also introduces the concept of ReLU as a modern alternative to the sigmoid function. The next video will cover how the network learns the appropriate weights and biases through training.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Inside Jeffrey Epstein's Network of Power
Bloomberg Originals

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth

Nvidia Crushes Earnings again — What Jensen Huang sees next for AI
CGTN America

Throwing out the first pitch for the Rockies for STEM Day!
Sick Science!