GPT in 600 Lines of Vanilla JavaScript: Summary
Key Concepts:
- Large Language Models (LLMs)
- GPT2 (specifically GPT2-small)
- Tokenization (including Byte Pair Encoding - BPE)
- Embeddings (Token & Position)
- Attention Mechanism (Multi-Head Attention)
- Multi-Layer Perceptron (MLP)
- Backpropagation (Gradient Descent)
- Language Head (Logits & Softmax)
- Reinforcement Learning from Human Feedback (RLHF)
- Mixture of Experts (MoE)
1. Introduction and Motivation:
- The talk aims to demystify LLMs and demonstrate that understanding their core mechanics doesn't require advanced linear algebra and calculus.
- The speaker built an Excel-based GPT2 implementation last year to teach the concepts, even to non-engineers like a CFO.
- This year, a vanilla JavaScript implementation of GPT2-small is used, accessible at
spreadsheetsarealluneed.ai/gpt2. - The goal is to provide web developers and full-stack JavaScript developers with a way to deeply understand how LLMs work under the hood without needing Python.
- The focus is on the why behind the code, building intuition with analogies and examples.
- Prerequisites:
- Motivation and curiosity
- Any STEM background
- Programming experience (JavaScript preferred)
- Basic awareness of linear algebra (matrix multiplication)
- Resources:
- JavaScript implementation of GPT2
- Chrome browser
- Discord server for questions
2. Setting Up the JavaScript Implementation:
- Go to
spreadsheetsarealluneed.ai/gpt2. - Download the GPT2-small CSV files from GitHub (model parameters). This is a large zip file containing the model weights.
- Drag and drop all the CSV files into the designated area on the webpage.
- The parameters are loaded into the browser's
index.dbdatabase (approximately 1.5 GB). - This allows the entire model to run locally in the browser, even without an internet connection.
- The interface resembles a Python notebook with cells that can be executed.
- Cells contain either JavaScript code or formulas that display results in a table format.
- Example:
promptToTokensis a DOM ID, and brackets are syntactic sugar to grab the DOM table with that ID. - Debugging: Full console debugging and stepping through the code using the browser's developer tools (using debugger statements).
3. Background on Large Language Models (LLMs):
- LLMs predict the next word/token in a passage of text.
- They are auto-regressive: the output is appended to the input and fed back into the model.
- Example: Given "Mike is quick he moves," the model predicts "quickly." This is then appended to generate the next word.
- LLMs transform the problem of predicting words into a math problem.
- The process involves:
- Mapping words to numbers (tokenization & embeddings).
- Number crunching (matrix operations).
- Mapping numbers back to words.
- The closeness of the number to the words gets weighted with probability distributions and then a random number generator is used to pick according to that distribution.
- A random number generator is used to sample from the probability distribution, introducing randomness in the output.
- Taking only the closest word to the prediction is called "greedy" or "temperature zero."
4. GPT2-small as a Foundation:
- GPT2 (2019) was considered too dangerous to release initially.
- It's the foundation of most modern LLMs, including chat GPT, Llama, Bard/Gemini, GPT4, and Claude.
- Luther AI research indicates that the fundamental recipe for building LLMs hasn't changed significantly since the transformer and OpenAI's GPT1 & GPT2.
- Understanding GPT2 provides 80% of the knowledge needed to understand state-of-the-art models.
5. Tokenization: Breaking Text into Subword Units:
- Tokenization splits input text into subword units called "tokens".
- Words can be composed of multiple tokens.
- Example: "reinjury" might be tokenized as "space re i n" + "jury".
- Word-based tokenization (one word = one token):
- Problem 1: Can't handle unknown or misspelled words (increases vocabulary size).
- Problem 2: Increases the vocabulary size, requiring more parameters and increasing model size.
- Character-based tokenization (one character = one token):
- Problem 1: Increases sequence length, requiring more memory and compute.
- Problem 2: Low semantic correlation; characters carry little meaning, requiring the model to work harder.
- Subword Tokenization (Byte Pair Encoding - BPE):
- A compromise between word-based and character-based tokenization.
- BPE has two phases:
- Learning Phase: Learns a vocabulary of tokens from a large corpus of text.
- Tokenization Phase: Translates input words into tokens from the learned vocabulary.
- BPE Learning Algorithm:
- Start with a vocabulary of individual characters.
- Count adjacent characters occurrences in the corpus.
- Merge the most frequent pair into a new token and add it to the vocabulary.
- Repeat steps 2 and 3 until a desired vocabulary size is reached.
- The goal is to find the most efficient way to represent the text.
- GPT2 tokenization uses a parsing process for breaking input into words based on OpenAI's released regex code. It then iterates to find the most efficient tokenization based on "vocab BPE" file.
- Tokenization as a Necessary Evil:
- Tokenization issues can cause problems, such as difficulty counting letters in words like "strawberry" due to different tokenizations based on casing, spaces, etc.
- These different token patterns make it harder for the model to understand.
- Word and character based tokenization aren't strict rules, and can sometimes be used.
- Tokenization can be applied to other things than just text, such as vision transformers that use patches of images as tokens or Whimo using trajectories in space.
6. Embeddings: Mapping Tokens to Semantic Meaning:
- Embeddings map tokens to a list of numbers (in GPT2-small, 768 numbers), capturing their semantic meaning.
- Analogy: Street addresses are like token IDs (identifiers), while embeddings are like square footage, bedrooms, and bathrooms (describe what's inside).
- Embedding values build a map where similar words are grouped together in a high-dimensional space.
- Word arithmetic (vector math) can be performed on embeddings to find relationships between words.
- Example:
king - man + woman = queen - Word2Vec: a popular word embedding technique that learned relationships like "France is to Paris as Italy is to Rome".
- Real-world embeddings have many more dimensions (GPT2 has 768), but the meaning of each column is unknown and uninterpretable.
- Despite the lack of interpretability, embeddings are useful for finding similarity.
- Training Embeddings:
- Embeddings are learned during training using backpropagation.
- The model is given a passage of text with the last token removed.
- It predicts the next word, and backpropagation adjusts the parameters (including embeddings) to get closer to the correct answer.
- Models learn from unsupervised text, discovering grammar, names, capitals, etc.
- Learning from Word Statistics (Distributional Hypothesis):
- Words that co-occur frequently have similar meanings.
- "You shall know a word by the company it keeps."
- Co-occurrence matrix: Count how often words co-occur within a window size.
- Embedding: A compressed co-occurrence matrix.
- Measuring Similarity (Cosine Similarity):
- Cosine similarity (angle between vectors) is used to measure the similarity of embeddings, not Euclidean distance.
- Cosine similarity captures relative co-occurrence better than Euclidean distance.
- The embeddings were learned during training so OpenAI gives this
model_WTEmatrix. In which is is vocabulary size tall and each row is simply the embedding for that token. - Embeddings can be applied to other things than just text, such as comparing words against images in CLIP.
7. Position Embeddings: Encoding Word Order:
- Word order matters in language (e.g., "The dog chases the cat" vs. "The cat chases the dog").
- Position embeddings add a sense of position to the token embeddings.
- Each position in the prompt has a slightly different offset in the embedding space.
- The original "Attention is All You Need" paper used sine and cosine formulas to generate position embeddings.
- GPT2 learns the position embeddings during training.
- GPT2 Implementation:
- A
model_WPmatrix (1024 rows for max context length, 768 columns for embedding dimension) contains the position offsets. - These offsets are added element-wise to the token embeddings to create position-aware embeddings.
- A
- Modern LLMs do not use these position embeddings, and now use ROPE (Rotary Position Embeddings).
8. Attention Mechanism: Letting Tokens Communicate:
- Attention allows tokens/words to "talk" to each other and convey their meaning.
- Tokens push and pull each other based on their relevance (distance) and importance (mass - value).
- Queries, keys, and values are used, but the "file cabinet" analogy doesn't fully capture the interaction.
- Attention shifts the position of a token in the embedding space to capture its specific meaning in context.
- Example: "moves" in the context of "quick" is shifted towards the "moves fast" area.
- Inside the code for the blocks, the steps are labelled with a step number for doing it in steps.
- Attention can be found in steps 4 through 9.
- The attention matrix (step 7) shows how much attention each word pays to every other word.
- In decoder-based transformers like GPT2, tokens can only look at previous tokens (upper triangle is zeroed out).
- Each row in the attention matrix sums up to one, representing the percentage of attention a word pays to others.
9. Multi-Layer Perceptron (MLP): Non-linear Transformation:
- MLP is another major operation inside each block/layer.
- It's a computational model inspired by the human brain (neurons and connections).
- Neuron Model:
- Inputs (x1...xn) are multiplied by weights (w1...wn), summed, and added to a bias term.
- The result is passed through an activation function (e.g., ReLU).
- ReLU: If the result is negative, output is 0; otherwise, the output is passed through.
- MLP Architecture:
- Neurons are arranged in columns (layers).
- Each node in a layer is fully connected to every node in the preceding layer.
- Hidden layers are between the input and output layers.
- Neural networks can be efficiently written as matrix multiplication.
- Universal Approximation Theorem: MLPs can approximate any function with enough neurons.
- MLPs are used to predict the embedding of the next token given the embedding of the current token.
- Backpropagation (Gradient Descent):
- An algorithm that adjusts the weights and biases of the MLP to improve accuracy.
- Analogy: A lost hiker on a foggy mountain uses the slope of the ground to find the way down.
- Calculus provides the slope (gradient) to move the parameters in the direction of lower error.
- GPT2's MLP has one hidden layer (4x the embedding dimension).
- Three steps: Applying weights and biases, applying GLU activation, and projecting back to the embedding dimension.
- Backpropagation is used for weights, biases, token embeddings, position embeddings, attention parameters (queries, keys, values), and layer normalization.
- Optimizing a model via back propagation is like a chef imitating a dish.
10. Iteration and Refining the Prediction:
- The model iteratively refines its prediction for the next token.
- GPT2-small iterates 12 times (more iterations in modern models).
- Each block performs identical operations but with different weights and parameters.
11. Language Head: From Embedding to Token:
- The language head converts the predicted token embedding back into a token.
- It takes the output of the last MLP in the last block.
- Steps:
- Apply layer normalization.
- Multiply the resulting embedding by the
model_wtematrix (vocabulary of embeddings). - The result is a column of "logits" or token scores representing the similarity of the predicted embedding to each token in the vocabulary.
- Apply softmax normalization to convert logits into a probability distribution (sums to one).
- Select the token with the highest probability.
- The Javascript implementation uses "greedy sampling" (temperature zero), always picking the highest probability token for consistency with other GPT2 implementations.
- Alternative sampling methods: Top-K sampling, nucleus/top-P sampling.
12. Chat GPT vs. GPT2: Training Differences:
- GPT2 is a base model trained to predict the next word from internet text.
- Chat GPT and InstructGPT are trained to be helpful assistants.
- Four-Step Pipeline:
- Pre-training: Imitate text on the internet (GPT2).
- Supervised Fine-tuning: Train on examples of ideal assistant responses (prompt and response pairs). Stanford Alpaca dataset is an example.
- Reward Model Training: Collect human preferences by comparing pairs of chosen and rejected responses to derive a scoring model.
- Reinforcement Learning from Human Feedback (RLHF): Use the scoring model to train the LLM to reinforce the nuanced preferences.
- Reinforcement Learning Analogy:
- Learning to play a game by exploring different paths and maximizing a score.
- Generating text is like walking a path through language, where some paths are more desirable than others.
13. Summary and Synthesis:
- Tokenization: Efficiently represents text through compression.
- Embeddings: Act as a recommendation system for words.
- Attention: Lets words communicate their context and refine recommendations.
- Neural Network: Learns to pull latent predictions out of the embeddings.
- Overall: The model learns to recommend the next word based on context, refined predictions, and a dictionary of embeddings.
- The goal of the workshop is to make mastery of LLMs feel within grasp.
14. Mixture of Experts (MoE)
- Mixture of Experts is a design implemented in the perceptron.
- It helps the model use more parameters without increasing the amount of compute.
- The perceptron is broken down in pieces to increase efficiency.
- It is implemented by selecting to only use a subset of the perceptron's thinking depending on which token comes in.
AI summaries can miss context or contain errors. Check important details against the original video.





