How LLMs work for Web Devs: GPT in 600 lines of Vanilla JS - Ishan Anand

AI EngineerAbout 11 min readJul 14, 2025Watch original
THE SUMMARYAI-generated

GPT in 600 Lines of Vanilla JavaScript: Summary

Key Concepts:

  • Large Language Models (LLMs)
  • GPT2 (specifically GPT2-small)
  • Tokenization (including Byte Pair Encoding - BPE)
  • Embeddings (Token & Position)
  • Attention Mechanism (Multi-Head Attention)
  • Multi-Layer Perceptron (MLP)
  • Backpropagation (Gradient Descent)
  • Language Head (Logits & Softmax)
  • Reinforcement Learning from Human Feedback (RLHF)
  • Mixture of Experts (MoE)

1. Introduction and Motivation:

  • The talk aims to demystify LLMs and demonstrate that understanding their core mechanics doesn't require advanced linear algebra and calculus.
  • The speaker built an Excel-based GPT2 implementation last year to teach the concepts, even to non-engineers like a CFO.
  • This year, a vanilla JavaScript implementation of GPT2-small is used, accessible at spreadsheetsarealluneed.ai/gpt2.
  • The goal is to provide web developers and full-stack JavaScript developers with a way to deeply understand how LLMs work under the hood without needing Python.
  • The focus is on the why behind the code, building intuition with analogies and examples.
  • Prerequisites:
    • Motivation and curiosity
    • Any STEM background
    • Programming experience (JavaScript preferred)
    • Basic awareness of linear algebra (matrix multiplication)
  • Resources:
    • JavaScript implementation of GPT2
    • Chrome browser
    • Discord server for questions

2. Setting Up the JavaScript Implementation:

  • Go to spreadsheetsarealluneed.ai/gpt2.
  • Download the GPT2-small CSV files from GitHub (model parameters). This is a large zip file containing the model weights.
  • Drag and drop all the CSV files into the designated area on the webpage.
  • The parameters are loaded into the browser's index.db database (approximately 1.5 GB).
  • This allows the entire model to run locally in the browser, even without an internet connection.
  • The interface resembles a Python notebook with cells that can be executed.
  • Cells contain either JavaScript code or formulas that display results in a table format.
  • Example: promptToTokens is a DOM ID, and brackets are syntactic sugar to grab the DOM table with that ID.
  • Debugging: Full console debugging and stepping through the code using the browser's developer tools (using debugger statements).

3. Background on Large Language Models (LLMs):

  • LLMs predict the next word/token in a passage of text.
  • They are auto-regressive: the output is appended to the input and fed back into the model.
  • Example: Given "Mike is quick he moves," the model predicts "quickly." This is then appended to generate the next word.
  • LLMs transform the problem of predicting words into a math problem.
  • The process involves:
    1. Mapping words to numbers (tokenization & embeddings).
    2. Number crunching (matrix operations).
    3. Mapping numbers back to words.
  • The closeness of the number to the words gets weighted with probability distributions and then a random number generator is used to pick according to that distribution.
  • A random number generator is used to sample from the probability distribution, introducing randomness in the output.
  • Taking only the closest word to the prediction is called "greedy" or "temperature zero."

4. GPT2-small as a Foundation:

  • GPT2 (2019) was considered too dangerous to release initially.
  • It's the foundation of most modern LLMs, including chat GPT, Llama, Bard/Gemini, GPT4, and Claude.
  • Luther AI research indicates that the fundamental recipe for building LLMs hasn't changed significantly since the transformer and OpenAI's GPT1 & GPT2.
  • Understanding GPT2 provides 80% of the knowledge needed to understand state-of-the-art models.

5. Tokenization: Breaking Text into Subword Units:

  • Tokenization splits input text into subword units called "tokens".
  • Words can be composed of multiple tokens.
  • Example: "reinjury" might be tokenized as "space re i n" + "jury".
  • Word-based tokenization (one word = one token):
    • Problem 1: Can't handle unknown or misspelled words (increases vocabulary size).
    • Problem 2: Increases the vocabulary size, requiring more parameters and increasing model size.
  • Character-based tokenization (one character = one token):
    • Problem 1: Increases sequence length, requiring more memory and compute.
    • Problem 2: Low semantic correlation; characters carry little meaning, requiring the model to work harder.
  • Subword Tokenization (Byte Pair Encoding - BPE):
    • A compromise between word-based and character-based tokenization.
    • BPE has two phases:
      1. Learning Phase: Learns a vocabulary of tokens from a large corpus of text.
      2. Tokenization Phase: Translates input words into tokens from the learned vocabulary.
    • BPE Learning Algorithm:
      1. Start with a vocabulary of individual characters.
      2. Count adjacent characters occurrences in the corpus.
      3. Merge the most frequent pair into a new token and add it to the vocabulary.
      4. Repeat steps 2 and 3 until a desired vocabulary size is reached.
    • The goal is to find the most efficient way to represent the text.
  • GPT2 tokenization uses a parsing process for breaking input into words based on OpenAI's released regex code. It then iterates to find the most efficient tokenization based on "vocab BPE" file.
  • Tokenization as a Necessary Evil:
    • Tokenization issues can cause problems, such as difficulty counting letters in words like "strawberry" due to different tokenizations based on casing, spaces, etc.
    • These different token patterns make it harder for the model to understand.
    • Word and character based tokenization aren't strict rules, and can sometimes be used.
    • Tokenization can be applied to other things than just text, such as vision transformers that use patches of images as tokens or Whimo using trajectories in space.

6. Embeddings: Mapping Tokens to Semantic Meaning:

  • Embeddings map tokens to a list of numbers (in GPT2-small, 768 numbers), capturing their semantic meaning.
  • Analogy: Street addresses are like token IDs (identifiers), while embeddings are like square footage, bedrooms, and bathrooms (describe what's inside).
  • Embedding values build a map where similar words are grouped together in a high-dimensional space.
  • Word arithmetic (vector math) can be performed on embeddings to find relationships between words.
  • Example: king - man + woman = queen
  • Word2Vec: a popular word embedding technique that learned relationships like "France is to Paris as Italy is to Rome".
  • Real-world embeddings have many more dimensions (GPT2 has 768), but the meaning of each column is unknown and uninterpretable.
  • Despite the lack of interpretability, embeddings are useful for finding similarity.
  • Training Embeddings:
    • Embeddings are learned during training using backpropagation.
    • The model is given a passage of text with the last token removed.
    • It predicts the next word, and backpropagation adjusts the parameters (including embeddings) to get closer to the correct answer.
    • Models learn from unsupervised text, discovering grammar, names, capitals, etc.
  • Learning from Word Statistics (Distributional Hypothesis):
    • Words that co-occur frequently have similar meanings.
    • "You shall know a word by the company it keeps."
    • Co-occurrence matrix: Count how often words co-occur within a window size.
    • Embedding: A compressed co-occurrence matrix.
  • Measuring Similarity (Cosine Similarity):
    • Cosine similarity (angle between vectors) is used to measure the similarity of embeddings, not Euclidean distance.
    • Cosine similarity captures relative co-occurrence better than Euclidean distance.
  • The embeddings were learned during training so OpenAI gives this model_WTE matrix. In which is is vocabulary size tall and each row is simply the embedding for that token.
  • Embeddings can be applied to other things than just text, such as comparing words against images in CLIP.

7. Position Embeddings: Encoding Word Order:

  • Word order matters in language (e.g., "The dog chases the cat" vs. "The cat chases the dog").
  • Position embeddings add a sense of position to the token embeddings.
  • Each position in the prompt has a slightly different offset in the embedding space.
  • The original "Attention is All You Need" paper used sine and cosine formulas to generate position embeddings.
  • GPT2 learns the position embeddings during training.
  • GPT2 Implementation:
    • A model_WP matrix (1024 rows for max context length, 768 columns for embedding dimension) contains the position offsets.
    • These offsets are added element-wise to the token embeddings to create position-aware embeddings.
  • Modern LLMs do not use these position embeddings, and now use ROPE (Rotary Position Embeddings).

8. Attention Mechanism: Letting Tokens Communicate:

  • Attention allows tokens/words to "talk" to each other and convey their meaning.
  • Tokens push and pull each other based on their relevance (distance) and importance (mass - value).
  • Queries, keys, and values are used, but the "file cabinet" analogy doesn't fully capture the interaction.
  • Attention shifts the position of a token in the embedding space to capture its specific meaning in context.
  • Example: "moves" in the context of "quick" is shifted towards the "moves fast" area.
  • Inside the code for the blocks, the steps are labelled with a step number for doing it in steps.
  • Attention can be found in steps 4 through 9.
  • The attention matrix (step 7) shows how much attention each word pays to every other word.
  • In decoder-based transformers like GPT2, tokens can only look at previous tokens (upper triangle is zeroed out).
  • Each row in the attention matrix sums up to one, representing the percentage of attention a word pays to others.

9. Multi-Layer Perceptron (MLP): Non-linear Transformation:

  • MLP is another major operation inside each block/layer.
  • It's a computational model inspired by the human brain (neurons and connections).
  • Neuron Model:
    • Inputs (x1...xn) are multiplied by weights (w1...wn), summed, and added to a bias term.
    • The result is passed through an activation function (e.g., ReLU).
    • ReLU: If the result is negative, output is 0; otherwise, the output is passed through.
  • MLP Architecture:
    • Neurons are arranged in columns (layers).
    • Each node in a layer is fully connected to every node in the preceding layer.
    • Hidden layers are between the input and output layers.
  • Neural networks can be efficiently written as matrix multiplication.
  • Universal Approximation Theorem: MLPs can approximate any function with enough neurons.
  • MLPs are used to predict the embedding of the next token given the embedding of the current token.
  • Backpropagation (Gradient Descent):
    • An algorithm that adjusts the weights and biases of the MLP to improve accuracy.
    • Analogy: A lost hiker on a foggy mountain uses the slope of the ground to find the way down.
    • Calculus provides the slope (gradient) to move the parameters in the direction of lower error.
  • GPT2's MLP has one hidden layer (4x the embedding dimension).
  • Three steps: Applying weights and biases, applying GLU activation, and projecting back to the embedding dimension.
  • Backpropagation is used for weights, biases, token embeddings, position embeddings, attention parameters (queries, keys, values), and layer normalization.
  • Optimizing a model via back propagation is like a chef imitating a dish.

10. Iteration and Refining the Prediction:

  • The model iteratively refines its prediction for the next token.
  • GPT2-small iterates 12 times (more iterations in modern models).
  • Each block performs identical operations but with different weights and parameters.

11. Language Head: From Embedding to Token:

  • The language head converts the predicted token embedding back into a token.
  • It takes the output of the last MLP in the last block.
  • Steps:
    1. Apply layer normalization.
    2. Multiply the resulting embedding by the model_wte matrix (vocabulary of embeddings).
    3. The result is a column of "logits" or token scores representing the similarity of the predicted embedding to each token in the vocabulary.
    4. Apply softmax normalization to convert logits into a probability distribution (sums to one).
    5. Select the token with the highest probability.
  • The Javascript implementation uses "greedy sampling" (temperature zero), always picking the highest probability token for consistency with other GPT2 implementations.
  • Alternative sampling methods: Top-K sampling, nucleus/top-P sampling.

12. Chat GPT vs. GPT2: Training Differences:

  • GPT2 is a base model trained to predict the next word from internet text.
  • Chat GPT and InstructGPT are trained to be helpful assistants.
  • Four-Step Pipeline:
    1. Pre-training: Imitate text on the internet (GPT2).
    2. Supervised Fine-tuning: Train on examples of ideal assistant responses (prompt and response pairs). Stanford Alpaca dataset is an example.
    3. Reward Model Training: Collect human preferences by comparing pairs of chosen and rejected responses to derive a scoring model.
    4. Reinforcement Learning from Human Feedback (RLHF): Use the scoring model to train the LLM to reinforce the nuanced preferences.
  • Reinforcement Learning Analogy:
    • Learning to play a game by exploring different paths and maximizing a score.
    • Generating text is like walking a path through language, where some paths are more desirable than others.

13. Summary and Synthesis:

  • Tokenization: Efficiently represents text through compression.
  • Embeddings: Act as a recommendation system for words.
  • Attention: Lets words communicate their context and refine recommendations.
  • Neural Network: Learns to pull latent predictions out of the embeddings.
  • Overall: The model learns to recommend the next word based on context, refined predictions, and a dictionary of embeddings.
  • The goal of the workshop is to make mastery of LLMs feel within grasp.

14. Mixture of Experts (MoE)

  • Mixture of Experts is a design implemented in the perceptron.
  • It helps the model use more parameters without increasing the amount of compute.
  • The perceptron is broken down in pieces to increase efficiency.
  • It is implemented by selecting to only use a subset of the perceptron's thinking depending on which token comes in.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.