How LLMs Actually Work
From tokenized text to trained parameters and next-token generation
Trace the full path from raw text to generated answer: how a tokenizer splits your prompt, how self-attention builds context-dependent representations, how next-token prediction and gradient descent turn data into parameters, and how post-training reshapes a base model into an assistant. You finish able to reason about token costs, sampling settings, context limits, and hallucination as consequences of the mechanism, not as mysteries.
See the Invisible
Interactive simulators visualise what's hidden from view.
Hands-On Labs
Step through executions tick by tick. Manipulate state.
Why, Not Just What
Understand the reasoning behind every design decision.
Quizzes & Cheatsheets
Verify your understanding and keep a quick reference handy.
Get Certified
Earn a shareable certificate to prove your deep expertise.
What's Covered
An LLM never sees your words. A tokenizer turns text into token IDs, an embedding lookup turns IDs into vectors, and a stack of transformer layers turns static embeddings into contextual representations. This pillar explains why token boundaries decide your context consumption and compute cost, and why the same word can mean different things to the model in different sentences.
You work through one causal self-attention update end to end: query, key, and value projections, scaled dot-product scores, softmax weights, causal masking, and value aggregation across multiple heads. Once you have computed it yourself, attention stops being a metaphor and becomes an operation you can reason about when context grows long or costs grow fast.
Training is next-token prediction with a score attached. You follow a training window from shifted input and target sequences through vocabulary logits, cross-entropy loss, backpropagation, and optimizer updates, then see how repeating that loop over massive data mixtures produces base-model capabilities without anyone programming them in.
A base model completes text; an assistant follows instructions. Supervised fine-tuning, reward models learned from human preferences, RLHF policy optimization, and DPO each reshape behavior in different ways. You learn what this alignment step changes, what it cannot guarantee about factual truth, and the trade-offs it introduces.
Generation is one token at a time: prefill, logits, temperature, top-k and top-p sampling, key-value caching, stop tokens. The same mechanism explains the failure modes, including lost-in-the-middle degradation, sampling nondeterminism, and hallucination as fluent unsupported content. You finish able to decide where output needs context controls or external verification.
The Curriculum
Comprehensive Lessons! Each with theory, interactive simulation, and quiz.
What an LLM Reads: Text to Tokens
From Embeddings to Contextual Representations
Self-Attention Under the Hood
Learning by Predicting the Next Token
How Loss Becomes Learning
From Base Model to Assistant
Generate One Token at a Time
Generation Limits and Reliability
This course in one line
Stop treating the model as a black box
Ready to see what's really happening?
All courses included with your subscription. Cancel anytime.