If you've read the "Attention Is All You Need" architecture diagrams and still don't feel like you get transformers, the Transformer Explainer from Georgia Tech's Polo Club is worth an hour of your time. It's an interactive visualization that runs a real GPT-2 model directly in your browser, so you can type a prompt and watch the model process it end to end.
The tool walks through the full pipeline: tokenization, token and positional embeddings, the self-attention mechanism (query, key, and value vectors), the multi-layer processing, and the final projection into a probability distribution over the next token. You can adjust the temperature setting and see how it reshapes those output probabilities, which makes an abstract sampling parameter tangible.
What makes this more than a diagram is that the numbers are live. Because an actual model is doing inference locally, the attention weights and activations you see correspond to the text you entered, not a canned example. That closes the gap between the math you read about and what a model does with your specific input.
For builders, this is a practical onboarding and teaching resource. Use it to explain to a non-ML stakeholder why context length and attention matter, to build intuition before you dig into a framework like PyTorch or Hugging Face Transformers, or to sanity-check your own mental model of how temperature and next-token prediction interact. The concepts scale up to modern GPT-class models even though the visualized model is the smaller GPT-2.
It pairs well with the underlying research and the group's other explainers (they've built similar interactive tools for CNNs and diffusion models). Start here to get the shape of the system, then read the original paper when the moving parts already make sense.