Tweet-GPT

A small language model with an interactive view of its predictions.

Role
Design & build
When
2026
Stack
  • PyTorch
  • ONNX
  • TypeScript

I built a small GPT in PyTorch to understand how a transformer turns earlier text into a prediction. This page makes part of that calculation visible: it shows the next-character probabilities and the attention weights as the model writes.

The model generates one character at a time. After the initial download, it runs entirely in your browser.

Loading Tweet-GPT…

Reading a prediction

At each step, the model assigns a score to every character in its vocabulary. Those scores become probabilities, and the generator selects a character from that distribution. The new character then becomes part of the input for the next step.

Pause generation to inspect the alternatives. The character that appears is only one of the possible next characters. The attention view shows how the model weights earlier positions while computing its representations. It helps expose the calculation, though it does not fully explain why a particular character was selected.

Training a model small enough to inspect

The architecture follows Andrej Karpathy's nanoGPT and teaching material. 1 It has about 850,000 parameters, four transformer layers, four attention heads, and a 128-character context window. Its vocabulary contains 152 characters and an end-of-tweet token.

I trained it on roughly 93,000 public tweets from TweetEval. 2 The best validation loss was about 1.60, compared with 5.03 for a uniform random prediction. The output is often awkward, but the model is small enough to run locally and inspect one step at a time.

I exported the weights to ONNX. The demo uses ONNX Runtime's WebAssembly backend, so generating text does not require a model API.

Cited works

  1. Andrej Karpathy. nanoGPT and Neural Networks: Zero to Hero.
  2. Cardiff NLP. TweetEval dataset.