Latent Fovea Open the visualizer

Is this really a model's latent space?

Yes. But only some of what you see is the raw data. The rest is either derived from it by a standard technique, or drawn to make it readable. This page separates the three.

The one-paragraph version

Every number on screen comes from the published weights of GPT-2, DistilGPT2 or MDLM, running for real in your browser. The point cloud is the model's own vocabulary: 50,257 tokens, each a vector of 768 numbers, its first latent space. You can't see 768 dimensions, so the cloud is a 3-D map of them, and like any map it distorts. When you type a query the model genuinely runs; every intermediate result is recorded and replayed slowly. The playback is slowed down, not the model.

Where the numbers come from

Nothing is simulated

The model runs once per token at full speed. The visualizer records everything it computed, then plays it back layer by layer at whatever speed you choose.

01 · Weights
The published modelOfficial releases from OpenAI, Hugging Face and Cornell, stored at half precision.
02 · Compute
A real forward passThe model's own maths, rebuilt from scratch and running in your browser.
03 · Replay
Recorded, then slowedAttention, neurons and predictions at every layer, replayed in the order they happened.

Checked against the standard reference implementations on the same weights. The built-in GPT-2 matches to 4 decimal places, and MDLM to about 0.00001. Models loaded from Hugging Face give the same top 10 next tokens, with probabilities within 0.0013.

Three kinds of thing on screen

Real, derived, illustrated

Real

The model's own numbers

Read straight out of the computation. No interpretation added.

  • The text it writes
  • Attention of every head, every layer
  • All 36,864 MLP neuron activations
  • How hard each head writes
  • A token's 768-number vector (the barcode)
  • Exact nearest neighbours, measured in 768-d
Derived

Real data through a standard lens

Established techniques applied to the real numbers. Useful, but they add assumptions.

  • The 3-D positions (UMAP)
  • The logit lens: "what it would say at layer N"
  • The local re-layout inside the fovea (PCA)
Illustrated

A picture of real data

Drawings that make the numbers readable. They are not objects inside the model.

  • The purple "thought path"
  • Attention arcs (averaged over heads)
  • Masked MDLM slots floating in space
  • Token colours (word, number…)

The most important caveat

The cloud is a map. The neighbour lines are the territory.

The real space has 768 dimensions. UMAP squeezes it into 3 so you can fly through it. Nearby points are usually close in the real space too, but big distances, directions and the overall shape mean very little. Run it again with a different random seed and the cloud looks different. The neighbour lines are measured in the full 768 dimensions, which is why they sometimes jump across the view: that jump is the map's distortion made visible.

token A token B
The map (3-D), schematic: two tokens that are neighbours can land far apart. The dashed line is the real neighbour link reaching across.
token A token B
The territory (768-d): the same two tokens, measured in the model's real space, sit next to each other. Look at a point and the fovea shows you this geometry.

Piece by piece

What each part of the screen is

Real

Attention heads (12 × 12 grid)

The actual attention weights each head assigns to each earlier token, for the token you are looking at. For MDLM, attention runs both ways, so a slot can attend to words on its right.

Real

MLP neurons

The actual activation of every neuron in every layer. MDLM's replay stores them at 8-bit precision to save memory, which is enough for colour.

Real

Barcode and neighbour lines

The token's actual 768 coordinates, and the tokens closest to it in the full space. The most faithful geometry on the page.

Derived

Logit lens column

The model's working state halfway through, pushed through its final output layer to ask "what would you say if you stopped here?". A standard interpretability technique, not something the model does itself. Trustworthy near the top, rough near the bottom.

Derived

3-D positions

UMAP of the vocabulary, computed once. Local neighbourhoods are roughly right; global layout is arbitrary.

Illustrated

The purple thought path

At each layer, the average map position of the model's top 10 guesses, weighted by how likely each is. A picture of what the model is leaning toward. It is not the model's internal state moving through space. That state is a 768-number vector at every layer; it is computed but not drawn as a point, because it doesn't live on the vocabulary map.

Illustrated

Attention arcs

Real weights, but averaged over the 12 heads of the current layer unless you pin one, and anything under 2% is hidden.

Illustrated

MDLM masked slots

A masked slot has no real position; it is drawn at the average of its guesses until it locks onto a word.

Illustrated

Colours

Word, subword, number, punctuation: a rule based on how the token is spelled. The model knows nothing about these categories.

Where people get confused

Two things to say clearly

01 · "FOVEATED" MEANS YOUR EYES

Foveation is how the viewer saves bandwidth, not the model's attention

Like a VR headset, the visualizer loads full detail only where you are looking. The model's attention is a separate, real thing, shown in the heads grid and the arcs. They share a metaphor; keep them apart.

02 · "LATENT SPACE" IS PLURAL

The cloud is one latent space. The model has more.

The cloud is the token embedding space: a fixed table and the model's entry point. In GPT-2 the same table also picks the output word, which is why plotting predictions on it makes sense. MDLM uses separate input and output tables, so placing its guesses on the input map is an approximation. The hidden states inside the model are latent spaces too, arguably the more interesting ones. They are fully computed and shown indirectly, through their strength, the logit lens and attention, but not as a cloud of their own.

If someone asks "so is it real?"

The one-line answer

The model and every number it produces are real. The cloud is an honest but distorted map of the model's vocabulary space. The glowing path is a picture of what it is leaning toward, not a literal trajectory.