Vision Transformers: ViT, Swin & DEiT
After this lesson, you will be able to:
- Slice an image into 16x16 patches, embed them as 'visual words,' and feed them to a standard transformer encoder — the entire ViT architecture in one breath
- Diagnose when ViT will lose to a CNN (small data, edge inference) versus when it wins decisively (300M+ images, multimodal models, foundation backbones)
- Pick the right ViT variant for the job — vanilla ViT for scale, DEiT when data is scarce, Swin when you need hierarchical CNN-like locality, MAE for self-supervised pretraining
- Read attention maps from a trained ViT to understand what the model is looking at, and use that to debug failures or build interpretable medical/scientific systems
Before You Start
Don't worry if "treat an image like a sentence" sounds weird — once you see the patch-embedding step, the rest of ViT is the same transformer encoder you already understand from BERT.
#Treating an Image Like a Sentence
The Vision Transformer (ViT) was introduced in 2020 by Google Research with a paper titled "An Image is Worth 16x16 Words." The deliberately provocative title summarizes the architecture exactly: take an image, slice it into patches, treat each patch as a word, and run BERT.
#The ViT Architecture in Three Steps
#Step 1: Patch Embedding
d_model-dim patch embedding.#Step 2: Add a CLS Token and Positional Encoding
[CLS] token to the sequence (its final hidden state will represent the whole image — used for classification). Then add a learned 1-D positional embedding to each token, including the CLS token.#Step 3: Run the Transformer Encoder
z_0 through L transformer encoder blocks (typically 12 for ViT-Base, 24 for ViT-Large, 32 for ViT-Huge). Each block is exactly the BERT block: multi-head self-attention → add & LayerNorm → FFN → add & LayerNorm.After the final layer, take the CLS token's hidden state, pass it through a linear head, and you have your image classifier. That's the entire architecture.
#Why ViT Was Surprising
- Trained on ImageNet alone (1.2M images): ResNet-152 beats ViT-Base.
- Pretrained on ImageNet-21k (14M images): roughly tied.
- Pretrained on JFT-300M (300M images): ViT beats every CNN on every benchmark.
You have 5,000 labeled medical images of skin lesions and need to build a classifier. Which is most likely to win?
#DEiT: Distilling ViT for the Rest of Us
The original ViT paper had an inconvenient asterisk: the JFT-300M dataset is private to Google. Without it, ViT was inferior to CNNs on every public benchmark — making the architecture a research curiosity for most of the field.
DEiT (Touvron et al. 2021, Facebook AI) fixed this with three tricks:
- Strong augmentation (RandAugment, mixup, cutmix, random erasing) — every transformer needs this for vision.
- Knowledge distillation from a CNN teacher — they added a special
[DIST]token next to[CLS], trained to match a RegNet teacher's output. The CNN's inductive bias gets transferred through the distillation loss. - Tuned hyperparameters for ViT specifically (learning rate, weight decay, augmentation strength).
Result: DEiT-B trained only on ImageNet-1k matches ViT-B trained on ImageNet-21k. The distillation token is dropped at inference.
#Swin: Hierarchical ViT with Shifted Windows
O(N² · d) per layer, which becomes painful at high resolution (e.g. 1024x1024 images give 4096 patches → 16M attention scores per head per layer).Swin Transformer (Liu et al. 2021) makes attention local:
- Divide the image into windows of, say, 7x7 patches.
- Inside each window, run normal multi-head self-attention. Cost:
O(N · M² · d)whereMis window size — linear inN. - Shift the windows by
M/2between alternating layers. Patches that were in different windows in layerkend up in the same window in layerk+1. Information crosses boundaries without ever needing global attention. - Patch merging between stages downsamples the resolution and increases channels — exactly like CNN downsampling stages.
#MAE: Self-Supervised Pretraining, Cheap and Effective
Supervised pretraining (training a ViT to classify ImageNet labels) requires labels. Labels are expensive. Worse, classification-pretrained features often don't transfer well to dense tasks.
Masked Autoencoder (He et al. 2021) uses self-supervision instead:
- Randomly mask 75% of the patches of an input image.
- Run the encoder (a standard ViT) on only the visible 25% of patches.
- A small decoder receives the encoder output plus learnable mask tokens at the masked positions and tries to reconstruct the missing pixels.
- After pretraining, throw away the decoder. The encoder is now a strong general-purpose backbone.
Why MAE works:
- The 75% masking ratio is much higher than BERT's 15%. Vision is much more spatially redundant than text — you can recognize a cat from any quarter of the image. With low masking, the model just interpolates from neighbors and learns nothing useful.
- The decoder is small (≤ 8 layers) and only used during pretraining. The encoder, which is the part you keep, processes only 25% of tokens — making pretraining 3-4x cheaper than supervised pretraining at the same model size.
- The result: ViT-Huge pretrained with MAE on ImageNet (no labels) outperforms ViT-Huge trained with supervised classification, on every downstream task.
#DINOv2: The 2023 Default Backbone
Today, when someone says "use a ViT backbone for X," they almost always mean a DINOv2 or SigLIP encoder. ResNet-50 has been retired from the same role.
#Try It Yourself: Patch Embedding from Scratch
Tests · Verify that the output shape is (batch_size, num_patches + 1, d_model) and that swapping cls_token / pos_embed for zeros breaks classification accuracy.
#Decision Rubric: Which Variant?
| Situation | Pick |
|---|---|
| Standard 224x224 image classification, decent compute | ViT-Base/16 with DEiT or MAE pretraining |
| <50K labeled images | DINOv2 or DEiT pretrained, fine-tune the head only first, then full model |
| Object detection, segmentation, high-resolution input | Swin Transformer — hierarchical, scales to high res |
| You want self-supervised pretraining on your own image data | MAE — cheap, strong, label-free |
| You need a frozen image embedder for downstream models | DINOv2 — features generalize like nothing else |
| You're building a multimodal LLM | SigLIP or DINOv2 — used by Gemini, Claude 3, GPT-4V |
| Edge deployment, mobile inference | Stick with MobileNet / EfficientNet-Lite — ViT is too big |
#Key Takeaways
- ViT treats images as sequences of patches. Slice into 16x16 patches, flatten, embed, add CLS token + position, run a standard transformer encoder. The architecture is BERT applied to image patches.
- ViT needs scale to win. At small data (≤1M images, ≤10K samples for fine-tuning) CNNs still beat ViT trained from scratch. ViT only wins when pretrained on >14M images. Always start from a pretrained backbone.
- DEiT solved ViT's data hunger. Knowledge distillation from a CNN teacher made ViT trainable on plain ImageNet. The distillation token is the trick.
- Swin brings CNN-like hierarchy back. Shifted local windows make attention linear in image size and bring back inductive bias for spatial structure. The default for detection and segmentation.
- Self-supervised pretraining (MAE, DINOv2) is the modern default. Frozen DINOv2 features outperform fine-tuned ResNet-50 features on most downstream tasks. Self-supervision, not labels, is now where ViTs get their strength.
#Quick Check
You have a 224x224 RGB image and patch size 16. How many tokens enter the transformer encoder?