Skip to content
Aryan.
← All work

03 · Computer vision, vision-language models

Few-Shot Adaptation of CLIP (Capstone)

Status
Completed (2026)
Role
Solo capstone project, UTS Bachelor of AI
Started
2026

PyTorch / CLIP / ConvNeXt / ViT

Elevator pitch

Reproduced and compared five state-of-the-art few-shot adaptation methods for CLIP: LP++, CoOp, PromptSRC, TaskRes, and PromptKD. Investigated generalisation under distribution shift and cross-dataset transfer. Best result: PromptKD reached 74.86% top-1 accuracy on Flowers102 at 16-shot, versus a 72.91% zero-shot CLIP baseline, with 606K trainable parameters over 5 epochs.

Why reproductions

Choosing one method and reporting on it is easy. Reproducing five means I actually understand each method's assumptions, hyperparameter sensitivity, and failure modes. This is the depth that matters when you're deciding which method to reach for on a new problem, or when you're reading a new paper and need to judge whether the claimed gains are real.

Setup

  • Teacher model (for PromptKD): ConvNeXt-Base
  • Student model: ViT-B-32-256
  • Dataset: Flowers102 (challenging fine-grained classification)
  • Shot count: 16-shot
  • Epochs: 5
  • Trainable parameters: 606K (student prompt tokens only)

Result

  • PromptKD: 74.86% top-1
  • Zero-shot CLIP baseline: 72.91% top-1
  • Gain over baseline: +1.95 points

Full comparison across the five methods, across multiple datasets, and under distribution shift is in the capstone report.

What I learned that isn't in the papers

  • Prompt-based methods are extremely sensitive to hyperparameter choices in ways the papers understate
  • Most of PromptKD's advantage comes from distillation off a larger teacher; the prompt scaffolding is secondary
  • LP++ is underrated: much cheaper than the prompt methods and gets close on many datasets. If you're compute-constrained, start there

Next

Legal Risk RAG (confidential) →