/

DisParQ describes an image with discrete part concepts, learned from a frozen self-supervised backbone without class labels or language. Each image patch is assigned to exactly one concept from a learnable dictionary, only a small set of concepts is active in each image, and quantized attributes capture how a concept varies from image to image.

PartImageNet

20,466 images · 16 concepts per image

CUB-200-2011

5,994 images · 16 concepts per image

Stanford Dogs

12,000 images · 16 concepts per image

Stanford Cars

8,144 images · 16 concepts per image

Oxford Flowers 102

1,020 images · 8 concepts per image