/

Stanford Dogs

A frozen DINOv2 ViT-S/14 backbone with a dictionary of 1,024 concepts, up to 16 per image, and 16 quantized attributes of 64 values each. Across its 12,000 training images, a typical image has 16.

Images

Browse 12,000 images Open an image to see its concepts and where else each one appears.

Concepts

All concepts Each card shows a concept's strongest match in four different classes.

Loading…