Multimodal

Open models and evaluations for images, audio, and music, from the early text-to-image work that became VQGAN-CLIP to current vision-language model suites.

Much of this work is community-driven: VQGAN-CLIP emerged from the EleutherAI Discord and helped start the open text-to-image era. More recent work includes the Robin suite of multi-scale vision-language models, studies of what those models systematically miss, and self-supervised representation learning for music.

Selected Papers

Newest first. Showing 12 of 18.

All 18 papers in this area