BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
Singh, Yadavalli, Arnett, and Warstadt. EMNLP Findings, 2026
Open models and evaluations for images, audio, and music, from the early text-to-image work that became VQGAN-CLIP to current vision-language model suites.
Much of this work is community-driven: VQGAN-CLIP emerged from the EleutherAI Discord and helped start the open text-to-image era. More recent work includes the Robin suite of multi-scale vision-language models, studies of what those models systematically miss, and self-supervised representation learning for music.
Newest first. Showing 12 of 18.
Singh, Yadavalli, Arnett, and Warstadt. EMNLP Findings, 2026
Rane, Vatsa, Pethe, Aktolun, Li, and Singh. Low-Resource Audio Codec @ ICASSP, 2026
Alam, Murali, Bharadwaj, Liu, Chung, Sharma, A, Kiran, Tam, and Vegesna. I Can't Believe It's Not Better @ ICLR, 2026
Bradshaw, Spangher, Biderman, and Colton. AI4Music @ NeurIPS, 2025
Kamachee, Casper, Ding, Yew, Reuel, Biderman, and Hadfield-Menell. arXiv, 2025
Zhou-Zheng, Backsund, Chan, Coventry, Eslami, Goel, Han, Soomro, and Wei. MIREX @ ISMIR, 2025
Bradshaw, Spangher, Fan, Biderman, and Colton. ISMIR, 2025
Longpre, Singh, Cherep, Tiwary, Materzynska, and 38 others. ICLR, 2025
Bradshaw, and Colton. ICLR, 2025
Roger, Humane, Kaplan, Gupta, Sun, Adamopoulos, Lim, Anthony, Fennell, and Rish. arXiv, 2025
Lemerle, Vanderbyl, Srivastav, Obin, and Roebel. arXiv, 2024
Crowson, Baumann, Birch, Abraham, Kaplan, and Shippole. ICML, 2024