A Vision-Language Model as a Teacher for Bird Vocalization Detection

· bioRxiv

1 min read Original article ↗

New Results

doi: https://doi.org/10.64898/2026.09.23.753955

Loading

Abstract

Time-frequency detection of bird vocalizations is an important step toward turning weakly annotated field recordings into usable data for studying avian communication. Supervised detectors trained on human expert annotations scale poorly across species and recording conditions, so we introduce a teacher-student setup in which the teacher, a vision-language model, labels bounding boxes on spectrograms of citizen-scientist recordings, and those labels train a student, a self-supervised bioacoustic encoder. We find that the student generally exceeds both the teacher and supervised models trained on human annotations and performs strongly on held-out datasets for both time-frequency and onset-offset localization. YOLO detectors trained on our teacher labels match those trained on human annotations, suggesting that VLM labels can substitute for costly expert labeling.

Competing Interest Statement

The authors have declared no competing interest.

Copyright 

The copyright holder for this preprint is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. All rights reserved. No reuse allowed without permission.