New Results
doi: https://doi.org/10.64898/2026.09.23.753955

Abstract
Time-frequency detection of bird vocalizations is an important step toward turning weakly annotated field recordings into usable data for studying avian communication. Supervised detectors trained on human expert annotations scale poorly across species and recording conditions, so we introduce a teacher-student setup in which the teacher, a vision-language model, labels bounding boxes on spectrograms of citizen-scientist recordings, and those labels train a student, a self-supervised bioacoustic encoder. We find that the student generally exceeds both the teacher and supervised models trained on human annotations and performs strongly on held-out datasets for both time-frequency and onset-offset localization. YOLO detectors trained on our teacher labels match those trained on human annotations, suggesting that VLM labels can substitute for costly expert labeling.
Competing Interest Statement
The authors have declared no competing interest.
Copyright
The copyright holder for this preprint is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. All rights reserved. No reuse allowed without permission.