Speech recognition
2nd MLC-SLM Challenge 2026
2nd Multilingual Conversational Speech Language Model Challenge 2026
Manual or gated
Multilingual Conversational Asr
Speaker Diarization
Speaker Attributed Asr
Acoustic Conversation Understanding
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_workshop_only_registration_agreement
- Code license
- not_specified
- License caution
- The official agreement limits the datasets to the 2026 MLC-SLM Workshop, prohibits redistribution and any other use, requires access controls, and requires return or destruction after termination. The challenge and baseline repositories have no detected license files, so their public visibility does not establish permission to reuse code. This second-edition release must not be conflated with the first MLC-SLM Eval annotation repository or its CC BY-SA 4.0 card label.
- Download notes
- The official challenge describes approximately 2,100 hours of two-speaker conversational training audio across 14 languages, plus approximately four development hours per language. Task 1 evaluates diarization and recognition with DER and time-constrained minimum-permutation WER/CER; Task 2 evaluates acoustic and semantic understanding through multilingual multiple-choice questions. The Task 1 system paper reports 150 development conversations across 21 language/accent categories and says evaluation references are not released. Access to training, development, and evaluation data requires challenge registration and acceptance of the data-use agreement. The helper downloads only public documentation, agreement, repository metadata, and paper metadata, then prints the manual registration path.
Safe-first helperscripts/download/mlc_slm_2nd_challenge.sh
Speaker, identity & emotion
ADD 2022
ADD 2022: The First Audio Deep Synthesis Detection Challenge
Safe-first helper
Audio Deepfake Detection
Low Quality Fake Audio Detection
Partially Fake Audio Detection
Synthetic Speech Detection
+2 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- The official challenge download page and every Zenodo record declare CC BY-NC-ND 4.0. Commercial use and distribution of modified material are prohibited; preserve attribution and review voice, source-recording, generator, and challenge-specific rights before redistribution.
- Download notes
- The public, ungated release covers low-quality fully fake speech, partially fake speech, and an audio generation-versus-detection game. Six official Zenodo records provide train/development, adaptation, and challenge test packages totaling approximately 49.5 GB. The helper downloads the challenge page, paper page, and all record metadata by default; archive files require an explicit opt-in and named record selection. The July 2026 traceback-translator paper evaluates ADD 2022 as a Mandarin continual-learning target but does not identify the exact track or split combination.
Safe-first helperscripts/download/add_2022.sh
Speaker, identity & emotion
ADD 2023
ADD 2023: The Second Audio Deepfake Detection Challenge
Safe-first helper
Audio Deepfake Generation
Audio Deepfake Detection
Partially Fake Audio Detection
Manipulation Region Localization
+3 more
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- The official download page and all eight Zenodo records declare CC BY-NC-ND 4.0. Commercial use and distribution of adaptations are prohibited. Track 1.2 evaluation includes participant-generated attacks, and other tracks derive from genuine and synthesized source speech; review source-media, voice, generator, and challenge terms before redistribution.
- Download notes
- ADD 2023 is one challenge family with four task tracks, not eight independent benchmarks. Eight official, ungated Zenodo records cover Track 1.1 fake generation evaluation, Track 1.2 fake detection train/development and two evaluation rounds, Track 2 manipulation localization train/development and evaluation, and Track 3 algorithm recognition train/development and evaluation. Together they are about 65.4 GB. The helper fetches only official pages and record metadata by default; archives require an explicit opt-in and named selection.
Safe-first helperscripts/download/add_2023.sh
Audio understanding, generation & events
ADQA-Bench
ADQA-Bench: Audio-Dependent Question Answering Evaluation Benchmark
Safe-first helper
Audio Question Answering
Audio Dependent Reasoning
Multiple Choice Question Answering
Shortcut Robustness
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_with_upstream_terms
- Code license
- not_applicable
- License caution
- The Hugging Face card declares Apache-2.0 for the release, but the dataset incorporates a portion of MMAU, MMAR, and MMSU plus newly annotated questions over other audio. Those component datasets and source recordings retain their own terms; review provenance and upstream media rights before redistribution or commercial use.
- Download notes
- The public, ungated DCASE 2026 Task 5 evaluation release contains 3,000 English multiple-choice questions with 3,000 WAV files spanning speech, music, and general sound understanding. Every item passed the organizers' four-stage Audio-Dependency Filtering process to reduce silent-audio and text-only shortcuts. The DCASE 2026 task-summary paper reports 1,607 development items, 14 participating teams, 36 ranked submissions, and two parameter-count tracks; its team results use the hidden 3,000-item evaluation set rather than the development set used for organizer baselines. The current release provides questions and choices without answers for challenge evaluation; the dataset card says answers will be released after the competition. The helper downloads official pages, the dataset card, API metadata, and the lightweight no-answer JSONL by default; the Hugging Face API reports approximately 2.94 GB of repository storage, so the audio snapshot requires explicit opt-in.
Safe-first helperscripts/download/adqa_bench.sh
Speech understanding & dialogue
ADReSS / ADReSSo
Alzheimer's Dementia Recognition through Spontaneous Speech Challenges
Manual or gated
Cognitive Impairment Detection
Alzheimers Dementia Classification
Mmse Score Regression
Cognitive Decline Prediction
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-3.0_with_password_protected_clinical_access
- Code license
- not_applicable
- License caution
- TalkBank says CC BY-NC-SA 3.0 governs its data unless otherwise indicated, prohibits incorporation into commercial products or large language models, and restricts DementiaBank membership to established researchers and clinicians or faculty-sponsored students. Password- protected data may not be shared with non-members or posted elsewhere. The recordings contain sensitive clinical and potentially identifiable speech; users must also follow the TalkBank Code of Ethics, NIH confidentiality protections, citation rules, and non-storage requirements for web processing.
- Download notes
- The official DementiaBank pages release age- and gender-balanced spontaneous-speech challenge sets after consortium approval. ADReSS 2020 provides enhanced full audio, normalized speech segments, transcripts, demographics, diagnosis labels, and MMSE scores for Alzheimer's-dementia classification and MMSE regression. ADReSSo 2021 is audio-only at evaluation time and adds longitudinal cognitive- decline prediction. The helper saves the public challenge pages, DementiaBank access rules, and the July 2026 cross-dataset evaluation paper, then prints the manual membership path; it never attempts to access password-protected participant recordings.
Safe-first helperscripts/download/adress_challenges.sh
Audio understanding, generation & events
AF-Reasoning-Eval
AF-Reasoning-Eval: Sound Reasoning Evaluation Benchmark
Safe-first helper
Audio Question Answering
Audio Reasoning
Audio Classification
Commonsense Reasoning
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC BY 4.0 metadata with mixed upstream audio terms
- Code license
- MIT
- License caution
- NVIDIA releases AF-Reasoning-Eval metadata under CC BY 4.0 and repository code under MIT. The 150 AQA items derive from Clotho-AQA, whose question-answer CSVs are MIT while Freesound audio retains per-file Creative Commons terms. The 7,227-item classification set derives from FSD50K, whose clips retain per-file CC0, CC-BY, CC-BY-NC, or CC Sampling+ terms in addition to the curated dataset's CC BY release.
- Download notes
- The helper downloads all four official JSON annotation files, totaling about 2.1 MB, plus the Sound-CoT README. It does not duplicate source audio. The AQA subset points to Clotho-AQA filenames and the classification subset points to FSD50K evaluation filenames; use those benchmarks' separate helpers and terms to obtain audio.
Safe-first helperscripts/download/af_reasoning_eval.sh
Audiovisual & cross-modal
Aff-Wild2
Aff-Wild2 Database
Manual or gated
Audio Visual Valence Arousal Estimation
Audio Visual Emotion Recognition
Facial Action Unit Detection
Facial Expression Classification
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_signed_eula_request_gated
- Code license
- not_applicable
- License caution
- The owner page requires a signed role-specific EULA and does not state a standard open-data license. Access, sharing, redistribution, commercial use, source-video rights, identifiable faces and voices, and derived artifacts remain controlled by the granted agreement and applicable privacy and media rights.
- Download notes
- Access is granted manually by the database owner after the appropriate signed EULA and an institutional-email request. Academic staff, PhD/postdoctoral researchers through a supervisor, industry users, and undergraduate/postgraduate students follow distinct owner-published procedures. The owner page describes 564 audiovisual videos, around 2.8 million frames, and 554 subjects with frame-level valence-arousal, expression, and action-unit annotations. The July 2026 evaluation paper instead reports 594 videos, so verify the revision delivered under the EULA.
Safe-first helperscripts/download/aff_wild2.sh
Speaker, identity & emotion
AffectDF
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
Safe-first helper
Emotional Speech Deepfake Detection
Audio Deepfake Detection
Synthetic Speech Detection
Audio Language Model Output Detection
+3 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_source_corpus_and_generation_model_terms
- Code license
- not_specified
- License caution
- The paper and Hugging Face card declare AffectDF CC BY-NC 4.0 for research and non-commercial use, subject to the original ESD, MSP-Podcast, and generation-model terms. The evaluation-output repository exposes no detected license, so public score files, checkpoints, and implementation material must not be assumed to share the dataset license. Review every incorporated source and model term before use or redistribution.
- Download notes
- The English benchmark contains 285,797 samples across five emotions and 21 TTS, VC, EVC, and large audio-language-model attack conditions. Its speaker-disjoint partitions contain 86,999 training, 23,330 development, and 175,468 test samples; the test set includes both acted ESD and spontaneous MSP-Podcast source speech. Protocol rows expose speaker, audio, attack, emotion, generation method and model, and real/spoof labels, supporting EER plus emotion-, attack-, and speaking-style robustness analysis. The paper also evaluates prompted Qwen2.5-Omni, Qwen3-Omni, and Voxtral and LoRA-tuned Voxtral. The helper saves owner documentation and release metadata by default. The approximately 1.0 MB protocol archive and approximately 40.2 GB of train/development/test audio are separate explicit opt-ins.
Safe-first helperscripts/download/affectdf.sh
Music
AI-Generated Cover Song Diagnostics
A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
Safe-first helper
Cover Song Generation Evaluation
Music Quality Assessment
Music Error Diagnosis
Music Feature Analysis
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_for_released_tables_no_audio
- Code license
- MIT
- License caution
- The repository's MIT license covers its software and associated documentation, including the released score, manifest, and feature tables. It does not grant rights to the absent source songs or generated cover audio. No raw audio is publicly released, and users must supply locally authorized files to rerun audio feature extraction.
- Download notes
- The public release covers 30 generated covers from five source songs and six systems. It includes anonymized source-song identifiers, system/file mappings, expert scores for melody, harmony, key, style, and arrangement/production, nine extracted features, and the analysis pipeline. The helper downloads these lightweight tables and official documentation by default. The paper and repository explicitly withhold raw audio because of source-song copyright and commercial-API licensing constraints; cloning the small repository does not provide audio.
Safe-first helperscripts/download/ai_cover_song_diagnostics.sh
Audio understanding, generation & events
AIR-Bench
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Safe-first helper
Audio Question Answering
Audio Instruction Following
Audio Language Model Evaluation
Speech Understanding
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- Apache-2.0
- License caution
- The HF dataset card lists CC BY-NC 4.0 and enumerates component sources with mixed upstream licenses, including MusicCaps, Clotho, Fisher, SpokenWOZ, Common Voice, IEMOCAP, acoustic-scene datasets, MUSIC-AVQA, FMA, MTG-Jamendo, NSynth, SLURP, VoxCeleb, LibriSpeech, CoVoST 2, Fake-or-Real, and VocalSound. Re-check component terms before redistribution or commercial use.
- Download notes
- The helper downloads the official GitHub README and Hugging Face dataset card by default. The full HF audio snapshot is about 45.9 GB, so it requires AIR_BENCH_DOWNLOAD_HF=1; cloning the evaluation repo is also opt-in.
Safe-first helperscripts/download/air_bench.sh
Speech recognition
AISHELL-1
AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR33 lists Apache License v2.0 and also describes the data as free for academic use; re-check upstream terms before redistribution or commercial use.
- Download notes
- OpenSLR hosts a 15 GiB speech/transcript archive plus a small supplementary resource archive with lexicon and speaker info. The helper saves the OpenSLR page and supplementary archive by default; the large corpus archive is an explicit opt-in.
Safe-first helperscripts/download/aishell_1.sh
Speech generation
AISHELL-3
AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
Safe-first helper
Text To Speech
Speech Synthesis
Multi Speaker Speech Synthesis
Mandarin Speech Synthesis
+1 more
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_specified
- License caution
- OpenSLR SLR93 lists Apache License v2.0 for the resource. The aishelltech external URL redirected but returned HTTP 429 during the 2026-07-09 check, so OpenSLR was used as the primary access page.
- Download notes
- OpenSLR lists one 19 GiB speech/transcript archive. The helper saves the official OpenSLR page by default and requires AISHELL3_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/aishell_3.sh
Speech recognition
AISHELL-4
AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
Safe-first helper
Meeting Transcription
Multi Channel Asr
Speech Enhancement
Speech Separation
+2 more
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- Apache-2.0
- License caution
- OpenSLR SLR111 lists the corpus under CC BY-SA 4.0. The official baseline repository is Apache-2.0. The arXiv paper is CC BY 4.0, which is separate from the share-alike dataset terms.
- Download notes
- The public OpenSLR release contains 211 real Mandarin meeting sessions totaling 120 hours, recorded with an eight-channel circular microphone array in small, medium, and large rooms. It provides accurate transcripts and speaker activity for meetings with four to eight speakers. The helper saves the OpenSLR page and baseline documentation by default. The approximately 5.2 GB test archive and 46 GB of training archives are explicit opt-ins. VibeVoice-ASR-BitNet section 3.1 and Table 4 evaluate the AISHELL-4 test set with CER.
Safe-first helperscripts/download/aishell_4.sh
Speech recognition
AliMeeting
AliMeeting: A Free Mandarin Multi-channel Meeting Speech Corpus
Safe-first helper
Meeting Transcription
Multi Channel Asr
Multi Speaker Asr
Speaker Diarization
+1 more
Access pathOpenSLR
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR119 lists AliMeeting under CC BY-SA 4.0. The corpus contains real Mandarin meetings with far-field microphone-array and near-field headset recordings; check challenge rules for benchmark submissions.
- Download notes
- The helper downloads small OpenSLR metadata by default. Corpus archives are large, about 73.24 GiB far-field train, 22.85 GiB near-field train, 3.42 GiB eval, and 8.90 GiB test, so they are explicit opt-ins.
Safe-first helperscripts/download/alimeeting.sh
Speech recognition
AMI
AMI Meeting Corpus
Safe-first helper
Meeting Speech Recognition
Multi Speaker Asr
Distant Speech Recognition
Meeting Understanding
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Official AMI pages say the corpus, signals, transcription, and some annotations are CC BY 4.0. OpenSLR SLR16 still lists an older modified CC BY-NC-SA v2.0 notice, so prefer the official AMI license page for current terms and re-check before redistribution.
- Download notes
- The helper downloads official annotation ZIPs by default. OpenSLR acoustic archives and the HF converted dataset are large, so audio downloads are explicit opt-ins.
Safe-first helperscripts/download/ami.sh
Representation & general suites
Androids Corpus
The Androids Corpus: A New Publicly Available Benchmark for Speech Based Depression Detection
Safe-first helper
Speech Based Depression Detection
Clinical Voice Assessment
Health Audio Classification
Paralinguistic Speech Analysis
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_academic_research_only_no_redistribution
- Code license
- not_specified
- License caution
- The official README limits the corpus to non-commercial academic research and prohibits redistribution, broadcasting, or making it publicly available anywhere else. No standalone code or data license file is present. The recordings expose depression/control labels plus age, gender, education, and identifiable voices, so users must also apply appropriate clinical-data, privacy, consent, and ethics review.
- Download notes
- The owner repository links a 3.69 GB archive containing 228 recordings from 118 native Italian speakers: 112 read-speech recordings and 116 spontaneous interview recordings, including 874 segmented interview clips, speaker metadata, turn timing, an openSMILE configuration, and official five-fold lists. The July 2026 voice-concept-bottleneck paper evaluates the 116-speaker interview subset with the official five-fold protocol. A later domain-generalization study also uses that interview subset and those participant-disjoint folds, comparing 10-, 20-, 30-, and 40-second segmentation before selecting 30-second windows; its linked implementation repository was unavailable on 2026-07-28. The helper saves owner documentation, repository metadata, and the primary Interspeech paper by default; archive download requires explicit acknowledgment of the restrictive terms.
Safe-first helperscripts/download/androids_corpus.sh
Music
ASAP
ASAP: A Dataset of Aligned Scores and Performances for Piano Transcription
Safe-first helper
Beat Tracking
Downbeat Tracking
Automatic Music Transcription
Score Performance Alignment
+3 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- Not specified in the source record.
- License caution
- ASAP's repository applies CC BY-NC-SA 4.0 to the dataset, including attribution, noncommercial, and share-alike conditions. Its ISMIR paper is separately CC BY 4.0. Reconstructed audio comes from MAESTRO, which also uses CC BY-NC-SA 4.0; retain both attributions and review the source release before redistribution.
- Download notes
- The public repository reports 236 distinct musical scores and 1,067 MIDI performances across 15 composers, with 519 performances also aligned to audio. It provides MusicXML and MIDI scores, performance MIDI, beat/downbeat, time-signature, and key-signature annotations, and alignment metadata. The helper downloads official documentation, repository metadata, and the approximately 420 KB metadata table by default. The approximately 45 MB combined annotation JSON and roughly 448 MB shallow repository clone are separate opt-ins. Audio is not redistributed by ASAP: obtain MAESTRO v2.0.0 separately and run the repository's initialize_dataset.py script to reconstruct the aligned and trimmed files. Music-JEPA section 4.4 uses ASAP audio in 4-second clips for beat tracking at 70 ms and 100 ms tolerances but does not release its exact probe configuration or derived clip manifest.
Safe-first helperscripts/download/asap.sh
Representation & general suites
ASD Benchmark
Unified DCASE 2020-2025 Task 2 Anomalous Sound Detection Representation Benchmark
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Audio Representation Evaluation
Domain Shift Robustness
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- source_release_terms_apply
- Code license
- MIT
- License caution
- The repository's MIT license covers benchmark software and documentation, not the underlying machine recordings. DCASE Task 2 dataset terms vary by year and source collection; review every official release before use or redistribution.
- Download notes
- The public benchmark supplies a unified evaluation pipeline and converted evaluation lists for the official DCASE 2020-2025 Task 2 releases. It preserves each year's split and scoring rules and supports frozen-embedding nearest-neighbor evaluation plus lightweight adaptation, reporting AUC and pAUC. The helper downloads official documentation and repository metadata by default; cloning the small code repository is opt-in. Obtain the six source datasets from their official DCASE pages. The helper intentionally does not fetch the repository's convenience Google Drive package because its separate redistribution terms are not stated.
Safe-first helperscripts/download/asd_benchmark.sh
Speaker, identity & emotion
ASVspoof 2015
Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015) Database
Safe-first helper
Speaker Verification Anti Spoofing
Spoofed Speech Detection
Synthetic Speech Detection
Unseen Attack Generalization
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata declares Creative Commons Attribution 4.0 International. Retain attribution and review the packaged license and source-speech provenance before redistribution or commercial use.
- Download notes
- The public Edinburgh DataShare release contains genuine speech from 106 speakers and synthetic speech from ten known and unknown text-to-speech and voice-conversion attacks, partitioned into training, development, and evaluation sets. The helper downloads official metadata, README, evaluation plan, summary paper, file descriptions, and extraction instructions by default. The approximately 2.1 MB protocol archive is a separate opt-in, and the three-part approximately 24.1 GB WAV archive requires ASVSPOOF2015_DOWNLOAD_AUDIO=1. Section 4.1 of arXiv:2607.21127 uses attacks S3 and S10 for expert calibration.
Safe-first helperscripts/download/asvspoof_2015.sh
Speaker, identity & emotion
ASVspoof 2017 V2
The 2nd Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2017) Database, Version 2
Safe-first helper
Speaker Verification Anti Spoofing
Presentation Attack Detection
Spoofed Speech Detection
Replay Attack Detection
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata explicitly declares Creative Commons Attribution-NonCommercial 4.0 International. The database uses genuine and replayed RedDots speech; retain attribution and review the packaged files and upstream RedDots conditions before use or redistribution.
- Download notes
- The official Version 2 release contains 42 speakers and genuine and replayed RedDots speech recorded across 179 replay sessions in 61 unique room, replay-device, and recording-device configurations. Training, development, and evaluation archives total approximately 1.4 GiB. The helper downloads the DataShare metadata, README, change log, instructions, evaluation plan, and primary papers by default; the approximately 104 KiB protocol archive and speech archives require separate explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2017.sh
Speaker, identity & emotion
ASVspoof 2019
ASVspoof 2019: The 3rd Automatic Speaker Verification Spoofing and Countermeasures Challenge database
Safe-first helper
Speaker Verification Anti Spoofing
Presentation Attack Detection
Spoofed Speech Detection
Synthetic Speech Detection
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- odc-by-1.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata lists Open Data Commons Attribution License and ships the ODC Attribution license text. The license text itself cautions that database contents can have separate rights; ASVspoof 2019 is derived from VCTK, so re-check component terms before redistribution.
- Download notes
- The DataShare record exposes README, license, evaluation plan, paper PDF, and LA/PA archives. The helper downloads small documentation/license files by default; LA is about 7.6 GiB and PA is about 17.7 GiB, so archives are explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2019.sh
Speaker, identity & emotion
ASVspoof 2021
ASVspoof 2021: Logical Access, Physical Access, and Speech Deepfake Challenge databases
Safe-first helper
Speaker Verification Anti Spoofing
Presentation Attack Detection
Spoofed Speech Detection
Synthetic Speech Detection
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_specified
- License caution
- The official ASVspoof 2021 page says the databases are available under an Open Data Commons Attribution Licence. Zenodo currently lists ODC-BY for LA and PA, and ODC-ODbL for DF; re-check active Zenodo metadata and component-source terms before redistribution. The GitHub baseline repository did not expose a detected license via the GitHub API on 2026-07-10.
- Download notes
- The helper downloads the evaluation plan, LA/PA/DF keys and metadata, Zenodo record metadata, and file-map metadata by default. Evaluation speech archives are large: LA is about 7.8 GB, PA is split into seven parts totaling about 47 GB, and DF is split into four parts totaling about 34.5 GB, so speech archives are explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2021.sh
Speaker, identity & emotion
ASVspoof 5
ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
Safe-first helper
Speaker Verification Anti Spoofing
Spoofed Speech Detection
Synthetic Speech Detection
Deepfake Speech Detection
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- odc-by-1.0_database_and_cc-by-4.0_bona-fide_audio
- Code license
- not_specified
- License caution
- The packaged license applies ODC Attribution 1.0 to the database and CC BY 4.0 to bona fide data. ODC-BY explicitly does not grant every right in individual contents; preserve Multilingual LibriSpeech provenance and review privacy, personality, and generated-voice rights. The baseline repository has no detected license.
- Download notes
- The public release contains 182,357 training, 142,134 development, and 681,872 evaluation utterances at 16 kHz from crowdsourced speech by roughly 2,000 speakers. It covers more than 20 spoofing attacks, seven adversarial attacks, codec conditions, countermeasure evaluation, and spoofing-robust speaker verification. The helper downloads official metadata, README, license, evaluation plan, paper page, and baseline README by default. The approximately 19.7 MiB protocol archive is a separate opt-in; the complete approximately 142.3 GB audio release remains on Zenodo and is not downloaded by the helper.
Safe-first helperscripts/download/asvspoof_5.sh
Audio understanding, generation & events
AudibleLight Eigenmike32
AudibleLight Eigenmike32-5 DCASE-STARSS23 Dataset
Safe-first helper
Sound Source Localization
Acoustic Source Tracking
Spatial Audio Understanding
Microphone Array Upsampling
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ambiguous-cc-tag-with-upstream-terms
- Code license
- CC-BY-4.0-release-archives
- License caution
- The Hugging Face card declares only the generic identifier "cc", which does not identify a specific Creative Commons license. Foreground events come from ESC-50, whose release is CC BY-NC 3.0, and simulated environments use Gibson meshes under their own database terms. The benchmark and generator Zenodo archives declare CC BY 4.0, but that does not override incorporated-data restrictions. Review all upstream terms before redistribution or commercial use.
- Download notes
- The public, ungated card describes one-minute synthetic spatial scenes, each rendered as five independent 32-channel Eigenmike captures at 24 kHz, with DCASE/STARSS23-format annotations using class label 0. Its stated counts conflict: it claims 121 scenes (111 training and ten evaluation) and 570 minutes, while the current Hub API inventory exposes 560 train and 30 test WAV files, or 590 capture-minutes. The helper downloads official cards, API records, paper metadata, and benchmark documentation by default. The Hub API reports about 57.1 GB of repository storage, so the audio snapshot requires AUDIBLELIGHT_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audiblelight_eigenmike32.sh
Speech understanding & dialogue
Audio Agent Bench Suite
AudioAgentBench: Evaluating Multi-Turn Voice Agents on Real-World Tasks
Safe-first helper
End To End Speech Dialogue
Spoken Question Answering
Audio Agent Tool Use
Multi Turn State Tracking
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_CC-BY-4.0_and_MIT_metadata
- Code license
- not_specified
- License caution
- The suite-level card declares CC BY 4.0, while each of the six component dataset cards currently declares MIT. No separate license file or evaluation-code repository is linked. Treat this metadata conflict as unresolved and confirm the intended terms with Arcada Labs before redistribution or commercial reuse.
- Download notes
- The official public suite comprises six English, multi-turn domains: conference assistance (75 turns), laptop sales (31), grocery ordering (30), dental appointments (25), event planning (29), and personal assistance (31), for 221 scripted turns total. Each domain releases user audio, transcripts, reference responses, knowledge-base context, tool schemas, expected function calls, and scoring labels. The card says two consenting voice actors recorded the inputs. The helper saves the suite and component cards plus API metadata by default; downloading all six snapshots (about 209 MB of current repository storage) requires AUDIO_AGENT_BENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audio_agent_bench_suite.sh
Audio understanding, generation & events
Audio Hallucination
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
Safe-first helper
Audio Hallucination Evaluation
Sound Event Existence Reasoning
Temporal Order Reasoning
Sound Source Attribute Reasoning
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_with_mixed_upstream_terms
- Code license
- no_detected_license
- License caution
- The three annotation cards and BEAF-Audio card state no license, and the official benchmark repository has no license file. Public access does not relicense AudioCaps/YouTube, ESC-50, VocalSound, CompA, or their source media. MIT in the separate adaptive-lalm-cd repository covers that paper's code only, not Audio Hallucination data or audio.
- Download notes
- The helper downloads the three small, pinned annotation Parquet files, official documentation, evaluation code, and live Hub metadata by default. The approximately 2.29 GB BEAF_Audio.tar object-existence archive is an explicit opt-in. Temporal-order and object-attribute audio must be obtained separately from the upstream CompA release.
Safe-first helperscripts/download/audio_hallucination.sh
Enhancement, separation & quality
Audio-Alpaca
Audio-Alpaca: A Preference Dataset for Aligning Text-to-Audio Models
Safe-first helper
Text To Audio Preference Modeling
Direct Preference Optimization
Text Audio Alignment
Temporal Event Alignment
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_apache-2.0_and_cc-by-nc-nd-4.0
- Code license
- cc-by-nc-nd-4.0
- License caution
- The Hugging Face dataset card declares Apache-2.0, while the linked official Tango repository applies CC BY-NC-ND 4.0 and does not explain whether that license excludes the dataset. Apply the more restrictive interpretation until the authors clarify scope. AudioCaps-derived captions, generated-audio/model terms, and any source-video rights also require separate review; neither license signal should be assumed to clear all upstream material.
- Download notes
- The public, ungated release contains 15,025 English prompt, chosen-audio, rejected-audio triplets across four construction strategies. Tango 2 generates candidate audio from AudioCaps training captions, perturbed prompts, and varied inference settings, then filters pairs with two CLAP models. Audio-Zero section 3.1 samples and filters 2,000 pairs for post-training but does not release its exact selection. The helper downloads official documentation and API metadata by default; the Hugging Face API reports approximately 9.71 GB of repository storage, so the audio snapshot requires AUDIO_ALPACA_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audio_alpaca.sh
Speech understanding & dialogue
AudioAgentSecurity
AudioAgentSecurity: Stealthy Concurrent Audio Prompt Injection against Multimodal LLM Agents
Safe-first helper
Audio Llm Safety Evaluation
Concurrent Audio Prompt Injection
Spoken Agent Tool Use Safety
Adversarial Audio Robustness
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_stated_auto_gated_access_conditions_apply
- Code license
- not_specified
- License caution
- Neither the Hugging Face API metadata nor the public GitHub repository declares a reusable license, and the gated dataset card was not publicly readable during verification. The arXiv perpetual non-exclusive license covers the article only. Do not infer rights for generated speech, source instructions adapted from MobileAgentBench, AndroidWorld, OSWorld, BrowserGym, and WebVoyager, physical recordings, model outputs, or third-party dependencies. Accept the live Hub terms and review upstream licenses and responsible security-research obligations before downloading or redistributing any assets.
- Download notes
- Section IV constructs 200 manually verified benign/malicious instruction pairs across eight agent scenarios, then applies ten inaudible, low-intelligibility, and semantic-confusion attacks. The paper reports 2,160 generated attack audio samples and evaluates eleven audio-capable agents in a sandboxed tool-trace environment with attack success and benign instruction-correctness metrics. It also defines CADV and prompt-defense comparisons, DEMAND-noise false-positive tests, physical distance/angle/overlap studies, and a 20-participant, 600-judgment perceptual study. The live auto-gated Hub snapshot reports 23,396,170,922 bytes and lists 200 benign WAVs, 1,997 mixed-attack WAVs, and three JSON manifests; this does not exactly match the paper's 2,160-attack count. The public repository releases the evaluator, attack generation and defense code plus selected distance and angle recordings, but not the paper's claimed generated-output logs. The repository README calls the suite AttackBench, so treat that as a release alias rather than a separate family. The helper fetches only primary documentation and live metadata by default. Repository cloning is opt-in; the large gated Hub snapshot additionally requires accepted access terms, authentication, an acknowledgement flag, and a separate download opt-in.
Safe-first helperscripts/download/audioagentsecurity.sh
Speech recognition
AudioBench
AudioBench: A Universal Benchmark for Audio Large Language Models
Safe-first helper
Audio Language Model Evaluation
Automatic Speech Recognition
Speech Translation
Spoken Question Answering
+4 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_upstream_terms
- Code license
- custom_noncommercial_unspecified_version
- License caution
- The paper introduces a suite of 8 tasks and 26 datasets, including 7 newly adapted or collected sets, while the maintained repository now supports more than 50 dataset configurations. The repository license file only says "Creative Commons NonCommercial" without a version, and its README states that each dataset remains under its respective license. Review every selected corpus and derived evaluation set before redistribution or commercial use.
- Download notes
- The helper downloads the official README, supported-dataset inventory, repository metadata, and license notice by default. Cloning the evaluation toolkit is opt-in because the repository is about 64 MB before Git history. It does not download the many upstream audio corpora, whose separate access paths and terms still apply.
Safe-first helperscripts/download/audiobench.sh
Enhancement, separation & quality
Audiobook Narration Appeal
Audio-Based Understanding of Audiobook Narration Appeal
Safe-first helper
Audiobook Appeal Prediction
Narration Quality Analysis
Paralinguistic Feature Analysis
Genre Conditioned Audio Modeling
+1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- Apache-2.0
- License caution
- The repository places its released CSV, supplementary material, and code under Apache-2.0. The CSV derives metadata and engagement counts from LibriVox and the Internet Archive; those services' terms and attribution requirements remain applicable. The license does not extend to separately hosted audiobook recordings or underlying texts. Verify per-item public-domain status in the intended jurisdiction and preserve source attribution. Proprietary Spotify engagement data is described only in aggregate and is not released.
- Download notes
- The public, ungated Spotify Research release contains one metadata row for each of 8,854 single-narrator English LibriVox audiobooks, covering 1,206 narrators and 65 genres. Fields include LibriVox and Internet Archive URLs, title grouping, narrator identifier, duration, genre, views, favorites, reviews, and days since publication. The paper uses time-normalized Internet Archive view rate as a noisy public proxy for narration appeal, evaluates global and genre-specific prediction, and ranks alternative narrations of the same text. The helper downloads the approximately 3.1 MB CSV, official documentation, license, paper, and supplementary material. Audiobook audio is not redistributed; users follow the released source URLs to public-domain LibriVox recordings. The paper's separate Spotify engagement analysis is proprietary and is not part of the public dataset.
Safe-first helperscripts/download/audiobook_narration_appeal.sh
Audio understanding, generation & events
AudioCaps
AudioCaps: Generating Captions for Audios in The Wild
Safe-first helper
Audio Captioning
Audio Language Modeling
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- academic_only
- Code license
- MIT
- License caution
- GitHub README says the code and dataset are free to use for academic purposes only. The repository has an MIT license, but the README adds the academic-use condition for repository material; re-check before redistribution or commercial use.
- Download notes
- Official CSVs contain captions, YouTube ids, and segment start times. Raw audio/video download requires the upstream form and is subject to AudioSet/YouTube availability and terms. The July 2026 AudioLDM pruning paper fine-tunes on the training split and evaluates generation from the 964-caption test split; its public code and checkpoints do not include generated test audio or per-item metric records.
Safe-first helperscripts/download/audiocaps.sh
Audio understanding, generation & events
AudioCards / ASFx Eval
AudioCards: Structured Metadata Improves Audio Language Models for Sound Design
Safe-first helper
Structured Audio Captioning
Structured Metadata Generation
Text Audio Retrieval
Sound Effect Classification
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_annotations_with_separate_Adobe_audio_terms
- Code license
- not_specified
- License caution
- Zenodo declares CC BY 4.0 for the released AudioCard CSV. The underlying Adobe sound effects are not included there: Adobe describes them as royalty-free but states that downloading and using them is governed by the Adobe Audition and related-software EULA. The July 2026 paper is CC BY-NC-SA 4.0, but that paper license does not release or license its absent four-field augmentation or perturbation artifacts.
- Download notes
- The public, ungated Zenodo release contains 499 CSV rows with filenames and 13 structured semantic, caption, and UCS fields. The original paper and project page describe 500 manually screened AudioCards, but the released CSV currently has 499 data rows. Pair its filename column with the separately downloaded Adobe Audition Sound Effects library to reproduce ASFx Eval. The July 2026 evaluation-framework paper uses 499 clips. Its Table 1 selects ten released semantic fields and adds five computed acoustic targets (LUFS, pitch, onset, offset, and a frequency profile), but says that augmented dataset "will" be released; those five acoustic annotations are not part of the current Zenodo file. The helper safely downloads the record metadata, project page, papers, Adobe landing page, and approximately 323 KB annotation CSV; it does not download the Adobe audio archives.
Safe-first helperscripts/download/audiocards.sh
Audio understanding, generation & events
AudioGrounding
Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events
Safe-first helper
Text To Audio Grounding
Temporal Audio Grounding
Sound Event Localization
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_on_Zenodo_with_upstream_terms
- Code license
- MIT
- License caution
- Zenodo declares CC BY 4.0 for the release and the official repository is MIT. The recordings derive from AudioCaps and AudioSet YouTube clips, so retain record attribution and review source-video rights, availability, and platform terms before redistribution or commercial use.
- Download notes
- The public, ungated v2 release contains 3,994 training, 488 validation, and 492 test clips derived from AudioCaps/AudioSet, with caption phrases aligned to one or more onset-offset intervals. The July 2026 GigaChat Audio report evaluates the combined 980 validation/test samples as a short-clip temporal-grounding benchmark and reports mIoU. The helper downloads the official repository documentation, Zenodo metadata, and approximately 5.2 MB of JSON annotations by default; the approximately 2.33 GiB audio archive requires explicit opt-in.
Safe-first helperscripts/download/audiogrounding.sh
Speech recognition
AudioMarathon
AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
Safe-first helper
Long Form Audio Understanding
Automatic Speech Recognition
Spoken Question Answering
Sound Event Detection
+7 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_mixed_upstream_terms
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for the benchmark and the GitHub repository is Apache-2.0. Its own component table additionally marks GTZAN research-only, VESUS academic-only, DESED mixed Creative Commons, and several other sources under separate terms. Apply the benchmark's noncommercial restriction and review every component's provenance, attribution, source-media, and redistribution conditions.
- Download notes
- The public release contains 6,567 English examples across ten long-context tasks, with audio between 90 and 300 seconds. It derives long-form items from LibriSpeech, RACE, Half-Truth, GTZAN, TAU, VESUS, SLUE, DESED, and VoxCeleb. The helper downloads official documentation, repository metadata, and the approximately 1.36 MB test CSV by default. The Hugging Face API reports approximately 58.75 GB of repository storage, so the full snapshot requires AUDIO_MARATHON_DOWNLOAD_HF=1; the evaluation repository clone is a separate opt-in.
Safe-first helperscripts/download/audio_marathon.sh
Speaker, identity & emotion
AudioMarkBench
AudioMarkBench: Benchmarking Robustness of Audio Watermarking
Safe-first helper
Audio Watermarking Robustness
Watermark Removal Detection
Watermark Forgery Detection
Adversarial Audio Robustness
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_upstream_terms_not_separately_specified
- Code license
- MPL-2.0
- License caution
- The repository's MPL-2.0 license covers its source code, but neither the README nor paper states a separate license for the released original, watermarked, or perturbed audio. AudioMarkData derives from Common Voice and the second corpus derives from CC BY 4.0 LibriSpeech; applicable Common Voice release terms, attribution requirements, speaker/privacy considerations, and rights in generated derivatives must be reviewed before reuse or redistribution.
- Download notes
- The public release evaluates AudioSeal/AudioSeal-B, Timbre, and WavMark against 12 no-box perturbation categories plus black-box and white-box adversarial attacks. AudioMarkData contains 20,000 five-second, 16 kHz Common Voice samples balanced for 25 languages, two reported biological-sex groups, and four age groups; the paper also samples 20,000 clips from LibriSpeech. The Drive folder releases original, watermarked, and perturbed audio, while the GitHub repository releases attack and evaluation code. The helper downloads official documentation, license text, repository metadata, and the paper by default; cloning the approximately 3.1 MB GitHub repository requires AUDIOMARKBENCH_CLONE_REPO=1. Google Drive audio remains a manual download so users can inspect its contents and upstream terms.
Safe-first helperscripts/download/audiomarkbench.sh
Speech understanding & dialogue
AudioMNIST
AudioMNIST: Exploring Explainable Artificial Intelligence for audio analysis on a simple benchmark
Safe-first helper
Spoken Digit Classification
Audio Classification
Speaker Metadata Analysis
Explainable Audio Ai
Access pathOfficial / other
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT
- Code license
- MIT
- License caution
- The repository-level LICENSE is MIT and GitHub API reports MIT. Confirm whether downstream use of recorded voices raises consent/privacy obligations beyond the code/data license.
- Download notes
- The repository contains about 30,000 spoken-digit WAV files from 60 speakers plus speaker metadata and Caffe examples. Because the GitHub repository is large, the helper downloads README/LICENSE metadata by default and requires AUDIO_MNIST_DOWNLOAD_REPO=1 before cloning the full repository.
Safe-first helperscripts/download/audio_mnist.sh
Audio understanding, generation & events
AudioSet
Audio Set: An ontology and human-labeled dataset for audio events
Safe-first helper
Audio Event Classification
Audio Tagging
Sound Event Detection
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- AudioSet dataset annotations/features are CC BY 4.0; the ontology is CC BY-SA 4.0. Original YouTube media remains subject to upstream availability and terms.
- Download notes
- Official release provides segment CSVs and precomputed 128-dimensional audio features; it does not redistribute original YouTube audio.
Safe-first helperscripts/download/audioset.sh
Audio understanding, generation & events
AudioSetCaps
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Safe-first helper
Audio Captioning
Audio Text Retrieval
Audio Language Pretraining
Synthetic Audio Question Answering
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- academic_research_only
- Code license
- not_specified
- License caution
- The Hugging Face card metadata labels the release CC BY 4.0, but the same official card explicitly allows only academic and research use. Apply the stricter research-only statement pending clarification. The GitHub repository has no detected license, and AudioSet, YouTube-8M, and VGGSound source-media rights and platform terms still apply.
- Download notes
- The public, ungated release provides synthetic captions for 6,117,099 ten-second clips sourced from AudioSet, YouTube-8M, and VGGSound, plus 18,414,789 intermediate question-answer pairs. The Hugging Face repository contains caption/Q&A CSVs rather than the full source audio and currently reports about 20.2 GB of storage. The helper downloads only official documentation and repository metadata by default; the large CSV files require AUDIOSETCAPS_DOWNLOAD_METADATA=1. The maintainers warn that AudioCaps and VGGSound evaluation examples overlap the release and should be filtered before training.
Safe-first helperscripts/download/audiosetcaps.sh
Audiovisual & cross-modal
AV-GC-AAD
Audiovisual, Gaze-controlled Auditory Attention Decoding Dataset KU Leuven
Safe-first helper
Auditory Attention Decoding
Attended Speaker Decoding
Attention Switch Detection
Eeg Speech Envelope Reconstruction
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-4.0
- Code license
- not_specified_for_crf_repository
- License caution
- The official Zenodo record declares CC BY-SA 4.0 for the released dataset. Copyrighted source speech and video are not included; the record provides only derived acoustic envelopes. The MATLAB baseline uses a custom KU Leuven license limited to academic users and internal, noncommercial research, and prohibits transfer without prior written agreement. The AAD-CRF repository currently has no detected license, so public source access does not establish code reuse rights. Human-subject consent, privacy, ethics, and re-identification constraints remain relevant.
- Download notes
- The versioned Zenodo release contains approximately 2.04 GB of preprocessed 128 Hz EEG/EOG, attended and unattended speech envelopes, trial metadata, condition IDs, and attention-side annotations for 13 consenting participants. The helper downloads only the official record metadata and 6 KB README by default; subject MAT files require explicit opt-in and a subject list. Original WAV and MP4 stimuli are excluded for copyright reasons. The 2026 CRF paper evaluates six trials per participant with one programmed spatial switch near five minutes, participant-dependent leave-one-trial-out folds, one-, two-, and four-second windows, accuracy, and switch delay.
Safe-first helperscripts/download/av_gc_aad.sh
Audiovisual & cross-modal
AV-SpeakerBench
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
Safe-first helper
Audio Visual Question Answering
Speaker Centric Audio Visual Reasoning
Audio Visual Temporal Grounding
Speech Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official project page, GitHub README, and Hugging Face card prose state CC BY-NC 4.0, while the Hugging Face card front matter incorrectly or inconsistently declares MIT. Use the more restrictive CC BY-NC 4.0 terms and re-check upstream source-video rights before redistribution or commercial use. GitHub reports no detected repository license.
- Download notes
- The public, ungated release contains 3,212 English multiple-choice questions plus aligned audio-only, visual-only, and audiovisual clips. The helper downloads official documentation by default; the Hugging Face API reports about 123 GB of repository storage, so the full snapshot requires AV_SPEAKERBENCH_DOWNLOAD_HF=1. Qwen3.5-Omni reports the benchmark in section 5.1.4, Table 7.
Safe-first helperscripts/download/av_speakerbench.sh
Audiovisual & cross-modal
AV-SyncBench
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Safe-first helper
Audio Visual Temporal Synchronization
Audio Visual Semantic Consistency
Global Offset Detection
Local Jitter Detection
+3 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- ModelScope declares MIT for the dataset repository, but the benchmark was collected from public online platforms and neither the paper nor card identifies per-video source rights. Treat those underlying rights and platform terms as separately controlling; do not infer that the paper license or GitHub placeholder licenses source media.
- Download notes
- The public ModelScope release contains seven split tar-gzip parts at revision 86c06579529a6e7b2cafb0dc386a50152a37fb98; its API reports approximately 68.2 GB of storage. The paper defines 38,390 samples derived from 3,269 three-to-thirteen-second videos across Voice, Music, and Sound, with three temporal and two semantic challenge types. Pairwise accuracy compares the original and perturbed audio against visual embeddings computed over non-overlapping 0.64-second chunks. The helper downloads paper, project, repository, and release metadata by default; the full ModelScope Git-LFS clone requires explicit opt-in. Evaluation code remains announced but unreleased. A focused 2026-08-09 recheck found the official GitHub repository unchanged at four files, with no push after 2026-03-21: its README still promises evaluation code later, and the tree contains only the project page, two images, and that README. The Hugging Face placeholder is likewise unchanged at revision 23329259e882f92acdfb8c0133e46b3a1c70cd0c and contains only a README and .gitattributes, so neither host releases an evaluation runner, model adapters, configurations, predictions, or per-item scores.
Safe-first helperscripts/download/av_syncbench.sh
Audiovisual & cross-modal
AVA Active Speaker
AVA Active Speaker: An Audio-Visual Dataset for Active Speaker Detection
Safe-first helper
Active Speaker Detection
Audio Visual Speaker Localization
Speaking Activity Detection
Face Track Classification
Access pathOfficial / other
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC BY 4.0
- Code license
- not_specified
- License caution
- Google's AVA download page states that all datasets listed there are CC BY 4.0. The GitHub repository has no detected license, so that statement should not be assumed to license repository code. Preserve attribution and review source-movie and hosting terms before redistributing media; video availability can change independently of the released annotations.
- Download notes
- The official v1.0 release associates visible face tracks with SPEAKING_AND_AUDIBLE, SPEAKING_BUT_NOT_AUDIBLE, or NOT_SPEAKING labels. Google reports 3.65 million labeled frames across approximately 39,000 face tracks; the current download page says dense labels cover 160 AVA movie clips that remained available on YouTube. The helper downloads official documentation and the 2.5 KB video-name manifest by default. Set AVA_ACTIVE_SPEAKER_DOWNLOAD_LABELS=1 for the approximately 23 MB train/validation annotation archives. It does not fetch the much larger source videos.
Safe-first helperscripts/download/ava_active_speaker.sh
Audiovisual & cross-modal
AVCap-Bench
Safe-first helper
Detailed Audio Visual Captioning
Atomic Audio Fact Coverage
Atomic Visual Fact Coverage
Audio Visual Alignment
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- CC BY 4.0 covers the article, not gated benchmark media or code. The AVCap-Codes card exposes no license field, so no benchmark or code license should be inferred. Apache 2.0 is declared separately for the training dataset and model; it does not automatically cover AVCap-Bench or override rights in source videos drawn from AVE, VGGSound, Condensed Movies, AVQA, Trailer30K, MPII-MVAD, and YouTube.
- Download notes
- AVCap-Bench is a 1,000-video held-out evaluation split fixed before training. Five annotators checked and corrected each detailed caption against visual facts, audio facts, and audio-visual temporal alignment. AVCap-Score creates 20 atomic questions per reference, answers them from a candidate caption, and uses a text judge to assign 1-5 semantic-equivalence scores before averaging by visual, audio, and joint dimensions. The owner code repository exposes a 2,580-file tree at revision `c57d273802944cec717902f26c9d041947960da3`, including the 1,000 videos, frozen test JSON, prompts, and three-stage evaluation runner, but requires manual Hugging Face approval. The separate approximately 842.5 GB AVCap-Dataset is the SFT/GRPO training release, not an additional benchmark. The helper downloads only public paper, collection, and repository metadata and does not attempt gated files.
Safe-first helperscripts/download/avcap_bench.sh
Audiovisual & cross-modal
AVDC
AVDC: Audio-Visual Decoupled Captions
Safe-first helper
Audio Visual Decoupled Captioning
Audio Only Captioning
Visual Only Captioning
Joint Audio Visual Captioning
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The Hugging Face card has no license field, and neither the dataset repository nor the official GitHub repository exposes a license file. The arXiv publication license does not license the released annotations, generated captions or reasoning traces, code, or source videos. ShareGPT4Video, Vript, and each source platform or media owner retain applicable terms; obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated Hugging Face release provides avdc_caption.json with decoupled audio, visual, and joint captions and omni_qa.json with questions, answers, reasoning steps, and split labels. The paper describes 10,000 long-form videos, 10,000 caption-derived QA pairs, and an AVDC-test split for visible versus invisible sound-event evaluation. The current JSON blobs total approximately 134 MiB, so the helper downloads only official documentation and API metadata by default and requires AVDC_DOWNLOAD_HF=1 for the annotation snapshot. Source videos are referenced by video ID but are not redistributed in the Hugging Face release; users must obtain applicable ShareGPT4Video- and Vript-sourced media separately. The training/evaluation repository is an additional AVDC_CLONE_REPO=1 opt-in.
Safe-first helperscripts/download/avdc.sh
Audiovisual & cross-modal
AVE
Audio-Visual Event Localization in Unconstrained Videos
Safe-first helper
Audio Visual Event Localization
Audio Visual Event Classification
Cross Modal Localization
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository does not expose a detected GitHub license and the README does not state a standalone dataset license. AVE is built from unconstrained videos, so source-video copyright, platform terms, and redistribution rights should be checked before use.
- Download notes
- The helper downloads the project page and official README by default and can clone the code repository. The dataset, precomputed audio features, and visual features are linked from Google Drive; download those manually or with a user-selected Drive tool after reviewing terms.
Safe-first helperscripts/download/ave.sh
Audiovisual & cross-modal
AVE-Compass
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Safe-first helper
Instruction Based Audio Video Editing
Joint Audio Visual Editing
Speech Editing
Audio Only Editing
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for the benchmark release. The evaluation repository has no LICENSE file or GitHub-detected license. Source-video filenames identify a mixture that includes web-video-derived clips, so users should retain attribution and review uploader, platform, privacy, and underlying media rights before redistribution or derivative use.
- Download notes
- The public, ungated release contains 145 curated source videos, 196 human-verified English editing instructions, 196 checklist JSON files with 2,688 fine-grained items, and 28 editing operation types across joint, speech, video-only, and audio-only branches. The helper downloads official documentation and repository metadata by default. Set AVE_COMPASS_DOWNLOAD_METADATA=1 for the lightweight instruction, checklist, and Dataset Viewer metadata files, or AVE_COMPASS_DOWNLOAD_HF=1 for the complete approximately 442 MB Hugging Face snapshot. The official project currently provides a citation but no public paper URL.
Safe-first helperscripts/download/ave_compass.sh
Audiovisual & cross-modal
AVHBench
A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Safe-first helper
Audio Driven Video Hallucination
Video Driven Audio Hallucination
Audio Visual Matching
Grounded Audio Visual Captioning
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository has no top-level license and GitHub detects no license. Public Google Drive access does not itself grant reuse or redistribution rights. AudioCaps, VALOR, YouTube-derived media, annotations, synthetic swapped videos, and voices retain their own copyright, platform, consent, privacy, and publicity constraints.
- Download notes
- The helper downloads the owner documentation, live repository metadata, and the 1,107,911-byte QA.json by default. That frozen annotation file contains 6,408 rows over 2,136 videos: 5,302 binary judgments and 1,106 captions. The owner-hosted video ZIP is a large explicit opt-in.
Safe-first helperscripts/download/avhbench.sh
Audiovisual & cross-modal
AVQA
AVQA: A Dataset for Audio-Visual Question Answering on Videos
Manual or gated
Audio Visual Question Answering
Multimodal Scene Understanding
Cross Modal Reasoning
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- noncommercial_or_permission_required
- Code license
- not_specified
- License caution
- The official page permits personal or classroom copying without fee only when it is not for profit or commercial advantage, and requires permission for broader copying, reposting, or redistribution. The repository has no license file or GitHub-detected license. Raw clips derive from VGGSound/YouTube, so source-media rights and availability also apply.
- Download notes
- The helper saves the official project page, repository README, and GitHub metadata, then prints the official OneDrive/Baidu download paths. It does not automate the combined archive or source-video retrieval. The public release provides QA annotations, a VGGSound-derived video manifest, raw-video access pointers, and large pre-extracted features.
Safe-first helperscripts/download/avqa.sh
Audiovisual & cross-modal
AVSBench
AVSBench: Audio-Visual Segmentation Benchmark
Manual or gated
Audio Visual Segmentation
Sounding Object Segmentation
Single Sound Source Segmentation
Multiple Sound Source Segmentation
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- Apache-2.0
- License caution
- The official project page licenses the AVSBench dataset published there under CC BY-NC 4.0, and the repository licenses the project under Apache-2.0. Source videos were collected from public YouTube material, so uploader rights, platform terms, and current video availability still apply; confirm whether the updated semantic release carries any additional terms during application.
- Download notes
- The benchmark covers semi-supervised Single Sound Source Segmentation (S4), fully supervised Multiple Sound Source Segmentation (MS3), and the later semantic-label AVSS task. The official project page publicly links the original AVSBench-object video-ID CSV and segmentation maps on Google Drive, while processed video/audio requires an email request. The repository directs users to the official application page for the updated object and semantic datasets. The helper saves lightweight official documentation, repository metadata, and license text, then prints these manual access paths; it does not automate Drive, email, or source-video retrieval.
Safe-first helperscripts/download/avsbench.sh
Audiovisual & cross-modal
AVSCapBench
AVSCapBench: Fine-Grained Audio-Visual Synergy Evaluation for Omni-Modal Video Captioning
Safe-first helper
Omni Modal Video Captioning
Audio Visual Captioning
Audio Event Captioning
Audio Visual Event Binding
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card and repository README declare CC BY-NC-SA 4.0, while the paper describes an academic-research-only restrictive release. The paper says clips come from YouTube, TikTok, and Video-MME, invokes fair use, and says only public URLs and timestamps are distributed, but the current Hugging Face repository lists 1,226 MP4 files. Treat the stricter academic/non-commercial interpretation as controlling and review source-platform, uploader, Video-MME, privacy, and copyright terms before downloading, redistributing, or publishing clips. The evaluation repository has no detected license.
- Download notes
- The public, ungated release contains 1,226 English video clips lasting 30 to 120 seconds, dense omni-modal captions, visual events, audio events separated into speech, music, and sound effects, and synergistic audio-visual events. The evaluation repository implements LLM-judged event recall. The helper downloads official documentation and repository metadata by default; set AVSCAPBENCH_DOWNLOAD_HF=1 for the approximately 19.6 GB Hugging Face snapshot.
Safe-first helperscripts/download/avscapbench.sh
Audiovisual & cross-modal
AVSD
Audio Visual Scene-Aware Dialog Dataset
Manual or gated
Audio Visual Dialogue
Video Question Answering
Multimodal Response Generation
Scene Understanding
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- unclear
- Code license
- MIT
- License caution
- The official repository is MIT-licensed, but it does not separately state that the MIT license covers the Google Drive dataset or underlying Charades videos. Treat dialog annotations and media rights conservatively, review any terms shown by the Drive/Charades access paths, and preserve source-video provenance before redistribution or commercial use.
- Download notes
- The CVPR paper introduces dialogs and final summaries for more than 11,000 Charades videos. The official DSTC7 repository reports 7,659 training, 1,787 validation, and 1,710 test dialogs and links the released challenge data through Google Drive. The helper saves the public repository documentation, license, and CVPR paper page, then prints the manual dataset and Charades media paths; it does not automate Google Drive or raw-video retrieval.
Safe-first helperscripts/download/avsd.sh
Audiovisual & cross-modal
AVUT
Audio-centric Video Understanding Benchmark without Text Shortcut
Safe-first helper
Audio Visual Question Answering
Audio Content Understanding
Audio Visual Alignment
Audio Event Localization
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the official repository nor Hugging Face card states a data or code license. The paper says AVUT contains only links to public YouTube videos and does not host or distribute video copies, while the current Hugging Face repository appears to include video files; users should review YouTube terms, source-video rights, and this discrepancy before downloading or redistributing media.
- Download notes
- The public, ungated release covers 2,662 English YouTube videos across 18 domains and 11,609 question-answer pairs in AV-Human and AV-Gemini. The helper downloads official documentation and four lightweight annotation JSON files by default; the Hugging Face API reports about 24.0 GB of repository storage, so the full snapshot requires AVUT_DOWNLOAD_HF=1. Qwen3.5-Omni reports AVUT in section 5.1.4, Table 7.
Safe-first helperscripts/download/avut.sh
Audiovisual & cross-modal
BackgroundMellow Cinematic Trailer Evaluation
BackgroundMellow: Narrative-Driven Cinematic Soundscape Generation Evaluation
Safe-first helper
Narrative To Audio Generation
Cinematic Soundscape Generation
Multi Track Audio Orchestration
Sound Event Coverage
+2 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The repository has no top-level LICENSE and GitHub detects no license; neither the public results sheet nor the paper states separate terms for the trailer annotations, generated prompts, cue manifests, experiment rows, or evaluation code. The arXiv distribution license covers the paper only. Underlying trailer audio/video remains subject to uploader, studio, copyright-holder, and YouTube terms, and generated outputs may carry additional model-service conditions. Obtain permission or clarification before redistribution or commercial use.
- Download notes
- The paper curates approximately 100 public YouTube movie trailers into 1,000 clips with generated story descriptions, sound categories, timings, and relative mix levels. Section 6 evaluates 40 sampled story prompts with retrieval-based sound-event coverage and temporal-IoU synchronization metrics. The public repository includes evaluation code, prompt-to-source mappings, cue manifests, comparison outputs, and aggregate ablation results; the linked public spreadsheet exports approximately 7.1 MB of row-level experiment results. The helper saves official documentation, repository metadata, the paper page, a lightweight evaluation mapping, and aggregate results by default. Exporting the spreadsheet requires BACKGROUNDMELLOW_DOWNLOAD_RESULTS=1, while cloning the approximately 121 MB repository requires BACKGROUNDMELLOW_CLONE_REPO=1. Source YouTube audio/video is referenced rather than redistributed as a clean benchmark archive.
Safe-first helperscripts/download/backgroundmellow_cinematic_trailer_eval.sh
Audiovisual & cross-modal
BAH
BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change
Manual or gated
Ambivalence Hesitancy Recognition
Multimodal Affect Recognition
Vocal Expression Analysis
Audio Visual Behavior Understanding
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- proprietary_research_only_eula
- Code license
- BSD-3-Clause
- License caution
- The ÉTS owner page expressly labels BAH proprietary and research-only. The current request process is limited to full-time faculty at an eligible university, higher-education institution, or equivalent organization; students and postdoctoral researchers cannot apply directly. The public repository's BSD-3-Clause license applies to code, not the gated recordings, transcripts, annotations, faces, or participant metadata. Review the signed EULA for storage, sharing, publication, retention, privacy, and downstream-use obligations.
- Download notes
- The release contains 1,427 videos totaling 10.60 hours from 300 participants across Canada, including 1.8 hours of annotated ambivalence/hesitancy moments. It provides raw videos, 16 kHz audio, timestamped transcripts, cropped and aligned faces, expert video- and frame-level labels and cues, participant metadata, and predefined participant-disjoint splits. Access is manual: an eligible full-time faculty member must submit the official form, list every team member, certify institutional eligibility, and sign the EULA. The helper saves only public owner, paper, repository, challenge, and request documentation, plus the ABAW 2026 calibration paper and implementation metadata, and never downloads participant data or model artifacts.
Safe-first helperscripts/download/bah.sh
Audio understanding, generation & events
Big Bench Audio
Artificial Analysis Big Bench Audio
Safe-first helper
Spoken Question Answering
Speech Reasoning
Audio Question Answering
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares MIT and the Xiaomi MiMo evaluator is Apache-2.0. The 1,000 English recordings contain verbatim questions from four BIG-Bench Hard tasks and were synthesized with 23 OpenAI-, Azure-, and AWS-provided voices; review inherited task terms and provider-generated-audio conditions rather than assuming the card resolves every component right.
Safe-first helperscripts/download/big_bench_audio.sh
Audiovisual & cross-modal
BioTalk-3D
BioTalk-3D: A Synchronized EEG, Audio, and Blendshape Dataset and Benchmark for 3D Facial Animation
Safe-first helper
Speech Driven 3d Facial Animation
Physiology Aware Facial Motion Reconstruction
Audio Physiology Facial Motion Alignment
Cross Subject Generalization
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_noncommercial_research_terms
- Code license
- not_released
- License caution
- DATA_LICENSE.md permits academic research, education, and non-commercial scientific evaluation, prohibits re-identification, biometric surveillance, clinical use, and redistribution, and requires citation. The repository contains documentation but no baseline code. Its README names additional data-card, ethics, and rendering documents that are not present in the current four-file repository tree. Review the packaged terms and institutional ethics requirements before use.
- Download notes
- BioTalk-3D is one synchronized multimodal dataset and benchmark family, not separate audio, electrophysiology, and facial-motion datasets. Version 1.0.0 reports 2,300 Chinese and English utterances from five anonymized participants and 20 sessions, totaling 13,120.016 seconds. The approximately 18.05 GB Baidu Netdisk package includes segmented WAV audio, 51-dimensional ARKit blendshapes, preprocessed session-level and segmented 64-channel scalp electrophysiological signals, rendering videos, and official train/validation/test splits. Raw acquisition files and original participant face videos are not released. The helper downloads only the lightweight repository documentation; obtain the data manually with extraction code em25.
Safe-first helperscripts/download/biotalk_3d.sh
Audio understanding, generation & events
BirdCLEF++ 2026
BirdCLEF++ 2026: Multi-label Animal Vocalization Detection in Pantanal Soundscapes
Manual or gated
Multi Label Animal Vocalization Detection
Passive Acoustic Monitoring
Soundscape Classification
Few Shot Species Detection
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- kaggle_competition_rules_and_source_terms_apply
- Code license
- Not specified in the source record.
- License caution
- The official LifeCLEF page does not state a standalone dataset license. Access requires a Kaggle account and acceptance of the competition rules. Focal recordings include Xeno-canto material and may retain recording-specific Creative Commons terms; soundscape recordings and annotations require review under the accepted competition rules. The paper's CC BY 4.0 license covers the article, not the competition audio.
- Download notes
- The official LifeCLEF page routes data access through the BirdCLEF 2026 Kaggle competition. The primary CLEF working paper reports 522.1 hours of audio: 344.5 hours of focal recordings across 206 species and 177.6 hours of Pantanal soundscapes from 23 sites. Only 1.03 hours (66 soundscape files and 739 unique five-second windows) are labeled; the remaining 176.5 soundscape hours are unlabeled. The task predicts 234 target species on five-second windows from 60-second soundscapes. The helper saves the official challenge page and paper metadata, then prints the manual Kaggle access path. It does not accept competition rules, authenticate, or download audio.
Safe-first helperscripts/download/birdclef_2026.sh
Speaker, identity & emotion
BLAB
BLAB: Brutally Long Audio Bench
Safe-first helper
Long Audio Question Answering
Word Localization
Named Entity Localization
Advertisement Localization
+5 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_cc_by_4.0_and_evaluation_only_metadata_with_upstream_video_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY 4.0 in structured metadata but also says the data should only be used for evaluation and not model training. The repository has no license file and GitHub detects no code license. BLAB does not redistribute the underlying YouTube audio; the paper says sources were Creative Commons-licensed, but each video's license, availability, and platform terms still apply. Resolve the card's conflicting reuse signals with the authors before training, redistribution, or commercial use.
- Download notes
- BLAB is one long-form audio-language benchmark with eight task configurations, not eight independent datasets. The paper reports 1,600 question-audio-answer items over more than 833 hours of YouTube-sourced audio, with clips averaging 51 minutes. The public, ungated Hugging Face release contains video URLs, audio-path fields, questions, answers, and task-specific timestamps rather than audio files. Its JSON metadata totals approximately 535 MB, so the helper downloads only official paper, repository, dataset-card, and API metadata by default; the full metadata snapshot and evaluation toolkit are separate opt-ins. Reconstructing inputs depends on each source video remaining available and on compliance with its terms.
Safe-first helperscripts/download/blab.sh
Enhancement, separation & quality
BVCC
BVCC: VoiceMOS Challenge 2022 Main-Track Dataset
Safe-first helper
Mean Opinion Score Prediction
Synthetic Speech Naturalness Assessment
Voice Conversion Quality Assessment
Text To Speech Quality Assessment
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_open_mixed_upstream_terms
- Code license
- not_separately_specified
- License caution
- Zenodo labels the release Other (Open), not a standard reusable data license. Its record explicitly prohibits redistribution of Blizzard Challenge samples and omits those files; Voice Conversion Challenge and ESPnet-TTS components retain their own terms. Treat ratings, metadata, audio, and scripts according to their component provenance, and do not infer broad redistribution or commercial rights from public access.
- Download notes
- The public VoiceMOS Challenge 2022 release provides unified MOS ratings and official train, development, and test splits for synthetic speech from past Voice Conversion and Blizzard Challenges plus ESPnet-TTS. The helper downloads official pages and Zenodo metadata by default. The approximately 273.4 MiB main-track archive, small out-of-domain package, and scoring package are separate opt-ins. Blizzard audio is intentionally absent from the archive; official scripts require users to obtain and preprocess that material under its original access terms. Section 4.2 of arXiv:2607.13477 constructs 60 BVCC test-set pairs with a human-MOS gap of at least 1.0 for its naturalness probe, but does not release the selected pair manifest.
Safe-first helperscripts/download/bvcc.sh
Speech generation
CapSpeech
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Safe-first helper
Style Captioned Text To Speech
Text To Speech With Sound Effects
Accent Captioned Text To Speech
Emotion Captioned Text To Speech
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_mixed_upstream_audio_terms
- Code license
- cc-by-nc-4.0
- License caution
- The dataset card, paper, and repository apply CC BY-NC 4.0 to CapSpeech resources. The release points to audio from Emilia, GigaSpeech, Common Voice, MLS, LibriTTS-R, VoxCeleb, EARS, Expresso, VCTK, VGGSound, FSDKaggle2018, ESC-50, and separate CapSpeech audio repositories; those recordings retain their own attribution, non-commercial, access, privacy, and media-rights constraints. Treat the CapSpeech license as covering its annotations and author contributions, not as overriding component-source terms.
- Download notes
- The public, ungated release contains more than 10 million machine-annotated and approximately 360,000 human-annotated English audio-caption records, with fixed pretraining and supervised-fine-tuning train, validation, and test splits across CapTTS, CapTTS-SE, AccCapTTS, EmoCapTTS, and AgentTTS. The main Hugging Face snapshot contains paths, transcripts, source labels, durations, and style captions rather than embedded source audio; the API reports approximately 4.31 GB compressed and 10.09 GB after processing. The helper downloads official documentation and API metadata by default, while the full metadata snapshot and code repository are separate opt-ins. The 2026 ProPS paper trains and evaluates prompt-conditioned speaker-profile distributions on CapSpeech's held-out splits.
Safe-first helperscripts/download/capspeech.sh
Audio understanding, generation & events
CASTELLA
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
Safe-first helper
Audio Moment Retrieval
Temporal Audio Grounding
Long Audio Retrieval
Audio Captioning
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_for_annotations_and_released_features
- Code license
- not_specified_for_audio_downloader
- License caution
- The annotation repository, its Hugging Face card, and the Zenodo feature record declare CC BY 4.0. That license covers the released annotations/features, not the underlying YouTube recordings. The separate raw-audio downloader repository has no detected license, and raw media remains governed by its owners and YouTube terms; verify availability and rights before downloading, redistribution, or commercial use.
- Download notes
- The public, ungated annotation release describes 1,862 real-world YouTube recordings split into 1,009 training, 213 validation, and 640 test items, with 3,925 human-written local captions and 11,308 temporal boundaries in English and Japanese. The Hugging Face mirror contains only the six lightweight annotation JSON files, not raw audio. The helper downloads those annotations plus official documentation and metadata by default. Precomputed MS-CLAP audio/text features are an approximately 2.78 GB Hugging Face opt-in (the Zenodo release is about 1.33 GB). Raw media must be reconstructed from YouTube IDs with the separate official downloader, subject to current availability and source-platform and recording rights; the helper only clones those tools when explicitly requested.
Safe-first helperscripts/download/castella.sh
Speech recognition
CDSD
CDSD: Chinese Dysarthria Speech Database
Manual or gated
Dysarthric Speech Recognition
Pathological Speech Recognition
Speech Intelligibility Assessment
Dysarthria Severity Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_signed_agreement_review_required
- Code license
- not_applicable
- License caution
- Access requires a signed owner agreement and approval. No standard open-data license is stated on the public access or paper pages; users must read the current agreement and retain its conditions rather than inferring rights from the paper's publication license. The corpus contains identifiable voices and health/disability information, so institutional ethics, privacy, storage, sharing, and permitted-purpose requirements need review before application and use.
- Download notes
- The owner describes 133 hours of Mandarin dysarthric speech from 44 speakers and reports a best benchmark character error rate of 16.4%. Access is not an anonymous archive download: applicants must download, complete, sign, and upload the owner's license agreement, then wait for approval by email. The helper saves only the official access page, blank agreement, paper page, and Crossref DOI metadata; it never submits an application or downloads clinical recordings.
Safe-first helperscripts/download/cdsd.sh
Audiovisual & cross-modal
CH-SIMS
CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of Modality
Safe-first helper
Multimodal Sentiment Analysis
Audio Sentiment Analysis
Visual Sentiment Analysis
Text Sentiment Analysis
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The MMSA repository is MIT-licensed, but neither the paper nor the repository README expressly applies that license to the CH-SIMS video and annotation files in the shared Drive folders. Treat dataset terms and source-media rights as unspecified and verify permitted use before redistribution or commercial use.
- Download notes
- The paper introduces 2,281 Chinese in-the-wild video segments with multimodal sentiment labels and separate text, audio, and visual annotations. The official MMSA repository provides shared Baidu and Google Drive folders containing raw video, processed features, and labels for CH-SIMS alongside MOSI and MOSEI. The helper downloads official paper, README, license, and repository metadata only; dataset files remain a manual Drive download, and cloning the toolkit is opt-in.
Safe-first helperscripts/download/ch_sims.sh
Audiovisual & cross-modal
CH-SIMS v2
CH-SIMS v2.0: A Fine-grained Multi-label Chinese Multimodal Sentiment Analysis Dataset
Safe-first helper
Multimodal Sentiment Analysis
Audio Sentiment Analysis
Visual Sentiment Analysis
Text Sentiment Analysis
+2 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official project page, paper, and repository state no dataset or code license, and the repository has no LICENSE file. Treat the videos, annotations, features, and code as all-rights-reserved unless the authors provide terms; also review source-media, speaker, privacy, and platform rights before reuse or redistribution.
- Download notes
- The official release extends and re-annotates CH-SIMS with 4,402 supervised segments carrying multimodal and unimodal sentiment labels plus 10,161 unlabeled segments for semi-supervised evaluation. It provides raw videos, extracted features, IDs, splits, and labels through separate Google Drive and Baidu folders. The helper downloads only the official homepage, repository README/API metadata, and arXiv metadata; Drive data remains manual and the code clone is opt-in.
Safe-first helperscripts/download/ch_sims_v2.sh
Music
ChartGenEval
ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
Safe-first helper
Music Generation Evaluation
Rhythm Game Chart Generation
Chart Audio Alignment
Chart Structure Evaluation
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_released_artifacts_corpus_not_released
- Code license
- Not specified in the source record.
- License caution
- The repository-level MIT license covers the released software and bundled artifacts. It does not grant rights to the absent chart/audio corpus; the README says potentially copyrighted community and commercial content is intentionally not redistributed. Users reproducing corpus-dependent experiments must supply a rights-compatible equivalent snapshot and review its chart, recording, and game-content terms.
- Download notes
- The public repository releases the NumPy-based evaluation toolkit, calibration artifact, corruption probes, plotting and reproduction scripts, and metric-only evaluation records. The paper validates seven output axes with nine held-out controlled-corruption tests across timing, audio response, structure, grammar, and human-reference gap measurements. The helper downloads official documentation and repository metadata by default; cloning the approximately 44 MB current file tree requires CHARTGENEVAL_CLONE_REPO=1. The underlying 3,880-chart calibration corpus and its audio are explicitly not distributed because they may contain copyrighted community and commercial material, so the released artifacts support verification of reported results but not full corpus-dependent reproduction.
Safe-first helperscripts/download/chartgeneval.sh
Speech recognition
CHILDES-Aligned
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling
Manual or gated
Asr
Child Speech Recognition
Long Form Speech Alignment
Forced Alignment
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0_with_talkbank_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC-SA 4.0 and its access agreement additionally limits use to non-commercial research, requires citation of the BEACON paper and every source CHILDES corpus used, incorporates the TalkBank Ground Rules, and prohibits audio redistribution. Access is manually reviewed. The paper's linked BEACON GitHub repository was not publicly reachable when checked, so no code license is claimed.
- Download notes
- The manually gated Hugging Face release contains a 413.3-hour general-purpose English child-speech configuration with corrected utterance timestamps and a quality-controlled 283-hour ASR-training configuration. The repository reports approximately 160.6 GB of storage. The helper prints the access steps by default and downloads a selected configuration only after the user has received access, authenticated with Hugging Face, and set CHILDES_ALIGNED_ACK_TERMS=1.
Safe-first helperscripts/download/childes_aligned.sh
Speech recognition
CHiME-4
CHiME-4 Challenge: Speech Separation and Recognition in Everyday Environments
Safe-first helper
Multichannel Speech Enhancement
Noise Robust Automatic Speech Recognition
Distant Speech Recognition
Microphone Array Generalization
Access pathLDC / licensed
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ldc_user_agreement
- Code license
- Apache-2.0
- License caution
- LDC2017S24 is controlled by the LDC User Agreement for Non-Members or applicable member terms and incorporates CSR-I WSJ0 material. The Apache-2.0 statement on the challenge download page covers the public baseline and supplementary annotations, not the speech and noise recordings. Do not redistribute licensed audio or infer open-data rights from the public documentation.
- Download notes
- The helper saves public challenge, data-layout, LDC-catalog, and recent evaluation documentation only. The official download page says the audio, six-channel annotations, acoustic simulation, and additional enhancement software are distributed through LDC2017S24 as part of the CHiME-3 package. LDC access requires an eligible membership or non-member license; the challenge page also says licensed WSJ0 access is required for the licensed portion. Public one- and two-channel annotations and Apache-2.0 baseline links are documented by the challenge, but they do not replace access to the evaluation audio.
Safe-first helperscripts/download/chime_4.sh
Speech recognition
CHiME-6
CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings
Safe-first helper
Distant Speech Recognition
Multi Speaker Asr
Speaker Diarization
Speech Separation
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR150 lists CC BY-SA 4.0. CHiME says CHiME-6 is a corrected-alignment version of CHiME-5 and recommends CHiME-6 for new work.
- Download notes
- The helper downloads OpenSLR transcriptions, floorplans, and license by default. Audio archives are large, about 97 GiB train, 11 GiB dev, and 12 GiB eval, so they are explicit opt-ins.
Safe-first helperscripts/download/chime_6.sh
Speech recognition
CHiME-7 DASR
The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios
Manual or gated
Distant Automatic Speech Recognition
Speaker Attributed Automatic Speech Recognition
Speaker Diarization
Meeting Transcription
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- mixed_manual_agreements
- Code license
- Apache-2.0
- License caution
- The benchmark is a protocol over three separately controlled corpora, not a single uniformly licensed download. Retain the CHiME, DiPCo, and task-specific LDC/Mixer 6 terms independently. The official baseline is part of ESPnet, whose repository is Apache-2.0 licensed.
- Download notes
- CHiME-7 DASR evaluates one system across revised CHiME-6, DiPCo, and challenge-specific Mixer 6 Speech partitions, ranking submissions by macro-averaged diarization-attributed WER across the three scenarios. The official ESPnet recipe generates the task layout, but it can automatically obtain only DiPCo. CHiME-5/CHiME-6 must be obtained through the CHiME license path, and the task's Mixer 6 release requires a separate LDC evaluation agreement; the challenge warns that this Mixer 6 version differs from LDC2013S03. The helper saves only public task, data, paper, and baseline documentation before printing the manual access steps.
Safe-first helperscripts/download/chime_7_dasr.sh
Audio understanding, generation & events
Clotho
Clotho: An Audio Captioning Dataset
Safe-first helper
Audio Captioning
Language Based Audio Retrieval
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_specified
- License caution
- Zenodo lists rights as Other (Attribution). Audio clips keep their original Freesound licenses, mostly Creative Commons with attribution, recorded in metadata CSVs. Captions are under the Tampere University license, mainly non-commercial with attribution.
- Download notes
- Clotho v2.1 audio archives total about 7.1 GiB; the helper downloads captions/metadata by default and makes audio opt-in.
Safe-first helperscripts/download/clotho.sh
Audio understanding, generation & events
Clotho-Moment
Clotho-Moment: Simulated Long-Audio Dataset for Language-Based Audio Moment Retrieval
Safe-first helper
Language Based Audio Moment Retrieval
Temporal Audio Grounding
Long Audio Retrieval
Audio Text Retrieval
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_on_hugging_face_card_with_upstream_terms
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares Apache-2.0 and the Lighthouse repository includes an Apache-2.0 license. The generated audio incorporates Clotho/Freesound foreground clips and Walking Tours/YouTube background audio, so the card does not erase component recording licenses, attribution requirements, or source-platform terms; review packaged provenance before redistribution or commercial use.
- Download notes
- The public, ungated release contains 51,240 one-minute synthetic English recordings split into 37,930 training, 5,741 validation, and 7,569 test samples. Each sample pairs a text query with a temporal boundary for a Clotho foreground event overlaid at a random interval on Walking Tours background audio. DCASE 2026 Task 6 uses it as a development dataset and advertises a 16.1 GB download. The helper downloads official documentation, license, and repository metadata by default; the audio WebDataset snapshot requires explicit opt-in. Hugging Face currently reports approximately 213 GB of repository storage including history, so users should verify available disk space and select only needed splits or shards.
Safe-first helperscripts/download/clotho_moment.sh
Audio understanding, generation & events
ClothoAQA
Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
Safe-first helper
Audio Question Answering
Audio Language Understanding
Machine Listening
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_applicable
- License caution
- Zenodo lists rights as Other (Attribution). The train/validation/test question-answer CSVs are MIT licensed by Tampere University. Audio files keep per-file Freesound licenses, mostly Creative Commons with attribution, recorded in clotho_aqa_metadata.csv.
- Download notes
- The helper downloads the QA split CSVs, metadata, and license by default. The 3.1 GiB audio_files.zip archive is opt-in because it contains Clotho/Freesound-derived audio.
Safe-first helperscripts/download/clotho_aqa.sh
Enhancement, separation & quality
CMI-RewardBench / CMI-Pref
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
Safe-first helper
Music Reward Model Evaluation
Music Preference Prediction
Music Quality Assessment
Text Music Alignment
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The CMI-Pref card and paper declare CC BY-NC-SA 4.0, and the evaluation repository contains Apache-2.0 code. The paper says some audio was generated through commercial APIs and describes a terms-aware release mechanism; generated-output and service terms may still apply. The composite CMI-RewardBench also incorporates PAM, MusicEval, and Music Arena, so their source licenses and music rights remain controlling for those subsets.
- Download notes
- The public, ungated CMI-Pref release contains 4,027 individual human preference votes over generated music, including a balanced 500-vote test split, 133.8 hours of English/Chinese material, and text, lyrics, reference-audio, musicality, alignment, confidence, and anonymized listener fields. CMI-RewardBench combines that test split with PAM, MusicEval, and Music Arena for music reward-model evaluation. The helper downloads the official cards, repository docs/configuration, the approximately 620 KB CMI-Pref test JSONL, and the approximately 4.8 MB composite test manifest by default. At repository revision 9235ef5c34106b63669958adfaa2eab61bd0e2cd, that manifest has 2,753 held-out rows: 500 CMI-Pref votes, 500 PAM clips, 413 MusicEval clips, and 1,340 filtered Music Arena pairs. The repository's separate 8.1 MB `all_train.jsonl` is baseline-tuning data, not another benchmark split, and the helper intentionally does not fetch it. The Hugging Face CMI-Pref snapshot is pinned by provenance at revision 5282fbe784e326014894b299cb22645cd7d56057. Its API reports approximately 15.0 GB of repository storage, so all MP3 assets require explicit opt-in. A July 2026 full-song generation report uses CMI-Reward as one evaluator on its separate 500-example multilingual test set; that paper does not release those 500 evaluation inputs.
Safe-first helperscripts/download/cmi_rewardbench.sh
Audiovisual & cross-modal
CMMA
CMMA: Benchmarking Multi-Affection Detection in Chinese Multi-Modal Conversations
Safe-first helper
Multimodal Sentiment Analysis
Multimodal Emotion Recognition
Multimodal Sarcasm Detection
Multimodal Humor Detection
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_cc_by_nc_4_0_style_terms
- Code license
- Not specified in the source record.
- License caution
- The repository's English agreement calls itself CC BY-NC 4.0 and permits attributed use, copying, distribution, display, and adapted material only for noncommercial purposes. It is custom license text, not a copy of the standard Creative Commons deed, and it says the original copyright of all television conversations remains with source owners. Review those third-party media rights before use or redistribution.
- Download notes
- The release contains 3,000 Chinese multi-party television conversations and 21,795 multimodal utterances with sentiment, emotion, sarcasm, humor, speaker, topic, and cross-task correlation annotations. The helper downloads the official README, data card, license, annotation guidance, text splits, and repository metadata by default. The public Google Drive CMMA.zip is approximately 12.66 GB and remains a manual download; set CMMA_CLONE_REPO=1 only to clone the approximately 5 MB metadata and text repository. CHARM section 4.1 evaluates the official 3,961-utterance test split.
Safe-first helperscripts/download/cmma.sh
Audiovisual & cross-modal
CMU-MOSEI
CMU-MOSEI: CMU Multimodal Opinion Sentiment and Emotion Intensity
Safe-first helper
Multimodal Sentiment Analysis
Multimodal Emotion Recognition
Audio Sentiment Analysis
Speech Emotion Recognition
+2 more
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The CMU Multimodal SDK and MultiBench repositories are MIT-licensed, but neither repository expressly applies that license to MOSEI annotations, processed features, or source YouTube media. Treat dataset terms as unspecified and review creator, platform, privacy, and media rights before reuse, redistribution, or commercial use.
- Download notes
- The official SDK describes more than 65 hours of annotated YouTube monologue video from more than 1,000 speakers and 250 topics. Each sentence has a sentiment score and six non-exclusive emotion scores for happiness, sadness, anger, surprise, disgust, and fear. The SDK publishes labels plus processed acoustic, visual, and language computational sequences; MultiBench provides an additional word-aligned processed package through Google Drive. The helper saves official documentation, dataset definitions, and repository metadata only. Cloning either toolkit is opt-in, and the Drive package remains a manual download.
Safe-first helperscripts/download/cmu_mosei.sh
Audiovisual & cross-modal
CMU-MOSI
CMU-MOSI: Multimodal Opinion-level Sentiment Intensity Dataset
Safe-first helper
Multimodal Sentiment Analysis
Audio Sentiment Analysis
Subjectivity Analysis
Audio Visual Sentiment Analysis
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- Both official code repositories are MIT-licensed, but neither license expressly grants rights to the MOSI annotations, processed features, or source media. Raw YouTube videos are not redistributed. Treat dataset terms as unspecified and review creator, platform, privacy, and media rights before reuse or redistribution.
- Download notes
- The original paper introduces 2,199 opinion segments from 93 English YouTube review videos with sentiment-intensity, subjectivity, visual, and acoustic annotations. The official CMU Multimodal SDK publishes labels and anonymized processed acoustic, visual, and language computational sequences, but explicitly does not share raw videos because of YouTube creator privacy. MultiBench provides an additional word-aligned processed release through Google Drive. The helper saves official documentation and repository metadata only; cloning either toolkit is opt-in and the Drive package remains a manual download.
Safe-first helperscripts/download/cmu_mosi.sh
Speech generation
CN-NewsTTS Bench
CN-NewsTTS Bench: A Target-Level Automatic Benchmark for Raw-Input Chinese News TTS Pronunciation
Safe-first helper
Chinese Text To Speech Evaluation
Pronunciation Accuracy Evaluation
Text Normalization Evaluation
Target Level Error Analysis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- MIT
- License caution
- CC BY 4.0 covers benchmark data, fixed ASR transcripts, results, documentation, and metadata; MIT covers repository code. The Zenodo audio consists of generated outputs from seven commercial TTS providers and is published as an evaluation artifact. The maintainers warn that reuse may remain subject to each provider or API's terms, so do not treat those audio archives as unrestricted speech-training data without a separate rights review.
- Download notes
- The public v0.1 release contains 200 development records and 800 public-test records with 1,240 auto-evaluable pronunciation targets, fixed transcripts from a three-ASR ensemble, target-level scoring code, and results for seven TTS products. The helper downloads official documentation, licenses, Zenodo metadata, the two small JSONL benchmark splits, schema, scorer, and checksums by default. The approximately 1.58 MB core archive and 1.72 MB full-transcript archive are separate opt-ins. The approximately 425 MB development-audio and 1.74 GB public-test-audio archives require explicit provider-terms acknowledgement and opt-in.
Safe-first helperscripts/download/cn_news_tts_bench.sh
Speaker, identity & emotion
Codec-SUPERB
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
Safe-first helper
Neural Audio Codec Evaluation
Speech Reconstruction
Audio Reconstruction
Music Reconstruction
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The Hugging Face card has no license field or provenance/rights statement. The repository README and badge call the project MIT, but the linked LICENSE file is absent and the GitHub API detects no license. Treat both data and code terms as unspecified until the maintainers publish authoritative license text, and verify the source field and upstream audio rights before reuse.
- Download notes
- The benchmark evaluates whether neural codecs preserve content, paralinguistics, speaker identity, and general audio information through downstream and signal-level metrics. The current official repository uses the public, ungated codec-superb-tiny release for regression runs: 6,000 rows split evenly across speech, audio, and music, with approximately 3.2 GB of downloads. The helper downloads official documentation by default; the dataset snapshot and repository clone are separate opt-ins.
Safe-first helperscripts/download/codec_superb.sh
Speaker, identity & emotion
Codecfake
The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
Safe-first helper
Codec Generated Speech Detection
Audio Deepfake Detection
Synthetic Speech Detection
Unseen Codec Generalization
+2 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0_with_source_corpus_terms
- Code license
- not_specified
- License caution
- Every official Zenodo record and the repository README declare CC BY-NC-ND 4.0 for Codecfake. The GitHub repository has no LICENSE file or detected code license, so the reuse terms for its implementation and checkpoints remain unspecified. VCTK, AISHELL-3, AudioCaps, and the generation systems retain their own terms; review all component conditions before use or redistribution.
- Download notes
- The paper defines a bilingual English/Mandarin family with 1,058,216 distinct real and codec-resynthesized samples derived from VCTK and AISHELL-3. Its C1-C7 protocol tests six seen codecs and one held-out codec; separate A1-A3 tracks test VALL-E, VALL-E X, and AudioGen outputs. The release arranges repeated real evaluation items inside condition directories, so repository directory counts should not be summed as a unique-sample total. The helper saves the official paper, repository documentation, and all six Zenodo metadata records by default. The approximately 172.7 GB release requires explicit opt-in and selection of individual records.
Safe-first helperscripts/download/codecfake.sh
Speech understanding & dialogue
CoDeTT
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
Safe-first helper
Turn Taking
Full Duplex Dialogue
Context Aware Turn Decision
Spoken Dialogue Intent Classification
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_specified
- License caution
- The Hugging Face card declares Apache-2.0 for the dataset, but the GitHub repository has no LICENSE file or detected license. The paper says real samples come from Candor and MagicData-RAMC and synthetic speech uses references from KeSpeech and Emilia; verify those upstream corpus and voice-data terms before redistribution or commercial use.
- Download notes
- The public, ungated release contains more than 300 hours of English and Chinese multi-turn dialogue for four turn-taking actions and 14 fine-grained intent scenarios across system-speaking and system-idle states. It mixes synthetic material with real conversational samples derived from Candor and MagicData-RAMC. The helper downloads official documentation and API metadata by default. The Hugging Face API reports approximately 51.1 GB of repository storage, so the single CoDeTT.lz4 archive requires CODETT_DOWNLOAD_HF=1.
Safe-first helperscripts/download/codett.sh
Audiovisual & cross-modal
CoMind
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
Safe-first helper
Audio Visual Social Reasoning
Joint Attention Estimation
Socially Conditioned Object Interaction Anticipation
Collaborative Handover Prediction
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official project page declares CC BY-NC 4.0, although its displayed license and terms links are currently placeholder anchors rather than a separate terms document. Treat the data as non-commercial, preserve attribution, and review privacy, voice, face, gaze, biometric, and participant-consent implications before reuse. The paper states that participants consented to release of identifiable video, but that does not remove downstream ethical obligations. No separate license notice is embedded in the official Python downloader; the paper itself uses arXiv's perpetual non-exclusive license.
- Download notes
- The official public release covers 41 hours of unscripted cooking collaboration across 80 sessions, with two synchronized egocentric cameras, two exocentric views, audio and WhisperX transcripts, gaze, hand tracking, camera trajectories, scene/object scans, and annotations. Its three benchmarks accept audio or transcribed speech from a ten-second context window for joint-attention estimation, socially conditioned object-interaction anticipation, and collaborative handover prediction. The helper saves the official page, paper metadata, first-party downloader, and annotation manifest by default. The approximately 5.0 MiB annotation JSON files are opt-in; large recording components remain available through the saved official downloader and are never fetched automatically.
Safe-first helperscripts/download/comind.sh
Speech recognition
Common Voice
Mozilla Common Voice
Manual or gated
Asr
Access pathOfficial / other
Upstream termsOpen / attribution signals
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC0-1.0
- Code license
- MPL-2.0
- License caution
- Common Voice data is CC0-1.0; cv-dataset metadata repo is MPL-2.0.
- Download notes
- The helper requires a per-release, per-language URL generated by Mozilla Data Collective and never stores credentials or a private generated URL. Set COMMON_VOICE_FILENAME when the signed URL does not expose a useful archive name.
Safe-first helperscripts/download/common_voice.sh
Audio understanding, generation & events
Common-Sense Facts Audio
Common-Sense Facts Audio: Spoken Fact-Completion Evaluation for Speech Language Models
Safe-first helper
Spoken Fact Completion
Factual Knowledge Retrieval
Cross Modal Pretraining Transfer
Speech Text Interleaving Ablation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_no_custom_terms_published
- Code license
- not_released
- License caution
- The Hub card declares only the generic "other" license tag and provides no custom terms or license file. The article's arXiv license does not license the curated facts, counterfactual annotations, timings, or Kokoro-generated audio. Confirm reuse and redistribution rights with the dataset owners.
- Download notes
- The public Hub release contains one 52,788,867-byte Parquet file with 281 English examples. Each row provides an incomplete prompt, a correct fact, a same-category counterfactual, all three audio renderings, and word-level prompt timings. The paper reports 282 examples across 13 categories, but the card declares 281 and does not identify the omitted item. Likelihood accuracy compares the correct and counterfactual spoken facts; prompt alignments also support layerwise transcription Recall@k analysis. The helper fetches official paper, project, card, API, and file-tree metadata by default; the approximately 52.8 MB Parquet snapshot requires explicit opt-in. The project marks code as coming soon, so no evaluation implementation is currently downloaded.
Safe-first helperscripts/download/common_sense_facts_audio.sh
Speaker, identity & emotion
CompSpoof V2
CompSpoof V2 Dataset
Safe-first helper
Component Level Audio Deepfake Detection
Speech Spoofing Detection
Environmental Sound Spoofing Detection
Audio Deepfake Detection
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_mixed_upstream_terms
- Code license
- not_specified
- License caution
- The Hugging Face card releases CompSpoof V2 under CC BY-NC 4.0 and requires compliance with every source dataset's terms. Components come from MLAAD, VCapAV, LibriTTS, EnvSDD, VGGSound, Common Voice, ASVspoof 5, English Conversation Corpus, AudioCaps, TUT acoustic-scene and sound-event data, and UrbanSound; those sources include CC0, attribution, noncommercial, ODC-By, GPL, and custom terms. The authors disclaim ownership of original audio. The baseline repository exposes no detected license, so do not infer code rights from the data card.
- Download notes
- The gated author-owned Hugging Face release contains 255,433 four-second clips (approximately 283 hours) across train, validation, evaluation, and test splits. Its five labels distinguish original audio and every bona fide/spoofed combination of speech and environmental components; evaluation and test add unseen generated speech, generated environmental audio, and codec transformations. The helper downloads public documentation, API metadata, paper pages, and baseline-repository metadata by default. The Hugging Face API reports approximately 130.0 GB of repository storage, so the snapshot requires login, license acknowledgement, and both COMPSPOOF_V2_ACK_TERMS=1 and COMPSPOOF_V2_DOWNLOAD_HF=1.
Safe-first helperscripts/download/compspoof_v2.sh
Audio understanding, generation & events
Concerto Accompaniment Benchmark
Concerto Accompaniment Benchmark for Score-Free Piano Concerto Accompaniment
Safe-first helper
Music Audio Alignment
Automatic Accompaniment
Score Free Accompaniment Generation
Downbeat Alignment
+1 more
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- MIT covers the public repository's code and released annotation/config files. It does not license the absent Music Minus One recordings or override the per-recording IMSLP terms listed in AudioDataSummary.csv, which include CC0, several CC variants, and public-domain status that may differ by jurisdiction. The repository does not state a license or public delivery path for the recorded solo-piano performances.
- Download notes
- The paper defines 150 alignment scenarios over four concerto movements, combining four recorded solo-piano performances, four commercial Music Minus One orchestra tracks, and eight IMSLP piano-orchestra mixes. The public repository provides the evaluation code, configuration tables, IMSLP source URLs, and measure-downbeat annotations. The helper saves those lightweight released artifacts by default and makes the repository clone opt-in. The commercial orchestra recordings are explicitly private and must be purchased separately. Although the paper calls the remaining data open source, the repository currently ignores audio and exposes no solo-piano recordings; do not infer a public audio download.
Safe-first helperscripts/download/concerto_accompaniment_benchmark.sh
Representation & general suites
COREFL
COREFL: Corpus of English as a Foreign Language
Safe-first helper
L2 Speaking Proficiency Assessment
Learner Speech Recognition
Spoken Written Modality Comparison
Cefr Level Prediction
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-ND-3.0-ES
- Code license
- not_applicable
- License caution
- The official terms require attribution, prohibit commercial use, and prohibit redistribution of remixed, transformed, or built-upon material. They additionally prohibit analysis or manipulation that could jeopardize confidentiality. Treat learner audio, transcripts, proficiency scores, L1, age, sex, and other participant metadata as human-subject data; minimize exports and do not use the helper to automate collection of participant records.
- Download notes
- COREFL v2.0 is an owner-hosted learner corpus with written and spoken English responses, manual speech transcriptions, CEFR placement scores, L1 and demographic metadata, task prompts, and native control subcorpora. The live statistics page reports 894 spoken documents and 264,306 spoken words across learner and native-control material; paired spoken/written responses from the same participant and prompt are a central design feature. The 2026 TTS-augmentation study uses a 510-response, 20.9-hour spoken subset, but does not release its exact split or generated derivatives. The helper saves official documentation and the paper only; filtered corpus exports require choosing criteria in the official interface and describing intended use.
Safe-first helperscripts/download/corefl.sh
Speech generation
CoVoMix2 Dialogue
CoVoMix2 Dialogue Test Set
Safe-first helper
Zero Shot Dialogue Synthesis
Multi Speaker Speech Generation
Overlapping Speech Generation
Speaker Consistency Evaluation
+1 more
Access pathOpenSLR
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_with_upstream_terms
- Code license
- not_applicable_or_not_specified
- License caution
- The test-set repository has no license file or GitHub-detected license. Its dialogue text derives from DailyDialog, whose authors distribute it under CC BY-NC-SA 4.0, and its acoustic prompts reference LibriSpeech test-clean, released under CC BY 4.0. Those upstream terms do not establish a license for the CoVoMix2 selection or annotations; seek clarification before redistribution or commercial use.
- Download notes
- The public repository contains 1,000 DailyDialog-derived dialogue transcript files and a 549 KB JSON manifest. Each row identifies two LibriSpeech test-clean prompt paths and supplies their transcriptions; prompt audio is not redistributed and must be obtained separately from the official LibriSpeech release. The helper downloads the paper, manifest, and GitHub metadata by default; cloning the small repository with all transcript files is explicit opt-in.
Safe-first helperscripts/download/covomix2_dialogue.sh
Speech recognition
CoVoST 2
CoVoST 2: Massively Multilingual Speech-to-Text Translation Corpus
Safe-first helper
Speech To Text Translation
Speech Translation
Multilingual Asr
Machine Translation
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Meta/GitHub list CoVoST data as CC0; the current Hugging Face card lists CC BY-NC 4.0, and current Mozilla Data Collective packages for CoVoST 2/Common Voice segments may be CC BY-NC 4.0.
- Code license
- cc-by-nc-4.0
- License caution
- Treat packaged mirrors conservatively and re-check the active source before redistribution. The GitHub license table also says Tatoeba evaluation sentences are CC BY 2.0 FR and Tatoeba speech has per-row licenses.
- Download notes
- The helper downloads the official CoVoST 2 translation TSV archives and split-generation script. CoVoST 2 rows match Common Voice 4 validated.tsv entries, so users must obtain Common Voice audio separately under the applicable upstream terms.
Safe-first helperscripts/download/covost2.sh
Audiovisual & cross-modal
CREMA-D
CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset
Safe-first helper
Speech Emotion Recognition
Audio Visual Emotion Recognition
Acted Emotional Speech
Crowd Sourced Emotion Annotation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- odbl-1.0
- Code license
- not_specified
- License caution
- The official README/LICENSE say the database is under the Open Database License 1.0 and individual contents are under the Database Contents License 1.0. GitHub reports license as NOASSERTION, so keep the explicit upstream text as authority.
- Download notes
- The helper downloads small README/license/CSV metadata by default. Full audio and video live in Git LFS and require about 7.55 GiB for a complete clone; upstream asks repository users to fill out the access/community form.
Safe-first helperscripts/download/crema_d.sh
Representation & general suites
Cross-Era
Cross-Era Dataset for Western Classical Music Style Analysis
Safe-first helper
Classical Music Style Classification
Historical Style Analysis
Global Key Detection
Chord Analysis
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The AudioLabs page requests citation but states no explicit license for its downloadable annotations, chord sequences, or chroma features. The supplementary Hugging Face card declares Apache 2.0 for its work-year metadata. The recent GitHub repository has no detected license and intentionally excludes datasets, models, embeddings, and outputs. Do not infer rights to the absent commercial recordings or use the supplementary metadata license to override upstream terms.
- Download notes
- Cross-Era contains 2,000 commercial classical-music recordings, balanced across piano and orchestra and five historical groupings. Because the recordings cannot be redistributed, the official release provides a 2,000-row annotation CSV, approximately 1 MB of extracted Chordino chord sequences, and an optional approximately 244 MB NNLS chroma archive instead. Expert global-key labels cover 1,200 Baroque, Classical, and Romantic works. A 2026 study adds an approximately 0.9 MB public work-year table with source URLs and uses all 2,000 works as a held-out representation target before evaluating chronological forecasts on a 480-work 1875-1940 subset. The helper fetches the official page, annotations, chord features, work-year metadata, and source cards by default; chroma features require explicit opt-in and commercial audio is never requested.
Safe-first helperscripts/download/cross_era.sh
Speech generation
CV3-Eval
CV3-Eval: CosyVoice 3 in-the-wild zero-shot speech synthesis benchmark
Safe-first helper
Zero Shot Text To Speech
Multilingual Voice Cloning
Cross Lingual Voice Cloning
Emotion Cloning
+6 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0 repository license; mixed upstream media rights
- Code license
- Apache-2.0
- License caution
- The repository root applies Apache-2.0, but the README says reference speech comes from Common Voice, FLEURS, EmoBox, and web-crawled real-world audio. Treat the repository license as insufficient to clear every source recording, and verify component provenance and rights before redistribution or commercial use.
- Download notes
- The official repository includes objective multilingual, cross-lingual, and emotion-cloning subsets plus subjective expressive, continuation, and Chinese-accent subsets. Qwen3.5-Omni section 5.2.3 calls the public cross-lingual subset both CV3-Eval and the Cross-Lingual benchmark, and reports mixed error rate over 12 source-target directions among Chinese, English, Japanese, and Korean. Qwen-Audio-3.0-TTS section 4.3 reuses those protocols and reports seven additional multilingual languages, but does not release a separate manifest or archive for that extension. The helper downloads the README and Apache-2.0 license by default; cloning the roughly 760 MiB repository, including evaluation audio and bundled scoring utilities/models, requires CV3_EVAL_CLONE_REPO=1.
Safe-first helperscripts/download/cv3_eval.sh
Audiovisual & cross-modal
DAIC-WOZ / E-DAIC
Distress Analysis Interview Corpus Wizard-of-Oz and Extended DAIC
Manual or gated
Speech Based Depression Detection
Psychological Distress Assessment
Multimodal Depression Classification
Clinical Interview Analysis
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- application_required_terms_not_publicly_specified
- Code license
- not_applicable
- License caution
- The owner page makes both clinical interview releases available upon request but does not publish reusable data-license terms on the public landing or application pages. Approval does not imply redistribution or commercial-use rights. The corpus contains identifiable voices, video, facial features, questionnaire responses, and mental-health labels; follow the approved agreement and institutional privacy, consent, security, and ethics requirements.
- Download notes
- USC ICT treats DAIC-WOZ and its extended E-DAIC release as one corpus lineage with separate application forms. DAIC-WOZ contains 189 Wizard-of-Oz interview sessions lasting 7–33 minutes, with participant audio, transcripts, facial features, and questionnaire responses. E-DAIC extends the corpus used by the AVEC 2019 depression challenge. The helper saves only public owner documentation and primary paper pages, then prints both application routes; it never submits a form, authenticates, or downloads sensitive interview data. The July 2026 LIWC study evaluates 189 DAIC-WOZ and 275 E-DAIC participants using pooled challenge labels under participant-level cross-validation, not the official AVEC test splits.
Safe-first helperscripts/download/daic_woz.sh
Audiovisual & cross-modal
Daily-Omni
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
Safe-first helper
Audio Visual Question Answering
Temporal Alignment
Cross Modal Reasoning
Audio Visual Event Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- GPL-3.0
- License caution
- The paper and Hugging Face card declare CC BY-NC-SA 4.0 for the benchmark, and GitHub reports GPL-3.0 for the repository. Videos are sampled from AudioSet, Video-MME, and FineVideo, so their upstream media rights and terms also require review.
- Download notes
- The public, ungated release contains 684 real-world videos and 1,197 English multiple-choice questions across six temporal audio-visual reasoning tasks. The helper downloads official documentation and qa.json by default; the Hugging Face API reports approximately 3.9 GB of storage, so Videos.tar requires DAILY_OMNI_DOWNLOAD_HF=1. Qwen3.5-Omni reports DailyOmni in section 5.1.4, Table 7.
Safe-first helperscripts/download/daily_omni.sh
Enhancement, separation & quality
DAPS
Device and Produced Speech Dataset
Safe-first helper
Speech Enhancement
Device And Environment Robustness
Professional Speech Production
Voice Conversion
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC_BY_NC_4_0
- Code license
- not_specified
- License caution
- The Zenodo record declares CC BY-NC 4.0 for the corpus. The public AudioCrossVerification repository has no license file and GitHub reports no detected license, so its notebooks and code must not be assumed reusable under the dataset license. Recorded human voices also warrant appropriate privacy and research-ethics handling.
- Download notes
- The helper downloads the owner page, Zenodo metadata, and public cross-verification repository documentation by default. Cloning the approximately 1.7 MiB notebook repository is optional. The single approximately 14.95 GiB daps.tar.gz corpus archive is a separate explicit opt-in and is not downloaded during metadata inspection.
Safe-first helperscripts/download/daps.sh
Audio understanding, generation & events
DataSED
DataSED: Dataset for Sound Event Detection of Environmental Noise
Safe-first helper
Sound Event Detection
Polyphonic Sound Event Detection
Monophonic Sound Event Detection
Temporal Audio Event Localization
+2 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- cc-by-4.0
- License caution
- Zenodo assigns CC BY-NC-SA 4.0 to the dataset archive. The separately published labeling tool is CC BY 4.0, while the Scientific Data descriptor is CC BY-NC-ND 4.0. The dataset combines field measurements and material from online repositories and says some sounds were manually added; review source-recording attribution and rights before redistribution rather than treating the compilation license as overriding upstream terms.
- Download notes
- The current public, ungated Zenodo record describes 717 mono 44.1 kHz WAV recordings totaling approximately 17.02 hours, with expert timestamp annotations for monophonic and polyphonic detection across 22 environmental classes. Its single archive is approximately 4.20 GiB, so the helper saves official metadata and provenance pages by default and requires DATASED_DOWNLOAD_AUDIO=1 for the archive. The 2025 data descriptor describes 712 files, while the updated Zenodo record and archive description say 717; use the downloaded release manifest as authoritative for the current version. The July 2026 active-learning paper creates 6,808 ten-second segments and a 5,446/681/681 split, but does not release its exact split indices.
Safe-first helperscripts/download/datased.sh
Audiovisual & cross-modal
DAVE
DAVE: Diagnostic Benchmark for Audio Visual Evaluation
Safe-first helper
Audio Visual Alignment
Multimodal Synchronization
Sound Absence Detection
Sound Discrimination
+3 more
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT label with upstream dataset terms
- Code license
- MIT stated in README; no repository LICENSE file
- License caution
- The Hugging Face card labels DAVE as MIT and the repository README says everything is MIT, but the GitHub repository has no LICENSE file. DAVE is built on EPIC-KITCHENS and Ego4D, and its card says it inherits their risks; verify both upstream datasets' access and media terms before redistribution or commercial use.
- Download notes
- The public, ungated Hugging Face release has EPIC-KITCHENS- and Ego4D-derived splits with seven diagnostic task views. The helper downloads the official cards, loader, and approximately 9 MB of JSON annotations by default. Media archives are excluded because the Hugging Face API reports about 113.3 GB of repository storage; DAVE_DOWNLOAD_HF=1 explicitly opts into the full snapshot.
Safe-first helperscripts/download/dave.sh
Audio understanding, generation & events
DCASE 2020 Task 2 ASD
DCASE 2020 Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Unsupervised Anomaly Detection
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- MIT
- License caution
- Each official Zenodo record declares CC BY-NC-SA 4.0, so commercial use is not authorized and adaptations must retain share-alike terms. The baseline repository's MIT license covers software, not benchmark audio. Preserve the packaged ToyADMOS and MIMII attribution and source notices.
- Download notes
- This is one challenge benchmark family assembled from subsets of ToyADMOS and MIMII, not a replacement release of either full source corpus. It covers toy car, toy conveyor, valve, pump, fan, and slide rail sounds. The development set supplies normal training audio and labeled normal/anomalous test audio; the additional set supplies normal training audio for the evaluation machine IDs; and the evaluation set contains the corresponding test audio. All recordings are approximately ten-second, single-channel, 16 kHz clips. The three official releases total approximately 13.23 GiB. The helper downloads official task, record, paper, and baseline metadata by default; archives require an explicit opt-in and part list.
Safe-first helperscripts/download/dcase2020_task2_asd.sh
Audio understanding, generation & events
DCASE 2021 Task 2 ASD
DCASE 2021 Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
Few Shot Domain Adaptation
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- mit
- License caution
- All three official Zenodo records declare CC BY-NC-SA 4.0, which prohibits commercial use and requires attribution and share-alike distribution. The official autoencoder baseline repository contains an MIT LICENSE. Preserve packaged attribution and machine/source provenance when redistributing derived material.
- Download notes
- This public challenge family introduced domain-shifted anomalous sound detection with seven machine types and three development sections plus three disjoint additional-training/evaluation sections per type. Each development section has about 1,000 normal source-domain and only three normal target-domain training clips, plus roughly 100 normal and 100 anomalous test clips in each domain. The final evaluation audio is public, but its normal/anomaly condition labels are withheld under the challenge protocol. The three official releases total approximately 15.25 GB. The helper downloads official task, record, and baseline metadata by default; archives require an explicit environment opt-in and part list.
Safe-first helperscripts/download/dcase2021_task2_asd.sh
Audio understanding, generation & events
DCASE 2022 Task 2 ASD
DCASE 2022 Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- mit
- License caution
- All three official Zenodo records declare CC BY 4.0. The official autoencoder baseline repository contains an MIT LICENSE. Preserve packaged attribution and machine/source provenance when redistributing derived material.
- Download notes
- The public domain-shift protocol covers seven machine types, with three development sections and three disjoint additional-training/evaluation sections per type. Each development section provides 990 normal source clips and ten normal target clips for training, plus 100 normal and 100 anomalous test clips; the final evaluation audio is public but its condition/domain labels remain hidden under the challenge protocol. The three official Zenodo releases total approximately 11.58 GiB. The helper downloads official task, record, and baseline metadata by default; archives require an explicit environment opt-in and part list. UD-ASD evaluates the development set in arXiv:2607.12576.
Safe-first helperscripts/download/dcase2022_task2_asd.sh
Audio understanding, generation & events
DCASE 2023 Task 2 ASD
DCASE 2023 Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
First Shot Learning
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- custom_dcase_challenge_license_v2.1
- License caution
- All three official Zenodo records declare CC BY 4.0. The current baseline repository points to a DCASE Challenge License v2.1 PDF, and the evaluator contains a license PDF with no SPDX-detected license; review those custom code terms separately from the dataset license.
- Download notes
- This is one first-shot challenge family, not three independent datasets. Seven development machine types are disjoint from the final additional-training/evaluation types, and each type has one section. Development sections provide 990 normal source-domain and ten normal target-domain training clips plus 100 normal and 100 anomalous test clips. Each final evaluation type has 200 test clips; the organizers subsequently released labels and an evaluator. The three official records total approximately 4.55 GB. The helper downloads official task, record, repository, and baseline-paper metadata by default; archives require an explicit opt-in and part list.
Safe-first helperscripts/download/dcase2023_task2_asd.sh
Audio understanding, generation & events
DCASE 2024 Task 2 ASD
DCASE 2024 Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
First Shot Learning
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_specified
- License caution
- All three official Zenodo records declare CC BY-NC-SA 4.0, which prohibits commercial use and requires attribution and share-alike distribution. The evaluator repository has no LICENSE file and GitHub reports no detected license, so its code terms remain unspecified.
- Download notes
- The public first-shot protocol uses seven development machine types and nine disjoint final-evaluation types, one section per type, with only normal source/target-domain audio available for training. The three official records contain approximately 2.24 GB of development, 1.99 GB of additional-training, and 398 MB of evaluation archives. Each final type has 200 test clips, and the organizers have released labels and an evaluator. The helper downloads official task, record, and evaluator metadata by default; archives require an explicit environment opt-in and part list. A July 2026 pseudo-label-distillation paper evaluates this release as part of its DCASE 2020-2025 Task 2 sweep.
Safe-first helperscripts/download/dcase2024_task2_asd.sh
Audio understanding, generation & events
DCASE 2024 Task 5
DCASE 2024 Task 5: Few-shot Bioacoustic Event Detection
Safe-first helper
Few Shot Bioacoustic Event Detection
Sound Event Detection
Animal Vocalization Detection
Five Shot Learning
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Both official Zenodo records declare Creative Commons Attribution 4.0 International. The benchmark combines multiple bioacoustic sources; retain the release attribution and review source-specific ethical or wildlife-recording constraints for downstream use.
- Download notes
- The official five-shot protocol provides the first five positive target events in each recording, then scores detection after the fifth event. The 2024 development release has 217 recordings: 174 training files covering 47 classes and 43 validation files covering seven classes. The official challenge reuses the 2023 evaluation release, which has 66 recordings across eight subsets. The helper downloads record metadata, class maps, and annotation-only archives by default. The current Zenodo files total approximately 20.4 GiB for development audio and 3.0 GiB for evaluation audio, so waveform archives require DCASE2024_TASK5_DOWNLOAD_AUDIO=1.
Safe-first helperscripts/download/dcase2024_task5.sh
Audio understanding, generation & events
DCASE 2024 Task 7 Sound Scene Synthesis
DCASE 2024 Task 7: Sound Scene Synthesis
Safe-first helper
Text To Audio Generation
Environmental Sound Scene Synthesis
Compositional Audio Generation
Audio Generation Quality Evaluation
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo declares CC BY 4.0 for the open-source dataset. The record says its audio is sourced from Freesound, so retain packaged source attribution and review per-clip provenance. The baseline repository has no LICENSE file and the task page states that it is mostly derived from the upstream AudioLDM repository; treat code terms as unspecified pending clarification.
- Download notes
- The public open-source release contains 310 manually composed four-second environmental sound scenes and corresponding structured text prompts. It uses only Freesound source audio and excludes the proprietary/private libraries present in the challenge reference data. The original protocol evaluates Fréchet Audio Distance with PANNs CNN14 Wavegram-Logmel embeddings plus listening tests for foreground fit, background fit, and audio quality. The challenge's 250 evaluation prompts and reference audios remain secret. The helper downloads official metadata and task documentation by default; the approximately 140 MiB public archive is an explicit opt-in.
Safe-first helperscripts/download/dcase2024_sound_scene_synthesis.sh
Enhancement, separation & quality
DCASE 2024 Task 9 LASS
DCASE 2024 Task 9: Language-Queried Audio Source Separation
Safe-first helper
Language Queried Audio Source Separation
Text Conditioned Audio Separation
Universal Sound Separation
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- All three official Zenodo records declare CC BY 4.0. Development audio remains in FSD50K and Clotho v2 under their own mixed or upstream terms, while validation/evaluation audio derives from Freesound; retain record attribution and inspect packaged per-clip provenance. The baseline repository has no LICENSE file or detected GitHub license and is largely derived from AudioSep, so its code terms are unspecified.
- Download notes
- The development release contains GPT-4-generated captions for the existing FSD50K development and evaluation clips; participants obtain the source FSD50K and Clotho v2 audio separately. The public validation release contains 3,000 synthetic mixtures built from 1,000 source clips with three captions per source. The evaluation release contains 3,000 additional synthetic mixtures plus 100 real overlapping Freesound clips annotated with two source queries each. Synthetic examples are scored with SDR; the real set uses listening tests for query relevance and overall quality. The helper downloads official task documentation, record metadata, and lightweight JSON/CSV annotations by default; approximately 1.14 GB of validation and evaluation audio is an explicit opt-in.
Safe-first helperscripts/download/dcase2024_lass.sh
Audio understanding, generation & events
DCASE 2025 Task 2 ASD
DCASE 2025 Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
First Shot Learning
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- custom_dcase_challenge_license_v2.1
- License caution
- All three official Zenodo records declare CC BY-NC-SA 4.0, which prohibits commercial use and requires attribution and share-alike distribution. The evaluator repository includes a DCASE Challenge License v2.1 PDF rather than an SPDX-style open-source license; review it before reuse.
- Download notes
- The public first-shot protocol trains only on normal machine sounds, tests source/target-domain generalization, and uses different machine types for development versus final evaluation. The development record has seven machine types and approximately 2.36 GB of archives; the approximately 1.98 GB additional-training record and 358 MB evaluation record cover eight different types. Each evaluation section has 200 test clips, and the organizers have released labels and an evaluator. The helper downloads official task, record, and evaluator metadata by default; archives require DCASE2025_TASK2_DOWNLOAD_ARCHIVES=1 and an explicit part list.
Safe-first helperscripts/download/dcase2025_task2_asd.sh
Audio understanding, generation & events
DCASE 2025 Task 5 AudioQA
DCASE 2025 Task 5: Multi-Domain Audio Question Answering
Manual or gated
Audio Question Answering
Temporal Audio Reasoning
Bioacoustic Question Answering
Multiple Choice Question Answering
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- mit_on_hugging_face_card
- Code license
- unspecified
- License caution
- The Hugging Face card metadata declares MIT, but the benchmark incorporates audio from Watkins Marine Mammal Sound Database, AudioSet, Mira, and other sources named by the organizers. Upstream audio and source-platform terms may be narrower and still apply; confirm them before redistribution or commercial use. No separate code license was found for the release scripts.
- Download notes
- The official English multiple-choice benchmark combines Bioacoustics QA, Temporal Soundscapes QA, and Complex QA (MMAU). The challenge page reports approximately 8.1K training and 2.4K development question-answer pairs; the released repository also includes the evaluation set. Access is public but auto-approved gated: users must sign in and provide basic identity and affiliation fields. The helper saves public challenge, paper, and repository API metadata, then prints the manual acceptance and authenticated download steps; it never downloads audio automatically.
Safe-first helperscripts/download/dcase2025_audioqa.sh
Audio understanding, generation & events
DCASE 2026 Task 1 HAC
DCASE 2026 Task 1: Heterogeneous Audio Classification
Safe-first helper
Heterogeneous Audio Classification
Hierarchical Audio Classification
Multimodal Audio Classification
Domain Generalization
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- All three official Zenodo records declare CC BY 4.0. The releases include per-sound Freesound license and uploader provenance, so users must also honor each source recording's terms. The official baseline repository has no detected license or LICENSE file; its code terms are unspecified.
- Download notes
- The task predicts 23 second-level Broad Sound Taxonomy categories and scores macro-averaged hierarchical F-score. Development uses the curated BSD10k-v1.2 release (about 11,000 sounds and 35 hours) and the noisier crowd-sourced BSD35k-CS release (about 35,000 sounds and 150 hours), both with text metadata and Freesound provenance. The public evaluation archive contains audio and metadata but intentionally omits labels. The helper downloads official pages, Zenodo records, READMEs, and approximately 7 MB of development metadata by default. Roughly 200 MB of CLAP features and 47 GB of audio/evaluation archives require separate explicit opt-ins.
Safe-first helperscripts/download/dcase2026_task1_hac.sh
Audio understanding, generation & events
DCASE 2026 Task 2 ASD
DCASE 2026 Task 2: Noise-Aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
First Shot Learning
+2 more
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- custom_dcase_challenge_license_v2.1
- License caution
- All three official Zenodo records declare CC BY-NC-SA 4.0, which prohibits commercial use and requires attribution and share-alike distribution. The shared baseline repository reports a non-standard DCASE Challenge License rather than an SPDX open-source license; review its terms separately.
- Download notes
- The public noise-aware first-shot protocol uses synchronized two-channel recordings: a near microphone emphasizes the target machine and a far microphone provides an environmental-noise reference. Development data contains seven machine types with normal training and labeled normal/anomalous test clips. Five different real-machine types appear in the additional-training and evaluation releases; evaluation labels are hidden by the challenge protocol. The three records total about 8.16 GB. The helper downloads official task, paper, repository, and Zenodo metadata by default; archives require an explicit environment opt-in and part list.
Safe-first helperscripts/download/dcase2026_task2_asd.sh
Audiovisual & cross-modal
DCASE2025 Task 3 Stereo SELD Dataset
DCASE2025 Task 3 Stereo Sound Event Localization and Detection Dataset
Safe-first helper
Sound Event Localization And Detection
Stereo Sound Source Localization
Audio Visual Sound Event Localization
Source Distance Estimation
+2 more
Access pathZenodo
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT
- Code license
- mixed
- License caution
- Zenodo metadata and the included LICENSE identify the dataset as MIT. Sony's data-generator repository is MIT. The official baseline repository has no detected GitHub license or LICENSE file, so its code terms are not specified. Because the release is derived from STARSS23 recordings of people and rooms, review the original privacy/provenance context before sensitive visual use.
- Download notes
- The public Zenodo v1.1.0 release contains 30,000 labeled development clips (41.7 hours) and 10,000 unlabeled evaluation clips (13.9 hours), each five seconds long, with 24 kHz stereo audio and aligned perspective video. It is derived from STARSS23 by sampling and converting its FOA audio and 360-degree video, and adds folded azimuth, source-distance, and onscreen/offscreen labels. The helper downloads official pages, record metadata, README, license, paper page, and generator documentation by default; the approximately 15.2 MB label archive is opt-in, while the approximately 27.6 GB audio/video release remains on Zenodo.
Safe-first helperscripts/download/dcase2025_stereo_seld.sh
Enhancement, separation & quality
DEMAND
DEMAND: Diverse Environments Multichannel Acoustic Noise Database
Safe-first helper
Speech Enhancement
Speech Denoising
Noise Robustness Evaluation
Multichannel Audio Processing
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-3.0
- Code license
- Not specified in the source record.
- License caution
- The record's structured license field says CC BY 4.0, but its owner-written description explicitly applies CC BY-SA 3.0 Unported to the audio and accompanying document. Apply the more specific attribution and share-alike terms pending clarification rather than treating the generic metadata field as controlling.
- Download notes
- The owner-authored Zenodo release provides real-world environmental noise as 16 synchronized single-channel WAV files per scene at both 16 kHz and 48 kHz. Its prose says this version contains 15 recordings, while the current file inventory exposes 18 named scenes and 36 scene archives; preserve that source inconsistency rather than silently choosing one count. The helper downloads official metadata and the approximately 86 KB technical description by default. Set DEMAND_SCENE to one documented scene and DEMAND_SAMPLE_RATE to 16k or 48k to fetch one checksum-published archive; the complete release is approximately 7.36 GB. Automatic Audio Equalization section 4.3 uses DEMAND at -5 to 30 dB SNR for a paper-specific 1,000-example noisy speech test, but does not release the exact derived test manifest or degradations.
Safe-first helperscripts/download/demand.sh
Speaker, identity & emotion
DementiaBank Pitt Corpus
DementiaBank English Pitt Corpus
Safe-first helper
Alzheimers Dementia Classification
Cognitive Impairment Detection
Speaker Privacy Evaluation
Speech Anonymization Evaluation
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-3.0_with_password_protected_clinical_access
- Code license
- not_applicable
- License caution
- TalkBank's default CC BY-NC-SA 3.0 terms prohibit commercial products and incorporation into large language models. Password-protected clinical data may not be shared with non-members or posted elsewhere. Users must preserve confidentiality, follow the TalkBank ethics and citation rules, and use non-retaining services if web processing is necessary. The recordings contain sensitive, potentially identifiable clinical speech.
- Download notes
- The official corpus index describes dementia and control recordings across four language tasks with transcripts and media. Access is password protected through DementiaBank membership; established researchers and clinicians may request access, while students require faculty sponsorship. The helper saves only public access, corpus-index, corpus-description, and ground-rules pages, then prints the manual membership path. It never authenticates or downloads clinical data.
Safe-first helperscripts/download/dementiabank_pitt.sh
Audio understanding, generation & events
DESED
Domestic Environment Sound Event Detection Dataset
Safe-first helper
Sound Event Detection
Audio Tagging
Weakly Supervised Sound Event Detection
Synthetic Soundscape Generation
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- MIT
- License caution
- Zenodo records for DESED real and synthetic list CC BY 4.0. The GitHub README says the Python code is MIT and that component datasets include license files at their roots; source media comes from AudioSet/YouTube, Freesound, MUSAN, SINS, and related sources, so re-check component terms before redistribution.
- Download notes
- The helper downloads the official repo plus Zenodo record JSON and small metadata/JAMS files by default. Real and synthetic audio archives are multi-GB and require explicit opt-in flags.
Safe-first helperscripts/download/desed.sh
Speech generation
Designed Vocalizations Dataset
Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion
Safe-first helper
Non Human Voice Conversion
Designed Vocalization Generation
Timbre Transfer
Sound Design Reproduction
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0-with-mixed-upstream-terms
- Code license
- not_applicable
- License caution
- The dataset card applies CC BY 4.0 to the compilation and authors' original metadata and annotations. Raw and designed clips retain their applicable source terms: VCTK and HiFi-TTS are CC BY 4.0, while Freesound clips are individually CC0 1.0, CC BY 3.0, or CC BY 4.0. Preserve the release NOTICE and per-row license, attribution, creator, and source fields when redistributing. The project does not publish a separate evaluation-code repository.
- Download notes
- The public, ungated release contains 237,574 mono 44.1 kHz WAV clips embedded in Parquet: 5,654 raw training sources, 226,160 non-parallel designed training clips, 120 test sources, and 5,640 aligned test references. The test protocol crosses source timbres seen or unseen during training with 40 seen and seven unseen effect presets. The helper downloads official documentation, API metadata, licensing notices, preset metadata, and the approximately 532 KB test-pair manifest by default. The full Hugging Face repository is approximately 37.1 GB and requires DESIGNED_VOCALIZATIONS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/designed_vocalizations.sh
Speaker, identity & emotion
DFADD
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
Safe-first helper
Audio Deepfake Detection
Synthetic Speech Detection
Diffusion Tts Detection
Flow Matching Tts Detection
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit_metadata_with_incorporated_source_terms
- Code license
- mit
- License caution
- The Hugging Face card and repository declare MIT, while the repository separately identifies VCTK bona fide speech as CC BY 4.0 and LJ Speech text as public domain. The generated subsets also depend on five official or unofficial TTS implementations and checkpoints. Do not assume the blanket MIT tag overrides source-corpus, voice, model, checkpoint, or generated-output terms; review every component before redistribution or commercial use.
- Download notes
- DFADD is one English audio-deepfake family spanning five diffusion or flow-matching TTS generators: Grad-TTS, NaturalSpeech 2, PFlow-TTS, StyleTTS 2, and Matcha-TTS. The paper reports 163,500 spoofed utterances paired with 44,455 VCTK bona fide utterances from 109 speakers. The current Hugging Face dataset-viewer snapshot contains 207,955 examples in train, validation, and test splits and reports approximately 28.6 GB of downloads. The repository says its separately packaged ZIP files were corrected in April 2025 for a Matcha-TTS audio/label mismatch and instructs users to merge VCTK_BONAFIDE into each generated subset. The helper downloads only the paper page, first-party documentation, license, and API metadata by default. The complete viewer snapshot or corrected ZIP packages require separate explicit opt-ins.
Safe-first helperscripts/download/dfadd.sh
Audio understanding, generation & events
DHAuDS
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
Safe-first helper
Test Time Adaptation
Audio Classification Robustness
Dynamic Corruption Robustness
Speech Command Classification
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_on_hugging_face_cards_with_upstream_terms
- Code license
- apache-2.0
- License caution
- All four Hugging Face cards declare Apache-2.0 and the code repository contains an Apache-2.0 LICENSE. The corrupted audio derives from Speech Commands V2, VocalSound, UrbanSound8K, ReefSet, QUT-NOISE, and DEMAND; those sources retain separate attribution, share-alike, non-commercial, or other terms. In particular, the Apache card labels should not be assumed to remove UrbanSound8K's non-commercial restriction or other upstream obligations.
- Download notes
- The public suite contains separately corrupted adaptation and evaluation sets derived from the held-out portions of Speech Commands V2, VocalSound, UrbanSound8K, and ReefSet. SC2-C, VS-C, and RS-C apply seven corruption categories at two severity levels; US8-C omits QUT-NOISE and DEMAND corruptions that overlap its target classes and uses four categories. The paper reports 908,196 derived samples in total and uses different random seeds for adaptation and evaluation corruptions. The helper downloads official documentation and repository metadata by default. The four Hugging Face repositories report about 50.0 GB of storage combined, so snapshots require DHAUDS_DOWNLOAD_HF=1 and an explicit DHAUDS_DATASETS selection.
Safe-first helperscripts/download/dhauds.sh
Speech recognition
Dialogs
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus for Dialog Assistants
Safe-first helper
Expressive Text To Speech
Conversational Text To Speech
Speech Emotion Classification
Automatic Speech Recognition
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- OpenRAIL
- Code license
- MIT
- License caution
- The dataset card links a custom OpenRAIL responsible-use license that permits use, modification, redistribution, and commercial use subject to its use-based restrictions. The paper and card state that performers gave written informed consent for public and commercial use. Review the complete LICENSE.md rather than treating OpenRAIL as an unrestricted permissive license. The linked VITS2 baseline repository is MIT.
- Download notes
- The public, ungated release contains 20.6 hours and 11,796 Russian utterances from face-to-face acted dialogues recorded in a studio by three professional performers. It provides transcripts, stress-marked text, speaker identifiers, and 12 style/emotion labels, with fixed 11,428/180/188 train, validation, and test splits. The paper evaluates the 188-item stratified test subset with six human-rated quality dimensions and trains a VITS2 expressive-TTS baseline. The helper downloads official documentation, API metadata, and the lightweight validation/test tables by default. The approximately 29.3 MB embedded- audio preview and 5.56 GB full Hugging Face snapshot are separate opt-ins.
Safe-first helperscripts/download/dialogs_ru.sh
Enhancement, separation & quality
Diamond Benchmark
Diamond Benchmark: 750 Real Degraded Speech Recordings for Evaluating Restoration Models
Safe-first helper
Speech Restoration
Speech Enhancement
Perceptual Speech Quality Evaluation
Content Preservation Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_unspecified
- Code license
- not_applicable
- License caution
- Hugging Face metadata labels the dataset "other", but the card provides no license text, source-corpus citation, consent statement, or redistribution terms. The manifest's `emolia_id` field appears to reference an upstream collection, but the card does not identify or license it. Treat the release as evaluation-only pending clarification and verify source-recording, speaker, transcript, and redistribution rights before reuse, especially for training or commercial purposes.
- Download notes
- The public, ungated release contains 750 English speech clips with real-world codec, bandwidth, noise, and clipping degradation, plus reference transcripts and speaker, duration, and sample-rate metadata. Its documented protocol combines DNSMOS-P.835 for perceptual quality with ASR character error rate for content preservation. The helper downloads the official card, API metadata, and approximately 197 KB manifest by default. The Hugging Face API reports approximately 340 MB of repository storage, so the audio snapshot is an explicit opt-in.
Safe-first helperscripts/download/diamond_benchmark.sh
Speaker, identity & emotion
DiffSSD
DiffSSD: A Diffusion-Based Dataset for Speech Forensics
Safe-first helper
Synthetic Speech Detection
Audio Deepfake Detection
Speech Forensics
Synthetic Speech Attribution
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0_with_incorporated_source_terms
- Code license
- not_applicable
- License caution
- The dataset card applies CC BY-NC-ND 4.0 only to the synthetic voice data and incorporates the requirements of LJ Speech, LibriSpeech, ChatGPT, and all ten TTS systems. The release does not include the real-speech files named by its split manifest. Review every incorporated source and service term before use or redistribution.
- Download notes
- The public, ungated release contains 70,000 English synthetic utterances from eight open and two commercial TTS systems. Its 94,226-row protocol combines those files with paths to 24,226 real LJ Speech and LibriSpeech utterances and defines 31,690 training, 7,923 validation, and 54,613 test rows. The helper downloads the official card, license, input-text CSVs, and approximately 8.2 MiB split manifest by default. The approximately 16.9 GiB synthetic-audio TAR requires DIFFSSD_DOWNLOAD_AUDIO=1; real speech must be obtained separately from its owner releases.
Safe-first helperscripts/download/diffssd.sh
Speaker, identity & emotion
DIHARD III
The Third DIHARD Speech Diarization Challenge
Manual or gated
Speaker Diarization
Speech Activity Detection
Overlapping Speech Diarization
Multisource Speech Diarization
Access pathLDC / licensed
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- ldc_user_agreement
- Code license
- not_applicable
- License caution
- LDC2022S12 and LDC2022S14 list the LDC User Agreement for Non-Members and are available through LDC membership/non-member access. Re-check active LDC terms and component source restrictions before use or redistribution.
- Download notes
- The LDC catalog records list web-download development and evaluation releases with user-agreement access. The helper only prints official access steps; it does not download LDC-controlled data.
Safe-first helperscripts/download/dihard_iii.sh
Music
Dilemmadata
Dilemmadata: A Symbolic Dataset for Music Research
Safe-first helper
Roman Numeral Analysis
Harmonic Analysis
Cadence Detection
Phrase Segmentation
+7 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0_metadata_with_upstream_terms
- Code license
- not_specified
- License caution
- The owner-provided .zenodo.json declares the dataset CC BY-NC-SA 4.0, but the GitHub repository has no LICENSE file and GitHub detects no license. AugmentedNet and Distant Listening source corpora retain independent terms. Apply the declared non-commercial share-alike terms to the processed release, review upstream rights, and do not infer a software license for processing code.
- Download notes
- The public v1.0 repository contains processed pitch-array TSVs, predefined AugmentedNet training, validation, and test splits, more than 40 Distant Listening subcorpora, processing scripts, summary tables, and column specifications. The helper downloads lightweight official documentation and repository metadata by default; cloning the approximately 84 MB GitHub repository and its processed symbolic data requires DILEMMADATA_CLONE_REPO=1. Git submodules point to the independent upstream corpora and are not initialized by the helper.
Safe-first helperscripts/download/dilemmadata.sh
Enhancement, separation & quality
DNS Challenge
Deep Noise Suppression Challenge
Safe-first helper
Speech Enhancement
Speech Denoising
Dereverberation
Personalized Speech Enhancement
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The repository legal notice says documentation/content are CC BY 4.0 and code is MIT. DNS training data includes component sources such as AudioSet, Freesound, VCTK, VocalSet, and multilingual speech, so component/source-media terms should be re-checked before redistribution or commercial use.
- Download notes
- DNS5 development and blind test sets are multi-GB archives; the full training resources are hundreds of GB compressed and about 1 TB unpacked. The helper saves official README/license/downloader-script files by default and makes data archives explicit opt-ins.
Safe-first helperscripts/download/dns_challenge.sh
Representation & general suites
Doppelganger
Doppelganger: Sound Effects and Their Synthetic Twins
Safe-first helper
Synthetic Real Sound Effect Retrieval
Audio Instance Matching
Audio Representation Evaluation
Synthetic Audio Detection
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The Hugging Face card's MIT tag covers manifests, embeddings, and generation logs, not every audio asset. Stable Audio Open twins use the Stability AI Community License; ElevenLabs twins are redistributed under the author's ElevenLabs license; real audio retains FSD50K, UrbanSound8K, Freesound, or DCASE 2023 Task 7 source terms, and restricted real sources remain reference-by-ID only.
- Download notes
- The public, ungated release pairs 10,420 verified real sound-effect references across 34 Universal Category System events with Stable Audio Open synthetic twins, plus a controlled seven-class DCASE 2023 Task 7 corpus and text-only ElevenLabs controls. Real recordings are referenced by source ID and are not redistributed in bulk. The helper downloads official documentation and repository metadata by default; cloning the roughly 10 MB code/manifests repository or downloading the approximately 8.48 GB Hugging Face release requires separate opt-ins.
Safe-first helperscripts/download/doppelganger.sh
Audio understanding, generation & events
DroneAudioDataset
DroneAudioDataset for Audio-Based Drone Detection and Identification
Safe-first helper
Acoustic Drone Detection
Drone Type Classification
Environmental Sound Classification
Session Grouped Generalization
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- Not specified in the source record.
- License caution
- DroneAudioDataset has no LICENSE file or GitHub-detected license, so public download access does not establish permission to redistribute or commercially reuse its author-recorded drone audio. Its README attributes negative clips to ESC-50 and Speech Commands; retain their provenance and applicable source terms. EchoHawk's MIT license covers its code and synthetic generator, not the separately hosted recorded dataset.
- Download notes
- The owner repository directly stores 23,408 WAV files in parallel binary and multiclass directory trees: each tree has 1,332 drone clips and 10,372 unknown clips, while the multiclass tree divides the drone clips equally between the repository labels `bebop_1` and `membo_1`. The two trees are alternate label organizations rather than independent recordings, so their totals must not be interpreted as 23,408 unique source clips. The unknown class includes clips derived from ESC-50 and Speech Commands, plus author-created silence. EchoHawk resamples and segments a working copy into 19,275 one-second examples and shows that adjacent slices from the same continuous recording can leak across naive clip-level splits. Its corrected protocol groups drone clips by 257 recording sessions and negatives by source clip or speaker. The helper downloads official documentation, repository metadata, and paper metadata by default; cloning either the approximately 281 MB compressed dataset repository or the separate evaluation toolkit requires an explicit opt-in.
Safe-first helperscripts/download/drone_audio_dataset.sh
Enhancement, separation & quality
DSD100
DSD100: The 2016 Signal Separation Evaluation Campaign Music Dataset
Safe-first helper
Music Source Separation
Singing Voice Separation
Music Demixing
Professionally Produced Music Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- source_specific_terms_not_stated_as_single_dataset_license
- Code license
- MIT_for_dsdtools
- License caution
- SigSep states that all 100 tracks derive from the Mixing Secrets Free Multitrack Download Library and directs users to that original source for usage rights; it does not declare one blanket DSD100 data license. Public download access therefore does not establish redistribution or commercial-use rights. The dsdtools parser has an MIT license, which does not cover the music or stems.
- Download notes
- DSD100 contains 100 full-length, 44.1 kHz stereo songs with mixture, drums, bass, vocals, and other stems, divided into 50-track train and 50-track test sets for the SiSEC 2016 professionally produced music task. It is one benchmark family rather than separate families for each stem or split. The helper saves the official dataset page and parser documentation by default. The live approximately 14.9 GB ZIP requires DSD100_DOWNLOAD_ARCHIVE=1; cloning the small parser toolkit separately requires DSD100_CLONE_PARSER=1.
Safe-first helperscripts/download/dsd100.sh
Speech understanding & dialogue
DuplexChat
DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling
Safe-first helper
Full Duplex Spoken Dialogue Modeling
Bilingual Spoken Dialogue Generation
Turn Taking Modeling
Backchannel Modeling
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_metadata_only_with_upstream_podcast_rights
- Code license
- MIT
- License caution
- The repository license explicitly applies MIT only to the software and reconstruction metadata. It grants no rights in podcast audio or RSS content, which remain with their respective publishers under per-show terms. The paper is CC BY 4.0, but that article license does not broaden rights in referenced recordings. Confirm source terms before reconstruction, training, redistribution, or commercial use.
- Download notes
- DuplexChat is one bilingual corpus family with separate English and Japanese manifests, not two independent benchmarks. The public, ungated Hugging Face release contains 15,304,412 English dialogue spans totaling 282,634 hours and 7,329,011 Japanese spans totaling 132,723 hours. It distributes approximately 791.5 MB of compressed URL-and-timing metadata and no podcast audio. The helper downloads official documentation, API metadata, and the lightweight count file by default; the full manifest snapshot and construction toolkit are separate opt-ins. Reconstructing speaker-separated stereo clips requires downloading source episodes whose URLs may expire, a GPU, and access to the gated DialogueSidon model.
Safe-first helperscripts/download/duplexchat.sh
Speech understanding & dialogue
Dynamic-SUPERB
Dynamic-SUPERB: A dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
Safe-first helper
Spoken Language Model Evaluation
Instruction Following
Speech Understanding
Audio Understanding
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_specified
- License caution
- GitHub API reported no repository license on 2026-07-09. The benchmark aggregates tasks from many component datasets; use the official task metadata and each upstream corpus license before redistribution, training, or commercial use.
- Download notes
- The helper downloads the official README and leaderboard documentation by default and clones the benchmark repository only with DYNAMIC_SUPERB_CLONE_REPO=1. The benchmark is collaborative and spans many speech, music, and general sound tasks, so underlying task data should be checked through each component source before use.
Safe-first helperscripts/download/dynamic_superb.sh
Speech recognition
Earnings-21
Earnings-21: A Practical Benchmark for ASR in the Wild
Safe-first helper
Asr
Named Entity Recognition
Long Form Speech Recognition
Financial Speech Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0_text
- Code license
- not_specified
- License caution
- The dataset README shows a CC BY-SA 4.0 badge, but LICENSE.md expressly covers only transcripts and associated text files used for alignment. The repository has no detected top-level license, so confirm audio rights before redistribution or commercial use.
- Download notes
- Earnings-21 contains 44 English-language earnings calls totaling about 39 hours, plus a representative 10-hour Eval-10 subset. The helper downloads official documentation and lightweight file/speaker metadata by default; sparse checkout of the approximately 770 MB media tree, transcripts, RTTMs, and bias lists is opt-in.
Safe-first helperscripts/download/earnings_21.sh
Speech recognition
Earnings-22
Earnings-22: A Practical Benchmark for Accents in the Wild
Safe-first helper
Asr
Accented Speech Recognition
Long Form Speech Recognition
Financial Speech Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_specified
- License caution
- The earnings22 README shows a CC BY-SA 4.0 license badge, and LICENSE.md states that transcripts and associated text files are CC BY-SA 4.0. The top-level GitHub repository does not expose a detected repository license, and audio is stored through Git LFS, so re-check upstream terms before redistribution or commercial use.
- Download notes
- Earnings-22 contains 125 English-language earnings-call files totaling about 119 hours. The helper downloads README, license, and metadata by default; sparse checkout of transcripts/media is opt-in, and Git LFS audio pull is a second explicit opt-in.
Safe-first helperscripts/download/earnings_22.sh
Speech recognition
Earnings25
Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance
Safe-first helper
Financial Domain Asr
Long Form Asr
Speaker Aware Asr Evaluation
Industry Aware Asr Evaluation
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_metadata_and_annotations_with_upstream_audio_terms
- Code license
- not_released
- License caution
- The paper applies CC BY 4.0 to transcripts, annotations, metadata, evaluation splits, and alignments, and the Zenodo record declares CC BY 4.0. The paper separately states that redistributed earnings-call audio remains subject to applicable original content-provider terms. Review those provider terms before redistributing, training on, or commercially using the recordings. No separate evaluation-code repository is linked in version 1.
- Download notes
- Earnings25 is one benchmark family with two complementary English evaluation tracks. Testset-full contains 498 hours of complete 2025 Q4 earnings calls from approximately 500 S&P 500 companies, preserving long-form conversational structure. Testset-segmented contains 46 hours in 290 industry-balanced, five-to-ten-minute segments sampled from more than 2,000 U.S. calls across 2025. Both tracks provide aligned transcripts plus speaker names and roles, industry labels, company identifiers, and call structure. The helper saves official paper and Zenodo metadata by default. The single approximately 12.0 GB ZIP requires explicit upstream-audio terms acknowledgement and opt-in.
Safe-first helperscripts/download/earnings25.sh
Speech generation
EARS
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
Safe-first helper
Speech Enhancement
Speech Dereverberation
Speech Denoising
Expressive Speech Modeling
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- CC-BY-NC-4.0_as_declared_in_repository_readmes
- License caution
- The official dataset repository supplies CC BY-NC 4.0 and applies it to the code and dataset. The benchmark README declares the same terms by linking that license, but its repository has no standalone license file. EARS-WHAM and EARS-Reverb also incorporate WHAM! noise and multiple RIR collections, so preserve every upstream attribution and noncommercial restriction rather than treating the EARS license as overriding component terms.
- Download notes
- EARS is one 100-hour, 48 kHz English speech family with 107 speakers, not a separate family for every derived track. The official paper and generation repository define an 86.8-hour training split, 1.7-hour validation split, and 3.7-hour test split, plus EARS-WHAM speech enhancement, EARS-Reverb dereverberation, and a 743-item, two-hour blind noisy-speech test whose clean references are withheld. Version 2 generation scripts are public for EARS-WHAM and EARS-Reverb. Safe defaults fetch only official documentation, transcripts, speaker statistics, licenses, and repository metadata. Benchmark cloning and speaker archives are opt-in; audio downloads can be limited with EARS_SPEAKERS and never run downloaded code.
Safe-first helperscripts/download/ears.sh
Audio understanding, generation & events
Eating Sound Collection
Eating Sound Collection: 20-class food-eating audio dataset
Safe-first helper
Eating Sound Classification
Acoustic Event Classification
Audio Classification
Robustness To Natural Audio Triggers
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- pddl_with_unresolved_source_media_rights
- Code license
- MIT
- License caution
- Kaggle labels the dataset ODC Public Domain Dedication and Licence (PDDL), and the baseline repository is MIT. The audio was extracted from 246 YouTube videos, however, and neither source-video licenses nor per-clip provenance are published on the dataset page. The Kaggle declaration should not be assumed to waive third-party copyright, performer, privacy, or platform rights in the recordings.
- Download notes
- The owner release contains 11,141 normalized eating-sound clips across 20 food categories, hand-cut from 246 YouTube videos. The helper saves public Kaggle metadata and baseline documentation by default. The approximately 6.27 GiB Kaggle archive requires explicit opt-in and an authenticated Kaggle CLI. The paper filters clips shorter than three seconds but does not release its exact retained-item or split manifest.
Safe-first helperscripts/download/eating_sound_collection.sh
Speech understanding & dialogue
EchoMind
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
Safe-first helper
Empathetic Spoken Dialogue Evaluation
Spoken Content Understanding
Vocal Cue Perception
Paralinguistic Reasoning
+4 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0_dataset_card_with_source_and_consent_caveats
- Code license
- not_specified
- License caution
- The Hugging Face dataset card declares CC BY-NC-SA 4.0. The GitHub repository has no detected root license, so do not infer a code license from bundled third-party model files. Environmental mixtures use AudioCaps sounds; synthesis also uses Doubao, GPT-4o-mini-TTS, and one voice cloned from a YouTube creator, while the human subset comes from two compensated speakers. The aggregate card label should not be assumed to override source-media rights, service terms, voice rights, consent constraints, or other upstream conditions. The arXiv license covers the article, not the separately released benchmark artifacts.
- Download notes
- EchoMind links three evaluation levels over shared, semantically neutral English scripts: content ASR and vocal-cue MCQs; MCQ reasoning that integrates lexical and acoustic evidence; and open-ended empathetic response generation. The paper reports 1,137 scripts and 3,356 synthetic target, neutral, and alternative-expression inputs across 39 vocal attributes, plus 1,453 human-recorded inputs from a balanced 491-script subset. Released artifacts include input and reference-response audio, script metadata, understanding/reasoning MCQs, model runners, and evaluation code. The Hugging Face API exposes 8,238 files and reports 7,399,956,191 bytes of storage, so the helper fetches only the paper, project/repository documentation, GitHub metadata, and Hugging Face metadata by default. Set ECHOMIND_DOWNLOAD_HF=1 to explicitly download the full ungated snapshot.
Safe-first helperscripts/download/echomind.sh
Speaker, identity & emotion
EMO-SUPERB
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
Safe-first helper
Speech Emotion Recognition
Speech Representation Evaluation
Speaker Independent Cross Validation
Standardized Dataset Partitioning
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_component_terms
- Code license
- not_specified
- License caution
- The official repository has no LICENSE file and GitHub reports no detected license, so the evaluation code and released partition/label files have unspecified terms. Each of the six underlying speech corpora retains its own access agreement or license; a public benchmark repository does not make their audio freely redistributable.
- Download notes
- The public repository provides the evaluation implementation, corpus adapters, and standardized speaker-independent partitions for IEMOCAP, CREMA-D, MSP-IMPROV, and BIIC-NNIME. EMO-SUPERB evaluates six corpora in total, also including MSP-Podcast and BIIC-Podcast, with separate primary/secondary-emotion settings for some corpora. The helper downloads official documentation by default and makes the approximately 23 MB GitHub repository clone opt-in. It does not fetch corpus audio; users must obtain each component through its official EULA, form, or repository path.
Safe-first helperscripts/download/emo_superb.sh
Music
EMOPIA
EMOPIA: A Multi-Modal Pop Piano Dataset for Emotion Recognition and Emotion-based Music Generation
Safe-first helper
Symbolic Music Emotion Recognition
Four Quadrant Emotion Classification
Arousal Classification
Valence Classification
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_official_CC-BY-4.0_and_CC-BY-NC-SA-4.0
- Code license
- MIT
- License caution
- Zenodo metadata declares CC BY 4.0, while the official repository README declares the dataset CC BY-NC-SA 4.0, prohibits commercial use, and asks users not to defame original music owners. Apply the stricter non-commercial share-alike terms pending clarification. Repository code is MIT. The release provides only YouTube identifiers and timestamps for source audio, whose copyright and platform terms remain separate.
- Download notes
- The public v1.0 release contains 1,087 transcribed MIDI clips with four-quadrant emotion labels, YouTube IDs, timestamps, and a notebook that reproducibly creates song-disjoint training, validation, and test splits with a fixed random state. Precomputed split CSVs are not included in the Zenodo ZIP. The original paper reports 1,087 clips from 387 songs; MIDI-RAE-JEPA section 3.1.3 evaluates 1,071 piano clips after its own preprocessing. The helper saves official documentation and metadata by default. The approximately 5.5 MB Zenodo ZIP is checksum-verified and requires EMOPIA_DOWNLOAD_DATA=1. Copyrighted source audio is not redistributed.
Safe-first helperscripts/download/emopia.sh
Audiovisual & cross-modal
EmoPrefer
EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
Safe-first helper
Multimodal Emotion Preference Prediction
Audio Visual Emotion Understanding
Pairwise Emotion Description Evaluation
Multimodal Judge Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_research_terms
- Code license
- Apache-2.0_official_and_MIT_audit
- License caution
- The official EmoPrefer subdirectory includes Apache-2.0 but its README also calls the service a non-commercial research preview. The gated MER2025 and MER2026 cards declare CC BY-NC 4.0 plus stricter academic-only, no-redistribution, and no-modification conditions that control files obtained through those gates. The MER2026 gate also prohibits mirroring and derived-file redistribution without written permission. The 2026 audit repository is MIT. Treat the narrower access terms as controlling where they conflict, and do not infer media rights from the public CSV release.
- Download notes
- The official repository publicly releases six small annotation tables: the original 574-pair EmoPrefer set, its V2 extension with 2,096 individual-annotator pairs, reverse-order variants for swap-consistency analysis, and variants exposing the two description-generator names for score calculation and shortcut auditing. The paired English audio/video comes from the separately gated MER2025 release and is not redistributed by the repository. MER2026 Track 3 now also packages the original and V2 audio/video, candidate/test tables, and a supplementary stage-2 set behind manual approval. The helper downloads the annotations, official documentation, licenses, challenge page, and public API metadata only; users must accept the applicable Hugging Face access conditions themselves to obtain media. The 2026 audit paper and code provide reproducible content-blind, counterfactual, audio-visual judge, and ODIN-style diagnostics but intentionally distribute no private data, model weights, predictions, or checkpoints.
Safe-first helperscripts/download/emoprefer.sh
Speech generation
EmoV-DB
The Emotional Voices Database: Towards Controlling the Emotional Expressiveness in Voice Generation Systems
Safe-first helper
Emotional Speech Synthesis
Expressive Tts
Speech Emotion Recognition
Voice Conversion
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_non_commercial
- Code license
- not_specified
- License caution
- The EmoV-DB license permits non-commercial research, teaching, scientific publication, and personal experimentation, and asks users to contact the dataset owner for commercial use. GitHub API reports license NOASSERTION/Other.
- Download notes
- OpenSLR SLR115 hosts per-speaker/per-emotion archives. The helper downloads OpenSLR/GitHub docs and license by default; speech archives are opt-in with EMOV_DB_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/emov_db.sh
Speech generation
EmphAssess
EmphAssess: A Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
Safe-first helper
Speech Resynthesis Emphasis Preservation
Speech To Speech Translation Emphasis Preservation
Word Level Emphasis Detection
Prosody Transfer Evaluation
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- CC-BY-NC-4.0_with_separately_licensed_dependencies
- License caution
- The owner README explicitly applies CC BY-NC 4.0 to the benchmark data and EmphaClass checkpoint, and the repository LICENSE is CC BY-NC 4.0. SimAlign and Whisper are MIT and WhisperX is BSD-4-Clause according to the repository. The benchmark speech uses synthetic Expresso voices; retain source attribution and do not assume the aggregate license overrides voice-model, recording, or other upstream rights.
- Download notes
- The public benchmark archive contains 3,652 synthetic English speech samples generated from 913 emphasis-annotated transcripts in four Expresso voices, plus tokenized text, emphasized-word indices, and voice identifiers. Evaluated systems resynthesize or translate every input; WhisperX transcribes and force-aligns the output, EmphaClass detects emphasized words, and SimAlign maps source emphasis to paraphrased or translated output. Precision, recall, and F1 are reported over the dataset. The helper saves the official paper, README, license, and repository metadata by default. The 226,646,154- byte dataset, 3,571,288,236-byte English EmphaClass checkpoint, and archived code repository require separate explicit opt-ins. The host publishes ETags but no cryptographic checksums, so the helper does not claim checksum verification.
Safe-first helperscripts/download/emphassess.sh
Audiovisual & cross-modal
EPIC-SOUNDS
EPIC-SOUNDS: A Large-Scale Dataset of Actions that Sound
Safe-first helper
Egocentric Audio Event Recognition
Sound Event Detection
Audio Visual Understanding
Action Sound Recognition
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The annotation README states all dataset files are published under Creative Commons Attribution-NonCommercial 4.0 International. Raw audio is derived from EPIC-KITCHENS-100 video recordings, so original dataset access terms and any HDF5 access approval should be checked before redistribution or commercial use.
- Download notes
- The helper downloads official docs and public annotation CSV files by default. Raw audio is not redistributed separately; the official README says to download EPIC-KITCHENS-100 videos and extract audio, or email the maintainers for access to an existing HDF5 file.
Safe-first helperscripts/download/epic_sounds.sh
Audio understanding, generation & events
ESC-50
ESC-50: Dataset for Environmental Sound Classification
Safe-first helper
Environmental Sound Classification
Audio Tagging
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-3.0
- Code license
- not_specified
- License caution
- ESC-10 subset clips are CC BY; ESC-50 as a whole is Creative Commons Attribution-NonCommercial. Per-clip Freesound attributions are in the repository LICENSE file.
Safe-first helperscripts/download/esc_50.sh
Audio understanding, generation & events
ESCUCHA
ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions
Safe-first helper
Spanish Speech Understanding
Audio Question Answering
Audio Reasoning
Multi Audio Comparison
+3 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The GitHub repository has no license file or detected GitHub license. The paper is CC BY 4.0 on arXiv, but that does not license the annotations, code, or source recordings. Review the rights and platform terms for each linked recording before downloading, redistribution, or commercial use.
- Download notes
- The public repository releases 1,000 Spanish questions as JSON and TSV, including 900 multiple-choice and 100 audio-instruction-following items. The helper downloads those approximately 2.2 MB of annotations plus the README and scorer by default; cloning the repository is opt-in. Audio is not redistributed: the release provides source URLs and a yt-dlp script for reconstructing up to 162.9 hours from public recordings, so availability can drift and source-platform terms apply.
Safe-first helperscripts/download/escucha.sh
Speech understanding & dialogue
Europarl-ST
Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates
Safe-first helper
Speech To Text Translation
Speech Translation
Multilingual Asr
Parliamentary Speech Translation
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The official README says the work carried out to construct Europarl-ST is released under CC BY-NC 4.0, while all rights of the underlying data belong to the European Union and respective copyright holders. Re-check EU/European Parliament reuse terms before redistribution or commercial use.
- Download notes
- The official page links release v1.1, which adds Romanian, Polish, and Dutch to German, English, Spanish, French, Italian, and Portuguese for 72 speech translation directions, plus a train-noisy set. The v1.1 archive is about 21 GB, so the helper downloads only the official page and README by default and requires EUROPARL_ST_DOWNLOAD_ARCHIVE=1 for the archive.
Safe-first helperscripts/download/europarl_st.sh
Speaker, identity & emotion
Fake-or-Real
Fake-or-Real (FoR) Dataset
Safe-first helper
Synthetic Speech Detection
Audio Deepfake Detection
Replay Robust Synthetic Speech Detection
Binary Real Fake Speech Classification
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_applicable
- License caution
- The official dataset page states no license, terms of use, or redistribution grant. Public, ungated downloadability is not a license. Obtain permission as needed and review the separate CMU Arctic, LJ Speech, VoxForge, generated-voice, and lab-recording rights before use or redistribution.
- Download notes
- The official York University page describes more than 195,000 real and synthetic speech utterances aggregated from contemporary TTS systems, CMU Arctic, LJ Speech, VoxForge, and lab recordings. It publishes four variants: original collected files; gender/class-balanced normalized audio; normalized files truncated to two seconds; and rerecorded two-second audio modeling transmission through a voice channel. The helper saves only the official landing page by default. The four archives total approximately 16.0 GiB and require both an explicit archive opt-in and a version selection.
Safe-first helperscripts/download/fake_or_real.sh
Audiovisual & cross-modal
FakeAVCeleb
FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
Manual or gated
Audio Visual Deepfake Detection
Synthetic Speech Detection
Face Swap Detection
Lip Sync Deepfake Detection
+1 more
Access pathOfficial / other
Upstream termsNot specified
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_specified_request_terms_apply
- Code license
- MIT
- License caution
- The repository explicitly describes its code as MIT, but states no separate dataset license. Its comparison table marks FakeAVCeleb as not rights-cleared with zero agreeing subjects; source clips come from celebrity YouTube videos. Author approval does not itself clear copyright, platform, voice, likeness, privacy, or publicity rights. Review the granted terms and local law before use or redistribution.
- Download notes
- The owner repository describes 500 real and 19,500 generated videos from 500 English-speaking celebrities, with real/fake audio and video combinations, lip-synced cloned speech, multiple face/audio generation methods, demographic metadata, and fine-grained labels. Dataset access is not a public direct download: applicants must submit the official Google form and, if approved, receive an author-provided download script. The helper prints the owner-controlled request path and never guesses or accepts a private archive URL.
Safe-first helperscripts/download/fakeavceleb.sh
Speech recognition
Fisher English
Fisher English Training Speech and Transcripts
Manual or gated
Automatic Speech Recognition
Conversational Speech Recognition
Telephone Speech Recognition
Speech Transcription
Access pathLDC / licensed
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_ldc_license
- Code license
- not_applicable
- License caution
- The LDC catalog pages list the LDC User Agreement for Non-Members and availability for Subscription/Standard Members and Non-Members. Re-check the current LDC agreement before use or redistribution.
- Download notes
- Fisher English is distributed by LDC after login/licensing. The helper only prints official access steps because the speech and transcripts are paid/licensed catalog releases, not publicly script-downloadable archives.
Safe-first helperscripts/download/fisher_english.sh
Speech recognition
FLEURS
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Safe-first helper
Asr
Speech Translation
Language Identification
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- HF dataset card lists CC BY 4.0.
Safe-first helperscripts/download/fleurs.sh
Speech recognition
Fleurs-SLU
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding
Safe-first helper
Multilingual Spoken Language Understanding
Spoken Topic Classification
Spoken Question Answering
Listening Comprehension
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-4.0_as_declared_on_hugging_face
- Code license
- not_specified
- License caution
- Both dataset cards declare CC BY-SA 4.0. The derived releases align FLEURS audio with FLORES, SIB-200, or Belebele annotations, so preserve attribution, share-alike obligations, and all source-dataset notices. The reproduction repository has no LICENSE file and GitHub reports no detected license, so no code reuse terms are claimed.
- Download notes
- Fleurs-SLU is one benchmark family with two aligned public releases. SIB-Fleurs maps FLEURS/FLORES utterances to seven SIB-200 topics for topical classification in 102 languages (692 hours). Belebele-Fleurs maps FLEURS speech to Belebele passages and questions for multilingual listening comprehension in 92 languages (944 hours), and can also support constructed long-form ASR. The Hugging Face API reports about 157.2 GB for SIB-Fleurs and 223.4 GB for Belebele-Fleurs. The helper downloads only official cards, Hub metadata, paper metadata, and repository documentation by default. A single explicit language-script configuration can be fetched with FLEURS_SLU_DATASET and FLEURS_SLU_CONFIG; it never requests the complete snapshots. MEUSLI section 4.1 evaluates ASR, speech translation, and seven-topic identification on five public SIB-Fleurs language subsets.
Safe-first helperscripts/download/fleurs_slu.sh
Speech recognition
FluencyBank
FluencyBank Shared Database for Fluency Research
Manual or gated
Word Level Timing
Speech Alignment
Disfluent Speech Recognition
Stuttering Research
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC_BY_NC_SA_3_0_with_password_protected_clinical_bank_rules
- Code license
- not_applicable
- License caution
- TalkBank applies CC BY-NC-SA 3.0 unless a corpus says otherwise and prohibits incorporating the data into commercial products or model parameters. Users must follow corpus citation requirements, the TalkBank Code of Ethics, confidentiality protections, local-storage restrictions for identifiable data, and non-retaining web-processing requirements. Individual corpora may impose additional terms.
- Download notes
- FluencyBank contains corpora from typically developing monolingual and bilingual participants, children and adults who stutter or clutter, and second-language learners. Most research data is password protected and restricted to approved FluencyBank consortium members, though the official corpus index also identifies a small number of open teaching or research collections. The helper saves only public documentation and prints the membership path; it never authenticates or downloads protected participant recordings.
Safe-first helperscripts/download/fluencybank.sh
Speech understanding & dialogue
Fluent Speech Commands
Fluent Speech Commands: A dataset for spoken language understanding research
Manual or gated
Spoken Language Understanding
Intent Classification
Slot Filling
Smart Home Voice Commands
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- Official Fluent.ai page says the dataset is released strictly for academic research only under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International, and not authorized for commercial use. No current public code repository was identified.
- Download notes
- The helper downloads the public dataset page and license PDF, then prints the manual Google Groups access path. The official page says the corpus contains 30,043 single-channel 16 kHz WAV utterances from 97 speakers with action, object, and location slot labels; do not commit granted links or downloaded audio.
Safe-first helperscripts/download/fluent_speech_commands.sh
Music
FMA
FMA: A Dataset For Music Analysis
Safe-first helper
Music Genre Classification
Music Auto Tagging
Music Recommendation
Artist Identification
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- Official GitHub README says metadata is CC BY 4.0, code is MIT, and audio is distributed under the license chosen by each artist because the dataset maintainers do not hold audio copyright. UCI lists the dataset record as CC BY 4.0, but per-track audio licenses should be checked before redistribution or commercial use.
- Download notes
- The helper downloads official README/license files by default. Metadata is 342 MiB; audio archives range from 7.2 GiB for fma_small to 879 GiB for fma_full, so metadata and audio are explicit opt-ins.
Safe-first helperscripts/download/fma.sh
Audio understanding, generation & events
FoleySet
FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset
Safe-first helper
Foley Sound Classification
Fine Grained Sound Event Classification
Audio Retrieval
Foley Sound Generation
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_from_cc0_source_audio
- Code license
- not_released
- License caution
- Zenodo declares the curated release CC BY 4.0. The paper says every incorporated Freesound waveform was selected under CC0 and retains uploader names, source IDs, URLs, tags, and descriptions for provenance. Preserve those fields and the dataset citation. No standalone baseline-code release was identified; the CC BY 4.0 arXiv paper license must not be assumed to grant software rights.
- Download notes
- The public, ungated Zenodo release contains 10,000 human-annotated Freesound-derived clips totaling approximately 9.5 hours. It defines fixed 8,000/1,000/1,000 train/validation/test splits without placing segments from one source recording in different splits, plus a two-level taxonomy of 9 major and 73 fine-grained Foley categories and one-shot/multi-shot labels. The paper benchmarks frozen PaSST features with linear probes on both classification levels. The helper downloads the official paper page and Zenodo metadata by default; the approximately 2.16 GB ZIP requires FOLEYSET_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/foleyset.sh
Audio understanding, generation & events
ForestIR
ForestIR: Physics-Informed Forest Sound Simulation for Array-Based Bioacoustic Remote Sensing
Safe-first helper
Forest Impulse Response Simulation
Bioacoustic Sound Source Localization
Microphone Array Design
Spatial Audio Robustness
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- bundled_media_terms_not_separately_specified
- Code license
- MIT
- License caution
- The repository has an MIT LICENSE and GitHub detects MIT, but neither the paper nor README states separate licenses or complete provenance terms for bundled bird vocalizations and recorded environmental noise. Do not assume the software license clears those recordings; verify source-media attribution, wildlife-recording, and field-recording rights before redistribution or commercial use. The request-only processed validation data may carry additional terms.
- Download notes
- The public repository provides a physics-informed forest impulse-response simulator, command-line and Python interfaces, measured and synthetic tree/microphone/source geometry presets, example bird calls, environmental noise recordings, and manifest-producing array rendering. The helper downloads official documentation, license, repository metadata, and the paper by default; cloning the approximately 31 MB repository requires FORESTIR_CLONE_REPO=1. The paper says processed site-recorded data needed to reproduce its main validation analyses must be requested from the authors, so the public repository must not be represented as a complete release of those field measurements.
Safe-first helperscripts/download/forestir.sh
Audiovisual & cross-modal
FriendBench
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Safe-first helper
Audio Visual Social Reasoning
Social Relationship Recognition
Familiar Stranger Classification
Multimodal Human Behavior Understanding
+3 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- cc-by-nc-4.0_declared_in_terms
- License caution
- The dataset card declares CC BY-NC 4.0 and says the benchmark inherits that license from Seamless Interaction. TERMS.md also applies CC BY-NC 4.0 to the original manifest, ratings, predictions, and evaluation code, but labels its own wording a draft awaiting confirmation against upstream terms. The release includes recorded human interactions, transcripts, relationship labels, and anonymized rater responses with bucketed demographic fields; users should retain attribution, limit use to non-commercial purposes, and review the live source corpus's privacy, consent, biometric, and responsible-use terms.
- Download notes
- The public, ungated validation release contains 96 participant-disjoint approximately 20-second dyadic clips, balanced across familiar versus stranger, early versus late, mixed versus same estimated-gender, and three recording-site conditions. Every item provides a 16 kHz mono audio mix, side-by-side audiovisual video, turn-level transcript, and binary familiarity answer key. The MINT 2026 paper evaluates 26 zero-shot models and roughly 90 human raters per modality with accuracy over answered trials, coverage, per-class recall, signal-detection sensitivity and response criterion, plus matched text, audio, and audiovisual comparisons. The release includes its evaluation script, model predictions, de-identified human ratings, and baseline table. The helper downloads the official card and API metadata by default. Set FRIEND_BENCH_DOWNLOAD_METADATA=1 for the lightweight manifest, ratings, predictions, baseline results, evaluator, terms, and citation, or FRIEND_BENCH_DOWNLOAD_HF=1 for the complete approximately 433 MB snapshot. The paper and card cover only the released `validation` config; possible future configs are not part of this benchmark version.
Safe-first helperscripts/download/friend_bench.sh
Audio understanding, generation & events
FSD50K
FSD50K: An Open Dataset of Human-Labeled Sound Events
Safe-first helper
Sound Event Classification
Sound Event Tagging
Audio Tagging
Machine Listening
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- not_specified
- License caution
- Zenodo says individual Freesound clips keep per-clip CC0, CC-BY, CC-BY-NC, or CC Sampling+ licenses, with clip-level license mappings in the metadata JSON files. FSD50K as a curated dataset is additionally released under CC BY, but the maintainers warn that a single global license is not straightforward because items have different licenses.
- Download notes
- The helper downloads the small documentation, ground-truth, and metadata ZIPs by default. Audio is split across about 24.7 GiB of dev archives plus about 6.2 GiB of eval archives, so audio download is an explicit opt-in.
Safe-first helperscripts/download/fsd50k.sh
Audio understanding, generation & events
FSDKaggle2018
FSDKaggle2018: Freesound General-Purpose Audio Tagging Challenge Dataset
Safe-first helper
Audio Tagging
Sound Event Classification
General Purpose Audio Tagging
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- not_specified
- License caution
- The Zenodo record license id is other-at. The record states FSDKaggle2018 as a curated dataset is CC BY, while individual Freesound clips retain per-clip Creative Commons licenses listed in train_post_competition.csv and test_post_competition_scoring_clips.csv. Kaggle competition rules and Freesound source terms should also be checked for challenge use.
- Download notes
- The helper downloads the Zenodo record JSON plus small documentation and metadata ZIPs by default. The audio archives are about 4.6 GiB combined, so audio download is an explicit opt-in.
Safe-first helperscripts/download/fsdkaggle2018.sh
Audio understanding, generation & events
FSDKaggle2019
FSDKaggle2019: Freesound Audio Tagging 2019 / Audio Tagging with Noisy Labels and Minimal Supervision
Safe-first helper
Audio Tagging
Sound Event Classification
Noisy Label Learning
Weakly Labeled Audio Classification
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- MIT
- License caution
- Zenodo reports license id other-at. The record says FSDKaggle2019 as a curated dataset is CC BY, while individual Freesound clips keep per-clip CC0, CC-BY, CC-BY-NC, or CC Sampling+ licenses and Flickr/YFCC noisy-train clips keep CC-BY or CC BY-SA licenses recorded in metadata CSVs. Kaggle/DCASE challenge rules and source media terms should also be checked for challenge or redistribution use.
- Download notes
- The helper downloads the Zenodo record JSON plus small documentation and metadata ZIPs by default. Full audio is about 25 GiB and includes a split noisy-train archive, so audio download is an explicit opt-in with part selection.
Safe-first helperscripts/download/fsdkaggle2019.sh
Speech understanding & dialogue
Full-Duplex-Bench
Full-Duplex-Bench: Evaluation Suite for Full-Duplex Spoken Dialogue Models and Voice Agents
Safe-first helper
Full Duplex Spoken Dialogue Evaluation
Turn Taking And Pause Handling
Interruption And Backchannel Handling
Overlap Robustness And Addressee Detection
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_by_version
- Code license
- CC-BY-NC-4.0_repository_license
- License caution
- The v1/v1.5 data README applies CC BY-NC 4.0 to v1 Candor/ICC-derived subsets and requires compliance with their upstream terms; it labels the owner-generated v1 synthetic subsets and v1.5 data MIT. The root repository carries CC BY-NC 4.0, despite GitHub reporting NOASSERTION. No separate license statement was found in the v2 README or for the external v3 Google Drive archive, so do not assume the repository license or the v1.5 MIT statement broadens rights in those assets. Review each version and source corpus before redistribution or commercial use.
- Download notes
- This is one evolving benchmark family, not four unrelated datasets. Versions 1 and 1.5 provide static offline audio tests: v1 covers pause handling, smooth turn-taking, backchannels, and interruption, while v1.5 adds 499 paired overlap trials across user interruption, backchannel, talking-to-other, and background-speech conditions. Its evaluator classifies respond, resume, uncertain, and unknown behavior from ASR transcripts and also measures stop/response latency and prosodic change. Version 2 instead runs dynamic multi-turn conversations between an examinee and automated examiner. Version 3 evaluates 100 human-recorded disfluent tool-use examples over 79 scenarios and 12 speakers. The helper downloads official documentation and repository metadata by default; cloning the approximately 47 MB repository is opt-in. Benchmark audio remains a manual Google Drive download. The JoyAI-Talker report, section 6, independently evaluates Joy-Duplex on the public v1.5 protocol; it does not define a separate benchmark.
Safe-first helperscripts/download/full_duplex_bench.sh
Enhancement, separation & quality
FUSS
FUSS: Free Universal Sound Separation Dataset
Safe-first helper
Universal Sound Separation
Arbitrary Sound Separation
Reverberant Source Separation
Sound Event Separation
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- Apache-2.0
- License caution
- The official license document and Zenodo record release FUSS as a whole under CC BY 4.0. Its input audio clips are CC0 Freesound files selected using prerelease FSD50K labels; those labels are not distributed with FUSS. Google Research's sound-separation code repository is Apache-2.0.
- Download notes
- The public Zenodo release provides train, validation, and eval mixtures containing one to four arbitrary sound sources, with dry and reverberant references, simulated room responses, and CC0 source clips. The helper saves official documentation, repository metadata, and the small license archive by default. Data archives are about 1.9 to 8.9 GB each and require explicit selection with FUSS_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/fuss.sh
Speech understanding & dialogue
GCM-Bench
GCM-Bench: A Dataset for Generative Context Mis-anchoring
Safe-first helper
Playback Relative Referent Anchoring
Interruption Context Consistency
Full Duplex Spoken Dialogue Evaluation
Access pathOfficial / other
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- Apache-2.0
- License caution
- The owner README states that both code and data are Apache 2.0, and the repository includes LICENSE, NOTICE, and third-party notices. This does not supersede DashScope/Qwen service terms or resolve all generated- speech, voice, model-output, privacy, or commercial-use questions for newly reproduced artifacts.
- Download notes
- The release contains all 108 16 kHz mono synthetic speech inputs, interruption metadata, frozen scenario and sample manifests, offline validation, construction, inference, ASR, and judging code. The helper downloads documentation, repository metadata, and the two lightweight JSON manifests by default; cloning the repository with roughly 130 MB of WAV files requires GCM_BENCH_CLONE_REPO=1. End-to-end evaluation also requires user-supplied DashScope/Qwen service credentials.
Safe-first helperscripts/download/gcm_bench.sh
Audio understanding, generation & events
Geo-ATBench
Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
Safe-first helper
Multi Label Audio Tagging
Environmental Sound Classification
Audio Geospatial Context Fusion
Context Aware Audio Tagging
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_upstream_terms
- Code license
- MIT
- License caution
- The Zenodo record declares CC BY 4.0 and the official repository is MIT. The paper says audio comes from Freesound and a prior geotagged-audio dataset, while contextual metadata derives from OpenStreetMap; preserve source attribution and review per-clip audio licenses and OSM attribution/database terms before redistribution.
- Download notes
- The public, ungated release contains 3,854 ten-second mono WAV clips totaling 10.71 hours, 28 fine-grained sound-event labels, three coarse event groups, and POI-derived geospatial semantic context over 11 OpenStreetMap feature categories. The helper downloads official documentation and repository/Zenodo metadata by default; the approximately 850 MB dataset archive is opt-in.
Safe-first helperscripts/download/geo_atbench.sh
Speech recognition
Ghana Speech Eval
Ghana Speech Eval: ASR Evaluation for 10 Ghanaian Languages
Safe-first helper
Automatic Speech Recognition
Multilingual Speech Recognition
Low Resource Speech Recognition
African Language Speech Recognition
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_claim_with_unresolved_upstream_terms
- Code license
- not_applicable
- License caution
- The Ghana Speech Eval card declares CC BY 4.0 and says this follows the source dataset. However, the current AfriSpeech public-v1 card and API metadata do not state a repository-level license. Preserve attribution, review the original collection and speaker-consent terms, and confirm upstream reuse rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains 9,967 read-speech clips with verbatim transcripts across Ahanta, Dagaare, Dangme, Ewe, Fante, Frafra/Gurene, Ga, Nzema, Sehwi, and Twi. It samples up to 1,000 clips per language from AfriSpeech's public African speech corpus, merges the source splits into one evaluation split, and filters clips to 3-15 seconds. The card explicitly reserves the fixed set for ASR evaluation rather than training. The helper downloads official cards and API metadata by default; the approximately 594 MB compressed benchmark snapshot requires GHANA_SPEECH_EVAL_DOWNLOAD_HF=1.
Safe-first helperscripts/download/ghana_speech_eval.sh
Speech recognition
GigaSpeech
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Manual or gated
Asr
Large Scale Speech Recognition
Text To Speech
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- gated_non_commercial_research_educational
- Code license
- Apache-2.0
- License caution
- HF terms restrict the database to non-commercial research and educational purposes; the GitHub code repo is Apache-2.0.
- Download notes
- HF access is gated and the official repo also asks users to fill out the Google Form first; full HF dataset size is about 2.6 TB.
Safe-first helperscripts/download/gigaspeech.sh
Speech recognition
GigaSpeechBench
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Safe-first helper
Automatic Speech Recognition
Speech To Text Translation
Multilingual Speech Recognition
Dialectal Speech Recognition
+6 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the Hugging Face repository metadata nor the official GitHub repository declares a license, and the GitHub repository has no LICENSE file. Do not infer reuse or redistribution rights from public, ungated access; obtain clarification from SpeechColab before redistribution or commercial use.
- Download notes
- The paper defines a 680-hour in-the-wild benchmark with five modules covering 14 low-resource language/region sets, six Chinese dialects, six English accents, 12 Chinese and English terminology domains, and child/older-adult speech; it also reports Chinese and English translation references for 11 languages. The current public, ungated Hugging Face repository contains the Low-Resource-Languages, CH-EN-Dialects, and Vertical-Domain modules, including audio archives and JSON metadata, but no separate age-group module was visible when checked. Translation result files are public, while the exact release coverage of translation references should be verified per metadata file. The helper downloads official documentation and repository/API metadata by default; the Hugging Face API reports approximately 86.3 GB of repository storage, so the dataset snapshot requires GIGASPEECHBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/gigaspeechbench.sh
Audio understanding, generation & events
GlobeAudio
GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models
Safe-first helper
Multilingual Audio Question Answering
Culturally Grounded Audio Understanding
Naturalistic Speech Reasoning
Prosodic And Paralinguistic Understanding
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_declared_on_dataset_card
- Code license
- not_applicable_no_evaluation_repository_linked
- License caution
- The dataset card declares CC BY-NC 4.0. The audio was clipped from publicly accessible YouTube media, however, and neither the card nor paper enumerates source-video URLs or demonstrates that every upstream rights holder authorized redistribution under that license. Treat the card declaration as the uploader's terms, not proof of cleared rights in third-party recordings; review provenance, privacy, platform terms, and intended non-commercial use before downloading or redistributing. The article itself uses arXiv's perpetual non-exclusive license.
- Download notes
- The owner-linked release provides 5,637 four-option MCQs with embedded 44.1 kHz audio across English, Russian, Chinese, Thai, Bengali, and Singlish configurations. Questions and distractors were written by native speakers from naturally occurring online audio; the paper reports 20-40 second clips, a 25.35 second mean, two-stage review, and 95.5% agreement on checked items. Evaluation uses exact option-label accuracy, with audio, transcript-only, blind, and translated-question ablations testing whether models use acoustic, prosodic, linguistic, and cultural evidence. The live Hub tree at revision `31015d8f0bbf746fb8f5f62a252b7c9ef04d22d7` contains 29 Parquet shards totaling 12,542,951,445 bytes. The card's aggregate 1,554,418,131-byte download-size field reflects only the Bengali configuration and understates the full snapshot. The helper therefore fetches only the official card and live API metadata by default and requires an explicit opt-in for the full dataset.
Safe-first helperscripts/download/globeaudio.sh
Speech recognition
Golos
Golos: Russian Dataset for Speech Research
Safe-first helper
Automatic Speech Recognition
Russian Speech Recognition
Far Field Speech Recognition
Language Modeling
Access pathOpenSLR
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_golos_license
- Code license
- not_specified
- License caution
- OpenSLR points to SberDevices' English and Russian PDF license files rather than an SPDX-style open license. Re-check those PDFs before redistribution, commercial use, or training release claims.
- Download notes
- OpenSLR SLR114 mirrors an 18 GiB Opus archive with Russian speech and transcripts, a 71 MiB QuartzNet acoustic model, and 4.7 GiB KenLM language models. The helper saves the OpenSLR page, README, checksums, and license PDFs by default and requires explicit opt-ins before downloading large archives.
Safe-first helperscripts/download/golos.sh
Music
GTZAN
GTZAN Genre Collection
Safe-first helper
Music Genre Classification
Music Information Retrieval
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The reachable Hugging Face dataset card does not state a data license. Treat redistribution and commercial use as unclear until the Marsyas/original dataset terms are confirmed.
- Download notes
- The Hugging Face card describes 1,000 30-second mono WAV tracks across 10 genres and provides a reproducible datasets loader. The helper saves the HF dataset card by default and requires GTZAN_DOWNLOAD_HF=1 before downloading the audio snapshot.
Safe-first helperscripts/download/gtzan.sh
Speech recognition
HALAS
HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
Safe-first helper
Automatic Speech Recognition
Asr Hallucination Detection
Asr Error Analysis
Span Level Error Detection
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- unknown_with_cc-by-sa-4.0_upstream
- Code license
- not_specified
- License caution
- The Hugging Face card's machine-readable license field is "unknown," and neither the GitHub nor Hugging Face release includes a standalone license file. The card says HALAS derives from CC BY-SA 4.0 Earnings-22 and describes the authors' annotations as intended for the same license unless otherwise specified, but it also tells users to consult the repository for authoritative terms. Treat annotation and code rights as unspecified pending an explicit license; Earnings-22 source terms continue to apply to separately obtained audio.
- Download notes
- The public, ungated release contains human-reviewed predictions from seven ASR systems for 3,611 English Earnings-22 segments, with utterance labels and character-span annotations for hallucination, looping, and hallucinated looping. Meeting-disjoint stratified splits contain 2,866 train and 745 test segments; the test split has a 22.6% hallucination rate and excludes audio at or below one second or below three words. The release contains annotations, predictions, corrected references, and source identifiers rather than redistributed audio; audio must be obtained separately from Earnings-22. The helper downloads official documentation, prompts, and repository metadata by default. The approximately 0.86 MB test CSV is opt-in, while the full approximately 7.2 MB Hugging Face snapshot requires HALAS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/halas.sh
Representation & general suites
HEAR
HEAR: Holistic Evaluation of Audio Representations
Safe-first helper
Audio Representation Evaluation
Speech Classification
Environmental Sound Classification
Music Classification
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- Apache-2.0
- License caution
- Zenodo record lists CC BY 4.0 but explicitly says component datasets have different open licenses and to inspect each dataset's LICENSE.txt; eval kit repo is Apache-2.0.
- Download notes
- The helper downloads the Zenodo record metadata and LICENSE.txt by default. Individual task archives are large and opt-in; the Zenodo record notes that TFDS-derived crema-d, GTZAN genre, and GTZAN music/speech tasks in this release were retracted because of a preprocessing bug.
Safe-first helperscripts/download/hear.sh
Speech generation
Hi-Fi TTS
Hi-Fi Multi-Speaker English TTS Dataset
Safe-first helper
Text To Speech
Speech Synthesis
Multi Speaker Speech Synthesis
High Fidelity Speech Synthesis
+1 more
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR109 lists CC BY 4.0. The corpus is based on public-domain LibriVox audiobooks and Project Gutenberg texts, but downstream users should still preserve attribution and check packaged metadata.
- Download notes
- OpenSLR lists one 39-41 GiB speech/text archive. The helper saves the official OpenSLR page by default and requires HIFITTS_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/hifitts.sh
Speaker, identity & emotion
HI-MIA
HI-MIA: A Far-field Text-Dependent Speaker Verification Database and the Baselines
Safe-first helper
Text Dependent Speaker Verification
Far Field Speaker Verification
Wake Word Speaker Verification
Speaker Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR85 lists Apache License v2.0. The paper says HI-MIA contains recordings from 340 people in real-room far-field settings and is extracted from AISHELL-WakeUp-1.
- Download notes
- OpenSLR SLR85 hosts the AISHELL Speaker Verification Challenge 2019 far-field text-dependent speaker verification data. The helper downloads the official page and filename mapping by default; train/dev/test archives are multi-GB and require HIMIA_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/hi_mia.sh
Representation & general suites
HyPoradise
HyPoradise: Generative Speech Recognition Error Correction with Large Language Models
Safe-first helper
Generative Speech Recognition Error Correction
N Best Hypothesis Correction
Zero Shot Llm Evaluation
Few Shot In Context Learning
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- Hugging Face declares MIT for v0 and the GigaSpeech subset, while the robust card and RobustGER repository declare Apache 2.0. These terms cover the released derived artifacts and code, not the absent source audio or every upstream corpus. Voice Memory separately declares CC BY-NC-SA 4.0 and is auto-gated; do not apply that later license to HyPoradise or apply HyPoradise's licenses to Voice Memory artifacts.
- Download notes
- HyPoradise v0 releases five-best ASR hypotheses paired with reference transcripts for 316,881 training and 17,383 test utterances across ten test conditions from nine source families: LibriSpeech clean/other, CHiME-4, WSJ, Switchboard, Common Voice, TED-LIUM 3, LRS2, ATIS, and CORAAL. The fixed protocol compares zero-shot prompting, few-shot in-context learning, full fine-tuning, and LoRA against 1-best and oracle reranking using WER. The separate Robust HyPoradise extension adds approximately 113,000 pairs for CHiME-4, VoiceBank-DEMAND, LibriSpeech-FreeSound, NOIZEUS, and RATS noise conditions. A GigaSpeech v1 subset supports later cross-modal error-correction work but is not a replacement for v0. The 2026 Voice Memory report reuses v0 and Robust test sets for LLM act-or-abstain evaluation, Recoverable Information Ratio, and Harmful Edit Rate. The helper saves cards, repository metadata, and papers by default; the approximately 384 MB v0, 40.8 MB GigaSpeech v1, and 115 MB Robust snapshots each require explicit opt-in.
Safe-first helperscripts/download/hyporadise.sh
Audiovisual & cross-modal
IEMOCAP
IEMOCAP: Interactive Emotional Dyadic Motion Capture Database
Manual or gated
Speech Emotion Recognition
Audio Visual Emotion Recognition
Multimodal Emotion Recognition
Dialogue Emotion Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_research_license
- Code license
- not_applicable
- License caution
- The official release page links a USC/SAIL data release form and Google request form. Treat access as manual/form-gated and re-check the signed release terms before redistribution, commercial use, or sharing derived copies.
- Download notes
- IEMOCAP is released by request after reading the license and submitting the official electronic release form; there is no public one-command archive URL.
Safe-first helperscripts/download/iemocap.sh
Speech understanding & dialogue
IFEval-Audio
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
Manual or gated
Audio Instruction Following
Speech Instruction Following
Structured Output Adherence
Semantic Correctness Evaluation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- apache-2.0_with_mixed_upstream_terms
- Code license
- cc_by_nc_unspecified_version
- License caution
- The Hugging Face card declares Apache-2.0, while the paper says clips are used as-is from Spoken SQuAD (CC BY-SA 4.0), TED-LIUM 3 (CC BY-NC-ND 3.0), MuChoMusic (CC BY-SA 4.0), WavCaps (academic use only), and AudioBench sources with inherited licenses. The AudioBench LICENSE says source code is Creative Commons NonCommercial without naming a version. Preserve the most restrictive source terms and verify per-clip provenance before redistribution or commercial use.
- Download notes
- The official release contains 280 English audio-instruction-answer triples spanning Content, Capitalization, Symbol, List Structure, Length, and Format requirements. It includes 240 speech triples and 40 music/environmental-sound triples, and reports Instruction Following Rate, Semantic Correctness Rate, and Overall Success Rate. The helper downloads public Hugging Face API metadata and official benchmark documentation by default; the approximately 45 MB compressed Hugging Face snapshot is opt-in and requires logging in and accepting the repository's access conditions.
Safe-first helperscripts/download/ifeval_audio.sh
Speech understanding & dialogue
IHBench
IHBench: Interruption Handling Benchmark
Safe-first helper
Post Interruption Recovery
Structured Workflow Dialogue
Multi Turn Spoken Dialogue
Speech Instruction Following
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY 4.0 for the released conversations, embedded synthesized audio, rubrics, and baseline responses. The evaluation repository is Apache 2.0. The paper says its per-conversation voices come from Common Voice, but does not identify the Common Voice release or voice records, the TTS system, or separate synthesis terms; retain attribution and review voice, privacy, and upstream conditions before redistribution or commercial use.
- Download notes
- The helper saves the dataset card, Hugging Face API metadata, paper page, toolkit README, license, and repository metadata by default. The 79,236-byte GPT-4o Audio baseline-response table is a separate opt-in; the 216,471,312-byte conversation table with embedded 16 kHz user audio is another opt-in. Cloning the evaluation toolkit is also opt-in. Full inference and LLM judging require model access and provider credentials; the helper does not collect or configure API keys.
Safe-first helperscripts/download/ihbench.sh
Speech recognition
IMDA National Speech Corpus
IMDA National Speech Corpus (NSC)
Safe-first helper
Singapore English Asr
Singlish Asr
Conversational Asr
Accent Aware Tts Evaluation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- singapore_open_data_licence
- Code license
- not_released
- License caution
- IMDA states that NSC is provided under the Singapore Open Data Licence, and the third-party mirror repeats that provenance. Review the current official licence and registration terms before use; the mirror does not create new rights or replace IMDA's authoritative release. The paper is CC BY 4.0, but that article licence does not release its curated 92-speaker selection or evaluation artifacts.
- Download notes
- IMDA NSC is one six-part Singapore English corpus family, not six independent benchmarks. IMDA describes an approximately 1.2 TB release of audio and transcripts, last expanded to all six parts in July 2021. Official access is free under the Singapore Open Data Licence but requires a registration form and Dropbox account; IMDA shares the folder by email. An ungated Mesolitica Hugging Face mirror is also public, explicitly identifies itself as a mirror, and reports per-part duration estimates. The 2026 Singlish TTS paper derives a 50-speaker, 55.50-hour in-domain selection from mirrored Part 3 and adds 42 held-out speakers, with a 60/20/20 per-speaker split for the in-domain condition. It does not release the selected speaker IDs, split manifests, prompts, generated speech, or score records, so that paper-specific evaluation protocol is not independently downloadable. The helper saves official and mirror documentation by default; the multi-terabyte mirror snapshot requires explicit opt-in.
Safe-first helperscripts/download/imda_nsc.sh
Speaker, identity & emotion
In-the-Wild
In-the-Wild: A Deepfake Detection Dataset
Safe-first helper
Audio Deepfake Detection
Synthetic Speech Detection
Cross Dataset Generalization
In The Wild Anti Spoofing
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_conflicting_official_signals_apache-2.0_vs_cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- The dataset owner's official page says the dataset and documentation are Apache License 2.0, while the current official Hugging Face card declares CC BY-SA 4.0. Treat the terms as unresolved and obtain clarification before redistribution or commercial use. Because the audio was collected from online videos of identifiable people, source copyright, platform terms, voice and likeness rights, and privacy obligations may also apply.
- Download notes
- The public, ungated release contains 37.9 hours of English recordings from 58 celebrities and politicians, including 17.2 hours of deepfakes collected from online videos. The owner page reports 20.8 hours of bona fide speech while the current Hugging Face card reports 20.7, apparently a rounding discrepancy. It is an evaluation-only cross-dataset benchmark without a prescribed training split. The helper downloads official pages, the dataset card, and Hub API metadata by default. The approximately 7.60 GiB ZIP archive requires IN_THE_WILD_DOWNLOAD_HF=1.
Safe-first helperscripts/download/in_the_wild_audio_deepfake.sh
Audiovisual & cross-modal
InCarEmo
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Manual or gated
Speech Emotion Recognition
Multimodal Emotion Recognition
Fatigue Detection
Driver Distraction Monitoring
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_academic_research
- Code license
- not_specified
- License caution
- The paper states that InCarEmo is licensed for non-commercial academic research and use, but neither the official repository nor linked Drive landing page provides a full license text. The repository contains no LICENSE file and no released code at the time checked. ArXiv's CC BY 4.0 license covers the paper, not the feature data.
- Download notes
- The paper describes synchronized Chinese RGB and infrared video, in-cabin audio, and dialogue text from 25 participants, with six emotion classes plus fatigue and distraction labels. It also defines an auxiliary English setting using translated text and synthesized English speech. Because of participant privacy, the official repository releases the dataset only as feature-level data through a public Google Drive folder, not as raw audio or video. The helper saves the public repository and paper metadata, then prints the manual Drive step; it does not download participant data.
Safe-first helperscripts/download/incaremo.sh
Speech recognition
Indic DiarBench
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Safe-first helper
Speaker Attributed Asr
Multilingual Speaker Diarization
Joint Diarization And Asr
Code Mixed Speech Recognition
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_source_media_rights_caveat
- Code license
- not_released
- License caution
- The official dataset card declares CC BY 4.0 for the release. The paper says the in-the-wild condition was curated from publicly available YouTube videos, but neither primary source separately analyzes the underlying recordings' copyright or platform terms. Retain attribution and review source-media rights before redistribution or commercial use. No separate baseline or evaluation code repository is linked in version 1.
- Download notes
- Indic DiarBench is one public test-only benchmark family spanning all 22 scheduled Indian languages. Its 1,164 recordings total about 108 hours across near-field meetings, far-field recordings, and in-the-wild audio, with human-corrected, time-aligned, speaker-attributed transcripts. It evaluates joint diarization and ASR with DER, cpWER, and WDER; the dataset card reports 485 meeting speakers, approximately 750 in-the-wild speakers, and 12.8% average overlap. The helper downloads official paper, card, and repository metadata by default. The Hugging Face API reports approximately 30.3 GB of repository storage, so the full snapshot requires INDIC_DIARBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/indic_diarbench.sh
Speech recognition
IndicContextEval
IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Safe-first helper
Contextual Asr
Multilingual Asr
Named Entity Recognition
Context Utilisation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card and official repository README declare the benchmark CC BY 4.0. The GitHub repository has no LICENSE file and GitHub detects no repository license, so do not assume that its evaluation outputs or forthcoming evaluation code use the same terms without clarification.
- Download notes
- The public, ungated release contains 16,884 natural-speech utterances totaling 55.93 hours from 555 speakers across Bengali, Gujarati, Hindi, Malayalam, Marathi, Odia, Telugu, and Urdu. It covers 23 professional domains and supplies seven controlled prompt levels: no context, language, structured metadata, natural-language description, English- script entities, native-script entities, and incorrect-domain entities. The helper downloads official documentation, repository metadata, the small published results table, and prompt-taxonomy supplements by default. The Hugging Face card reports a 6.48 GB current download and its API reports approximately 19.64 GB of repository storage including history, so embedded audio requires INDIC_CONTEXT_EVAL_DOWNLOAD_HF=1.
Safe-first helperscripts/download/indic_context_eval.sh
Speech generation
InstructTTSEval
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
Safe-first helper
Controllable Text To Speech
Speech Instruction Following
Acoustic Parameter Control
Descriptive Style Control
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- not_specified
- License caution
- The Hugging Face dataset card lists MIT, while the GitHub repository has no detected license. The paper says reference audio was curated from NCSSD plus movies, TV dramas, variety shows, and other film/television sources; it also says the dataset is solely for academic and research use. Treat the card license as insufficient to clear third-party media rights, and verify source terms before redistribution or commercial use.
- Download notes
- The helper downloads the official GitHub README and Hugging Face dataset card by default. The public, ungated Hugging Face snapshot contains English and Chinese Parquet splits with embedded reference audio and uses about 1.8 GB, so data download requires INSTRUCT_TTS_EVAL_DOWNLOAD_HF=1; cloning the evaluation repository is separately opt-in.
Safe-first helperscripts/download/instruct_tts_eval.sh
Audiovisual & cross-modal
InterPet4D
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
Safe-first helper
Audio Conditioned Motion Generation
Human Pet Interaction Modeling
Multimodal Motion Generation
Animal Behavior Understanding
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_noncommercial_terms
- Code license
- not_applicable
- License caution
- The current public Hugging Face card declares CC BY-NC 4.0 and says participants consented to data release. The paper's ethical statement instead says users must agree to a data-use agreement prohibiting redistribution, surveillance, and biometric identification, but no such agreement or click-through gate is exposed on the current ungated dataset page. Apply the stricter paper restrictions pending clarification, do not redistribute the corpus, and do not infer that arXiv's CC BY 4.0 paper license governs the released participant recordings.
- Download notes
- The public, ungated v1 release contains 227 approximately 17-20-second egocentric clips from 113 interaction sessions involving 13 dogs and about 23 human participants. It provides synchronized 48 kHz stereo MP3 audio, MERT embeddings, and human-hand, human-body, dog-skeleton, and SMAL motion parameters. The helper downloads the official dataset card and API metadata by default; the Hugging Face API reports about 10.7 GB of repository storage, so the full snapshot requires INTERPET4D_DOWNLOAD_HF=1. The paper also describes 12-view and egocentric RGB video, but the current Hugging Face file inventory releases audio and motion artifacts rather than those raw videos.
Safe-first helperscripts/download/interpet4d.sh
Representation & general suites
JapanEEG
JapanEEG: A 1000-hour EEG-EMG-audio dataset of Japanese speech production
Safe-first helper
Overt Speech Neural Decoding
Eeg Speech Representation Learning
Eeg Emg Audio Alignment
Speech Artifact Modeling
+3 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC0-1.0
- Code license
- CC0-1.0
- License caution
- The versioned OpenNeuro dataset_description.json declares CC0 and the companion repository uses CC0-1.0. Cite the dataset DOI and paper despite the waiver. Treat voice, EEG, EMG, demographic, and transcript data as potentially identifying, and do not infer that CC0 eliminates privacy, consent, ethics, or institutional-review responsibilities.
- Download notes
- OpenNeuro version 1.0.0 provides approximately 955 GB in BIDS format. The paper reports 1,020 hours of synchronized scalp EEG, facial EMG, and 48 kHz speech audio from three healthy native Japanese speakers, recorded longitudinally with three 62-128-channel EEG systems. The release includes event timing, ASR transcriptions, participant and channel metadata, and overt, covert, and listening conditions. The helper downloads only lightweight official metadata, documentation, and provenance; obtain selected files or the full corpus through the versioned OpenNeuro page. The companion repository's current README and downloader refer to older ds007602/ds007600 identifiers, so they must not be used as a substitute for the paper's ds007808 release.
Safe-first helperscripts/download/japaneeg.sh
Speech recognition
JASMIN-CGN
JASMIN-CGN: Extension of the Spoken Dutch Corpus with Speech of Elderly People, Children and Non-Natives
Manual or gated
Automatic Speech Recognition
Child Speech Recognition
Accented Speech Recognition
Non Native Speech Recognition
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_signed_license_noncommercial
- Code license
- not_applicable
- License caution
- The official non-commercial product is free but requires a signed license and account login; the owner page directs commercial users to a separate commercial product. Review the current agreement before use or redistribution.
- Download notes
- The Dutch Language Institute distributes the approximately 115-hour Dutch/Flemish speech corpus after account login and a signed license agreement. Its official page says the corpus contains read speech and human-machine dialogues from adolescents, non-native speakers, and seniors, with WAV audio plus TXT and TextGrid annotations. The helper prints the official access steps only; it does not bypass login or download corpus audio. The 2026 evaluation paper selects 120 human-machine-interaction test utterances from children, older adults, and Flemish speakers and also reports ASR results on the full corresponding test sets.
Safe-first helperscripts/download/jasmin_cgn.sh
Speech generation
JVS
Japanese versatile speech corpus
Manual or gated
Speaker Identification
Speaker Verification
Multi Speaker Speech Synthesis
Speaking Style Robustness
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_research_noncommercial_and_personal_use
- Code license
- not_applicable
- License caution
- The owner page permits the audio for academic research, noncommercial research including within commercial organizations, and personal use; commercial use requires contacting the owners. Audio redistribution is prohibited except for small illustrative excerpts such as about ten files. Tags are CC BY-SA 4.0, while texts derive from JSUT and retain its terms. Review voice and biometric-data considerations.
- Download notes
- The official public Google Drive archive is approximately 3.5 GB and contains about 30 hours of 24 kHz studio speech from 100 professional Japanese speakers. Each speaker has 100 shared parallel100 readings, 30 speaker-specific nonpara30 readings, ten whispered utterances, and ten falsetto utterances, plus metadata tags and alignments. The helper saves the official project and paper pages, then prints the manual Drive path; it does not automate the large archive. A July 2026 voice-clone-attribution paper uses parallel100 and a JVS-varied mixture of nonpara30, falset10, and whisper10 as speaker-identification controls, and uses six jvs001 parallel100 utterances as Seed-VC input.
Safe-first helperscripts/download/jvs.sh
Audiovisual & cross-modal
K9-Bench
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
Safe-first helper
Audio Visual Question Answering
Long Video Understanding
Animal Behavior Understanding
Temporal Reasoning
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_custom_gated_terms
- Code license
- not_specified
- License caution
- The dataset card declares CC BY-NC 4.0, while its required access terms additionally prohibit commercial use, sharing, redistribution, and derivative datasets without written permission. The authors disclaim ownership of the linked YouTube media; original uploader copyright, YouTube terms, and availability remain applicable. The evaluation repository has no LICENSE file, so its code terms are unspecified.
- Download notes
- The auto-approved gated Hugging Face release contains 4,744 English five-choice question-answer pairs over 907 domestic-dog YouTube videos. Its approximately 1.8 MB parquet table includes source video URLs, questions, options, categories, and answer keys but does not mirror the videos. The five task categories cover posture, action sequence, context, cause-effect, and interaction analysis. The project also evaluates audio-plus-video input and finds modest gains for action sequences. The helper saves public first-party documentation by default; after accepting the dataset terms and authenticating with Hugging Face, an explicit acknowledgment can fetch only the lightweight parquet metadata. It never downloads the linked YouTube videos.
Safe-first helperscripts/download/k9_bench.sh
Speech recognition
KeSpeech
KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects
Manual or gated
Mandarin Asr
Dialect Asr
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_no_adaptations_no_distribution
- Code license
- not_specified
- License caution
- Downloading data means accepting the custom license in dataset_license.md.
Safe-first helperscripts/download/kespeech.sh
Speech generation
Kimi-Audio-GenTest
Kimi-Audio Generation Test Set for Speech Conversation
Safe-first helper
End To End Speech Conversation
Spoken Instruction Following
Speech Emotion Control
Speech Speed Control
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_declared_by_dataset_card
- Code license
- MIT
- License caution
- The Hugging Face card declares MIT for the dataset, but it does not separately document prompt-audio provenance or recording rights. The toolkit's MIT license applies to its code and does not license third-party model outputs or absent human judgments. Review voice, content, and downstream model terms before redistribution or reuse.
- Download notes
- The owner release contains 191 Chinese spoken instructions and a metadata JSONL with transcript, ability, and audio filename. Its six released ability labels are speed (12), accent (51), explicit emotion (60), empathy or implicit emotion (48), storytelling (10), and tongue twister (10). Section 6.2.4 and Table 7 compare Kimi-Audio with GPT-4o, Step-Audio-chat, GLM-4-Voice, and GPT-4o-mini using 1-5 human ratings for speed, accent, emotion, empathy, and style control. The helper downloads official documentation and the 28,488-byte prompt manifest by default; the full 191-WAV Hub snapshot and toolkit clone require separate opt-ins.
Safe-first helperscripts/download/kimi_audio_gentest.sh
Audio understanding, generation & events
Kitchen20
Kitchen20 Kitchen Sound Dataset
Safe-first helper
Environmental Sound Classification
Kitchen Sound Classification
Domestic Audio Classification
Access pathOfficial / other
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- Apache-2.0
- License caution
- The official repository provides an Apache License 2.0 file covering the released repository. The class-source table identifies source recordings and offsets but does not provide a per-recording license inventory, so users should preserve provenance and review source-audio rights before redistribution or commercial use.
- Download notes
- The official public repository's kitchen20.csv lists 1,070 audio clips in 20 kitchen classes across nine folds. The recent TriA evaluation uses the 800 clips in folds 1-5: folds 1-3 train, fold 4 validates, and fold 5 tests. The helper downloads the approximately 67 KB split metadata, class-source table, license, and repository metadata by default. Cloning the approximately 325 MB repository, including WAV audio and baseline code, requires explicit opt-in.
Safe-first helperscripts/download/kitchen20.sh
Speech recognition
KSC2
KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus
Safe-first helper
Automatic Speech Recognition
Kazakh Speech Recognition
Code Switched Speech Recognition
Low Resource Asr
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- cc-by-4.0
- License caution
- The ISSAI owner page and corpus DOI state CC BY 4.0 and explicitly permit academic and commercial use with attribution. The Hugging Face card's machine-readable MIT tag conflicts with the owner terms; use CC BY 4.0 for the corpus. The linked recipe repository is also marked CC BY 4.0, but component broadcast, podcast, parliament, KSC, and KazakhTTS2 source rights should still be reviewed for redistribution.
- Download notes
- The public, ungated Hugging Face release contains around 1,200 hours and more than 600,000 transcribed Kazakh utterances, including Kazakh-Russian code-switching. It is stored as ten multipart archive files totaling about 80.8 GB. The helper downloads the official pages, dataset card, API metadata, and repository README/license by default; the full snapshot requires KSC2_DOWNLOAD_HF=1. GigaAM Multilingual section 3.3 names KSC2 as public fine-tuning data, not as one of its Common Voice/FLEURS evaluation sets.
Safe-first helperscripts/download/ksc2.sh
Speaker, identity & emotion
KVoiceBench / KOpenAudioBench / KMMAU
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
Safe-first helper
Korean Spoken Question Answering
Speech Instruction Following
Speech Reasoning
Speech Safety Evaluation
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_Apache-2.0_and_CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The KVoiceBench and KOpenAudioBench cards declare Apache-2.0; KMMAU declares CC BY-NC-SA 4.0. These benchmarks adapt prior benchmark content or source speech corpora, so the declared repository licenses may not replace upstream terms. Review VoiceBench/OpenAudioBench provenance and the KSS, KMSAV, and Seoul Corpus conditions before redistribution or commercial use. Raon-Eval code is Apache-2.0.
- Download notes
- The three public, ungated Hugging Face releases contain 12,345 Korean test samples: 7,306 KVoiceBench spoken-QA items transferred from VoiceBench, 2,835 KOpenAudioBench items transferred from OpenAudioBench, and 2,204 KMMAU audio-understanding items derived from KSS, KMSAV, and Seoul Corpus. The helper downloads official cards, repository metadata, and the paper page by default. Full snapshots require KOREAN_SPEECHLM_BENCHMARKS_DOWNLOAD_HF=1 because Hugging Face reports approximately 4.64 GB, 608 MB, and 4.49 GB of repository storage respectively.
Safe-first helperscripts/download/korean_speechlm_benchmarks.sh
Speech recognition
L2-ARCTIC
L2-ARCTIC: A Non-native English Speech Corpus
Manual or gated
Asr
Accented Speech Recognition
Mispronunciation Detection
Accent Conversion
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- Official homepage states the corpus is released under CC BY-NC 4.0; usage outside that license requires contacting the TAMU dataset owner. The current release covers 24 non-native English speakers plus suitcase-corpus material, while the Interspeech 2018 paper describes the initial v1.0 release.
- Download notes
- Official access requires reviewing the CC BY-NC 4.0 license terms and submitting the download form with name, email, and affiliation. The project sends a generated Google Drive link by email, so the helper prints manual access steps rather than storing or using private generated URLs.
Safe-first helperscripts/download/l2_arctic.sh
Enhancement, separation & quality
L3DAS21
L3DAS21 Challenge: Machine Learning for 3D Audio Signal Processing
Safe-first helper
Three Dimensional Speech Enhancement
Speech Enhancement
Sound Event Localization And Detection
Sound Source Localization
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- DataCite lists CC BY 4.0 for the Zenodo dataset. The generated mixtures incorporate LibriSpeech speech and FSD50K sound events, so preserve LibriSpeech attribution and review FSD50K's per-clip Creative Commons terms, including non-commercial clips. The official GitHub repository exposes no LICENSE file or GitHub-detected license, so clarify code terms before reuse or redistribution.
- Download notes
- The public Zenodo V1 release is a 65-hour corpus of multi-source, multi-perspective B-format Ambisonics audio generated from impulse responses measured at 252 positions with two first-order Ambisonics microphones. Task 1 contains more than 30,000 spatial speech mixtures for 3D speech enhancement, with clean monaural speech targets. Task 2 contains 900 one-minute soundscapes for sound-event localization and detection, with up to three simultaneous events and 100 ms activity, class, and Cartesian-location targets. Both tasks have one- and two-microphone tracks. The helper downloads official pages, repository metadata, DataCite metadata, and the paper by default; cloning the code and downloading selected large archives are separate opt-ins.
Safe-first helperscripts/download/l3das21.sh
Enhancement, separation & quality
L3DAS22
L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
Safe-first helper
Three Dimensional Speech Enhancement
Speech Enhancement
Sound Event Localization And Detection
Sound Source Localization
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- Kaggle's official API lists CC BY 4.0 for L3DAS22. The generated mixtures incorporate LibriSpeech speech and FSD50K sound events; preserve LibriSpeech attribution and check FSD50K's per-clip Creative Commons licenses, including non-commercial clips. The official GitHub repository exposes no LICENSE file or GitHub-detected license, so clarify code terms before redistribution.
- Download notes
- The public Kaggle release contains 105,757,713,362 bytes of multi-source, multi-perspective B-format Ambisonics audio. Task 1 has more than 40,000 spatial speech mixtures totaling nearly 90 hours at 16 kHz, with up to three overlapping background noises and clean monaural speech targets. Task 2 has 900 30-second soundscapes totaling 7.5 hours at 32 kHz, with 14 office-relevant event classes, up to three overlaps, and 100 ms activity and Cartesian-location targets. Both tasks provide one- and two-microphone tracks using one or two first-order Ambisonics microphones. The helper downloads official pages, repository metadata, the paper, documentation, and Kaggle API metadata by default; cloning the code and downloading the 105.8 GB dataset are separate opt-ins. Full data download requires the Kaggle CLI and account credentials.
Safe-first helperscripts/download/l3das22.sh
Audio understanding, generation & events
LAT-Bench
LAT-Bench: Long-form Audio Temporal Awareness Benchmark
Safe-first helper
Long Form Audio Understanding
Temporal Audio Reasoning
Dense Audio Captioning
Temporal Audio Grounding
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_claim_with_upstream_media_terms
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for LAT-Bench, while the GitHub repository is Apache-2.0. The release references externally hosted source audio rather than relicensing or redistributing it; review each recording's rights and platform terms before retrieval, redistribution, or commercial use.
- Download notes
- The public, ungated release provides Chinese and English metadata plus task JSONL files for a human-verified 40-hour benchmark, with 25 hours of Chinese and 15 hours of English audio up to 30 minutes long. Dense Audio Captioning, Temporal Audio Grounding, and Targeted Audio Captioning annotations are included. Audio is not bundled; the metadata records source URLs, so availability can drift and source-platform terms apply. The helper downloads the approximately 2.6 MB annotations, official documentation, and repository metadata by default; cloning the evaluation repository is opt-in.
Safe-first helperscripts/download/lat_bench.sh
Representation & general suites
LFR Benchmarking Dataset Factory
Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities
Safe-first helper
Voice Type Classification
Addressee Classification
Child Vocal Maturity Classification
Child Speech Recognition
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_source_specific_and_participant_consent_terms
- Code license
- MIT
- License caution
- The repository's MIT license covers the benchmark-generation software and bundled metadata, not the underlying child recordings or annotations. Each contributing corpus retains its own access, consent, ethics, and reuse conditions. The paper is CC BY-SA 4.0; that publication license does not authorize redistribution of the sensitive source audio or derived clips.
- Download notes
- This is one benchmark-construction family with four derived tracks, not 27 separately released benchmark families. The public pipeline standardizes long-form child-centered recordings in ChildProject and DataLad, maps heterogeneous human annotations, and creates child-disjoint train, validation, and test splits for voice type, addressee, vocal maturity, and orthographic transcription tasks. The paper inventories 27 contributing corpora across more than 18 languages and 14 countries: six source corpora are described as public, while 21 require access controls. No fixed all-corpora benchmark archive is publicly downloadable. Safe defaults fetch only the paper, repository documentation, license, API metadata, and the lightweight manual annotation-column index; cloning the pipeline is opt-in and never fetches source recordings.
Safe-first helperscripts/download/lfr_benchmarking_factory.sh
Speech recognition
Libri-Light
Libri-Light: A Benchmark for ASR with Limited or No Supervision
Safe-first helper
Automatic Speech Recognition
Self Supervised Speech Representation
Semi Supervised Asr
Unsupervised Speech Representation
+1 more
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The GitHub repository says Libri-Light code is MIT, but the reviewed README and data-preparation docs do not state a standalone data license. The paper describes the data as derived from open-source LibriVox audiobooks, so re-check per-source audiobook terms and attribution requirements before redistribution.
- Download notes
- The helper downloads official README/license/data-preparation/evaluation docs by default. Limited-supervision finetuning data is about 0.6 GiB, ABX item data is small, and unlabeled audio ranges from 35 GiB small to 321 GiB medium and 3.05 TiB large, so all data archives require explicit opt-in.
Safe-first helperscripts/download/libri_light.sh
Speech recognition
LibriCSS
LibriCSS: Continuous Speech Separation Dataset
Safe-first helper
Continuous Speech Separation
Overlapped Speech Recognition
Multi Channel Asr
Speaker Diarization
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_upstream_basis_data_license_not_separately_stated
- Code license
- MIT
- License caution
- The repository LICENSE applies MIT to the software and documentation and identifies the source LibriSpeech corpus as CC BY 4.0, but it does not state a separate license for the replayed LibriCSS recordings. Preserve LibriSpeech attribution and confirm recording-level rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains ten approximately one-hour sessions made from LibriSpeech utterances replayed through eight loudspeakers and captured with a seven-channel circular microphone array in an office meeting room. Each session has six ten-minute mini-sessions spanning 0%-40% overlap, including separate short-gap and long-gap 0% conditions. The official repository provides preparation, Kaldi ASR, and WER-scoring tools for utterance-wise and continuous-input evaluation. The helper downloads official documentation and repository metadata by default; the direct Google Drive archive is 6,407,297,637 bytes (about 5.97 GiB) and requires LIBRICSS_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/libricss.sh
Speech recognition
LibriHeavy
LibriHeavy: A 50,000-Hour ASR Corpus with Punctuation, Casing, and Context
Safe-first helper
Automatic Speech Recognition
Long Form Speech Recognition
Contextual Speech Recognition
Punctuation And Casing Prediction
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The official repository LICENSE and Hugging Face card declare Apache 2.0 for the released code and metadata. LibriHeavy does not repackage the audio, and that license does not supersede Libri-Light or the underlying LibriVox audiobook terms. Preserve source attribution and verify audiobook-specific rights before redistribution or commercial use.
- Download notes
- LibriHeavy aligns and labels the Libri-Light audiobook audio with punctuation, casing, and preceding-text context. It provides nested 509-, 5,042-, and 50,794-hour training subsets plus 22.3-hour dev, 10.5-hour test-clean, and 11.5-hour test-other sets with no overlapping training speakers or books. The helper downloads official documentation and API metadata by default. Selected compressed manifests require LIBRIHEAVY_DOWNLOAD_MANIFESTS=1; the Hub inventory is approximately 15.9 GB because it includes multi-gigabyte large manifests and duplicate raw variants. Audio is not stored in the Hub repository and must be obtained separately through Libri-Light. The July 2026 Apple report names LibriHeavy for within-speaker ABX codec evaluation but does not identify its split or release its ABX items.
Safe-first helperscripts/download/libriheavy.sh
Enhancement, separation & quality
LibriMix
LibriMix: An Open-Source Dataset for Generalizable Speech Separation
Safe-first helper
Speech Separation
Speech Enhancement
Noisy Speech Separation
Source Separation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_derived
- Code license
- MIT
- License caution
- The LibriMix repository license is MIT for code/scripts. Generated data is derived from LibriSpeech, which OpenSLR lists as CC BY 4.0, plus WHAM noise; re-check WHAM terms and cite all upstream components before redistribution or commercial use.
- Download notes
- The helper clones the official generator/metadata repository by default. Running generation downloads LibriSpeech clean subsets and WHAM noise, then creates many mixtures; the upstream README estimates about 430 GiB for Libri2Mix plus 332 GiB for Libri3Mix, with additional source storage, so generation requires an explicit opt-in and an external storage directory.
Safe-first helperscripts/download/librimix.sh
Speaker, identity & emotion
LibriSeVoc
LibriSeVoc: LibriTTS Self-Vocoding Dataset for Neural Vocoder Artifact Detection
Safe-first helper
Audio Deepfake Detection
Synthetic Speech Detection
Vocoder Artifact Detection
Neural Vocoder Identification
+2 more
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_with_libritts_cc_by_4.0_upstream
- Code license
- mit
- License caution
- The author repository's MIT license explicitly covers the repository, while neither the primary paper nor download instructions state a separate license for the 45 GB audio archive. The author-owned Hugging Face metadata-only page declares CC BY-SA 4.0, but that does not clearly license the separately hosted archive. Bona fide speech is derived from LibriTTS, whose source release is CC BY 4.0. Treat archive rights as unspecified and preserve LibriTTS attribution and upstream terms; clarify generated-audio redistribution and commercial use with the authors.
- Download notes
- LibriSeVoc is one English self-vocoding benchmark family, not six independent datasets. It pairs 13,201 LibriTTS utterances with one reconstruction from each of WaveNet, WaveRNN, WaveGrad, DiffWave, MelGAN, and Parallel WaveGAN. The paper reports 92,407 clips at 24 kHz, spanning 34.92 hours of bona fide and 208.74 hours of synthesized speech, with non-overlapping 60/20/20 train, development, and test splits of 55,440, 18,480, and 18,487 clips. Evaluation covers binary detection and vocoder identification, intra- and cross-dataset EER, plus resampling/noise robustness. The official archive is public and approximately 45.0 GB. The helper downloads only the paper page, author documentation, repository license/API metadata, and lightweight dataset-card metadata by default; the archive and toolkit clone are separate opt-ins.
Safe-first helperscripts/download/librisevoc.sh
Speech recognition
LibriSpeech
LibriSpeech ASR corpus
Safe-first helper
Asr
Speech Reconstruction
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR12 and HF dataset card list CC BY 4.0.
- Download notes
- The helper downloads the OpenSLR landing page and checksums by default; corpus archives require LIBRISPEECH_DOWNLOAD_ARCHIVES=1. Qwen3-TTS section 4.1.2 evaluates tokenizer reconstruction on all 2,620 utterances in test-clean with PESQ, STOI, UTMOS, and WavLM-based speaker similarity. Qwen-Audio-VAE sections 4.1-4.3 use LibriSpeech for speech reconstruction and ablation evaluation.
Safe-first helperscripts/download/librispeech.sh
Speech generation
LibriTTS
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Safe-first helper
Text To Speech
Speech Synthesis
Voice Cloning
Multi Speaker Speech Synthesis
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR60 lists CC BY 4.0. LibriTTS is derived from LibriSpeech, which in turn derives from LibriVox audio and Project Gutenberg text.
- Download notes
- OpenSLR lists seven archives totaling tens of GiB; the helper saves the official OpenSLR page by default and requires LIBRITTS_DOWNLOAD_ARCHIVES=1 before downloading archives.
Safe-first helperscripts/download/libritts.sh
Speech recognition
LibriWASN
LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices
Safe-first helper
Meeting Transcription
Continuous Speech Separation
Speaker Diarization
Multi Channel Asr
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- Zenodo declares the LibriWASN release CC BY 4.0 and includes the license text; the official tools repository is MIT. The recordings replay LibriSpeech/LibriCSS material, so preserve that provenance and review upstream speech-data terms when redistributing derived data.
- Download notes
- The public release contains 20 hours of meeting-like recordings in two rooms with approximately 200 ms and 800 ms reverberation times. Five smartphones and four microphone arrays provide 29 asynchronous audio channels, with 0%-40% overlap conditions and ground-truth diarization. The same LibriSpeech sentences and speakers as LibriCSS were replayed. The helper downloads official metadata, README/license files, paper page, and repository documentation by default. Twelve per-room and per-overlap ZIP archives total about 55.8 GB and require explicit opt-in; the official repository's full downloader also fetches LibriCSS as a transcription/reference dependency.
Safe-first helperscripts/download/libriwasn.sh
Speech recognition
Live Gurbani Captioning Benchmark v1
Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan
Safe-first helper
Closed Vocabulary Singing Captioning
Live Audio Tracking
Lyrics Alignment
Sung Scripture Identification
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_annotations
- Code license
- MIT
- License caution
- The repository LICENSE applies CC BY 4.0 to ground-truth annotations in test/ and the committed baselines, and MIT to the software and documentation. The four underlying Kirtan recordings are referenced by YouTube ID rather than bundled; the repository license does not grant rights to those recordings, and platform terms, uploader rights, performer rights, and local cultural or legal considerations remain separate.
- Download notes
- The public v1 repository contains ground-truth timelines for four hand-reviewed Sikh Kirtan recordings, each evaluated from 0%, 33%, and 66% start offsets, giving 12 cases and approximately 57 minutes of scored audio. It also releases a standard-library Python scorer, visualization and annotation tools, and empty, shifted, and perfect baselines. The primary metric is one-second frame accuracy with a one-second boundary collar. The helper downloads official documentation and repository metadata by default; cloning the small repository with annotations and evaluation code is opt-in. Audio is not redistributed: the official README identifies four YouTube video IDs and provides yt-dlp/ffmpeg preparation commands, so availability and source-media rights must be checked separately.
Safe-first helperscripts/download/live_gurbani_captioning_v1.sh
Speech generation
LJSpeech
The LJ Speech Dataset
Safe-first helper
Text To Speech
Speech Synthesis
Single Speaker Speech Synthesis
Asr
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- public_domain
- Code license
- not_applicable
- License caution
- Official page says text, audio, and annotations are public domain; the HF mirror lists unlicense.
- Download notes
- The official archive is about 2.6 GiB. The helper saves the official dataset page by default and requires LJSPEECH_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/ljspeech.sh
Audiovisual & cross-modal
LLP
Look, Listen, and Parse: Audio-Visual Video Parsing Dataset
Safe-first helper
Audio Visual Video Parsing
Temporal Event Localization
Audio Event Detection
Visual Event Detection
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- unclear
- License caution
- The official repository does not contain a license file and GitHub reports no detected license. Its README says GPLv3 but links to a GPLv3 license in an unrelated repository, so neither that statement nor the linked file establishes clear terms for LLP annotations or code. Obtain clarification before redistribution or commercial use; source videos also retain their original rights and YouTube terms.
- Download notes
- The official repository releases lightweight weak-label train/validation/test CSVs and dense audio and visual event annotations for validation and test. The full CSV contains 11,849 ten-second YouTube segment references; the published split files contain 10,000 train, 649 validation, and 1,200 test rows. The helper downloads documentation and all annotation CSVs by default. Extracted VGGish, ResNet-152, and R(2+1)D features remain a manual Google Drive download, while raw videos must be reconstructed from referenced YouTube segments subject to availability and platform terms.
Safe-first helperscripts/download/llp.sh
Audio understanding, generation & events
LOCATA
LOCATA: IEEE-AASP Challenge on Acoustic Source Localization and Tracking
Safe-first helper
Sound Source Localization
Acoustic Source Tracking
Multi Source Localization
Moving Source Localization
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ODC-By-1.0
- Code license
- not_specified
- License caution
- Zenodo identifies the final dataset release as Open Data Commons Attribution, and the official corpus page says both LOCATA and its VCTK speech material use Open Data Commons terms. Neither official MATLAB repository exposes a LICENSE file or a GitHub-detected license, so clarify software terms before redistribution.
- Download notes
- The open final release contains corrected development and evaluation datasets with close-talking speech, distant multichannel recordings from four microphone-array configurations, and OptiTrack ground truth for sources and microphones. Its six tasks span static and moving, single- and multi-source scenarios with static or moving arrays. The dev.zip and eval.zip archives total about 19.3 GB, so the helper downloads only official pages, documentation, Zenodo metadata, and tool metadata by default.
Safe-first helperscripts/download/locata.sh
Audiovisual & cross-modal
LRRo
LRRo: A Lip Reading Data Set for the Under-resourced Romanian Language
Safe-first helper
Visual Speech Recognition
Lip Reading
Isolated Word Recognition
Low Resource Romanian Speech
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- The Zenodo record declares CC BY 4.0. Wild LRRo was collected from Internet videos and the release contains identifiable speaker imagery, so attribution, privacy, likeness, and any source-media rights should still be reviewed for redistribution, biometric use, or deployment. The documentation-only GitHub repository has no detected license.
- Download notes
- The public, ungated Zenodo release contains Wild LRRo, with more than 20 hours, over 35 speakers, 1,100 word instances, and a 21-word vocabulary, plus Lab LRRo, with more than 5 hours, 19 speakers, 6,400 word instances, and a 48-word vocabulary. Both isolated-word collections provide train, validation, and test subsets. VSRo-200 section 4.6 evaluates transfer on both LRRo subsets. The helper saves the official Zenodo record and repository documentation by default; the 291,504,956-byte archive is opt-in with LRRO_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/lrro.sh
Audiovisual & cross-modal
LRS2-BBC
The Oxford-BBC Lip Reading Sentences 2 Dataset
Manual or gated
Audio Visual Speech Recognition
Visual Speech Recognition
Lip Reading
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- non-commercial_academic_access_agreement
- Code license
- not_applicable
- License caution
- BBC R&D restricts LRS2 to non-commercial research by universities, reputable academic institutions, and relevant public organizations; companies and independent researchers are not eligible. The signed agreement controls use, and the BBC asks researchers to obtain approval before publishing sample images. Review the current agreement and BBC broadcast-content rights before use.
- Download notes
- The official VGG page describes 144,482 pre-train/train/validation/test utterances from BBC television, with date-disjoint validation and test broadcasts, and links a password-protected 50 GB package plus public split file links. Access requires submitting the BBC LRS2 Data Sharing Agreement from an official academic address; approved researchers receive a countersigned agreement and password. The helper saves the official pages, agreement, and paper metadata only and does not request credentials or download the corpus.
Safe-first helperscripts/download/lrs2.sh
Audiovisual & cross-modal
LRS3-TED
LRS3-TED: A Large-Scale Dataset for Visual Speech Recognition
Manual or gated
Audio Visual Speech Recognition
Visual Speech Recognition
Lip Reading
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- unspecified
- Code license
- not_applicable
- License caution
- No LRS3 data license is stated on the currently accessible official landing page or paper record. TED/TEDx source-video rights remain with their owners; obtain the official terms before use and do not infer reuse rights from third-party mirrors.
- Download notes
- The paper introduces more than 400 hours of aligned face tracks, audio, subtitles, and word boundaries from TED and TEDx for visual and audio-visual speech recognition. The official VGG landing page still identifies LRS3 and links its dataset page, but that linked page returned HTTP 404 when checked on 2026-07-22. The helper preserves the live official landing page and paper metadata and reports the unavailable official download route; it does not substitute an unofficial mirror.
Safe-first helperscripts/download/lrs3.sh
Audiovisual & cross-modal
LVOmniBench
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Manual or gated
Audio Visual Question Answering
Long Video Understanding
Cross Modal Reasoning
Temporal Localization
+2 more
Access pathHugging Face
Upstream termsNot specified
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The paper says every source video carries a Creative Commons license, but neither the official dataset card nor repository states a benchmark-level license or records each video's exact Creative Commons variant in public documentation. The GitHub repository has no LICENSE file or detected license. Review the access form, per-video rights, attribution requirements, and YouTube terms before reuse or redistribution; the paper's CC BY 4.0 license covers the article, not automatically the dataset or evaluation code.
- Download notes
- The gated Hugging Face release contains 275 English Creative Commons-licensed YouTube videos totaling 140 hours, with durations from 10 to 90 minutes, plus 1,014 manually authored four-option question-answer pairs. Questions span nine categories and require joint reasoning over speech, music, or sound with visual evidence. The helper downloads public repository documentation and API metadata by default. The Hugging Face API reports approximately 187.4 GB of repository storage, so the complete snapshot requires approval, authentication, and LVOMNIBENCH_DOWNLOAD_HF=1. OmniReasoner evaluates LVOmniBench in its main audio-visual table and reports duration-wise tool-use behavior on the benchmark.
Safe-first helperscripts/download/lvomnibench.sh
Enhancement, separation & quality
Lyra-SA
Lyra Lab Singing Assessment Dataset
Manual or gated
Singing Quality Assessment
Full Song Singing Assessment
Singing Score Prediction
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- not_applicable
- License caution
- The official page states CC BY-NC 4.0 for non-commercial use, requires attribution to the source page and notice, reserves copyright to Tencent Music Entertainment Group, and requires separate permission for commercial use. The recordings are WeSing user performances of copyrighted songs; rely on the official authorization and application terms rather than inferring broader music or performance rights from the Creative Commons label.
- Download notes
- The official Tencent Music Lyra Lab page describes 1,000 complete mobile-karaoke recordings: 100 user covers for each of 10 Chinese songs, with no repeated singer. The release includes 44.1 kHz 16-bit mono WAV audio, listener-provided overall singing scores, rough singer labels, timed lyrics, and reference MIDI. Access is application-based: users must complete the official form and accept its terms, after which Lyra Lab says it emails a download link within three days. The helper saves official documentation and the recent SongSQA paper, then prints the manual application path; it never guesses or bypasses an emailed archive URL.
Safe-first helperscripts/download/lyra_sa.sh
Speaker, identity & emotion
M6
M6: Multi-Generator, Multi-Domain, Multi-Lingual and Cultural, Multi-Genres, Multi-Instrument Machine-Generated Music Detection Databases
Safe-first helper
Machine Generated Music Detection
Synthetic Music Detection
Music Generator Attribution
Cross Domain Robustness
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- Not specified in the source record.
- License caution
- The Hugging Face repository has no dataset card, license metadata, or license file. The article is CC BY 4.0, but that publication license does not license the archive. M6 combines AI outputs with human music drawn from GTZAN, FMA, COSIAN, MISD, and online services; the article discusses non-commercial and redistribution conditions but does not provide a per-file rights manifest. Review every component corpus, generator, service, composition, performance, and recording term before redistribution or commercial use.
- Download notes
- The public, ungated release contains 13,493 WAV examples across six machine-generated-music detection settings spanning instruments, multilingual and culturally varied songs, genres, long-form music, and general music. The paper reports 4,299 human and 9,194 AI-generated examples and fixed 64/16/20-percent train/validation/test proportions within each setting. The helper downloads the official article, DOI metadata, Hugging Face repository metadata, file inventory, and LFS rules by default. The sole data archive is approximately 30.2 GB and requires M6_DOWNLOAD_HF=1. A July 2026 follow-up uses an unreleased balanced 1,000-sample, ten-genre subset for aesthetics-bias analysis.
Safe-first helperscripts/download/m6_music_detection.sh
Audio understanding, generation & events
MACS
MACS: Multi-Annotator Captioned Soundscapes
Safe-first helper
Audio Captioning
Audio Tagging
Multi Annotator Audio Labeling
Acoustic Scene Captioning
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other-nc
- Code license
- not_applicable
- License caution
- Zenodo lists MACS as Other (Non-Commercial), and LICENSE.txt grants experimental non-commercial use with attribution to Tampere University/Machine Listening Group. Audio files come from TAU Urban Acoustic Scenes 2019, whose Zenodo record also lists Other (Non-Commercial).
- Download notes
- The helper downloads the MACS annotations, competence scores, license, and TAU Urban Acoustic Scenes 2019 docs/metadata by default. The upstream TAU 2019 audio shards are large, so audio download is an explicit opt-in.
Safe-first helperscripts/download/macs.sh
Music
MADB
MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
Safe-first helper
Music Aesthetic Assessment
Music Quality Assessment
Music Score Regression
Multimodal Music Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_upstream_audio_rights
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 and research-only use, but also warns that audio may remain subject to original copyright restrictions. No license file is present in the official GitHub repository. Apply the non-commercial dataset terms and independently review MuChin, generator-service/output, and unidentified online-source rights before redistributing or using audio.
- Download notes
- The paper describes 9,999 tracks rated by 30 trained annotators, with about ten ratings per track across ten perceptual dimensions and an overall score plus comments and tags. The public, ungated Hugging Face release includes all annotations and 1,730 Suno/Levo audio tracks; its card points to the separate MuChin repository for another 4,400 tracks and says remaining tracks came from diverse online sources. The helper downloads official documentation and repository metadata by default, makes the approximately 69 MB annotation tables opt-in, and requires MADB_DOWNLOAD_HF=1 for the approximately 18.6 GB Hugging Face snapshot.
Safe-first helperscripts/download/madb.sh
Representation & general suites
MAEB
MAEB: Massive Audio Embedding Benchmark
Safe-first helper
Audio Embedding Evaluation
Audio Text Retrieval
Audio Classification
Audio Clustering
+6 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_component_terms
- Code license
- Apache-2.0
- License caution
- Apache-2.0 covers the MTEB software and benchmark registry, not the underlying MAEB task datasets. The 30 tasks draw on multiple speech, music, and environmental-audio sources with their own licenses, access controls, attribution requirements, and media rights. Review every selected task's metadata and upstream dataset terms before downloading, redistributing, or using it commercially.
- Download notes
- The public MTEB registry defines the 30-task MAEB beta suite across speech, music, environmental sound, and audio-text evaluation in more than 100 languages. It includes retrieval, classification, clustering, pair classification, reranking, multilabel classification, and zero-shot classification tasks. The helper downloads official documentation, repository metadata, the Apache-2.0 license, and the lightweight benchmark registry by default; cloning the approximately 55 MB MTEB source repository is opt-in. It does not fetch component datasets, which MTEB tasks acquire separately and which can be large or restricted.
Safe-first helperscripts/download/maeb.sh
Music
MAESTRO
MAESTRO: MIDI and Audio Edited for Synchronous TRacks and Organization
Safe-first helper
Automatic Music Transcription
Piano Transcription
Music Synthesis
Symbolic Music Modeling
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_applicable
- License caution
- Official Magenta page says the dataset is made available by Google LLC under Creative Commons Attribution Non-Commercial Share-Alike 4.0.
- Download notes
- The helper downloads v3.0.0 CSV/JSON metadata by default. The MIDI-only archive is about 56 MiB and the full WAV+MIDI archive is about 101 GiB, so both are explicit opt-ins.
Safe-first helperscripts/download/maestro.sh
Audio understanding, generation & events
MAESTRO Real
MAESTRO Real: Multi-Annotator Estimated Strong Labels
Safe-first helper
Sound Event Detection
Soft Label Sound Event Detection
Multi Annotator Label Aggregation
Long Form Sound Event Detection
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_noncommercial
- Code license
- not_applicable
- License caution
- The packaged Tampere University license permits copying and use only for experimental and non-commercial purposes with the full copyright notice and source acknowledgment. It explicitly prohibits commercial use, including selling or distributing results or content achieved through use of the dataset.
- Download notes
- The public development release contains 49 real-life recordings from five acoustic scenes with crowdsourced soft strong labels for 17 classes. Recordings are three to five minutes long and come from subsets of TUT Sound Events 2016 and 2017. DCASE 2024 Task 4 combines MAESTRO Real with DESED and evaluates 11 MAESTRO classes; its separate 26-file MAESTRO evaluation set is not included in the public development archive. The helper downloads official metadata, README, license, and the sub-megabyte annotation archive by default; the approximately 2.43 GiB audio archive requires explicit opt-in. The official Zenodo description reports 189 minutes 52 seconds total while its packaged README reports 97 minutes 4 seconds, so verify duration against the downloaded files rather than assuming either figure.
Safe-first helperscripts/download/maestro_real.sh
Speech recognition
MAGICDATA Mandarin Chinese Read Speech Corpus
Safe-first helper
Automatic Speech Recognition
Speaker Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR lists CC BY-NC-ND 4.0 and says the corpus is freely published for non-commercial or academic use. Re-check the current official page before redistribution or commercial use.
- Download notes
- The helper downloads the OpenSLR page and small metadata archive by default. Speech archives are large, including about 52 GiB train, 1.0 GiB dev, and 2.2 GiB test, so archive download is an explicit opt-in.
Safe-first helperscripts/download/magicdata_mandarin.sh
Music
MagnaTagATune
MagnaTagATune: A Music Annotation Benchmark from the TagATune Game
Safe-first helper
Music Auto Tagging
Music Annotation
Music Similarity
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-3.0
- Code license
- GPL-3.0
- License caution
- The original TagATune details page says the data is CC BY-NC-SA 3.0 except scripts released under GPL v3, enabling non-commercial research redistribution. Audio clips are Magnatune excerpts with artist/album URLs for purchase or commercial licensing.
- Download notes
- City University MIRG hosts metadata, annotations, comparisons, Echo Nest features, and three 1 GiB MP3 split archives. The helper downloads CSV metadata by default and makes features/audio explicit opt-ins.
Safe-first helperscripts/download/magnatagatune.sh
Speech recognition
MCIF
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
Safe-first helper
Multimodal Crosslingual Instruction Following
Audio Visual Instruction Following
Automatic Speech Recognition
Speech To Text Translation
+9 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card and paper declare MCIF data CC BY 4.0, and the official repository licenses evaluation and construction code under Apache-2.0. The source talks, papers, slides, voices, and presenter videos come from the ACL Anthology's CC BY 4.0 ACL 2023 release; preserve benchmark and source attribution and consider voice, likeness, and conference-media privacy obligations in downstream redistribution.
- Download notes
- MCIF is a professionally annotated test benchmark built from 100 ACL 2023 scientific talks: 21 core presentations totaling two hours support all 13 tasks, and 79 additional talks expand long-form summarization to about 10 hours overall. It provides English speech, video, and transcripts with prompts and references in English, German, Italian, and Mandarin Chinese. Fixed- and mixed-prompt configurations each expose 362 long-form and 1,560 short-form examples; the underlying media is shared rather than duplicated across configurations. The helper downloads the official cards, API metadata, four small Parquet manifests, and eight compressed reference files by default. Cloning the evaluation code is optional, and the public Hugging Face media snapshot is a separate opt-in because the repository is approximately 7.58 GiB.
Safe-first helperscripts/download/mcif.sh
Speaker, identity & emotion
MCR-Bench
MCR-Bench: Modal Conflict Resolution Benchmark for Large Audio-Language Models
Safe-first helper
Cross Modal Conflict Resolution
Audio Text Robustness
Audio Question Answering
Speech Emotion Recognition
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_upstream_terms
- Code license
- apache-2.0
- License caution
- The repository contains an Apache-2.0 LICENSE, although its README badge says MIT. Neither statement clearly relicenses the released benchmark archive or embedded source audio. ClothoAQA/Clotho, MELD's copyrighted Friends clips, and VocalSound retain their own terms, so verify item-level provenance and upstream permissions before redistribution or commercial use.
- Download notes
- The public Google Drive release contains approximately 3,000 English samples across audio question answering, speech emotion recognition, and vocal-sound classification. Each audio item is paired with faithful, adversarial, irrelevant, and neutral text conditions to measure whether an audio-language model follows contradictory text instead of audio evidence. The benchmark derives its task audio from ClothoAQA, MELD, and VocalSound. The helper downloads official documentation and repository metadata by default and prints the manual Drive path; cloning the documentation-only repository is opt-in.
Safe-first helperscripts/download/mcr_bench.sh
Enhancement, separation & quality
MECAT
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Safe-first helper
Fine Grained Audio Understanding
Audio Captioning
Audio Question Answering
Speech Understanding
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-3.0
- Code license
- Apache-2.0
- License caution
- Both Hugging Face cards declare CC BY 3.0 for the released audio and annotations, while the repository code is Apache 2.0 and the arXiv paper is CC BY 4.0. MECAT is derived from a selected subset of ACAV100M; preserve attribution and review the source dataset and underlying media terms before redistribution or commercial use.
- Download notes
- MECAT is one benchmark family with two public, ungated WebDataset releases rather than two independent families. MECAT-Caption has 20,052 training and 20,052 test clips, each annotated for systematic, content-specific, and acoustic-environment captioning. MECAT-QA uses the same eight speech/music/general-sound mixture domains and has 100,260 question-answer pairs in each split across perception, analysis, and reasoning tracks. The DATE metric and baseline evaluation code are in the repository. The helper downloads first-party cards, API metadata, paper, license, and repository documentation by default; the approximately 16.2 GB caption and 42.4 GB QA snapshots require separate explicit opt-ins.
Safe-first helperscripts/download/mecat.sh
Enhancement, separation & quality
MedleyDB
MedleyDB: A Multitrack Dataset for Annotation-Intensive MIR Research
Safe-first helper
Music Information Retrieval
Melody Extraction
Music Source Separation
Instrument Recognition
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- MIT
- License caution
- The official downloads page says MedleyDB is free for non-commercial research use only and identifies the dataset as Creative Commons Attribution-NonCommercial-ShareAlike 4.0. It also asks users not to republish the dataset in full or in part without consent, even though redistribution is technically allowed under the license. The GitHub tooling repository is MIT.
- Download notes
- The helper saves official pages and repository license/README by default. It checks the Zenodo request records only with MEDLEYDB_CHECK_ZENODO=1, downloads the public sample archive only with MEDLEYDB_DOWNLOAD_SAMPLE=1, and clones the annotation/metadata/tooling repo only with MEDLEYDB_CLONE_REPO=1. Full MedleyDB and MedleyDB 2.0 audio require requesting access through the official Zenodo records.
Safe-first helperscripts/download/medleydb.sh
Audiovisual & cross-modal
MELD
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
Safe-first helper
Speech Emotion Recognition
Multimodal Emotion Recognition
Dialogue Emotion Recognition
Sentiment Analysis
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- gpl-3.0
- Code license
- GPL-3.0
- License caution
- GitHub and the Hugging Face dataset card list GPL-3.0. MELD clips are derived from the Friends TV series, so downstream users should re-check media rights and any fair-use/research assumptions before redistribution or commercial use.
- Download notes
- The helper downloads official project, README, license, and dataset-card metadata by default. Raw audio/video tarballs and extracted feature/model tarballs are opt-in because they are larger and include TV-derived media clips.
Safe-first helperscripts/download/meld.sh
Music
MeloBottleneck Evaluation Suite
MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck
Safe-first helper
Melody Skeleton Extraction
Ornament To Backbone Reduction
Variation To Theme Alignment
Gongche Skeleton Alignment
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The supplementary repository has no LICENSE file and GitHub detects no license. The paper's arXiv license does not grant reuse rights for the released code, benchmark derivatives, MIDI files, or experiment outputs. TAVERN, Jiugong, and the seven training-corpus sources retain their own terms; obtain author and upstream clarification before redistribution or commercial use.
- Download notes
- The public supplementary repository releases the paper's three evaluation families inside a 4.76 MB codebase ZIP: synthetic out-of-distribution Main O2B splits, a 619-sequence TAVERN V2T set, and a 20-sequence Jiugong O2G set, together with preprocessing code, split metadata, result records, and symbolic MIDI/pickle artifacts. The paper reports 2,600 O2B sequences, 12,700 TAVERN notes, and 3,100 Jiugong notes. The helper downloads the official README and GitHub repository metadata by default; set MELOBOTTLENECK_DOWNLOAD_CODEBASE=1 to fetch the checksum-recorded code/data ZIP. The separate 55.7 MB RAR result bundle and 9.8 MB demo archive are not fetched.
Safe-first helperscripts/download/melobottleneck_eval.sh
Audiovisual & cross-modal
MER2023
MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
Manual or gated
Multimodal Emotion Recognition
Discrete Emotion Classification
Dimensional Emotion Recognition
Modality Robustness
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_additional_gated_terms
- Code license
- not_specified
- License caution
- The current Hugging Face card declares CC BY-NC 4.0 and limits access to academic research. Its gate also prohibits handing the database or derived labeling files to third parties and prohibits modification without written consent. The challenge paper describes a separate EULA with academic-only, no-editing, and no-upload conditions. The baseline repository has no LICENSE file. Clips were collected from movies and television, so underlying media, performer, privacy, and platform rights remain separate.
- Download notes
- The request-gated release extends CHEAVD and provides 3,373 labeled Train&Val clips, 411 MER-MULTI test clips, 412 corrupted-modality MER-NOISE test clips, and MER-SEMI with 834 labeled plus 73,148 unlabeled clips. MER-MULTI evaluates joint discrete-emotion and valence prediction, MER-NOISE tests robustness to noisy audio and blurred video, and MER-SEMI evaluates semi-supervised discrete emotion recognition. The helper saves public paper, repository, and Hugging Face API metadata only. The current Hugging Face repository is about 140 GB, password-protected, and requires approval, so benchmark files are left as a manual download.
Safe-first helperscripts/download/mer2023.sh
Audiovisual & cross-modal
MER2024
MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
Manual or gated
Multimodal Emotion Recognition
Discrete Emotion Classification
Semi Supervised Emotion Recognition
Modality Robustness
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_additional_gated_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 and non-commercial use. Its gate prohibits transfer of the database or derived labeling files to third parties and modification without written consent. The challenge page additionally limits the dataset to academic research and prohibits uploading samples. The MER2024 README displays an Apache-2.0 badge but links to a license under the later MER2025 directory; the repository root and MER2024 directory contain no applicable LICENSE file, so the MER2024 code license is recorded as unspecified. Source-video, performer, privacy, and platform rights remain separate.
- Download notes
- This request-gated extension of MER2023 provides 5,030 labeled Train&Val clips and 115,595 unlabeled clips. MER-SEMI evaluates 1,169 annotated clips from the unlabeled pool, MER-NOISE evaluates 1,170 clips with additive audio noise and image blur, and MER-OV evaluates free-form emotion labels. Table 1 of the paper reports 322 MER-OV samples, while the surrounding prose says 332; the index preserves that primary-source discrepancy. The current Hugging Face tree is approximately 218.4 GB and requires approval. The helper saves only public paper, project, repository, and Hugging Face API metadata.
Safe-first helperscripts/download/mer2024.sh
Music
Million Song Dataset
Million Song Dataset (MSD)
Safe-first helper
Music Recommendation
Collaborative Filtering
Music Auto Tagging
Music Similarity
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_noncommercial_mixed
- Code license
- gpl_version_unspecified
- License caution
- The packaged LICENSE applies archived Echo Nest API terms allowing non-commercial research/use with attribution and requiring separate permission for other purposes. It identifies MusicBrainz track years as public domain and MusicBrainz tags/counts as CC BY-NC-SA 2.0. Complementary lyrics, cover-song, tag, recommendation, and other datasets have their own terms. The repository describes its code only as GNU Public License without specifying a version; MSD distributes no listenable audio.
- Download notes
- The core public release contains Echo Nest-derived audio features and metadata for one million tracks, not listenable audio. The Taste Profile subset contains 48,373,586 user-song-play-count triplets for 1,019,318 users and 384,546 songs. The helper downloads official pages, terms, and the approximately 100 KiB tag vocabulary by default; the approximately 500 MB Taste Profile archive, 1.8 GB 10,000-song feature subset, and code clone require separate opt-ins. The July 2026 CCBR paper uses an unreleased processed selection of 31,046 items and 166,188 users plus separately obtained audio and cached descriptions.
Safe-first helperscripts/download/million_song_dataset.sh
Speech recognition
MInDS-14
MInDS-14: Multilingual and Cross-Lingual Intent Detection from Spoken Data
Safe-first helper
Spoken Language Understanding
Intent Classification
Automatic Speech Recognition
Multilingual Speech Understanding
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Hugging Face dataset card lists CC BY 4.0 and exposes 14 spoken e-banking intents across 14 language varieties. No separate code license was identified for the dataset card.
- Download notes
- The helper downloads the Hugging Face dataset card by default. Dataset snapshots include audio and are opt-in; choose one locale such as en-US or all with MINDS14_CONFIG.
Safe-first helperscripts/download/minds14.sh
Speech generation
Ming-Freeform-Audio-Edit
Ming-Freeform-Audio-Edit: Free-form instruction-based speech editing benchmark
Safe-first helper
Instruction Based Speech Editing
Semantic Speech Editing
Speech Content Deletion
Speech Content Insertion
+6 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_with_upstream_terms
- Code license
- not_specified
- License caution
- The Hugging Face dataset card declares Apache-2.0, but the released audio is derived from multiple upstream corpora whose terms still apply. Seed-TTS Eval states no data license, LibriTTS is CC BY 4.0, and GigaSpeech uses its own agreement/access conditions. The evaluation repository has no license file or detected GitHub license; review each source corpus and clarify annotation/code rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains Chinese and English source speech, natural-language editing instructions, and metadata for semantic deletion, insertion, and substitution plus acoustic emotion, dialect, speed, pitch, and volume changes. The paper's section 6.3 constructs the semantic set from 896 Chinese and 655 English Seed-TTS test samples and reports separate basic and full instruction versions; its acoustic sets also use Seed-TTS audio. The current dataset card additionally names LibriTTS and GigaSpeech as source corpora. The helper downloads official documentation and repository metadata by default; the Hugging Face API reports approximately 1.07 GB of repository storage, so audio and annotations require explicit opt-in.
Safe-first helperscripts/download/ming_freeform_audio_edit.sh
Speech recognition
MIR-1K vocal
MIR-1K
Safe-first helper
Singing Voice Transcription
Singing Voice Separation
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_on_official_page
- Code license
- not_applicable
- License caution
- Official MIR Lab page has direct downloads but no visible license statement; its MIR-1K.rar URL returned 404 on 2026-07-09. Figshare mirror lists CC BY 4.0.
Safe-first helperscripts/download/mir_1k_vocal.sh
Speaker, identity & emotion
MixFake
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
Safe-first helper
Audio Deepfake Detection
Synthetic Speech Detection
Synthetic Music Detection
Synthetic Environmental Sound Detection
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_source_corpus_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY 4.0 for the release, but contains no explanatory license text. The GitHub repository has no LICENSE file or detected license, so its code terms are unspecified. MixFake derives material from ASVspoof 2019 LA, FMA-Medium, Sonics, FakeMusicCaps, and EnvSDD; review all source-corpus, generated-media, voice, and redistribution terms in addition to the dataset-card declaration.
- Download notes
- MixFake is one benchmark family with isolated-source and mixed-source evaluation conditions, not separate benchmark families for its foreground and background tasks. The paper reports 252,500 clips totaling 673.69 hours: 510.59 hours of isolated speech, music, and environmental sound plus 163.10 hours of mixtures. Train, development, and evaluation partitions cover real/fake foreground-background combinations at -5, 0, 5, 10, 15, and 20 dB SNR. The helper downloads the official paper, repository documentation, and Hugging Face card/API metadata by default. The public snapshot is split across 67 7-Zip volumes totaling approximately 66.7 GiB and therefore requires MIXFAKE_DOWNLOAD_DATA=1; cloning the released baseline and score files is a separate opt-in.
Safe-first helperscripts/download/mixfake.sh
Speech recognition
MLC-SLM Eval
Multilingual Conversational Speech Language Model Challenge Eval Ground Truth
Safe-first helper
Multilingual Conversational Asr
Speaker Diarization
Speaker Attributed Asr
Long Form Speech Recognition
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0_annotation_card_label
- Code license
- not_specified
- License caution
- The Hugging Face card labels the released ground-truth repository CC BY-SA 4.0. It does not publish a separate license file or establish terms for the absent evaluation recordings. The official baseline repository has no detected license, so do not assume the annotation label licenses its code. Obtain the audio and its terms from the challenge owners before attempting full benchmark reproduction.
- Download notes
- The public, ungated Hugging Face release contains approximately 6.55 MB of oracle segmentation, speaker labels, and transcriptions for the challenge's 32-hour Eval-1 and 32-hour Eval-2 sets. It covers English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese; English additionally spans five accent groups. The repository currently contains 224 annotation files and no audio. Challenge participants previously received evaluation recordings, but the official paper, dataset card, and baseline repository provide no current public audio URL. The helper downloads official documentation and repository metadata by default; the lightweight annotation snapshot requires MLC_SLM_EVAL_DOWNLOAD_HF=1 and does not include audio. VibeVoice-ASR evaluates the MLC-Challenge set, while VibeVoice-ASR-BitNet section 3.1 and Table 4 report six MLC language subsets.
Safe-first helperscripts/download/mlc_slm_eval.sh
Speech recognition
MLS
MLS: A Large-Scale Multilingual Dataset for Speech Research
Safe-first helper
Multilingual Asr
Asr
Language Modeling
Limited Supervision Asr
Access pathOpenSLR
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR94 lists CC BY 4.0. MLS is derived from LibriVox audiobooks and provides public-file-hosted archives plus MD5 checksums.
- Download notes
- OpenSLR links original FLAC and compressed OPUS archives for English, German, Dutch, French, Spanish, Italian, Portuguese, and Polish; archives range from about 1.6 GiB to multiple TiB, so the helper downloads only the page and checksums by default and requires MLS_DOWNLOAD_ARCHIVES=1 for audio.
Safe-first helperscripts/download/mls.sh
Enhancement, separation & quality
MMAE
MMAE: A Massive Multitask Audio Editing Benchmark
Safe-first helper
Instruction Based Audio Editing
Multi Round Audio Editing
Multi Hop Audio Editing
Speech Editing
+4 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The Hugging Face card does not declare a dataset license or identify licenses for the source audio. The official GitHub repository has an MIT LICENSE for its code, but that must not be assumed to license the benchmark audio. Review source-media rights and obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated release contains 2,000 high-fidelity input samples across sound, speech, music, and mixtures, organized by six complexity levels, two granularities, and eight operation types. Its 17,741 rubric criteria evaluate instruction following and context consistency. The helper downloads official documentation and repository metadata by default; cloning the evaluation repository is opt-in, and the Hugging Face API reports approximately 4.43 GB of repository storage, so the audio snapshot requires MMAE_DOWNLOAD_HF=1.
Safe-first helperscripts/download/mmae.sh
Speech generation
MMAG
MMAG: A Multi-Control Mixed Audio Generation Benchmark
Safe-first helper
Text To Mixed Audio Generation
Speech Music Sound Effect Composition
Voice Prompt Conditioned Generation
Speaker Identity Control
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_noncommercial_card_wording_conflict
- Code license
- no_detected_license
- License caution
- The Hugging Face metadata and card label the dataset CC BY 4.0, but the card's license paragraph also limits it to non-commercial research use; treat that inconsistency conservatively and seek clarification for commercial use. The repository README repeats CC BY 4.0 but has no detected license file. Source clips are selected from AudioCaps, VGGSound, and MECAT test sets, whose dataset, platform-media, and recording terms continue to apply; the aggregate label should not be assumed to override those upstream rights.
- Download notes
- The owner release contains three test configurations: 3,974 main-set clips with detailed overall captions; 691 voice-cloning cases with 2-5-second speaker prompts; and 1,828 timestamp-control cases with precise speech and foreground-event boundaries. The timestamp set is selected from the main set, and the voice-cloning set is a dedicated capability subset, so these are tracks of one benchmark family rather than separate family counts. The helper downloads owner documentation, live metadata, and all three manifests (about 3.2 MB total) by default. The approximately 1.06 GB audio and prompt snapshot is opt-in.
Safe-first helperscripts/download/mmag.sh
Audio understanding, generation & events
MMAR
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Safe-first helper
Audio Question Answering
Audio Reasoning
Chain Of Thought Reasoning Evaluation
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for its original metadata and audio snapshot. The GitHub repository has no license file or detected GitHub license, so do not assume that declaration separately licenses the later MMAR-Rubrics annotations or evaluation code. Original in-the-wild video/audio rights and platform terms also apply.
Safe-first helperscripts/download/mmar.sh
Audio understanding, generation & events
MMAU
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Safe-first helper
Audio Question Answering
Audio Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- HF cards: MMAU-test-mini is cc-by-nc-4.0; MMAU-test is mit
- Code license
- Apache-2.0
- License caution
- The two HF dataset cards list different licenses; re-check before redistribution.
Safe-first helperscripts/download/mmau.sh
Audio understanding, generation & events
MMAU-Pro
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Safe-first helper
Audio Question Answering
Audio Reasoning
Long Audio Understanding
Multi Audio Reasoning
+5 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0, but the paper says almost all audio was sourced from in-the-wild recordings and the spatial subset reuses EasyCom. Review source-media and EasyCom terms before redistribution or commercial use. The official GitHub repository has no license file or detected license, so evaluator code terms are unspecified.
- Download notes
- The public, ungated release contains 5,305 expert-authored multiple-choice and open-ended QA instances spanning 49 skills across speech, environmental sound, music, and their mixtures. It includes multiple-audio, spatial, instruction-following, and up-to-10-minute long-form cases. The helper downloads official documentation, repository metadata, the evaluator, and Hugging Face metadata by default; the Hugging Face API reports approximately 47.5 GB of repository storage, so the audio and test snapshot requires MMAU_PRO_DOWNLOAD_HF=1.
Safe-first helperscripts/download/mmau_pro.sh
Speech generation
MMGenre
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Safe-first helper
Singing Voice Synthesis
Genre Conditioned Singing Voice Synthesis
Singing Voice Genre Alignment
Score Conditioned Singing Voice Synthesis
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_generated_source_caveat
- Code license
- cc-by-4.0
- License caution
- The dataset card, dataset LICENSE, and repository LICENSE declare CC BY 4.0. The benchmark was derived from music generated with Suno V4.5 and then source-separated and automatically aligned; the release license does not independently resolve any terms or rights attached to the generation service or generated source content, so review those conditions for downstream use.
- Download notes
- The public, ungated Hugging Face release contains 3,152 aligned Chinese singing-voice and symbolic-score segments from 148 generated songs, totaling about 4.36 hours across 10 major genres and 26 subgenres. The paper body says 27 subgenres, but the dataset card identifies this as a counting error and treats the released 26-subgenre taxonomy as authoritative. The helper downloads official documentation and API metadata by default, can clone the lightweight score/code repository, and requires explicit opt-in for the approximately 5.54 GB Hugging Face snapshot.
Safe-first helperscripts/download/mmgenre.sh
Audiovisual & cross-modal
MMOU
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
Safe-first helper
Audio Visual Question Answering
Long Video Understanding
Cross Modal Reasoning
Temporal Grounding
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0 metadata label; non-commercial research restriction in paper
- Code license
- not_applicable
- License caution
- Both Hugging Face cards declare Apache-2.0, but the MMOU paper says the dataset is released solely for non-commercial research. Apply the stricter non-commercial restriction pending clarification. Videos were collected from public web platforms including YouTube, so source copyright, platform terms, availability, and any per-video rights also apply.
- Download notes
- The public, ungated NVIDIA release contains 20,000 English questions over 11,877 long-form web videos. The 5,000-item Test Mini split includes labels for local evaluation; answers for the main 15,000-item split are withheld for evaluator submission. The helper downloads official cards and API metadata by default, makes the approximately 48 MB question files opt-in, and requires a separate opt-in for the approximately 322.8 GB community-hosted video snapshot. Audio-Visual Flamingo evaluates MMOU as an omni-modal benchmark.
Safe-first helperscripts/download/mmou.sh
Speech understanding & dialogue
MMSU
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Safe-first helper
Spoken Language Understanding
Speech Reasoning
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- not_specified
- License caution
- HF dataset card lists MIT; GitHub code repo did not expose a detected license.
Safe-first helperscripts/download/mmsu.sh
Enhancement, separation & quality
MoisesDB
MoisesDB: A Dataset for Source Separation Beyond 4-Stems
Safe-first helper
Music Source Separation
Multi Stem Source Separation
Instrument Source Separation
Singing Voice Separation
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- cc-by-nc-sa-4.0
- License caution
- The official repository applies CC BY-NC-SA 4.0 to MoisesDB and its packaged loader/evaluation materials, and the Music AI page limits the dataset to non-commercial research use. Attribution and ShareAlike obligations apply; commercial use is not permitted by this release.
- Download notes
- The official Music AI research page provides the dataset through its browser download flow. The release contains 240 songs by 47 artists across 12 high-level genres, totaling 14 hours, 24 minutes, and 46 seconds, with mixtures and a hierarchical stem/source taxonomy extending beyond the common four-stem setup. The helper saves the official page plus repository README and license by default; it does not fetch the large audio archive, and cloning the loader/evaluation repository is opt-in.
Safe-first helperscripts/download/moisesdb.sh
Audiovisual & cross-modal
Movie Gen Audio Bench
Movie Gen Audio Bench: Video-Conditioned Sound Effect and Music Generation
Safe-first helper
Video To Audio Generation
Text And Video To Audio Generation
Video Conditioned Sound Effect Generation
Joint Sound Effect And Music Generation
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- not_applicable
- License caution
- The repository applies Creative Commons Attribution-NonCommercial 4.0 to Movie Gen Bench. Commercial use is prohibited and attribution is required. The release contains generated videos and generated audio; users should still review prompt, depicted-subject, music, and other third-party rights that the repository license may not control. No separate scoring-code release is identified.
- Download notes
- Meta's public repository releases two JSONL prompt manifests for 527 Movie Gen Video-generated test videos: one for sound effects and one for joint sound effects plus background music. The benchmark supports video-to-audio and text-plus-video-to-audio evaluation and includes human-reviewed sound and music captions. The helper downloads the repository documentation, license, and both approximately 338 KB total prompt manifests by default. The media archives include Meta's non-cherry-picked generated audio and are separate explicit opt-ins: the SFX archive is approximately 8.05 GiB and the SFX+music archive is approximately 8.08 GiB. Flowley section 4.3.5 evaluates the public benchmark as a zero-shot distribution-shift test.
Safe-first helperscripts/download/movie_gen_audio_bench.sh
Speech generation
MS-SNSD
Microsoft Scalable Noisy Speech Dataset
Safe-first helper
Speech Enhancement
Speech Denoising
Noisy Speech Synthesis
Subjective Speech Quality Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The README says Microsoft provides the datasets as-is under the original terms received. It lists PTDB-TUG under ODbL 1.0, Edinburgh/VoiceBank material under its DataShare license, selected Freesound noise as CC0, and DEMAND as CC BY-SA 3.0. Re-check component terms before redistribution or commercial use.
- Download notes
- The official GitHub repository contains clean speech, noise, noisy test, and clean test directories plus scripts for generating noisy speech at configurable SNRs. The helper saves README/license/generator files by default; cloning the large repository is an explicit opt-in.
Safe-first helperscripts/download/ms_snsd.sh
Audiovisual & cross-modal
MSMD
Multimodal Sheet Music Dataset
Safe-first helper
Audio Sheet Music Retrieval
Sheet Music Audio Retrieval
Piece Identification
Score Following
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_official_CC-BY-4.0_and_CC-BY-NC-SA-4.0
- Code license
- CC-BY-NC-SA-4.0_repository_wide
- License caution
- Both Zenodo records declare CC BY 4.0, while the official MSMD repository contains a repository-wide CC BY-NC-SA 4.0 LICENSE. This conflict is not resolved by the release pages; conservatively follow the noncommercial share-alike terms and verify intended rights with the maintainers before commercial use or redistribution. Preserve attribution and inspect the licenses of the underlying Mutopia scores. CODA's MIT license covers its evaluation code, not MSMD data.
- Download notes
- MSMD is one multimodal benchmark family rather than separate retrieval and score-following datasets. The original release describes 497 classical pieces and 344,742 notehead-to-audio/MIDI alignments, with score images, notation graphs, MIDI-derived performance features, and composer-aware splits; a 2019 correction removed 12 pieces with unreliable da-capo alignments and updated the split files. The approximately 9.56 GB original Zenodo archive contains preprocessed data without rendered audio. A later approximately 1.92 GB score-following package contains 354 training, 19 validation, and 94 test pieces as NPZ/WAV pairs. The helper downloads official repository, paper, release, and recent-evaluation metadata by default; either large archive requires an explicit part selection.
Safe-first helperscripts/download/msmd.sh
Speaker, identity & emotion
MSP-Podcast
MSP-Podcast: A Large Naturalistic Speech Emotional Dataset
Manual or gated
Speech Emotion Recognition
Dimensional Emotion Recognition
Categorical Emotion Recognition
Speaker Independent Emotion Recognition
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_academic_license
- Code license
- not_applicable
- License caution
- The owner page currently calls the release an Academic License and requires an institution-signed FDP data-transfer agreement. Although the page says source podcasts were chosen under permissive licenses, the signed corpus agreement controls access and reuse; review it directly before commercial use, redistribution, or sharing copies.
- Download notes
- Version 2.0 contains 264,705 naturalistic podcast speaking turns totaling 409 hours, with speaker-independent train, development, and three test partitions. It provides categorical emotion and activation, dominance, and valence labels; Test3 releases audio but withholds labels, speaker information, transcripts, and alignments for web-based evaluation. Access is free but institution/form-gated: an authorized institutional representative must sign the official academic agreement and send it to the corpus owner. There is no public archive URL.
Safe-first helperscripts/download/msp_podcast.sh
Enhancement, separation & quality
MSRBench
MSRBench: A Benchmarking Dataset for Music Source Restoration
Safe-first helper
Music Source Restoration
Music Source Separation
Singing Voice Restoration
Robust Audio Restoration
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- MIT
- License caution
- The paper and Hugging Face card declare CC BY-NC 4.0 for MSRBench, so commercial use is not authorized by this release. The MSRKit repository's MIT license covers the evaluation and baseline code, not the benchmark audio. Preserve attribution and review any third-party degradation-tool terms separately.
- Download notes
- MSRBench is one public benchmark family and the validation set for the MSR Challenge 2025. It contains 250 professionally mixed clips and unprocessed targets for each of eight instrument classes. Every item has the original mixture plus 12 analog, environmental, conventional codec, and neural-codec degradation conditions. The Hub repository is approximately 28.4 GB and the card reports 28.7 GB after extraction. The helper downloads only official pages, cards, API metadata, paper metadata, and repository documentation by default. Set MSRBENCH_STEM to exactly one available class to fetch its ZIP; cloning the lightweight MIT toolkit is a separate opt-in.
Safe-first helperscripts/download/msrbench.sh
Speech understanding & dialogue
MSU-Bench
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Safe-first helper
Multi Speaker Conversation Understanding
Speaker Identification
Speaker Attribute Recognition
Speaker Verification
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_upstream_terms
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY-NC 4.0 and limits the release to non-commercial academic research because audio derives from third-party film/TV, telephone, meeting, and podcast sources. The repository LICENSE applies MIT only to code and explicitly gives the benchmark data separate non-commercial academic-research terms; review each source corpus and media right before redistribution. The paper reports approximately 731 hours of source corpora, not 731 released hours.
- Download notes
- The public, ungated release contains 2,847 English and Chinese four-choice QA items over 241 multi-speaker audio clips, including 2,223 human-verified items across 16 tasks. The helper downloads official documentation, repository metadata, and the approximately 5.8 MB test JSONL by default; the Hugging Face API reports approximately 2.5 GB of repository storage, so the audio and annotations snapshot requires MSU_BENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/msu_bench.sh
Speech understanding & dialogue
MSWC
Multilingual Spoken Words Corpus
Safe-first helper
Keyword Spotting
Spoken Term Search
Multilingual Speech Classification
Forced Alignment
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- MLCommons and the HF card list CC BY 4.0. MSWC is derived from crowd-sourced sentence-level audio, including Common Voice, so preserve source attribution and re-check the active source terms for downstream redistribution.
- Download notes
- The official MLCommons page and HF card describe 50 languages, more than 340,000 keywords, 23.4 million 1-second examples, and over 6,000 hours. The helper saves official docs and the HF card by default; direct per-language audio, splits, and alignments are opt-in with MSWC_DOWNLOAD_ARCHIVES=1 because high-resource language audio archives can be many GiB.
Safe-first helperscripts/download/mswc.sh
Speech recognition
mTEDx
The Multilingual TEDx Corpus for Speech Recognition and Translation
Safe-first helper
Multilingual Asr
Speech To Text Translation
Speech Translation
Sentence Level Alignment
Access pathOpenSLR
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- OpenSLR SLR100 lists CC BY-NC-ND 4.0. The corpus is derived from TEDx Talks; respect TED/TEDx source terms as well as the corpus license.
- Download notes
- OpenSLR SLR100 hosts ASR-only language archives, speech-translation language-pair archives, IWSLT 2021 test sets, and a small French talk gender annotation CSV. Archives are multi-GiB, so the helper downloads only the OpenSLR page and small CSV by default and requires MTEDX_DOWNLOAD_ARCHIVES=1 for archive downloads.
Safe-first helperscripts/download/mtedx.sh
Music
MTG-Jamendo
MTG-Jamendo Dataset for Automatic Music Tagging
Safe-first helper
Music Auto Tagging
Music Genre Classification
Musical Instrument Recognition
Music Mood Theme Recognition
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- Apache-2.0
- License caution
- The GitHub README says repository metadata is CC BY-NC-SA 4.0, code is Apache-2.0, audio files keep individual Creative Commons licenses listed in audio_licenses.txt, and the dataset is made available solely for non-commercial research and academic use unless Jamendo grants separate authorization.
- Download notes
- The helper clones or updates the official metadata/scripts repository by default and saves Zenodo record metadata. The upstream downloader can fetch very large archives: raw_30s audio is about 508 GiB, raw_30s audio-low is about 156 GiB, and autotagging_moodtheme audio-low is about 46 GiB, so media downloads require MTG_JAMENDO_DOWNLOAD_MEDIA=1.
Safe-first helperscripts/download/mtg_jamendo.sh
Representation & general suites
MuChin
MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
Safe-first helper
Multimodal Music Understanding
Colloquial Music Description
Professional Music Description
Structured Lyric Generation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_metadata_with_restricted_commercial_audio_use_and_upstream_song_rights
- Code license
- MIT_for_v1_repository
- License caution
- The V1 repository and Hugging Face card declare MIT, but the owner separately says scholars should obtain song audio legally, use it for academic purposes only, and not use it for commercial model training without copyright-holder authorization. Treat those express audio restrictions and underlying song, recording, lyric, artist, and album rights as controlling caveats; MIT must not be read as relicensing commercial music. V2 repeats the same audio restrictions and is gated.
- Download notes
- MuChin V1 is a public, ungated 1,000-song benchmark with paired amateur and professional Chinese descriptions, musical sections, rhyme structures, lyrics, and metadata. The IJCAI paper evaluates five pretrained music encoders through fixed shallow tag predictors and uses a weighted six-part Gestalt-based protocol for structured-lyric generation by four LLM families. The helper downloads official cards, repository metadata, and license text by default; the approximately 20.9 MB V1 annotation archive and 3.59 GB V1 MP3 archive are separate opt-ins. The expanded 6,066-song V2 is auto-gated on Hugging Face, about 31.3 GB, overlaps V1 on 724 songs, and requires explicit opt-in plus accepted access. Qwen-Audio-3.0-Gen-Preview Table 9 reuses MuChin for VAE reconstruction but does not identify the version or item manifest, so that derived protocol is not represented as a new split.
Safe-first helperscripts/download/muchin.sh
Speaker, identity & emotion
MUGEN
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
Safe-first helper
Multi Audio Understanding
Audio Grounding
Speech Understanding
Speaker And Demographic Understanding
+5 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_with_mixed_upstream_terms
- Code license
- MIT
- License caution
- The 35 Hugging Face dataset cards do not declare a data license. The repository's MIT license identifies its subject as software and associated documentation, so it should not be assumed to relicense embedded audio. The paper describes public corpora, specialized academic corpora, Mozilla Data Collective speech, and synthesized speech as sources; review each task's upstream corpus and generation terms before redistribution or commercial use.
- Download notes
- The public, ungated release provides 35 separate Hugging Face task repositories with 1,750 five-way audio-grounding problems and 9,250 audio clips across seven dimensions. Ten tasks add a reference clip, producing six concurrent audio inputs. The helper downloads official documentation, repository metadata, the license, and the Hugging Face collection inventory by default; set MUGEN_DOWNLOAD_TASK to one of the documented task names to explicitly download that task. The current cards report about 10.7 GB of compressed files across all task repositories and also expose reduced-candidate splits for input-scaling analysis.
Safe-first helperscripts/download/mugen.sh
Speech understanding & dialogue
MultiChallenge Audio
Speech adaptation of MultiChallenge for multi-turn spoken dialogue evaluation
Safe-first helper
Multi Turn Spoken Dialogue
Speech To Text Dialogue
Speech To Speech Dialogue
Instruction Retention
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit_claim_with_upstream_and_synthetic_voice_caveats
- Code license
- Apache-2.0
- License caution
- The MiMo evaluation-set card declares MIT but also says Xiaomi does not own the included datasets and directs users to upstream terms. Scale AI's 266-example source text release is CC BY 4.0. The report says the filtered conversations were rendered with a commercial TTS model and 250 voices, but it does not identify the provider or state separate rights for those generated voices; review those conditions before redistribution or commercial use.
Safe-first helperscripts/download/multichallenge_audio.sh
Speech generation
MultiRef-Compass
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Safe-first helper
Multi Reference Audio Video Generation
Audio Visual Consistency Evaluation
Voice Timbre Similarity
Speech Lip Synchronization
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY 4.0 and the evaluation repository is MIT. The paper describes assets from Pexels and Freesound plus generated images and voices, so preserve per-asset attribution and review source-platform and generation-service terms; the dataset-level tag should not be assumed to erase those underlying rights.
- Download notes
- The public release contains 350 English structured prompts across four boards, image references, and input reference videos for board 4. The helper downloads the dataset card, API metadata, CSV/JSONL manifests, and repository license by default. Cloning the evaluation toolkit and downloading the approximately 1.96 GB Hugging Face snapshot are separate opt-ins. Although the paper evaluates explicit audio-reference conditions and the toolkit accepts reference_audio_path, the current dataset schema exposes only image and video reference columns and includes no audio files; do not claim those paper scenarios are fully reproducible from this snapshot.
Safe-first helperscripts/download/multiref_compass.sh
Music
MulTTiPop
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
Safe-first helper
Automatic Music Transcription
Multitrack Music Transcription
Audio Midi Alignment
Music Information Retrieval
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_for_released_midi_and_metadata
- Code license
- not_applicable
- License caution
- The dataset card declares CC BY 4.0 and says its aligned MIDI is adapted from the CC BY 4.0 Lakh MIDI Dataset. Source audio is not licensed or redistributed; users are instructed to obtain only the referenced segments and use MulTTiPop for evaluation rather than training. YouTube availability, platform terms, and commercial-song rights remain applicable.
- Download notes
- The public, ungated release contains aligned multitrack MIDI and metadata for 572 commercial-pop segments (about 3.5 hours), divided into artist-disjoint development and test splits. It does not redistribute source audio; metadata provides YouTube video identifiers and segment timestamps. The helper downloads the dataset card, API metadata, and lightweight split manifests by default, while the full MIDI/metadata snapshot requires MULTTIPOP_DOWNLOAD_HF=1.
Safe-first helperscripts/download/multtipop.sh
Speaker, identity & emotion
MUSAN
MUSAN: A Music, Speech, and Noise Corpus
Safe-first helper
Voice Activity Detection
Music Speech Discrimination
Speech Music Noise Classification
Speaker Recognition Augmentation
Access pathOpenSLR
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR17 lists Attribution 4.0 International (CC BY 4.0). The paper describes MUSAN as music, speech, and noise recordings released under a flexible Creative Commons license.
- Download notes
- OpenSLR lists the corpus archive as 11 GiB. The helper downloads the OpenSLR landing page by default and requires MUSAN_DOWNLOAD_ARCHIVE=1 for the large archive.
Safe-first helperscripts/download/musan.sh
Enhancement, separation & quality
MUSDB18
MUSDB18: A Corpus for Music Separation
Safe-first helper
Music Source Separation
Singing Voice Separation
Audio Source Separation
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other-non-commercial
- Code license
- MIT
- License caution
- Zenodo lists "Other (Non-Commercial)" and the official pages state educational/academic use only, with commercial use requiring copyright-holder permission. Track sources include DSD100/Mixing Secrets, MedleyDB CC BY-NC-SA 4.0, Native Instruments stems, and Easton Ellises/Heise CC BY-NC-SA 3.0 material.
- Download notes
- The helper saves the official SigSep and Zenodo pages by default. The 4.7 GiB compressed STEMS archive and 22.7 GiB uncompressed HQ archive require explicit terms acknowledgement and opt-in.
Safe-first helperscripts/download/musdb18.sh
Audiovisual & cross-modal
MUSIC-AVQA
MUSIC-AVQA: Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Safe-first helper
Audio Visual Question Answering
Multimodal Scene Understanding
Spatiotemporal Reasoning
Music Understanding
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- conflicting: GitHub API/LICENSE file reports MIT, while the README License section mentions GPLv3
- License caution
- No standalone dataset license was found on the official page or README on 2026-07-10. Raw musical-performance videos and extracted features should be treated as upstream media with terms to verify before redistribution or commercial use.
- Download notes
- The helper downloads the official project page, README/LICENSE, and public JSON QA annotations by default. Raw videos and large feature files are hosted through Google Drive and Baidu Drive links on the project page/README; they are not downloaded automatically.
Safe-first helperscripts/download/music_avqa.sh
Audiovisual & cross-modal
MUSIC-AVQA-R
MUSIC-AVQA-R: Robust Audio-Visual Question Answering
Safe-first helper
Audio Visual Question Answering
Robustness Evaluation
Long Tail Reasoning
Question Paraphrase Robustness
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- No LICENSE file or standalone dataset terms were found in the official repository on 2026-07-25. The derived questions and evaluation code should not be assumed open-licensed; original MUSIC-AVQA media rights and its unspecified dataset license remain applicable.
- Download notes
- The helper downloads official documentation and the GitHub tree manifest by default. Set MUSIC_AVQA_R_CLONE_REPO=1 to clone the approximately 90 MB repository containing the three public JSON test files and evaluation code. Training/validation annotations and source videos remain in the separately indexed MUSIC-AVQA release.
Safe-first helperscripts/download/music_avqa_r.sh
Audiovisual & cross-modal
MUSIC-AVQA-v2.0
MUSIC-AVQA-v2.0: A Balanced Dataset for Unbiased Audio-Visual Question Answering
Safe-first helper
Audio Visual Question Answering
Dataset Bias Evaluation
Balanced Evaluation
Music Understanding
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- research_purpose_only_statement_for_youtube_videos
- Code license
- Not specified in the source record.
- License caution
- The repository-wide GPL-3.0 file covers the packaged repository, while its README separately says the manually collected YouTube videos are for research purposes only. Upstream YouTube and original MUSIC-AVQA media rights remain applicable; do not infer that GPL-3.0 re-licenses third-party video content.
- Download notes
- The helper downloads the official README, GPL-3.0 license, GitHub tree manifest, and lightweight additional-video manifest by default. Set MUSIC_AVQA_V2_CLONE_REPO=1 to clone the repository with the biased and balanced QA JSON splits. The much larger 1,040-video archive remains a manual Dropbox download.
Safe-first helperscripts/download/music_avqa_v2.sh
Audio understanding, generation & events
MusICA-MetaBench
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music
Safe-first helper
Music Perception Question Answering
Audio Music Understanding
Symbolic Music Understanding
Sheet Music Understanding
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_CC-BY-4.0_and_CC-BY-SA-4.0_by_source
- Code license
- GPL-3.0-or-later
- License caution
- ChoraleBricks-derived benchmark instances are CC BY 4.0, while ChoralSynth-derived instances are CC BY-SA 4.0. Original code, question templates, ontology, and configurations are GPL-3.0-or-later. Model outputs and inference logs are provided without a separate license. Source recordings and scores are downloaded separately and retain the upstream dataset terms; select the license by each benchmark file's documented source rather than treating the repository as uniformly licensed.
- Download notes
- The public repository releases the evaluation and benchmark-generation pipeline, question templates, ontology, configurations, inference logs, and pre-generated benchmark instances. The paper's main ChoraleBricks instance and its ChoralSynth validation instance each contain 300 five-option questions balanced across audio, symbolic, and sheet-image modalities; 20% use "none of the other options" as the correct answer. Items test pitch, rhythm, and harmony perception. The helper downloads official documentation, licenses, configs, and both sub-500 KB TSV instances by default. Cloning the approximately 49 MB GitHub repository is opt-in. Source audio and scores are not stored in that repository and must be obtained from ChoraleBricks or ChoralSynth under their respective terms.
Safe-first helperscripts/download/musica_metabench.sh
Audio understanding, generation & events
MusicCaps
MusicCaps: A Dataset of Music Captions
Safe-first helper
Music Captioning
Text To Music Evaluation
Music Understanding
Audio Language Modeling
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- The Hugging Face card and Kaggle-converted metadata list CC BY-SA 4.0 for the annotation CSV. The referenced media are 10-second clips from AudioSet/YouTube, so original media copyright, platform terms, and availability still apply.
- Download notes
- The public release is a CSV of 5,521 music-text pairs with YouTube IDs, segment timestamps, AudioSet labels, aspect lists, and musician-written captions. Raw audio is not mirrored by the dataset and must be reconstructed from YouTube/AudioSet subject to upstream availability and terms.
Safe-first helperscripts/download/musiccaps.sh
Music
MusicNet
MusicNet: A Dataset for Music Transcription and Multi-label Classification
Safe-first helper
Music Transcription
Multi Label Music Classification
Note Onset Labeling
Musical Instrument Recognition
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo lists CC BY 4.0 for the MusicNet release. The record says audio recordings are Creative Commons licensed and Public Domain performances from the Isabella Stewart Gardner Museum, the European Archive Foundation, and Musopen, with per-recording provenance described in the metadata.
- Download notes
- The helper saves the Zenodo record JSON and 44 KiB metadata CSV by default. Reference MIDI files are a small opt-in download, while the full audio/label archive is about 10.3 GiB and requires MUSICNET_DOWNLOAD_AUDIO=1.
Safe-first helperscripts/download/musicnet.sh
Audiovisual & cross-modal
MUStARD
MUStARD: Multimodal Sarcasm Detection Dataset
Safe-first helper
Multimodal Sarcasm Detection
Spoken Sarcasm Detection
Audio Visual Pragmatics
Dialogue Context Understanding
+1 more
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The official GitHub repository includes an MIT license and the Hugging Face card declares MIT. MUStARD audiovisual clips are excerpted from Friends, The Big Bang Theory, The Golden Girls, and Sarcasmaholics Anonymous; neither machine-readable label establishes that those third-party recordings are relicensed. Review copyright, fair-use, and redistribution constraints before downloading or using the media, especially commercially.
- Download notes
- The official release contains 690 audiovisual utterances from four English-language comedy sources, binary sarcasm labels, dialogue context and speakers, and fixed five-fold indices. The helper downloads the official documentation, repository metadata, 375 KB annotation JSON, and 28 KB fold file by default. The public Hugging Face repository is approximately 4.63 GB and includes raw audiovisual clips plus pre-extracted BERT and ResNet features, so it requires explicit opt-in. CHARM section 4.1 evaluates all 690 items for text calibration and uses stratified five-fold cross-validation for audio-text late fusion.
Safe-first helperscripts/download/mustard.sh
Enhancement, separation & quality
NISQA
NISQA Speech Quality Corpus
Safe-first helper
Speech Quality Assessment
Mean Opinion Score Prediction
Non Intrusive Speech Quality
Multidimensional Speech Quality
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The GitHub README says the corpus is provided under the original terms of the source speech and noise samples, generally non-commercial research with some subsets permitting commercial use; individual README/license files inside the archive should be checked. The Zenodo record reports license id other-at, and model weights are CC BY-NC-SA 4.0.
- Download notes
- The helper downloads the official README, corpus wiki markdown, model-weight license, and Zenodo record JSON by default. The full NISQA_Corpus.zip archive is about 15.9 GB and requires NISQA_DOWNLOAD_CORPUS=1.
Safe-first helperscripts/download/nisqa.sh
Audio understanding, generation & events
Nonspeech7k
Nonspeech7k dataset: Classification and analysis of human non-speech sound
Safe-first helper
Human Vocal Sound Classification
Paralinguistic Classification
Domestic Safety Monitoring
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0 per explicit record description
- Code license
- not_applicable
- License caution
- Zenodo's structured license field says CC BY 4.0, but the same official record description says the dataset is allowed only for non-commercial academic research under CC BY-NC-SA 4.0. Apply the stricter explicit statement pending author or repository clarification. The record asks users to acknowledge Freesound, YouTube, and Aigei as sources; preserve provenance and review source-recording rights before redistribution.
- Download notes
- The public author-owned Zenodo release contains 7,014 strongly single-label-annotated mono WAV clips at 32 kHz, lasting 0.5-4 seconds. Its fixed split has 6,289 training and 725 test clips. The helper downloads Zenodo metadata, the two small annotation CSV files, and the YouTube provenance list by default. The approximately 2.54 GB train and test archives require explicit license acknowledgment and opt-in. TriA section 4.1 further divides the official training set 9:1 for training and validation while preserving the 725-clip test set.
Safe-first helperscripts/download/nonspeech7k.sh
Speech generation
NonverbalTTS
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
Safe-first helper
Expressive Text To Speech
Nonverbal Vocalization Synthesis
Paralinguistic Fidelity Evaluation
Emotion Conditioned Speech Generation
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The dataset card's prose says annotations are CC BY-NC-SA 4.0 and audio follows the original VoxCeleb and Expresso licenses, while its machine-readable front matter declares Apache-2.0 for the dataset. Treat the more specific prose and upstream source restrictions as controlling pending clarification; do not assume the Apache tag re-licenses recordings or annotations for commercial use.
- Download notes
- The public English release contains about 17 hours of speech derived from VoxCeleb and Expresso, with text-aligned labels for ten nonverbal vocalization types and eight emotion categories. The paper reports speaker-disjoint train, development, and test splits; OpenSTBench section 4.2 evaluates all 359 test samples for English-to-Chinese speech-translation paralinguistic fidelity, using text-tag positions to approximate event locations. The helper downloads official documentation and repository metadata by default. The approximately 108 MiB test Parquet is a separate opt-in, and the full approximately 3.9 GiB Hugging Face snapshot requires NONVERBAL_TTS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/nonverbal_tts.sh
Speech recognition
NOTSOFAR-1
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
Safe-first helper
Distant Automatic Speech Recognition
Speaker Attributed Automatic Speech Recognition
Speaker Diarization
Continuous Speech Separation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- Microsoft applies CC BY 4.0 to the released data and MIT to the baseline repository. The repository notes that challenge-only Dev-set-2 is not part of the current open release and is restricted to publications about systems developed during the challenge; use the current public subsets for new work.
- Download notes
- The public English recorded-meeting release currently documents 237 meetings averaging six minutes across 30 conference rooms, with 4-8 attendees and 35 speakers. It provides train, dev, an 80-meeting eval-small set matching the challenge evaluation set, and a 129-meeting eval-full set, with ground truth available for both evaluation releases. The original challenge and paper describe roughly 280 meetings; the post-challenge open release removes Dev-set-2 for legal and quality reasons and adds eval-full, so benchmark versions must be reported. Single-channel and known-geometry seven-channel tracks use speaker-attributed tcpWER for ranking. The helper downloads only official documentation and license files by default; use the repository's versioned download utilities for selected audio subsets. The Hugging Face repository has very large historical storage and must not be snapshotted wholesale.
Safe-first helperscripts/download/notsofar_1.sh
Music
NSynth
NSynth: Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
Safe-first helper
Audio Synthesis
Music Synthesis
Musical Instrument Modeling
Timbre Modeling
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- Apache-2.0
- License caution
- Official Magenta dataset page lists the dataset under CC BY 4.0; Magenta code is Apache-2.0.
- Download notes
- The full TFDS download is about 73 GiB. The helper saves the official dataset page by default and requires NSYNTH_DOWNLOAD_ARCHIVES=1 before downloading split archives.
Safe-first helperscripts/download/nsynth.sh
Speech recognition
Nyra Verbatim Speech Benchmark
Nyra Verbatim Speech Benchmark for Controllable Verbatim Automatic Speech Recognition
Safe-first helper
Automatic Speech Recognition
Verbatim Speech Recognition
Intended Speech Recognition
Disfluency Detection
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_apache_2_0_and_unspecified
- Code license
- stated_MIT_without_repository_LICENSE_file
- License caution
- The English Hugging Face card declares Apache-2.0 and identifies its AMAAI Lab DisfluencySpeech source, which also declares Apache-2.0. The German dataset card declares its license as unknown and provides no separate grant, so do not assume the English terms apply. The evaluator README says MIT, but the repository currently has no LICENSE file or GitHub-detected license; confirm terms with the owner before reuse that depends on a formal code grant.
- Download notes
- The helper downloads owner-controlled dataset cards, API metadata, and evaluator documentation by default. The English and German Hugging Face snapshots are explicit opt-ins and total approximately 1.17 GB in current repository files. Cloning the evaluator and its cached model predictions is a separate opt-in.
Safe-first helperscripts/download/nyra_verbatim_speech_benchmark.sh
Audiovisual & cross-modal
Omni-Cloze
Omni-Cloze: Omni Detailed Captioning Benchmark
Safe-first helper
Detailed Audio Captioning
Detailed Visual Captioning
Detailed Audio Visual Captioning
Cloze Question Answering
+1 more
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the official Hugging Face card nor the GitHub repository states a data or code license, and GitHub reports no detected license. The release includes video files without documented source-media provenance or reuse terms; obtain clarification and review media rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains 2,320 audio-visual files across nine main domains and 47 subcategories, with approximately 70,000 fine-grained cloze blanks. The helper downloads official documentation and evaluation scripts by default; the 25.1 MB JSONL metadata and Hugging Face snapshot are separate opt-ins. The current repository files total about 6.1 GB, while the Hugging Face API reports about 11.3 GB of repository storage including history. Qwen3.5-Omni evaluates detailed audio-visual captioning on Omni-Cloze in section 5.1.4, Table 7.
Safe-first helperscripts/download/omni_cloze.sh
Audiovisual & cross-modal
OmniBench
OmniBench: Towards the Future of Universal Omni-Language Models
Safe-first helper
Omni Modal Question Answering
Audio Visual Question Answering
Tri Modal Reasoning
Speech Understanding
+2 more
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official dataset card has no license field, and the official repository has no LICENSE file or GitHub-detected license. The project page's footer links CC BY-SA 4.0 but does not clearly state that it covers the benchmark data, code, or component media. Obtain clarification and review source-image/audio rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains 1,142 four-choice questions that jointly pair an image, audio, and text prompt across seven task types. Audio spans speech, sound events, and music. The helper downloads official documentation and repository metadata by default; the Hugging Face card reports approximately 1.26 GB of downloads, so the media-bearing Parquet snapshot requires OMNIBENCH_DOWNLOAD_HF=1. OPOD evaluates OmniBench as its omni-modal benchmark and reports accuracy in the Experiments benchmark block of arXiv:2607.20918.
Safe-first helperscripts/download/omnibench.sh
Audiovisual & cross-modal
OmniGAIA
OmniGAIA: Towards Native Omni-Modal AI Agents
Safe-first helper
Omni Modal Question Answering
Audio Visual Reasoning
Multi Hop Reasoning
Tool Use
+2 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0
- Code license
- MIT
- License caution
- The Hugging Face card declares Apache-2.0 and the GitHub repository declares MIT. The benchmark curates media from FineVideo, LongVideoBench, LongVideo-Reason, COCO 2017, and other Hugging Face sources; verify component media rights and attribution requirements before redistribution or commercial use.
- Download notes
- The public, ungated release contains 360 English test tasks with audio, image, and video inputs across nine domains. The helper downloads official documentation and the lightweight test metadata JSON by default; the Hugging Face API reports about 9.9 GB of repository storage, so the full media snapshot requires OMNIGAIA_DOWNLOAD_HF=1. Qwen3.5-Omni reports OmniGAIA without a thinking prompt or answer-tag formatting in section 5.1.4, Table 7.
Safe-first helperscripts/download/omnigaia.sh
Speech recognition
Omnilingual ASR Corpus
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
Safe-first helper
Multilingual Asr
Low Resource Asr
Zero Shot Asr
Speech Representation Learning
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- Not specified in the source record.
- License caution
- Meta's dataset card and repository README declare the corpus CC BY 4.0; repository code and models are Apache 2.0. The current empty MoLGE derivative repository has no dataset card or verifiable license, and any future derivative would retain the source-recording terms. The corpus contains human voices and spontaneous responses, so attribution, privacy, consent, and local-law considerations remain relevant even though access is ungated.
- Download notes
- The commissioned corpus provides spontaneous speech and transcripts for 348 languages, with train, development, and test splits where available. The paper reports 3,350 hours; the current Hugging Face repository exposes 349 configurations and approximately 491 GB of storage. The helper downloads only the official dataset card, Hub API metadata, repository documentation, and current Hub metadata for the MoLGE segmented derivative by default. A single explicit language-script configuration from Meta's canonical release can be requested with OMNILINGUAL_ASR_DOWNLOAD_CONFIG=1; it never requests the complete snapshot. The MoLGE paper links a derivative in which the authors use MMS-FA to segment the 2,342 hours of Omnilingual ASR data used in their experiments into roughly 30-second chunks. As checked on 2026-07-28, however, that Hugging Face repository contained only .gitattributes in its visible tree: no card, audio, manifest, or license was available. Its API reported approximately 26.1 GB of used storage, but that unexplained backend field does not provide downloadable files or establish a usable release. Treat it as an announced, currently metadata-only preprocessing variant rather than a distinct benchmark. Omnilingual ASR models cover more than 1,600 languages, but that model coverage must not be confused with this 348-language public corpus.
Safe-first helperscripts/download/omnilingual_asr_corpus.sh
Audiovisual & cross-modal
OmniRetriever-Bench
OmniRetriever-Bench: 12-Direction Audio-Video-Text Retrieval Benchmark
Safe-first helper
Audio Text Retrieval
Audio Video Retrieval
Audio Video Text Retrieval
Cross Modal Retrieval
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_research_use
- Code license
- not_specified
- License caution
- The Hugging Face card labels the annotations Apache-2.0 but adds a biometric-identification, profiling, and surveillance prohibition, while the paper calls the release a custom research-use license. Treat the additional restriction as nonstandard and controlling pending a standalone license file. Underlying TikTok media remains owned by its uploaders and governed by platform terms. The evaluator repository has no detected license.
- Download notes
- The public, ungated release contains a 923,982-byte CSV with 3,782 held-out English-captioned audio-video-text triples and evaluates six single-modal plus six dual-modal retrieval directions. The helper downloads the official dataset card, CSV annotations, evaluator README, and paper page. Media is not redistributed; each row contains a TikTok source URL and clip interval, so availability depends on the source platform and users must obtain media themselves under applicable platform and uploader terms.
Safe-first helperscripts/download/omniretriever_bench.sh
Audiovisual & cross-modal
OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Manual or gated
Audio Visual Question Answering
Audio Visual Reasoning
Long Video Understanding
Temporal Reasoning
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- conflicting_CC-BY-NC-SA-4.0_and_CC-BY-NC-ND-4.0
- Code license
- not_specified
- License caution
- The official GitHub README says CC BY-NC-SA 4.0, while the Hugging Face card metadata declares CC BY-NC-ND 4.0 and the access form additionally requires non-commercial research use and no redistribution without permission. Apply the stricter gated terms pending clarification. The authors explicitly do not own the raw-video copyrights, so source-media rights remain separate; the GitHub repository has no detected license for its evaluation code.
- Download notes
- The gated Hugging Face release contains 628 English videos lasting from several seconds to 30 minutes and 1,000 manually verified question-answer pairs. Every question requires complementary audio and visual evidence and includes atomic step-by-step reasoning annotations; the paper reports 762 speech, 147 sound, and 91 music questions across 13 reasoning types. The helper downloads public repository documentation and API metadata by default. The Hugging Face API reports approximately 114 GB of repository storage, so the complete snapshot requires approval through the dataset questionnaire, authentication, and OMNIVIDEOBENCH_DOWNLOAD_HF=1. OmniReasoner evaluates OmniVideoBench in its main audio-visual benchmark table and reports duration- and audio-type-specific results.
Safe-first helperscripts/download/omnivideobench.sh
Speech recognition
Opencpop-test
Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis
Manual or gated
Singing Voice Transcription
Mandarin Singing Voice
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- The official page limits the dataset to non-commercial use under CC BY-NC-ND 4.0, retains corpus copyright with the Opencpop Team, and directs commercial users to contact the team. The repository has no detected license. The license-page path is spelled /liscense/.
- Download notes
- The official corpus contains 3,756 Mandarin singing utterances from 100 songs and about 5.2 hours of studio audio. Its fixed test split is five randomly selected songs. Qwen3.5-Omni section 5.1 evaluates singing transcription on Opencpop-test and reports CER in Table 5. The helper saves the official landing, download, license, test-set, and paper pages by default; corpus access still requires the owner Google Form and emailed instructions.
Safe-first helperscripts/download/opencpop_test.sh
Music
OpenMIC-2018
OpenMIC-2018: An Open Dataset for Multiple Instrument Recognition
Safe-first helper
Musical Instrument Recognition
Music Auto Tagging
Multi Label Audio Classification
Music Information Retrieval
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo says Spotify AB releases the dataset under CC BY 4.0 and includes full license terms in the archive. The included metadata contains licenses for each audio recording, so check per-track metadata before redistribution or commercial use.
- Download notes
- The Zenodo archive is about 2.6 GiB and contains 10-second OGG clips, VGGish features, crowd-sourced labels, metadata with per-recording licenses, and train/test partitions. The helper saves the Zenodo record JSON and official README by default and requires OPENMIC_DOWNLOAD_ARCHIVE=1 before downloading the full archive.
Safe-first helperscripts/download/openmic_2018.sh
Speech generation
OpenSTBench
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
Safe-first helper
Speech To Text Translation Evaluation
Speech To Speech Translation Evaluation
Streaming Speech Translation Evaluation
Translation Quality Evaluation
+6 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other
- Code license
- MIT_with_CC-BY-SA-4.0_adapted_components
- License caution
- The paired-set card marks the dataset license as "other" and says users must comply with the original data and synthesis-component terms. LibriTTS is CC BY 4.0, but neither the repository's MIT license nor the paper's CC BY-NC-SA 4.0 license establishes standalone reuse terms for all translated metadata or Qwen3-TTS-generated reference audio. OpenSTBench's original code is MIT, while adapted SimulEval latency components are CC BY-SA 4.0. MSLT, RAVDESS, MCAE-SPPS, NonverbalTTS, and SynParaSpeech retain their own access and license terms.
- Download notes
- OpenSTBench provides a public evaluation package for offline and streaming speech-to-text and speech-to-speech translation. The paper's section 4.2 evaluates separate public source datasets for translation, speech quality, emotion, paralinguistics, temporal consistency, and latency, and constructs a 300-sample, 35-speaker LibriTTS-based paired set for speaker preservation. That paired set is public and ungated on Hugging Face and includes original and prompt LibriTTS audio, translated text, and Qwen3-TTS-synthesized reference speech. The helper downloads the paper, repository documentation, license notices, dataset card, and Hugging Face API metadata by default. Cloning the evaluation toolkit is opt-in, and downloading the approximately 511 MiB paired-set snapshot requires OPENSTBENCH_DOWNLOAD_PAIRED_SET=1. The helper does not fetch the other component datasets.
Safe-first helperscripts/download/openstbench.sh
Audiovisual & cross-modal
OV-MERD
OV-MERD: Open-Vocabulary Multimodal Emotion Recognition Dataset
Manual or gated
Open Vocabulary Multimodal Emotion Recognition
Audio Visual Emotion Understanding
Acoustic Emotion Cue Reasoning
Free Form Multi Label Emotion Prediction
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_additional_gated_terms
- Code license
- Apache-2.0_with_noncommercial_notice
- License caution
- The paper and Hugging Face card identify OV-MERD as CC BY-NC 4.0. MER2025's gated terms further limit use to academic research and non-commercial purposes, prohibit distribution of the dataset or derivative annotation/label files to third parties, and prohibit modification without prior written consent. The OV-MER repository includes Apache-2.0 code terms but also calls the service a non-commercial research preview. Source clips derive from MER2023 movie and television media, so underlying media rights remain separate.
- Download notes
- OV-MERD extends a consented subset of MER2023 movie and television clips with human-checked acoustic and visual clues, merged multimodal descriptions, and open-vocabulary emotion labels. The paper reports 236 emotion categories, one to nine labels per sample (most have two to four), and clips that are mostly one to four seconds long. The current official repository points to the gated MER2025 release, which bundles OV-MERD label and description tables with audio, video, subtitles, and face features for the broader challenge corpus. The Hugging Face API reports approximately 442 GB of repository storage. The helper downloads only public official documentation, repository metadata, the paper landing page, and the Hugging Face API response; users must request access and download any benchmark files manually.
Safe-first helperscripts/download/ov_merd.sh
Speech recognition
Pansori-TEDxKR
Pansori-TEDxKR: Korean Speech Corpus Generated from Korean Language TEDx Talks
Safe-first helper
Korean Asr
Speech Transcription
Subtitle Aligned Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- OpenSLR lists Creative Commons BY-NC-ND 4.0. The corpus is derived from TEDx talks, so downstream use should also respect TED/TEDx source terms and original media rights.
- Download notes
- OpenSLR SLR58 hosts about 3 hours of Korean TEDx speech from 41 speakers with subtitle-boundary segmentation and manual alignment checks. The helper downloads the OpenSLR page, about page, info, and checksum by default; the 174 MB corpus archive is opt-in.
Safe-first helperscripts/download/pansori_tedxkr.sh
Speaker, identity & emotion
ParaPairAudioBench
ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge
Safe-first helper
Paralinguistic Pairwise Judgment
Audio Language Model Judge Evaluation
Speaking Style Comparison
Speech Rate Comparison
+4 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_non_commercial_and_gated
- Code license
- not_specified
- License caution
- The benchmark repository has no license file or detected GitHub license. Its README lists SVC as non-commercial academic research, EARS and Expresso as CC BY-NC 4.0, and LibriTTS as CC BY 4.0. Treat the pair annotations and builder code as rights-unspecified, obtain SVC approval for the affected age/gender rows, and preserve each source corpus's terms.
- Download notes
- The official repository releases pairwise JSON annotations and swapped-order variants for speech rate, emphasis, and style, plus source-pair metadata and builders for age and gender. The paper reports 5,175 pairs across five criteria, including tie and same-/cross-transcript conditions. The helper downloads official documentation and repository metadata by default; cloning the approximately 6 MB repository is opt-in and does not fetch underlying audio. Age and part of gender require manually approved SVC access; the remaining source audio comes from EARS, Expresso, and LibriTTS under their own download terms.
Safe-first helperscripts/download/parapair_audio_bench.sh
Speaker, identity & emotion
PartialEdit
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
Safe-first helper
Partial Speech Deepfake Detection
Partial Speech Deepfake Localization
Neural Speech Editing Detection
Codec Artifact Analysis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_upstream_terms_and_partial_release
- Code license
- not_applicable
- License caution
- Zenodo declares CC BY 4.0 for the released record. The audio is derived from VCTK, whose official release is also CC BY 4.0, but users should preserve both provenances and review neural-editor output terms. The license does not make the withheld Audiobox-derived E3/E4 subsets public or grant rights to reconstruct them.
- Download notes
- The official Zenodo release contains the E1 (VoiceCraft), E1-Codec, E2 (SSR-Speech), and E2-Codec subsets derived from VCTK, plus the E1/E2 protocol CSV and modified-text metadata. The four audio archives total approximately 21.9 GB. The project and Zenodo description state that E3 (Audiobox-Speech) and E4 (Audiobox) cannot be released under Audiobox's license. The helper downloads official pages and Zenodo record metadata by default; the approximately 7.7 MB protocol/text metadata and large audio archives are separate opt-ins. SALMONN-2 section IV-E and Table VIII evaluate temporal spoof localization on PartialEdit using mean intersection over union.
Safe-first helperscripts/download/partialedit.sh
Speech recognition
PazaBench
PazaBench: A Benchmark for Automatic Speech Recognition on Low Resource Languages
Safe-first helper
Automatic Speech Recognition
Low Resource Asr
Multilingual Asr
Asr Efficiency
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The Hugging Face Space card declares MIT, although its visible tree has no standalone LICENSE file. That declaration does not relicense the 11 source dataset groups: their displayed terms range from CC0 and CC BY to CC BY-SA, CC BY-NC-SA, gated CC BY-NC, Apache-2.0, and mixed terms. Review each upstream dataset card and access agreement before use, redistribution, or commercial deployment.
- Download notes
- The public Microsoft Research Africa leaderboard currently reports WER, CER, and inverse real-time factor across 61 African languages and 53 ASR/language models. It evaluates 16 kHz mono speech from 11 named public or community dataset groups, including African Next Voices, ALFFA, FLEURS, Common Voice 23.0, WAXAL, Naija Voices, and TWB Voice. The public Space exposes the leaderboard, submission interface, dataset inventory, and implementation, but its repository does not contain a standalone unified audio snapshot, frozen item manifest, or result CSV; obtain evaluation audio from each named upstream provider under that provider's access terms. The helper saves official documentation, Space metadata, implementation metadata, and the dataset inventory only; it does not download source corpora or model weights.
Safe-first helperscripts/download/pazabench.sh
Audio understanding, generation & events
PhysioNet/CinC 2016 Heart Sound
Classification of Heart Sound Recordings: The PhysioNet/Computing in Cardiology Challenge 2016
Safe-first helper
Abnormal Heart Sound Detection
Phonocardiogram Classification
Cardiac Screening
Clinical Audio Classification
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ODC-By-1.0
- Code license
- mixed_entry_specific_terms
- License caution
- PhysioNet licenses the released files under the Open Data Commons Attribution License v1.0. Submitted challenge systems carry their own entry-specific licenses. Preserve dataset and PhysioNet attribution; the public license does not remove clinical-audio privacy, ethics, or re-identification responsibilities.
- Download notes
- The helper saves the official challenge and license pages by default. The approximately 190 MB training.zip archive is public but requires explicit opt-in. The current PhysioNet file tree exposes six training directories, A-F, totaling 3,240 recordings; challenge prose also retains an earlier description of 3,126 recordings in five databases. The challenge's subject-disjoint test set remains private.
Safe-first helperscripts/download/physionet_cinc_2016_heart_sound.sh
Speech generation
PodEval
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
Safe-first helper
Podcast Generation Evaluation
Long Form Audio Generation Evaluation
Dialogue Naturalness Evaluation
Speech Quality Evaluation
+3 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_for_real_pod_manifest_and_linked_audio
- Code license
- MIT
- License caution
- The repository's MIT license covers the software and accompanying documentation, but no separate data license is stated for the Real-Pod manifest. The linked podcast recordings remain hosted by third parties and retain creator, publisher, platform, voice, music, and other media rights. The maintainers direct users to follow legal and ethical rules and use the reference data for research and education; public links do not grant redistribution or commercial-use rights.
- Download notes
- The public framework evaluates podcast generation across text, speech, and audio using objective metrics, LLM-based judging, and structured listening tests. Its Real-Pod reference manifest contains 51 topics across 17 categories, with one publicly accessible Apple Podcasts episode link per topic. The repository does not redistribute those recordings. The helper downloads official documentation, the small JSON manifest, license, repository metadata, and arXiv metadata by default; cloning the approximately 12 MB MIT toolkit is opt-in and still does not download podcast audio.
Safe-first helperscripts/download/podeval.sh
Audiovisual & cross-modal
POLY-SIM 2026
Polyglot Speaker Identification with Missing Modality 2026 Grand Challenge
Safe-first helper
Closed Set Speaker Identification
Audio Visual Speaker Recognition
Audio Only Speaker Recognition
Cross Lingual Speaker Recognition
+1 more
Access pathOfficial / other
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the official repository nor the results/evaluation-plan papers state a data or code license. The release derives from MAV-Celeb interviews, talk shows, and television debates sourced from YouTube; uploader copyright, likeness/privacy considerations, and platform terms remain applicable. Confirm reuse and redistribution rights before downloading or publishing derived data.
- Download notes
- The English-Urdu challenge release provides raw paired speech/face data, precomputed features, and CSV files for train, development, and hidden-label test splits. The paper reports 6,850 English and 12,706 Urdu samples across the three splits. The helper downloads official documentation only by default, optionally clones the small baseline repository, and leaves Drive-hosted data as a manual download.
Safe-first helperscripts/download/polysim_2026.sh
Music
POP909
POP909: A Pop-song Dataset for Music Arrangement Generation
Safe-first helper
Symbolic Music Generation
Piano Arrangement Generation
Melody Conditioned Accompaniment Generation
Beat Tracking
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_repository_license_with_underlying_music_rights_caveat
- Code license
- MIT
- License caution
- The official repository includes an MIT LICENSE and GitHub detects MIT, but neither the README nor paper separately analyzes rights in the transcribed arrangements or underlying popular compositions. Original commercial-song audio is not redistributed. Retain the repository notice and review composition, arrangement, performance, and jurisdiction-specific rights before redistribution or commercial use.
- Download notes
- The public repository contains MIDI arrangements for 909 popular songs, including melody, bridge, and piano-accompaniment tracks, multiple arrangement versions, and aligned beat, chord, key, and tempo-derived annotations. The paper describes approximately 60 hours of symbolic performances created by professional musicians. Original commercial-song audio is not included. The helper downloads official documentation, license, repository metadata, and the song index by default; cloning the approximately 47 MB GitHub repository and its MIDI/annotation files requires POP909_CLONE_REPO=1.
Safe-first helperscripts/download/pop909.sh
Speech recognition
Preference-ASR
Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs
Safe-first helper
Instruction Following Asr
Preference Aware Asr
Inverse Text Normalization
Disfluency Preservation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-4.0_author_annotations_with_source_corpus_terms
- Code license
- Apache-2.0
- License caution
- The repository license applies CC BY-SA 4.0 to author-created instruction, preference-text, and category annotations and Apache 2.0 to the downloader and normalizer. Audio, original transcripts, and related fields retain their source-corpus terms. The current released rows reference LibriSpeech and AMI (CC BY 4.0), Common Voice (CC0), and VoxPopuli (CC0 with European Parliament source notice). The license also documents Earnings-22, GigaSpeech, and SPGISpeech for the broader construction pipeline; GigaSpeech and SPGISpeech remain gated and restricted even though no rows from them appear in the current public manifest.
- Download notes
- The official repository now publishes a 2,822-row English JSONL manifest with audio paths, source transcripts, natural-language instructions, preferred transcripts, and category labels. It contains 1,113 case, 906 entity, 516 disfluency, 228 normalization, and 59 standard items drawn from VoxPopuli, LibriSpeech, AMI, and Common Voice. The repository README supersedes the paper's earlier 3,210-preference-plus-335-standard composition and describes this released manifest as the test set. The helper downloads the paper, official documentation, license, repository metadata, and the 1.6 MB manifest by default; cloning the evaluation normalizer and upstream audio downloader is opt-in. Source audio is not redistributed and must be obtained separately under each corpus's terms.
Safe-first helperscripts/download/preference_asr.sh
Speech recognition
Primewords Chinese Corpus Set 1
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR lists Attribution-NonCommercial-NoDerivatives 4.0 International and describes the corpus as free for academic use. Re-check upstream terms before redistribution or commercial use.
- Download notes
- OpenSLR hosts a 9.0 GiB Mandarin speech/transcript archive recorded from 296 native Chinese speakers. The helper saves the OpenSLR page by default and makes the full archive an explicit opt-in.
Safe-first helperscripts/download/primewords_chinese.sh
Representation & general suites
PROCESS-2
PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
Manual or gated
Cognitive Impairment Classification
Mild Cognitive Impairment Detection
Dementia Detection
Cognitive Score Regression
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_research_only_data_use_agreement
- Code license
- Apache-2.0
- License caution
- The participant data agreement prohibits re-identification, external linkage for identification, redistribution, public hosting, inclusion in public model-training datasets, surveillance, profiling, discriminatory uses, and biometric-identification training. It also requires secure storage, authorized-personnel access, deletion on completion or request, citation, and omission of identifiable speech excerpts from publications. The Hugging Face card uses the machine tag "other"; Apache-2.0 applies to the separate analysis code and Zenodo software archive, not to the clinical dataset.
- Download notes
- The helper saves the public Hugging Face API response, official analysis-repository documentation, repository metadata, code license, and Zenodo software metadata. It never authenticates, submits the access form, or downloads participant speech, transcripts, or clinical metadata.
Safe-first helperscripts/download/process_2.sh
Enhancement, separation & quality
PVQD
Perceptual Voice Qualities Database
Safe-first helper
Clinical Voice Quality Assessment
Pathological Voice Assessment
Perceptual Voice Rating
Cape V Prediction
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- Mendeley Data v4 and its DataCite DOI record declare CC BY 4.0. Preserve attribution and modification notices. The recordings contain human voices and clinical voice-quality information, and the release includes demographics; lawful and ethical handling of identifiable and health-related data remains necessary even though access is ungated.
- Download notes
- The public, ungated Mendeley Data v4 release contains 296 mono 44.1 kHz, 16-bit WAV recordings of sustained /a/ and /i/ vowels and six CAPE-V sentences, plus 13 XLSX files with demographics and experienced clinicians' CAPE-V and GRBAS ratings. The complete 310-file release is approximately 514.5 MiB. The helper downloads the official dataset page, DataCite DOI record, and live Mendeley file manifest by default. PVQD_DOWNLOAD_ANNOTATIONS=1 fetches the approximately 0.6 MiB PDF/XLSX documentation and labels; PVQD_DOWNLOAD_ALL=1 explicitly downloads the complete release including identifiable clinical voice recordings. The July 2026 voice-concept bottleneck paper uses an 80:20 speaker split and derives per-utterance segments with voice activity detection; those split and segment artifacts are not part of PVQD itself.
Safe-first helperscripts/download/pvqd.sh
Audiovisual & cross-modal
QIVD
Qualcomm Interactive Video Dataset: Can Vision-Language Models Answer Face to Face Questions in the Real-World?
Manual or gated
Situated Audio Visual Question Answering
Real Time Audio Visual Understanding
Spoken Query Understanding
When To Answer Prediction
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_publicly_specified_account_terms_apply
- Code license
- not_applicable
- License caution
- No public dataset license text was found on the official landing page or in the paper on 2026-07-21. The paper says crowd contributors signed consent permitting research and commercial use of their video and audio, but that consent statement is not itself a downstream dataset license. Review and retain any terms shown in Qualcomm's account/download flow before use or redistribution.
- Download notes
- The official Qualcomm release page describes 2,900 short English video files across 13 semantic categories. Each clip contains raw audio with a spoken question, an annotated transcription, a text answer, and a timestamp indicating when enough context is available to answer. The helper saves the public landing page, then prints the manual Qualcomm account/download path because the release flow is JavaScript-driven and may present account-specific terms. Qwen3.5-Omni calls it Qualcomm IVD and evaluates audio-query interaction in section 5.1.4, Table 7.
Safe-first helperscripts/download/qivd.sh
Enhancement, separation & quality
QualiSpeech
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
Safe-first helper
Speech Quality Assessment
Multidimensional Speech Quality
Speech Quality Reasoning
Natural Language Quality Description
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0-with-upstream-restrictions
- Code license
- not_separately_specified
- License caution
- The official card declares CC BY-NC-SA 4.0 for QualiSpeech. Component audio retains upstream conditions from BVCC, NISQA, and GigaSpeech; in particular, Blizzard Challenge data may not be redistributed and is not included in the Hugging Face release. The card does not state a separate software license for its download and merge scripts.
- Download notes
- The public, ungated release provides official train, validation, and test CSVs with seven numerical quality scores, four specific descriptions, and natural-language quality reasoning. The paper reports 10,558 train, 2,167 validation, and 1,852 test utterances from BVCC, recent TTS systems, GigaSpeech, and NISQA. The helper downloads the dataset card, scripts, split CSVs, and BVCC filename lists by default. Its approximately 1.46 GiB non-BVCC WAV archive is opt-in; BVCC audio is deliberately excluded and must be obtained separately under its original terms before running the official merge script. SALMONN-2 section IV-E and Table VIII evaluate QualiSpeech using Pearson correlation coefficient.
Safe-first helperscripts/download/qualispeech.sh
Audiovisual & cross-modal
RAVDESS
The Ryerson Audio-Visual Database of Emotional Speech and Song
Safe-first helper
Speech Emotion Recognition
Audio Visual Emotion Recognition
Acted Emotional Speech
Emotional Song Recognition
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_applicable
- License caution
- Zenodo lists CC BY-NC-SA 4.0 for the dataset and says commercial licenses are available separately. The linked PLOS ONE paper itself is CC BY, but that article license is not the dataset license.
- Download notes
- The helper downloads the Zenodo metadata JSON by default. The audio-only speech and song ZIPs are about 215 MB and 198 MB respectively; video archives are much larger and are not downloaded by default.
Safe-first helperscripts/download/ravdess.sh
Speaker, identity & emotion
REAL-TSE Challenge
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Manual or gated
Target Speaker Extraction
Online Target Speaker Extraction
Offline Target Speaker Extraction
Real Conversational Speech Separation
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- access_restricted_terms_not_publicly_specified
- Code license
- MIT
- License caution
- The official evaluation repository is MIT, but that code license does not license the challenge audio. The public challenge page restricts DEV/EVAL use to validation or final evaluation, forbids training and fine-tuning, and says access was limited to registered teams; it does not state a standalone dataset license. DEV and EVAL-1 derive from AISHELL-4, AliMeeting, AMI, DiPCo, and CHiME-6, whose upstream terms also remain applicable.
- Download notes
- The challenge reports 6,991 Mandarin/English mixture-enrollment trials over 2,309 real conversational mixtures and 11.3 hours of mixture audio. DEV has 1,991 REAL-T-derived pairs; EVAL-1 and EVAL-2 contain 2,000 seen and 3,000 unseen pairs without public references. The July 2026 SonicAGI paper documents both tracks and reports second and fifth place; the MERL submission reports first place in offline Track 2 and audits DNSMOS and speaker-similarity metric fragility on DEV. Neither paper releases challenge data, models, training pipelines, outputs, or attack artifacts. The public helper saves official pages, repository metadata, and the papers only. Organizers distributed password-protected data by email exclusively to registered teams, registration closed on May 31, 2026, and no public dataset URL is currently provided.
Safe-first helperscripts/download/real_tse.sh
Audio understanding, generation & events
RealDESED
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Safe-first helper
Sound Event Detection
Domestic Sound Event Detection
Polyphonic Sound Event Detection
Temporal Audio Event Localization
+3 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc0_or_cc-by_per_audio_file_and_cc-by-4.0_annotations
- Code license
- MIT
- License caution
- The Zenodo record is open and lists CC BY 4.0 at record level, but its description and official repository clarify that each audio recording and corresponding metadata row uses the per-file license recorded in metadata.csv, either CC0 or CC BY; remaining metadata and annotations are CC BY 4.0. Preserve creator attribution for CC BY recordings and consult metadata.csv before redistribution. The baseline repository is MIT.
- Download notes
- The public release contains 5,710 real-home recordings (37.85 hours) from 652 participants, with 64,430 temporal annotations across 15 domestic event classes. Collector-disjoint train, validation, and test splits contain 3,704, 999, and 1,007 recordings. All validation and test labels and 1,108 training files (29.91%) received additional review. The protocol reports PSDS1 and PSDS2, including macro-averaged variants, and analyzes annotation aggregation, long-form inference, and recording-condition metadata. The helper saves official Zenodo/GitHub metadata, documentation, and the two lightweight collection and annotation guides by default; the three audio archives total approximately 8.74 GB and require explicit opt-in.
Safe-first helperscripts/download/realdesed.sh
Speaker, identity & emotion
RealMAN
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
Safe-first helper
Multichannel Speech Enhancement
Speech Source Localization
Dynamic Speaker Localization
Variable Array Generalization
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- unspecified
- License caution
- The official repository README declares the dataset CC BY 4.0, but the repository has no detected standalone license file and GitHub reports no repository license. Treat CC BY 4.0 as covering the dataset release, preserve attribution and notices, and verify terms separately for the baseline code and any downstream derived artifacts.
- Download notes
- The public, ungated release contains 83.7 hours of 32-channel speech recorded in 32 scenes and 144.5 hours of background noise recorded in 31 scenes, with direct-path speech, transcriptions, source locations, and scene and speaker metadata. The repository lists approximately 531.4 GB of training data, 27.5 GB of validation mixtures, 39.3 GB of test mixtures, 158 GB of raw validation/test recordings, and 129 MB of dataset information; the Hugging Face API currently reports about 812.0 GB of repository storage. The helper downloads only official documentation and API metadata by default, while the complete snapshot requires explicit opt-in. A July 2026 geometry-aware enhancement paper evaluates RealMAN after resampling to 8 kHz and fixed four-second test segments, and separately tests array generalization on CHiME-4.
Safe-first helperscripts/download/realman.sh
Speech understanding & dialogue
RealSI
RealSI: Open Benchmark for Simultaneous Interpretation in Real-world Scenarios
Safe-first helper
Simultaneous Speech To Text Translation
Simultaneous Speech To Speech Translation
Long Form Speech Translation
Streaming Latency Evaluation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- CC-BY-4.0_repository_license
- License caution
- The repository declares the dataset CC BY 4.0, but its README also says the authors do not own the copyright in the source videos and describes the release as annotations plus public video links for educational and informational use. The current repository additionally contains WAV derivatives. Treat CC BY 4.0 as covering author-created annotations, and independently review source-video copyright, platform terms, and the README disclaimer before using or redistributing audio.
- Download notes
- The official public repository contains timestamped Chinese-English and English-Chinese transcripts and translations for 20 natural, approximately 3-8 minute recordings across ten domains. The release totals 95 minutes 29 seconds and 778 utterance segments (431 En-to-Zh and 347 Zh-to-En), and its current tree also includes 20 WAV files. SimulS2ST-Omni section 4.1 and appendix D.3 reuse RealSI for sentence-level and long-form streaming S2TT/S2ST evaluation. The helper downloads official documentation, repository metadata, and all 20 lightweight JSON annotation files by default. The repository's WAV payload is about 351 MiB, so cloning the repository requires REALSI_CLONE_REPO=1.
Safe-first helperscripts/download/realsi.sh
Speech understanding & dialogue
RedVox
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
Safe-first helper
Multilingual Speech Llm Safety
Speech Llm Fairness Evaluation
Stereotype And Harmful Request Robustness
Modality Conditioned Safety
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_manual_custom_terms_not_publicly_readable
- Code license
- announced_Apache_2_0_but_not_released
- License caution
- The Hub card declares only "other" and requires manual approval; its README and dataset-info files return HTTP 401 before approval. Review the customized research-use and intended-use terms in the access flow. The article's CC BY 4.0 license does not cover participant recordings, and MUSAN-derived silence, ambient-noise, and babble controls retain upstream terms. Do not infer Apache-2.0 code terms while the linked repository remains unavailable.
- Download notes
- The manually gated owner release supplies the planned 3,414-entry subset from 26 consenting speakers in English, French, Italian, Spanish, and German. It evaluates open-ended safety, stereotype fairness, and response relatedness for spoken harmful requests, non-speech-audio-plus-text requests, and matched text-only controls. Public Hub metadata exposes five language configurations, 856 WAV files, five metadata JSONL files, and about 414.5 MB of storage. The helper fetches the paper, live API metadata, and a paginated public file listing by default; dataset files require owner approval, authentication, acknowledgement of the terms, and explicit opt-in. The separate full 6,118-entry collection from 52 participants remains private and is not represented as public. The paper-linked code repository still returns HTTP 404, so the released family currently has no verified public evaluator.
Safe-first helperscripts/download/redvox.sh
Music
RUBATO
RUBATO: A Multi-Version Benchmark for Robust Music Transcription and Analysis
Safe-first helper
Automatic Music Transcription
Beat Tracking
Downbeat And Measure Tracking
Local Key Estimation
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- not_separately_specified
- License caution
- Zenodo labels the deposition CC BY 3.0, but metadata_versions.csv assigns per-recording terms that include CC0, CC BY, CC BY-SA, CC BY-ND, CC BY-NC, CC BY-NC-SA, CC BY-NC-ND, ambiguous "CC"/"CC0?", and EEF. Treat the per-recording field as controlling and review it before redistribution or commercial use; scripts inside the archive have no separate license statement.
- Download notes
- The open Zenodo v0.3 release contains 566 versions of 15 musical works (about 42.9 hours), including 22.05 kHz mono audio, aligned score MIDI/MuseScore/PDF/images, performance video, note/beat/measure/local-key/structure annotations, and audio-to-score warping paths. The helper downloads the 83 KB version metadata and Zenodo API record by default; the approximately 6.26 GB archive requires RUBATO_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/rubato.sh
Audio understanding, generation & events
RUL-MuchoMusic
RUL-MuchoMusic from RUListening / Are You Really Listening?
Safe-first helper
Music Question Answering
Perceptual Music Understanding
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- RUL repo/HF card: mit; upstream MuChoMusic dataset: CC BY-SA 4.0
- Code license
- MIT
- License caution
- RUL-MuchoMusic derives from MuChoMusic; check upstream audio/source terms too.
Safe-first helperscripts/download/rul_muchomusic.sh
Speech recognition
S-DiverSe
S-DiverSe: Spanish Diverse Speech
Safe-first helper
Automatic Speech Recognition
Pathological Speech Recognition
Spanish Speech Recognition
Asr Robustness
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository has no license file or detected GitHub license. The paper is CC BY 4.0 on arXiv, but that does not license the annotations, reconstruction code, or linked source recordings. The manifest includes health-condition metadata and potentially identifiable speech/transcripts; review consent, privacy, research ethics, source rights, and platform terms before use or redistribution.
- Download notes
- The public repository releases a TSV manifest for 444 manually transcribed Spanish segments totaling 3.2 hours from 22 speakers with amyotrophic lateral sclerosis, Parkinson's disease, or post-stroke effects. Metadata includes speaker ID, sex, condition, intelligibility, source URL, timestamp, and duration. Audio is not redistributed; the repository provides a yt-dlp/ffmpeg reconstruction script for public video sources, whose availability and platform terms can change. The paper evaluates corpus-level and condition-specific WER after lowercasing, punctuation removal, digit expansion, and retention of filled pauses. The helper downloads annotations, documentation, and reconstruction code only; cloning the repository is opt-in.
Safe-first helperscripts/download/s_diverse.sh
Audio understanding, generation & events
SAKURA
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Safe-first helper
Audio Language Model Benchmark
Audio Question Answering
Audio Reasoning
Paralinguistic Reasoning
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_with_mixed_upstream_terms
- Code license
- not_specified
- License caution
- Neither the GitHub repository nor the four Hugging Face dataset cards publishes an explicit license. The release derives audio from Common Voice 17.0, CREMA-D, MELD, ESC-50, and the animal-sound dataset of Sasmaz and Tek; their separate terms and source-media constraints remain applicable. Public access is not permission for unrestricted redistribution or commercial use.
- Download notes
- The public, ungated release provides four 500-audio tracks: Gender, Language, Emotion, and Animal. Each audio item has paired single-hop perception and multi-hop reasoning questions, yielding 4,000 question instances across eight sub-tracks. The repository includes WAV files, CSV/JSON annotations, golden answers, GPT-4o judge prompts, and accuracy code. The helper fetches official documentation and repository/Hugging Face metadata by default; set SAKURA_DOWNLOAD_TRACK to GenderQA, LanguageQA, EmotionQA, AnimalQA, or all to explicitly download audio, or set SAKURA_CLONE_REPO=1 to clone the roughly 220 MiB repository.
Safe-first helperscripts/download/sakura.sh
Speaker, identity & emotion
SALMon
SALMon: A Suite for Acoustic Language Model Evaluation
Safe-first helper
Acoustic Consistency Evaluation
Acoustic Semantic Alignment
Speaker Consistency
Speaker Gender Consistency
+5 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official repository and Hugging Face card license the SALMon dataset under CC BY-NC 4.0 because some source datasets are non-commercial. The release derives speech or acoustics from Expresso, VCTK, LJSpeech, FSD50K, and EchoThief and also includes Azure TTS output, so preserve component provenance and review upstream terms. The evaluation repository has no standalone license file and GitHub reports no detected code license.
- Download notes
- The public, ungated release contains 1,600 positive/negative audio pairs across eight configurations for acoustic consistency and acoustic-semantic alignment. It covers speaker identity, speaker gender, sentiment, background sound, and room impulse response, using a likelihood-ranking protocol for speech language models. The helper saves official documentation and repository metadata by default; set SALMON_DOWNLOAD_HF=1 for the approximately 562 MB Hugging Face snapshot. Google Drive provides the same benchmark as raw WAV files.
Safe-first helperscripts/download/salmon.sh
Audiovisual & cross-modal
SAVEE
Surrey Audio-Visual Expressed Emotion Database
Manual or gated
Speech Emotion Recognition
Audio Visual Emotion Recognition
Speaker Independent Emotion Classification
Explainable Speech Emotion Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_free_for_research_registration_required
- Code license
- not_applicable
- License caution
- The official download page says SAVEE is available free of charge for research purposes, but publishes no standard license, redistribution grant, or commercial-use terms. Treat access as research-only and request clarification from the owner before redistribution or use outside research. The recordings contain identifiable actor voices and faces, so privacy, likeness, and research-ethics considerations remain applicable.
- Download notes
- The owner page describes 480 British-English audiovisual utterances from four male actors, with 120 utterances per speaker across anger, disgust, fear, happiness, sadness, surprise, and neutral. It reports 44.1 kHz audio and 60 fps video. Access requires registration, after which approved users log in to the password-protected data directory. The helper saves only public owner documentation and never authenticates or downloads the corpus.
Safe-first helperscripts/download/savee.sh
Representation & general suites
SEABAD
SEABAD: Southeast Asian Bird Activity Detection
Manual or gated
Bird Activity Detection
Binary Bird Presence Detection
Passive Acoustic Monitoring
Tropical Bioacoustic Detection
+1 more
Access pathZenodo
Upstream termsNot specified
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- MIT_claimed_in_readme_no_license_file
- License caution
- Zenodo's structured record declares CC BY 4.0 for the compilation, while the official repository says individual positive clips retain their Xeno-Canto licenses, including CC0, CC BY-SA, CC BY-NC-SA, and CC BY-NC-ND, and negative clips retain each source dataset's terms. Use the included provenance metadata and comply per recording; do not infer that the record-level license removes noncommercial, share-alike, no-derivatives, or attribution requirements. The repository README calls the curation code MIT, but the repository currently has no LICENSE file and GitHub detects no license, so that code statement should be clarified before reuse.
- Download notes
- Zenodo v1.0.0 releases 50,000 balanced three-second, 16 kHz mono WAV clips: 25,000 bird-present clips spanning 1,677 Southeast Asian species and 25,000 bird-absent clips. Fixed stratified train, validation, and test splits contain 40,000, 5,000, and 5,000 clips. Positive recordings derive from Xeno-Canto; negatives derive from BirdVox-DCASE-20k, Freefield1010, Warblr, FSC-22, ESC-50, and DataSEC. The helper downloads official Zenodo, paper, and repository metadata by default. The single approximately 3.87 GiB mybad.zip archive requires explicit source-terms acknowledgment and an audio opt-in. DrongoNet section 7.1 evaluates the held-out 5,000-clip test split across five seeds and reports AUC, accuracy, recall, and F1.
Safe-first helperscripts/download/seabad.sh
Audiovisual & cross-modal
Seamless Interaction
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Safe-first helper
Dyadic Audiovisual Interaction Modeling
Audiovisual Behavioral Motion Generation
Conversational Gesture Generation
Active Listener Motion Generation
+6 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- CC-BY-NC-4.0_repository_wide
- License caution
- The Hugging Face card and repository LICENSE declare CC BY-NC 4.0, including attribution and noncommercial restrictions; the repository does not provide a separate software license for its Python toolkit. The release contains identifiable audiovisual recordings, participant metadata, relationship information, transcripts, and behavioral features. The card documents occasional timestamp and participant-ID errors and says hundreds of hours of between-interaction "meta time" remain unreleased. Review the official data policy and applicable privacy, consent, biometric, and human-subject requirements before reuse; public download access does not remove those obligations.
- Download notes
- Seamless Interaction is one source dataset and modeling/evaluation family, not separate benchmark families for its improvised and naturalistic configurations or annotation types. The public, ungated release reports more than 4,000 hours from more than 4,000 participants, with 4,065.04 hours and 64,739 dyadic interactions split into improvised and naturalistic train, development, and test partitions. It provides synchronized participant video, 48 kHz denoised audio, time-aligned transcripts and VAD, body and facial features, movement representations, relationship/context metadata, and limited first- and third-party internal-state and behavior annotations. The paper's section 6 and appendix B propose automatic metrics plus separate human Dyadic Body and Dyadic Face Protocols, while the dataset card does not define one mandatory leaderboard protocol. The paper also describes participant-disjoint private test partitions (40.75 naturalistic hours and 15.32 improvised hours) held out for future benchmarking; those private assets are not publicly released. The helper fetches only official cards, API/repository metadata, and license text by default. An explicit metadata opt-in downloads about 9.3 MB of CSV indexes; a separate opt-in clones the roughly 36 MB toolkit. Dataset files are organized as approximately 50 GB batches and the complete release is reported as roughly 27 TB, so the helper deliberately does not wrap a full snapshot download. Use the official browser or toolkit to select individual interactions or batches.
Safe-first helperscripts/download/seamless_interaction.sh
Speech generation
Seed-TTS Eval
Seed-TTS objective zero-shot speech generation evaluation set
Safe-first helper
Zero Shot Text To Speech
Voice Cloning
Zero Shot Voice Conversion
Speech Intelligibility Evaluation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository has no LICENSE file and GitHub reports no detected license. The objective set selects 1,000 English samples from Common Voice and 2,000 Mandarin samples from DiDiSpeech-2, so verify both component-source terms before redistribution or commercial use. Public access does not imply an open license.
- Download notes
- The helper downloads the official README and lightweight evaluation scripts by default; cloning the evaluation repository is opt-in. The public objective EN/ZH test set is linked through Google Drive and must be downloaded manually. The Seed-TTS paper states that the 100-sample-per-language subjective set is not released because of copyright restrictions.
Safe-first helperscripts/download/seed_tts_eval.sh
Speech generation
Silent Speech EMG
Silent Speech EMG: Facial Electromyography Recordings for Voicing Silent Speech
Safe-first helper
Emg To Speech Synthesis
Silent Speech Recognition
Articulatory To Acoustic Mapping
Open Vocabulary Speech Generation
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- Zenodo declares the corpus CC BY 4.0 and the official repository declares its code MIT. Preserve creator attribution and citation. Because the release contains biometric facial-muscle signals and a participant's voice, downstream reuse should also consider privacy, biometric-data, and research-ethics obligations even though the official record is public.
- Download notes
- The official Zenodo release contains one approximately 3.92 GB emg_data.tar.gz archive with facial EMG and audio from silent and vocalized English speech by one speaker. The official repository describes approximately 19 hours of open-vocabulary recordings and supplies fixed evaluation manifests, preprocessing, synthesis, direct recognition, and ASR-scored evaluation code. The helper downloads Zenodo metadata, official README/LICENSE files, and paper pages by default. The corpus archive requires SILENT_SPEECH_EMG_DOWNLOAD_ARCHIVE=1 and is verified against the official MD5 checksum.
Safe-first helperscripts/download/silent_speech_emg.sh
Speech generation
SILMA Open-source Arabic TTS Benchmark
Open-source Arabic TTS Benchmark
Safe-first helper
Speech Synthesis
Text To Speech
Arabic Speech Synthesis
Dialectal Speech Synthesis
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_declared_at_space_level
- Code license
- Apache-2.0_declared_at_space_level
- License caution
- The Space metadata declares Apache-2.0, but the repository contains no separate license file and does not document the provenance or reuse terms of its Arabic prompts. Generated clips may also remain subject to the licenses and acceptable-use terms of the evaluated TTS models. Treat the Space-level declaration as insufficient to resolve all prompt, voice, and model-output rights before redistribution or commercial use.
- Download notes
- SILMA's public, ungated Hugging Face Space provides fixed prompts and generated model audio for direct listening comparisons in Modern Standard Arabic, Egyptian Arabic, and Saudi Arabic. The current release has 10 MSA prompts across four systems, five Egyptian prompts across five systems, and five Saudi prompts across three systems. SILMA says this release deliberately prioritizes auditory assessment because WER, CER, speaker similarity, and UTMOS do not fully capture Arabic speech nuances; it does not publish aggregate human ratings or a formal automatic scoring protocol. The helper downloads the official README, application source, three prompt CSVs, and Space API metadata by default. Cloning the approximately 29.6 MB Space, including generated evaluation audio, requires SILMA_ARABIC_TTS_CLONE_SPACE=1.
Safe-first helperscripts/download/silma_open_source_arabic_tts.sh
Speaker, identity & emotion
SingFox
SingFox: A Multi-Lingual Singfake Detection Corpus
Manual or gated
Singing Voice Deepfake Detection
Multilingual Singfake Detection
Cross Dataset Generalization
Synthetic Source Verification
+2 more
Access pathOfficial / other
Upstream termsNot specified
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_specified_request_restricted_with_upstream_media_terms
- Code license
- not_specified
- License caution
- The Zenodo record does not declare a reusable data license, and the GitHub repository has no LICENSE file or GitHub-detected license. The arXiv manuscript is CC BY-NC-ND 4.0, which licenses the paper rather than the corpus. SingFox incorporates real singing and source material from multiple upstream datasets and generated outputs from several models; access approval does not by itself grant redistribution or commercial-use rights. Review the signed corpus documentation and every upstream source and model term before use.
- Download notes
- SingFox is one multilingual singing-deepfake benchmark family with six related tracks, not six independent dataset releases. T1 covers 14 widely spoken non-Indic languages; T2 covers six Indic languages; T3 evaluates five music types; T4 combines T1 and T2 as the recommended 20-language global evaluation; T5 mixes real and fake vocals and instruments as alternative-fake conditions; and T6 defines closed-set and open-set source-verification trials over T4. The paper reports 113,802 clips, more than 126.32 hours, 20 languages, and about 1,150 singers across the family. Track totals overlap because T4 combines T1 and T2 and T6 reuses T4 test material. The helper saves the paper, repository documentation/API metadata, and Zenodo record metadata, then prints the owner-request steps; it cannot fetch the restricted audio.
Safe-first helperscripts/download/singfox.sh
Representation & general suites
Single-Item Kawaii Measure Validation Data
Validating the Single Item Kawaii Measure
Safe-first helper
Voice Cuteness Rating
Subjective Voice Evaluation
Paralinguistic Perception Modeling
Psychometric Measure Validation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- Not specified in the source record.
- License caution
- The public workbook does not display a data license or reuse notice. The paper's CC BY-NC-ND 4.0 license does not automatically license the spreadsheet, participant responses, or source voice/game-character stimuli. Treat participant identifiers as study pseudonyms, minimize unnecessary processing, and obtain author clarification before redistribution or commercial use. Separately sourced voice and character media retain their original rights.
- Download notes
- The paper validates a one-item kawaii/cuteness rating across nine Japanese participant datasets covering voice assistants and video-game characters. It reports 1,228 total participant records from 967 unique participants and evaluates convergent, known-groups, and cross-context validity. The author-linked public Google workbook exports to an approximately 1.0 MB XLSX with participant-ID, dataset-summary, analysis, and nine study-data worksheets. It contains ratings and derived analysis rather than a bundled audio corpus. The helper saves the public paper by default; exporting the participant-level workbook requires explicit opt-in and acknowledgment that reuse terms are unspecified.
Safe-first helperscripts/download/single_item_kawaii_measure.sh
Enhancement, separation & quality
SingMOS-Pro
SingMOS-Pro: A Comprehensive Benchmark for Singing Quality Assessment
Safe-first helper
Singing Quality Assessment
Singing Mos Prediction
Singing Voice Synthesis Evaluation
Singing Voice Conversion Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_upstream_source_terms
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY 4.0 and the official predictor repository is MIT. SingMOS-Pro includes ground truth and outputs from singing synthesis, conversion, resynthesis, and song-generation systems built from 12 source datasets. The paper and card do not provide a per-file license inventory, so retain source-corpus, performer, composition, model-output, and service terms rather than assuming the card clears all embedded audio rights.
- Download notes
- The public, ungated release contains 7,981 Chinese and Japanese singing clips totaling 11.15 hours, generated by 41 models across 12 source datasets. At least five experienced annotators rated every clip for overall MOS, and 4,155 clips additionally have lyrics and melody scores. The helper downloads official documentation, API metadata, split definitions, and system metadata by default. The approximately 11.6 MB sample/rating annotations require a separate opt-in, while the Hugging Face API reports approximately 2.83 GB of repository storage for the full audio snapshot.
Safe-first helperscripts/download/singmos_pro.sh
Enhancement, separation & quality
Slakh2100
Slakh2100: The Synthesized Lakh Dataset
Safe-first helper
Music Source Separation
Multi Instrument Automatic Transcription
Music Information Retrieval
Synthetic Multitrack Music
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- MIT
- License caution
- The official Slakh page states that Slakh2100 and Flakh2100 are licensed under Creative Commons Attribution 4.0 International. The slakh-utils repository is MIT. Slakh is synthesized from Lakh MIDI Dataset v0.1, so keep source MIDI attribution/provenance in downstream use.
- Download notes
- The helper downloads the official Slakh page and utility README/LICENSE by default. Zenodo currently hosts the full Slakh2100 record plus a tiny prototyping subset; the helper can save the Zenodo landing pages with SLAKH_CHECK_ZENODO=1, but archive file selection should be made from the live Zenodo records because the full corpus is large.
Safe-first helperscripts/download/slakh2100.sh
Speech recognition
SLUE
SLUE: Spoken Language Understanding Evaluation
Safe-first helper
Spoken Language Understanding
Automatic Speech Recognition
Named Entity Recognition
Named Entity Localization
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- SLUE-VoxPopuli is CC0; SLUE-VoxCeleb is CC BY 4.0; HF dataset metadata also advertises cc0-1.0 and cc-by-4.0 tags.
- Code license
- MIT
- License caution
- SLUE redistributes curated subsets of VoxPopuli and VoxCeleb plus task annotations. The VoxCeleb license notice says original and cropped video copyrights remain with the original owners.
- Download notes
- The helper downloads official toolkit docs and component license files by default. Hugging Face dataset snapshots are opt-in because they contain audio-derived benchmark data; use SLUE_DATASETS to choose slue, slue-phase-2, or both.
Safe-first helperscripts/download/slue.sh
Speech understanding & dialogue
SLURP
SLURP: A Spoken Language Understanding Resource Package
Safe-first helper
Spoken Language Understanding
Intent Classification
Slot Filling
Semantic Entity Labeling
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Textual annotations are CC BY 4.0; Zenodo-hosted audio is non-commercial (Zenodo license id other-nc; GitHub README states CC BY-NC 4.0).
- Code license
- not_specified
- License caution
- GitHub README separates textual data and audio licensing and says a less-strict audio license may be available by contacting the dataset owner. No standalone code license file was found in the repo via raw GitHub on 2026-07-09.
- Download notes
- The helper clones or updates the official annotation/code repository by default and downloads Zenodo LICENSE.txt. Audio archives are about 3.9 GiB real plus 2.8 GiB synthetic, so audio download is an explicit opt-in.
Safe-first helperscripts/download/slurp.sh
Speech recognition
SmartGlasses Challenge 2026
SLT 2026 SmartGlasses Challenge: Egocentric Speech Interaction on AI Glasses
Manual or gated
Time Stamped Speaker Attributed Asr
Meeting Transcription
Speaker Diarization
Spoken Language Understanding
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- access_restricted_terms_not_publicly_specified
- Code license
- not_specified
- License caution
- The challenge page does not publish a standalone dataset license and reserves organizer control over the participation terms. The public evaluation repository has no LICENSE file and GitHub reports no detected license, so its availability must not be treated as permission to redistribute the toolkit or corpus. Obtain permission from the organizers before reusing data or code beyond the challenge.
- Download notes
- The official challenge covers dyadic conversations and multi-party meetings recorded with a four-channel microphone array on smart glasses. Across train, development, and test, Track 1 reports 518 sessions and 44.95 hours, while Track 2 reports 196 sessions and 62.03 hours. Each track evaluates time-stamped speaker-attributed ASR with tcpWER and multiple-choice spoken-language understanding; public reference answers are limited to development data. The helper saves the official challenge page, public evaluation-toolkit documentation and metadata, and the July 2026 system paper. Corpus access required registration and agreement to challenge rules, download links were emailed to participating teams, registration closed in June 2026, and no current public corpus URL is provided.
Safe-first helperscripts/download/smartglasses_challenge_2026.sh
Enhancement, separation & quality
SOMOS
SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis
Safe-first helper
Mean Opinion Score Prediction
Synthetic Speech Naturalness Assessment
Text To Speech Quality Assessment
Cross Corpus Speech Quality Evaluation
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_applicable
- License caution
- Zenodo declares CC BY-NC-SA 4.0 and restricts the dataset to non-commercial research use with citation and same-terms redistribution. Synthetic audio was generated from a public-domain LJ Speech voice using multiple TTS models, while evaluation texts also draw from Blizzard Challenge, LJ Speech, Wikipedia, and general web sources; preserve provenance and review source-text terms for redistribution or derivative releases.
- Download notes
- The public, ungated release contains 20,000 synthetic utterances, 100 natural utterances, and 374,955 crowdsourced naturalness ratings. It provides carefully designed 70/15/15 train, validation, and test partitions with unseen systems, listeners, and texts. The helper saves the official project page, Zenodo metadata, and primary paper page by default; the approximately 3.70 GiB archive requires SOMOS_DOWNLOAD_ARCHIVE=1. A July 2026 cross-corpus MOS study uses all 20,100 SOMOS samples in its 19-dataset leave-one-dataset-out evaluation.
Safe-first helperscripts/download/somos.sh
Enhancement, separation & quality
Song Describer Dataset
The Song Describer Dataset: A Corpus of Audio Captions for Music-and-Language Evaluation
Safe-first helper
Music Captioning
Text To Music Generation Evaluation
Music Text Retrieval
Music Codec Reconstruction
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-4.0
- Code license
- MIT
- License caution
- Zenodo and the repository declare CC BY-SA 4.0 for the dataset and MIT for code. The audio originates from MTG-Jamendo and retains per-track Creative Commons licenses recorded in audio_licenses.txt; preserve attribution and apply each track's terms in addition to the dataset license.
- Download notes
- The public, ungated release contains 706 approximately two-minute MTG-Jamendo tracks with 1,106 crowdsourced English captions. Its human-validated evaluation subset contains 546 tracks and 746 captions. The helper downloads the official annotations, per-track audio-license list, metadata, dataset documentation, and Zenodo record by default; the approximately 3.09 GiB audio archive requires explicit opt-in. Qwen-Music section 4.2.2 evaluates codec reconstruction on all 546 validated tracks. Qwen-Audio-VAE sections 4.1-4.2 also use the dataset for music reconstruction evaluation.
Safe-first helperscripts/download/song_describer.sh
Enhancement, separation & quality
SongBench
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
Safe-first helper
Full Song Quality Assessment
Fine Grained Song Aesthetics Assessment
Text To Song Generation Evaluation
Music Reward Model Evaluation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- LICENSE.txt resembles MIT but adds an academic-only restriction and a complete prohibition on commercial or production use, so it is not standard MIT. The paper's arXiv license does not license the absent corpus. That corpus includes commercial-generator outputs and 1,000 copyrighted songs; source-song, composition, recording, generated- output, service, voice, lyric, and prompt rights remain relevant. The repository names MuQ as an MIT dependency, whose terms apply separately.
- Download notes
- The public repository releases a MuQ-based seven-output evaluator, configuration, approximately 96.1 MiB checkpoint, and 100 bilingual lyric/style test prompts. It scores vocal, instrumental, melody, structure, arrangement, mixing, and musicality quality. The helper downloads code, terms, configuration, prompts, and official metadata by default; the checkpoint requires explicit opt-in. The paper's core 11,717-song, 683.5-hour expert-rated corpus, 95:5 train/ID split, 352-song OOD test set, source audio, annotations, and exact manifests are not public. Qwen-Audio-3.0-Gen-Preview Table 7 uses the evaluator on an unspecified small set whose items and generated audio are likewise unreleased.
Safe-first helperscripts/download/songbench.sh
Enhancement, separation & quality
SongEval
SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
Safe-first helper
Song Aesthetics Assessment
Music Quality Prediction
Full Song Generation Evaluation
Human Preference Modeling
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC-SA 4.0 and the GitHub toolkit includes Apache-2.0. The paper says the audio includes outputs from five open and commercial song generators plus real and deliberately poor examples; generated-output, service, and any underlying music rights may still apply, and the release does not provide per-item provenance in metadata.jsonl. Review those rights before redistributing audio or relying on the card license alone.
- Download notes
- The public, ungated release contains 2,399 complete English and Chinese songs (about 140 hours) spanning nine mainstream genres. Sixteen musically trained annotators rated coherence, memorability, vocal breathing and phrasing naturalness, structural clarity, and overall musicality on five-point scales. The helper downloads official cards, API metadata, the approximately 1.27 MB rating JSONL, and toolkit documentation by default. The Hugging Face API reports approximately 16.1 GB of repository storage, so fetching all MP3 files requires explicit opt-in.
Safe-first helperscripts/download/songeval.sh
Audio understanding, generation & events
SongFormBench
Safe-first helper
Music Structure Analysis
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- cc-by-4.0
- License caution
- Dataset card lists CC BY 4.0. Audio reconstruction notes reference HarmonixSet and BigVGAN resources.
Safe-first helperscripts/download/songformbench.sh
Audiovisual & cross-modal
Sonic Seasoning
Sonic Seasoning: a Multi-Source Perceptual Dataset of Taste-Evoking Sounds
Safe-first helper
Taste From Audio Regression
Taste Conditioned Music Retrieval
Music Representation Evaluation
Crossmodal Audio Perception
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for the compilation, ratings, splits, and redistributed clips, and the repository code is Apache-2.0. Audio provenance varies: the release combines pre-existing music, MusicGen outputs, and stimuli from 13 prior studies. The card limits redistributed clips to non-commercial research and instructs users to cite the originating studies; review underlying music, publication-stimulus, performer, and generated-output rights in addition to the compilation license.
- Download notes
- The public, ungated Hugging Face release contains 377 uniformly encoded WAV clips with normalized sweet, bitter, salty, sour, and spicy ratings; some subsets also provide temperature and emotion annotations. Its fixed split column contains 269 train, 68 validation, and 40 test items. The paper evaluates ten frozen audio encoders with a shared multi-task regression protocol and uses a 309-item pool for taste-conditioned retrieval. The helper downloads the official card, repository docs, API metadata, and approximately 34 KB ratings/path Parquet file by default. The Hugging Face API reports approximately 797 MB of repository storage, so the audio snapshot and the approximately 642 KB code repository are separate opt-ins. The repository README still calls the training dataset private, which conflicts with the current public, ungated Hugging Face release; the helper follows the live dataset state.
Safe-first helperscripts/download/sonic_seasoning.sh
Audio understanding, generation & events
SONYC-UST-V2
SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context
Safe-first helper
Audio Tagging
Urban Sound Tagging
Multilabel Sound Classification
Spatiotemporal Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo v2.3 lists CC BY 4.0 and the record README says the SONYC-UST dataset is offered under Creative Commons Attribution 4.0 International. Challenge rules also restrict private external data for reproducible task submissions.
- Download notes
- The helper downloads Zenodo record JSON plus README, annotations, taxonomy, and unpack script by default. Audio is split across 19 archive shards totaling about 12.8 GiB, so audio download is an explicit opt-in.
Safe-first helperscripts/download/sonyc_ust_v2.sh
Audio understanding, generation & events
Soroll-IA
Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring
Safe-first helper
Multi Label Audio Tagging
Weakly Supervised Sound Event Classification
Industrial Sound Monitoring
Real World Environmental Audio Classification
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- Kaggle and the official benchmark README declare CC BY-NC 4.0, which prohibits commercial use and requires attribution. The benchmark repository has no LICENSE file or GitHub-detected license, so its code terms are unspecified.
- Download notes
- The public Kaggle release contains 7,396 FLAC clips (approximately 22 hours and 2.17 GB) recorded by two fixed sensing nodes in the Port of Valencia, with 26 weakly labeled industrial sound classes. It provides two ground-truth variants: a permissive non-cross-validated annotation set and a conservative set requiring agreement from at least two-thirds of annotators, plus five-fold assignments. The helper downloads official metadata, paper, and benchmark documentation by default; the audio and annotations require explicit opt-in and an authenticated Kaggle CLI.
Safe-first helperscripts/download/soroll_ia.sh
Speech generation
SoulX-Singer-Eval / GMO-SVS
SoulX-Singer-Eval: A Benchmark and Evaluation Suite for Zero-Shot Singing Voice Synthesis
Safe-first helper
Zero Shot Singing Voice Synthesis
Score Conditioned Singing Voice Synthesis
Singing Voice Editing
Cross Lingual Singing Voice Synthesis
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_upstream_and_consent_caveats
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC 4.0, while the evaluation code is Apache-2.0. GMO-SVS incorporates GTSinger, M4Singer, and Opencpop, whose corpus, performer, composition, recording, and redistribution terms still apply. The paper says recruited Mandarin singers consented to academic open-source use and English prompts were sliced from Mixing Secrets; those narrower provenance and consent conditions must not be broadened by the aggregate card license. Review upstream terms before redistribution, derivatives, voice cloning, or commercial use.
- Download notes
- The public, ungated release combines GMO-SVS, an 802-sample evaluation selection derived from the official Opencpop and M4Singer test splits plus GTSinger, with 100 prompt segments from 50 unseen Mandarin and English singers in SoulX-Singer-Eval. It releases word- and phone-level prompt and target manifests, source and prompt WAV files, and an Apache-2.0 evaluation suite for WER or CER, speaker similarity, F0 errors, MCD, SingMOS-Pro, and Sheet-SSQA. The helper downloads the card, API metadata, eight lightweight annotation JSONLs, and evaluation-suite documentation by default. Repository cloning is optional, and the approximately 888 MiB Hugging Face snapshot is a separate audio opt-in. SemBridge Section 4 uses this protocol to evaluate transfer from its pretrained speech-token interface into score-conditioned singing.
Safe-first helperscripts/download/soulx_singer_eval.sh
Audio understanding, generation & events
Spatial LibriSpeech
Spatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning
Safe-first helper
Sound Source Localization
Source Distance Estimation
Room Acoustics Estimation
Direct To Reverberant Ratio Estimation
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- Apple's copyrights in the dataset are CC BY 4.0. Its license preserves upstream terms for LibriSpeech, Microsoft DNS noise, AudioSet, and CC0 Freesound material and says Apple makes no representations about those upstream rights; review component provenance before redistribution.
- Download notes
- The public Apple release provides more than 650 hours of 16 kHz first-order ambisonic speech and optional distractor noise, synthesized from LibriSpeech across more than 200,000 acoustic conditions and 8,000 synthetic rooms. Labels cover 3D source position and speaking direction, room geometry, C50, DRR, EDT, T20, and T30. The helper downloads official documentation by default; the approximately 365 MiB metadata Parquet file and individual FLAC samples are explicit opt-ins. The README still describes raw 19-channel audio as forthcoming and directs users to contact Apple rather than exposing a public download.
Safe-first helperscripts/download/spatial_librispeech.sh
Speech recognition
Speak & Improve Corpus 2025
The Speak & Improve Corpus 2025: An L2 English Speech Corpus for Language Assessment and Feedback
Manual or gated
Spoken Language Assessment
L2 English Proficiency Scoring
Non Native Speech Recognition
Spoken Grammatical Error Correction
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_noncommercial_research_and_education_only
- Code license
- not_applicable
- License caution
- The official agreement grants a non-exclusive, non-transferable right for non-commercial research and education only. It prohibits sublicensing and sharing any part of the corpus, even with another licensee; publishing corpus-derived items or statistics requires prior Cambridge University Press & Assessment approval, apart from permitted excerpts under 100 words. The owner also instructs users not to submit the corpus to services that retain data for training.
- Download notes
- The official Cambridge release contains about 315 hours and more than 45,000 utterances from over 7,000 L2 English test submissions, with holistic CEFR-aligned proficiency scores and fixed training, development, and evaluation sets. About 55 hours additionally have manual transcripts, disfluencies, and language-error corrections. Access requires accepting the owner license and completing the registration form; the portal does not expose a reusable direct archive URL, so the helper prints the manual steps only. The July 2026 shortcut-reliance paper evaluates all four open-speaking task types on 39,490 training, 5,616 development, and 3,209 evaluation utterances.
Safe-first helperscripts/download/speak_and_improve_2025.sh
Speech understanding & dialogue
SPEARBench
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Safe-first helper
Streaming Speech To Speech Evaluation
Conversational Naturalness Evaluation
Response Latency Evaluation
Interruption Evaluation
+4 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_derived_from_seamless_interaction
- Code license
- MIT
- License caution
- The GitHub repository's MIT license covers the website and helper code. Neither the paper nor project page states a separate license for the downloadable benchmark audio, which is extracted from Seamless Interaction; public access does not establish redistribution or commercial-use rights, so verify the source corpus and package terms before reuse.
- Download notes
- The public project page links a browser-mediated SharePoint package containing 5,419 selected English question-answer dialogues from the Seamless Interaction development and test sets (37.33 hours including contexts, questions, and human answers). The helper downloads only official documentation, leaderboard metadata, an example submission CSV, and inference instructions; obtain the audio package manually from the project page and keep its directory structure intact. The live repository publishes inference wrappers, aggregate and per-subset metric CSVs, reports, and plots for evaluated systems, but not the complete evaluation pipeline claimed in the paper.
Safe-first helperscripts/download/spearbench.sh
Speech understanding & dialogue
Speech Commands
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Safe-first helper
Keyword Spotting
Limited Vocabulary Speech Recognition
Audio Classification
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Google Research blog and Hugging Face dataset card list Creative Commons BY 4.0; HF card also asks users not to try to identify speakers.
- Download notes
- TensorFlow Datasets reports v0.02 as about 2.37 GiB download / 8.17 GiB extracted; avoid accidental full downloads in automated checks.
Safe-first helperscripts/download/speech_commands.sh
Speaker, identity & emotion
Speech DF Arena
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
Safe-first helper
Audio Deepfake Detection
Synthetic Speech Detection
Cross Dataset Generalization
Detector Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- component_specific_with_sonar_terms_unresolved
- Code license
- not_specified
- License caution
- The toolkit repository has no LICENSE file and GitHub reports no detected license. Its protocol CSVs expose absolute path strings from the authors' evaluation setup and do not license the underlying audio. Each component corpus retains its own terms; SONAR's terms remain unspecified. Treat public source access as distinct from permission to redistribute data, code, checkpoints, or score files.
- Download notes
- Speech DF Arena is one cross-dataset evaluation protocol and leaderboard, not a new redistribution of its component corpora. Phase I evaluates detectors on 14 public evaluation subsets spanning In the Wild, ASVspoof 2019/2021/2024, Fake or Real, CodecFake, ADD 2022/2023, DFADD, LibriSeVoc, and SONAR. The toolkit standardizes CSV protocols with absolute file paths and spoof/bonafide labels, 16 kHz input, EER, pooled EER, accuracy, and F1. Its public repository contains model adapters and protocol CSVs but no component audio. The helper downloads only the paper page, repository documentation/API metadata, and leaderboard metadata by default. Cloning the approximately 400 MB toolkit is an explicit opt-in; acquire every component dataset from its own official source.
Safe-first helperscripts/download/speech_df_arena.sh
Speech recognition
Speech-MASSIVE
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
Safe-first helper
Multilingual Spoken Language Understanding
Intent Classification
Slot Filling
Zero Shot And Few Shot Transfer
+4 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The owner repository and dataset card explicitly declare the audio dataset CC BY-NC-SA 4.0, which requires attribution, restricts commercial use, and applies share-alike to adaptations. Apache-2.0 covers the released first-party code, not the audio or inherited MASSIVE annotations under a different grant. Confirm the upstream MASSIVE terms and Hugging Face gating conditions before redistribution or derived releases.
- Download notes
- Speech-MASSIVE records native-speaker readings of a portion of the textual MASSIVE corpus in Arabic, German, Spanish, French, Hungarian, Korean, Dutch, Polish, European Portuguese, Russian, Turkish, and Vietnamese. Each language has 2,033 validation and 2,974 test items; French and German also have 11,514-item full training splits, while every language has a 115-item few-shot selection. Labels retain 18 domains, 60 intents, and 55 slots. The paper establishes cascaded and end-to-end SLU protocols under zero-shot and few-shot settings, plus ASR, language-identification, and speech-translation baselines. The owner repository releases training/evaluation code and supplementary metrics. The helper saves lightweight cards, current Hub metadata, and the code license by default; the roughly 23.7 GB main snapshot, 35.8 GB auto-gated test snapshot, and code clone require explicit opt-in.
Safe-first helperscripts/download/speech_massive.sh
Speech generation
SpeechEditBench
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
Safe-first helper
Instruction Guided Speech Editing
Content Editing
Speaker Editing
Emotion Editing
+5 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_with_upstream_terms
- Code license
- Apache-2.0
- License caution
- The repository and Hugging Face card release contributor-authored code, documentation, metadata, and benchmark assets under Apache-2.0, while requiring compliance with source-corpus terms. The paper's appendix identifies mixed upstream conditions: CC BY sources, Apache-2.0 sources, CC BY-NC and CC BY-NC-SA sources, StoryTTS research-only restrictions, MagicData-RAMC custom terms, and the IEMOCAP access agreement. Apply those component restrictions to affected rows and audio rather than treating the aggregate label as overriding them.
- Download notes
- The public, ungated v1.1 release contains 4,700 English and Chinese source-instruction pairs and 5,400 audio files across seven atomic editing tasks and a compositional split. Evaluation separately measures target success, lexical-content preservation, and joint success. The helper downloads official documentation, release metadata, and the eight sample JSONL files by default. The Hugging Face API reports approximately 3.75 GB of repository storage, so audio requires SPEECH_EDIT_BENCH_DOWNLOAD_HF=1; cloning the roughly 3.2 MB evaluation repository is a separate opt-in.
Safe-first helperscripts/download/speech_edit_bench.sh
Speech understanding & dialogue
SpeechEQ
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models
Safe-first helper
Spoken Dialogue Emotional Intelligence
Paralinguistic Reasoning
Multi Turn Speech Understanding
Acoustic Multiple Choice
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The Hugging Face card has no license field, and neither the dataset repository nor the official code repository exposes a license file. The paper itself is CC BY-NC-SA 4.0, but that publication license must not be assumed to license the released benchmark audio, annotations, or code. Obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated English release contains 2,265 six-turn dialogues totaling 42 hours 23 minutes across 15 EQ-i 2.0 subscales. Evaluation selects between high- and low-EQ acoustic renditions of identical text at turns four and six, testing pitch, energy, rate, pauses, and sustained conversational context. The helper downloads official documentation and repository metadata by default. The five Parquet shards contain embedded audio and total approximately 2.45 GB, so the full Hugging Face snapshot requires SPEECHEQ_DOWNLOAD_HF=1.
Safe-first helperscripts/download/speecheq.sh
Audio understanding, generation & events
SpeechJBB
SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech
Safe-first helper
Spoken Jailbreak Evaluation
Multilingual Speech Safety
Code Switched Speech Robustness
Pseudo Word Obfuscation Robustness
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_no_aggregate_terms_published
- Code license
- no_license_file_detected
- License caution
- The article is CC BY 4.0 and Appendix E reports MIT for JailbreakBench, CC BY-SA 4.0 for MGSM and FLEURS-SLU SIB, and CC BY 4.0 for FLEURS. Those source licenses do not establish terms for the translated and code-switched prompts, pseudo-word annotations, or XTTS-rendered audio. The Hub card says only "other," and the GitHub repository has no license file. Confirm artifact rights with the owners before reuse or redistribution.
- Download notes
- The public Hub contains 11,618 WAV files and reports approximately 6.82 GB of storage. The core safety benchmark has 7,500 files: 1,500 clean benign, 1,500 clean harmful, and 1,500 each at 10%, 30%, and 50% pseudo-word insertion. It covers five monolingual and ten code-switched settings derived from 100 harmful and 100 benign JailbreakBench prompts. The remaining 4,118 WAV files are Speech-MGSM, FLEURS ASR, and FLEURS-SLU comprehension controls. The helper downloads official cards and repository metadata by default; the full audio snapshot and the roughly 44 MB GitHub repository are separate opt-ins. The Hub publishes audio only, without a structured prompt/label manifest; its card also contains stale namespace and split examples, so use the eight declared test configurations in McGill-NLP/SpeechJBB.
Safe-first helperscripts/download/speechjbb.sh
Speech generation
SpeechParaling-Bench
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
Safe-first helper
Paralinguistic Speech Generation
Spoken Instruction Following
Fine Grained Voice Control
Intra Utterance Paralinguistic Variation
+3 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_declared_by_dataset_card
- Code license
- no_detected_license
- License caution
- The Hugging Face card declares Apache-2.0 for the dataset. The GitHub repository has no detected license, so the dataset declaration should not be assumed to license its software, third-party model generations, or judge records. The arXiv and project-site publication licenses do not resolve those artifact rights.
- Download notes
- The release contains 1,001 aligned Chinese-English queries, yielding 2,002 input WAVs across 691 paralanguage-control, 120 dynamic-variation, and 190 situational-adaptation cases per language. It covers 101 features in 13 dimensions. The helper saves owner documentation and live repository/Hub metadata by default; the two compact metadata manifests, approximately 1.66 GB Hub snapshot, and code/result clone are separate opt-ins.
Safe-first helperscripts/download/speechparaling_bench.sh
Speech understanding & dialogue
SpeechRole
SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
Safe-first helper
Speech Role Playing
Speech Dialogue Evaluation
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- not_specified
- License caution
- HF cards list MIT; GitHub repo has no separate detected license.
Safe-first helperscripts/download/speechrole.sh
Audiovisual & cross-modal
SpEmoC
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
Manual or gated
Speech Emotion Recognition
Audio Visual Emotion Recognition
Multimodal Emotion Recognition
Cross Dataset Emotion Generalization
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_academic_eula
- Code license
- not_specified
- License caution
- The official agreement limits use to academic research, education, scientific publication, and other non-commercial research; prohibits redistribution and sharing download links; and leaves copyright in the source movie and television clips with their respective owners. The public benchmark repository has no LICENSE file or detected GitHub license, so code and public split/metadata terms remain unspecified.
- Download notes
- The benchmark contains 30,000 refined clips curated from 306,544 raw speaking segments across 3,100 English-language movies and television series, with aligned audio, visual, and text modalities and a near-balanced seven-emotion label distribution. The official project says the full dataset is available only after a requestor and faculty advisor or principal investigator sign the access agreement and submit it from an institutional email address. The helper saves public project, paper, repository, and agreement metadata, then prints the manual application steps; it never downloads restricted media.
Safe-first helperscripts/download/spemoc.sh
Speech recognition
SPGISpeech
SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
Manual or gated
Asr
Fully Formatted Transcription
Financial Speech Recognition
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- gated_academic_research_internal_use
- Code license
- not_specified
- License caution
- HF terms say the content is for academic research purposes and internal use only, prohibit redistribution without prior written consent, and include additional restrictions on creating competing databases/products and identifying individuals.
- Download notes
- Hugging Face access requires logging in and accepting Kensho terms. The HF card lists split sizes from 11 GiB for dev/test to 530 GiB for the L training subset; the helper refuses to download until SPGISPEECH_ACK_TERMS=1 is set.
Safe-first helperscripts/download/spgispeech.sh
Enhancement, separation & quality
SpInt
SpInt: A Spanish Speech Intelligibility Dataset
Safe-first helper
Speech Intelligibility Assessment
Objective Intelligibility Metric Evaluation
Speech Enhancement Evaluation
Spanish Speech
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_for_released_SpInt_artifacts
- Code license
- Not specified in the source record.
- License caution
- Zenodo declares CC BY 4.0 for the released SpInt package. The record explicitly withholds the original clean Spanish Matrix Test recordings to respect their licensing conditions, so the release is not a standalone audio corpus and CC BY 4.0 must not be extended to those absent recordings.
- Download notes
- The public Zenodo v1.0 release provides behavioral intelligibility labels for 5,148 processed Spanish utterances, per-stimulus and listener-response metadata, complex speech-enhancement masks, noise signals, and a reconstruction script. The clean Spanish Matrix Test recordings are deliberately excluded because of their separate license; users must obtain that corpus independently to reconstruct the stimuli. The helper downloads the official record, README, reconstruction script, and approximately 2.7 MB JSON metadata by default. The approximately 807 MiB noise and 3.08 GiB mask archives require explicit opt-in.
Safe-first helperscripts/download/spint.sh
Speaker, identity & emotion
SpoofCeleb
SpoofCeleb: Speech Deepfake Detection and SASV in the Wild
Manual or gated
Speech Deepfake Detection
Synthetic Speech Detection
Spoofing Robust Speaker Verification
In The Wild Anti Spoofing
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_source_media_rights
- Code license
- not_applicable
- License caution
- The official project and Hugging Face tag state CC BY 4.0, while the project explicitly says copyright in the human speech files remains with the original video owners. Access is granted only after a request and agreement to Hugging Face terms. Treat the Creative Commons label as covering the released compilation and author contributions, not as clearance of every underlying video, voice, likeness, privacy, or generated-speech right.
- Download notes
- The gated author-owned Hugging Face release contains more than 2.5 million bona fide and synthetic utterances from 1,251 VoxCeleb1 speakers. Its 23 TTS attacks and speaker-disjoint train, validation, and evaluation partitions support both speech deepfake detection and spoofing-robust speaker verification protocols. The helper downloads the official project page, paper pages, and Hugging Face API metadata by default. The API reports approximately 268.3 GB of repository storage, so the snapshot requires author approval, local Hugging Face authentication, explicit acceptance of the terms, and both SPOOFCELEB_ACK_TERMS=1 and SPOOFCELEB_DOWNLOAD_HF=1. Section 4.1 of arXiv:2607.21127 evaluates balanced TTS attacks from all SpoofCeleb splits but does not publish its clipped row selection.
Safe-first helperscripts/download/spoofceleb.sh
Audio understanding, generation & events
SpurAudio
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification
Safe-first helper
Few Shot Audio Classification
Environmental Sound Classification
Shortcut Learning Evaluation
Background Shift Robustness
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_mixed_upstream_terms
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY 4.0 for the released mixtures, but the benchmark derives audio from five upstream datasets with separate terms. In particular, ESC-50 and UrbanSound8K include non-commercial restrictions; review all component licenses before redistribution or commercial use. The evaluation repository is MIT.
- Download notes
- The public, ungated Hugging Face release contains 16,381 WAV files in train, validation, and test splits and reports approximately 7.69 GB of repository storage. It mixes foreground events from ESC-50, UrbanSound8K, VocalSound, WILD DESED, and USM with unrelated background textures to measure IID-versus-OOD shortcut reliance in 1-shot and 5-shot classification. The helper downloads official documentation and repository metadata by default; the audio snapshot is opt-in.
Safe-first helperscripts/download/spuraudio.sh
Speech recognition
ST-CMDS
ST-CMDS-20170001_1: Free ST Chinese Mandarin Corpus
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR38 lists Creative Commons BY-NC-ND 4.0 and asks users to cite the data as "ST-CMDS-20170001_1, Free ST Chinese Mandarin Corpus." Re-check upstream terms before redistribution or commercial use.
- Download notes
- OpenSLR hosts an 8.2 GiB archive with cellphone-recorded Mandarin speech, transcriptions, and metadata from 855 speakers and 102,600 utterances. The helper saves the OpenSLR page by default and requires ST_CMDS_DOWNLOAD_ARCHIVE=1 for the large archive.
Safe-first helperscripts/download/st_cmds.sh
Speech understanding & dialogue
StanceBench
StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
Safe-first helper
Interpersonal Stance Evaluation
Audio Llm Judge Evaluation
Conversational Speech Understanding
Paralinguistic Understanding
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_for_bundled_benchmark_metadata
- Code license
- MIT
- License caution
- The repository currently contains an MIT LICENSE covering its software and bundled materials, although its README still says the project license is pending. Seamless Interaction audio remains under Meta's CC BY-NC 4.0 terms and is not redistributed. Treat the current LICENSE file as authoritative for the repository while preserving the upstream noncommercial restriction for source audio.
- Download notes
- StanceBench is a distinct evaluation protocol derived from the Improvised subset of Seamless Interaction, not a second release of its audio. The repository provides nine stance definitions, role/category mappings, a 9,817-row conversation mapping, clip-construction and model evaluation code, and analysis tooling. S0-S5 evaluate single-speaker clips; S6-S8 evaluate a target response with partner context. Safe defaults fetch the paper, repository documentation, current license, questions, and role mappings. The approximately 8.8 MB interaction mapping is opt-in, and source audio must be obtained separately from Meta's official release.
Safe-first helperscripts/download/stancebench.sh
Audio understanding, generation & events
STAR-Bench
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
Safe-first helper
Foundational Acoustic Perception
Pitch Perception
Loudness Perception
Duration Perception
+6 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- conflicting_MIT_and_Apache-2.0_signals
- License caution
- The Hugging Face card declares the dataset CC BY-NC 4.0 and labels use as research-only. Preserve the terms and provenance of Clotho, FSD50K, STARSS23, and internet-sourced audio. The repository's LICENSE file is MIT, but its README badge says Apache-2.0 and describes both data and code as research-only; clarify the intended software terms before redistribution or commercial use.
- Download notes
- The public, ungated v1.0 release contains 2,353 English multiple-choice questions: 951 for foundational perception, 900 for temporal reasoning, and 502 for spatial reasoning. Its current metadata revises the v0.5 questions reported in the paper. Foundational audio is synthesized; temporal tasks draw on Clotho and FSD50K, while spatial tasks use STARSS23 and in-the-wild audio. The helper downloads official documentation, repository metadata, and the approximately 2 MB of JSON question metadata by default. The 2.74 GB audio archive and evaluation repository are separate opt-ins.
Safe-first helperscripts/download/star_bench.sh
Speech understanding & dialogue
Step-Audio-R1.5 Benchmarks
Step-Audio-R1.5 Open Paralinguistic Evaluation Benchmarks
Safe-first helper
Audio Captioning
Paralinguistic Speech Understanding
Spoken Question Answering
Speech Dialogue
+1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_repository_declaration_with_upstream_media_voice_and_privacy_caveats
- Code license
- Apache-2.0
- License caution
- The root repository declares Apache 2.0. Step-Caption includes curated YouTube and Bilibili clips, while Step-DU and Step-SPQA contain recorded or synthesized voices and sensitive inferred traits such as age and gender. Review source-platform terms, media rights, consent, privacy, biometric, and publicity restrictions separately; do not assume the repository license clears third-party content or voice rights.
- Download notes
- The official package releases 905 Step-Caption, 87 Step-DU, and 550 Step-SPQA audio examples with JSONL metadata and prompts. The helper fetches pinned documentation, annotations, prompts, and repository-tree metadata by default. Set STEP_AUDIO_R1_5_DOWNLOAD_REPO=1 to clone the complete approximately 1.06 GB object snapshot, including 1,542 audio files.
Safe-first helperscripts/download/step_audio_r1_5_benchmarks.sh
Audiovisual & cross-modal
StoryAD-QA
StoryAD-QA: Narrative-Comprehension Evaluation for Long-Form Audio Description
Safe-first helper
Long Form Audio Description Evaluation
Narrative Comprehension
Multiple Choice Question Answering
Context Conditioned Reasoning
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- not_specified_pending_finalization
- License caution
- The repository's LICENSE is a placeholder that says release terms will be finalized and merely recommends MIT or Apache-2.0 for code and CC BY-NC 4.0 or another author-approved license for annotations. Those recommendations are not grants. The release contains no movie video, audio, frames, subtitles, or scripts; clip identifiers derive from ten MAD-Eval movies in LSMDC, and users must obtain lawful access to the underlying copyrighted media separately.
- Download notes
- The official ECCV 2026 repository releases 2,572 manually verified, five-option question-answer pairs across two tracks: 1,609 segment-only questions over 30-, 60-, 120-, and 240-second windows, and 963 context-conditioned questions using 30, 60, or 90 seconds of preceding context plus a 30-second target clip. The repository includes full annotation CSVs, question-only files, answer keys with rationales, prompts, and a local accuracy scorer. The paper reports 2,574 retained questions, while the repository README says its public release contains 2,572 after validation and cleanup; this entry uses the released-file total. The helper saves official documentation, license notice, summary, evaluator, and repository metadata by default; cloning the approximately 8 MB repository is opt-in.
Safe-first helperscripts/download/storyad_qa.sh
Audiovisual & cross-modal
StreamArena
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
Manual or gated
Continuous Audio Visual Understanding
Real Time Audio Visual Question Answering
Long Horizon Multimodal Memory
Proactive Interaction
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- conflicting_CC-BY-4.0_and_CC-BY-NC-4.0_with_upstream_video_rights
- Code license
- announced_Apache-2.0_but_repository_unavailable
- License caution
- Appendix C.7 says the annotations are CC BY 4.0, while the live Hugging Face card declares CC BY-NC 4.0; apply the stricter non-commercial terms pending clarification. The paper says original YouTube videos are not redistributed, but the dataset card says it ships about 300 GB of MP4, metadata, and subtitles under a research fair-use rationale. Neither Creative Commons declaration establishes rights to third-party source videos, music, speech, subtitles, or archived copies. Review per-video rights and platform terms before downloading, use, or redistribution.
- Download notes
- The public, ungated release contains 243 hour-scale videos averaging 88.8 minutes and 3,646 bilingual open-ended tasks. Its 263 real-time perception, 877 historical-retrospection, 1,732 external-tool, and 774 proactive-interaction questions carry causal query and evidence timestamps; the paper evaluates raw-audio and ASR-assisted omni-modal systems and separately measures factual accuracy and proactive timing. The helper downloads the project page, live Hub metadata, and four lightweight annotation/metadata JSONL files by default. The dataset card reports roughly 300 GB of 243 per-video tar files, so media require STREAMARENA_DOWNLOAD_VIDEOS=1. The paper promises an Apache-2.0 code release, but its linked GitHub repository returned HTTP 404 on 2026-08-08; the helper therefore does not guess an alternate code URL.
Safe-first helperscripts/download/streamarena.sh
Speech generation
StyleSet
StyleSet: A Spoken Language Benchmark for Speaking-Style-Related Speech Generation
Safe-first helper
Voice Style Instruction Following
Fine Grained Prosody Control
Intra Utterance Style Control
Non Verbal Vocalization Control
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_declared_by_dataset_card_with_IEMOCAP_and_generated_audio_caveats
- Code license
- not_applicable_no_code_release_linked
- License caution
- The dataset card says NTU Speech Lab releases StyleSet under MIT, and the paper says the authors obtained IEMOCAP-author consent to redistribute the adapted role-playing contexts. The release excludes original IEMOCAP recordings, whose access and use terms remain separate. Prompt speech was synthesized with GPT-4o-audio, so the aggregate declaration should not be assumed to override applicable service-output, voice, privacy, or source-dialogue rights. No terms can be inferred for absent model generations, judge records, or human ratings.
- Download notes
- The owner release contains two 20-item test configurations. Voice style instruction following pairs a base repetition prompt with a fine-grained instruction spanning emotion, volume, pace, emphasis, pitch, within-utterance change, and non-verbal cues. Role playing supplies paired audio prompts and contexts adapted from IEMOCAP for two sides of a dialogue; it does not redistribute the original IEMOCAP recordings. The paper generates a 20-turn dialogue, inserts two-second inter-turn silences, crops at one minute, and evaluates style on a five-point scale plus realism as a binary decision. The voice track uses a separate five-point text-and-style-adherence rubric. Audio-language-model judges sample five verdicts per item and are compared with four human evaluators using Pearson correlation. The live card declares 148,908,359 compressed bytes for role playing and 35,109,899 bytes for voice instruction following. The helper saves the card and live release metadata by default; either released test configuration is an explicit opt-in.
Safe-first helperscripts/download/styleset.sh
Speech recognition
SUPERB
SUPERB: Speech Processing Universal PERformance Benchmark
Safe-first helper
Speech Representation Evaluation
Phoneme Recognition
Automatic Speech Recognition
Keyword Spotting
+11 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- Apache-2.0
- License caution
- S3PRL is mostly Apache-2.0, with repository notes indicating some Facebook-authored files are CC BY-NC. SUPERB tasks use multiple external datasets, so each component corpus must be downloaded and licensed through its own official source.
- Download notes
- SUPERB is a benchmark suite over multiple upstream corpora. The helper downloads official documentation/license files by default and only clones the S3PRL toolkit with SUPERB_CLONE_TOOLKIT=1; underlying corpora such as LibriSpeech, Speech Commands, VoxCeleb, and IEMOCAP keep their own access paths and licenses.
Safe-first helperscripts/download/superb.sh
Representation & general suites
Surge Pitch Dataset
Pitch Audio Dataset (Surge synthesizer)
Safe-first helper
Musical Pitch Classification
Musical Pitch Ranking
Synthesizer Preset Classification
Audio Representation Evaluation
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- Zenodo declares CC BY 4.0 for the released dataset. The paper's publication license and the Surge synthesizer and preset licenses are separate; retain the dataset citation and review synthesizer/preset terms if regenerating or redistributing modified renders.
- Download notes
- The public, ungated Zenodo release contains 3.4 hours of four-second sounds generated from 2,084 human-authored Surge presets. Each preset is rendered at MIDI pitches 21 through 108 with velocity 64, a three-second note-on duration, and RMS normalization. The helper saves official record metadata and the paper by default; the approximately 7.58 GB tar archive requires explicit opt-in. NABEATs section 4.1 uses this release for downstream pitch classification under clean and constructed noisy conditions.
Safe-first helperscripts/download/surge_pitch.sh
Audiovisual & cross-modal
SVHalluc
SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
Safe-first helper
Speech Vision Semantic Hallucination
Speech Vision Temporal Hallucination
Global Semantic Alignment
Fine Grained Semantic Alignment
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_research_purposes_notice
- Code license
- no_standard_open_source_license
- License caution
- The repository and dataset card deliberately use a custom notice, not a standard open license. They state that the release is for research purposes, transfer no ownership of third-party video, and require users to follow SVHalluc, YouCook2, YouTube, and source-video terms. The article's arXiv license covers only the paper. Obtain clarification before redistribution, modification, or commercial use.
- Download notes
- The ungated test release contains 2,405 balanced multiple-choice rows referencing 872 MP4 files: 1,422 semantic items over GSA, FGSA, and CSMB, and 983 temporal items over TA, TF, and CMTB. The manifest is approximately 771 KB and all referenced paths are present in the approximately 6.26 GB Hub snapshot. The helper downloads primary documentation, terms, live repository metadata, the dataset card, and the complete question-answer manifest by default. Media require SVHALLUC_DOWNLOAD_HF=1; cloning the small owner repository is a separate opt-in. The release documents option-letter accuracy but supplies no model adapters or general response parser.
Safe-first helperscripts/download/svhalluc.sh
Speech recognition
Switchboard
Switchboard-1 Release 2 conversational telephone speech corpus
Manual or gated
Automatic Speech Recognition
Conversational Speech Recognition
Telephone Speech Recognition
Speaker Recognition
Access pathLDC / licensed
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_ldc_license
- Code license
- not_applicable
- License caution
- LDC catalog pages list membership/licensing terms and web-download access. Re-check the current LDC agreement before use or redistribution; do not treat benchmark recipes or transcripts as granting rights to the underlying audio.
- Download notes
- Switchboard-1 Release 2 and the 2000 HUB5 English Evaluation Speech set are distributed by LDC after login/licensing. The helper only prints official access steps because the corpus and standard evaluation audio are not publicly script-downloadable.
Safe-first helperscripts/download/switchboard.sh
Audiovisual & cross-modal
SyncBench
SyncBench: Causal-Semantic Audio-Visual Synchronization Evaluation for Generative Models
Safe-first helper
Audio Visual Synchronization Evaluation
Audio Visual Generation Evaluation
Video To Audio Evaluation
Causal Semantic Alignment
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The official code repository is MIT, but the Hugging Face dataset has no card, license tag, or license file. The paper is CC BY 4.0, which does not license the released generated videos or prompts. Model-output provider terms and any rights in prompt/source content may also apply; obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated Hugging Face repository contains 1,185 generated MP4s across six model directories plus two small evaluator-score JSON files, and reports approximately 12.9 GB of repository storage. The paper's section 4.4 defines SyncBench as 185 curated prompts across five audio-visual domains, while the current release has up to 200 numbered clips per model and does not include a dataset card or prompt manifest. The helper downloads official documentation, repository metadata, Hugging Face metadata, and the two lightweight score files by default; all videos require SYNCBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/syncbench.sh
Speaker, identity & emotion
SynSFX
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
Manual or gated
Non Speech Audio Deepfake Detection
Synthetic Sound Effect Detection
Unseen Generator Robustness
Cross Domain Audio Forensics
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- academic_research_only
- Code license
- not_released
- License caution
- The official release page labels SynSFX "Academic research only" but does not publish a full standalone dataset license or redistribution terms. The authentic partition incorporates AudioCaps, Clotho, ESC-50, TACoS, and WavCaps material, so their source-media and per-clip terms remain applicable. The arXiv paper uses arXiv's perpetual non-exclusive publication license, which does not license the dataset. No official evaluation-code release was linked when checked.
- Download notes
- The official release page provides a direct private-storage download route for the academic-research-only corpus. The paper reports 43,374 clips totaling 178 hours: 16,922 authentic clips from AudioCaps, Clotho, ESC-50, TACoS, and WavCaps, plus 26,452 clips synthesized by seven text-to-audio systems. It also defines a 1,890-prompt controlled subset shared across all seven generators and train, validation, in-domain test, and unseen-generator test protocols. The official page rounds the duration to approximately 180 hours and inconsistently lists 26,460 synthetic clips, despite retaining the 43,374 total; this entry uses the internally consistent paper counts. The helper saves the official page and paper by default. Because the uncompressed-WAV archive is large and the publisher states research-only access, downloading it requires both SYNSFX_ACK_RESEARCH_ONLY=1 and SYNSFX_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/synsfx.sh
Audio understanding, generation & events
Synth-DoPaCo
Synth-DoPaCo: Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
Safe-first helper
Audio To Soap Note Generation
Long Form Audio Summarization
Clinical Conversation Understanding
Speech To Text Generation
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY 4.0 for the released train/dev dataset, and the official metrics repository is Apache-2.0. The audio and conversations are synthetic, but the hidden challenge test set includes acted and realistic recordings that are not covered by the public dataset release or its license statement. Do not infer access or redistribution rights for those withheld recordings.
- Download notes
- The public, ungated WebDataset release contains 7,200 training and 400 development English doctor-patient dialogues with 16 kHz mono Opus audio, transcripts, generation metadata, and reference SOAP notes. The dataset paper describes the complete 8,800-conversation, approximately 1,329-hour corpus; the public challenge repository currently exposes only the 7,600 train/dev examples. The helper downloads official pages, cards, API metadata, rules, and metrics repository metadata by default. The Hugging Face API reports about 14.6 GB of repository storage, so the full snapshot is opt-in. BeTraC's 875-item blind test set is withheld and cannot be fetched by the helper.
Safe-first helperscripts/download/synth_dopaco.sh
Speech recognition
Tadabur
Tadabur: A Large-Scale Quran Audio Dataset
Safe-first helper
Quranic Speech Recognition
Arabic Speech Recognition
Reciter Identification
Word Level Alignment
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_source_and_cultural_caveats
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0, limits use to research and education, and adds ethical and cultural expectations for respectful Qur'anic use. Audio was collected from public Qur'anic repositories and archives, but the release does not provide per-recording source-license provenance; confirm source rights before redistribution. The linked repository has no standalone LICENSE file, so no code license is claimed.
- Download notes
- The public, ungated Hugging Face release contains more than 365,000 verse-level Arabic recitation examples totaling over 1,400 hours from more than 600 reciters, with simple and Uthmani text plus automatically derived word timestamps. It exposes one training split rather than a fixed held-out evaluation split. The helper downloads official documentation and repository metadata by default; the Hugging Face API reports approximately 1.94 TB of repository storage, so the audio-bearing snapshot requires TADABUR_DOWNLOAD_HF=1.
Safe-first helperscripts/download/tadabur.sh
Audio understanding, generation & events
TAU Spatial Sound Events 2019
TAU Spatial Sound Events 2019: Ambisonic and Microphone Array
Safe-first helper
Sound Event Localization And Detection
Sound Source Localization
Sound Event Detection
Spatial Audio Understanding
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_TAU_noncommercial
- Code license
- custom_TAU_noncommercial
- License caution
- Both Zenodo records use a custom Tampere University license permitting experimental non-commercial use with attribution and prohibiting commercial use. The source events derive from the DCASE 2016 Task 2 isolated-event dataset, so preserve that provenance. The baseline repository applies closely matching custom experimental/non-commercial terms to its code.
- Download notes
- The public DCASE 2019 Task 3 release has 400 development and 100 evaluation scenes of one minute each at 48 kHz, in matching 4-channel first-order Ambisonic and tetrahedral-microphone formats. Scenes use stationary sources from 11 classes, real impulse responses measured at 504 azimuth-elevation-distance combinations across five indoor locations, natural ambient noise, and zero or up to two overlapping events. Version 2 includes temporal and azimuth/elevation labels for both development and evaluation audio. The helper downloads official pages, papers, record metadata, READMEs, and license files by default; the approximately 10.1 GB audio release remains on Zenodo, while the roughly 490 KB label archives are an explicit opt-in.
Safe-first helperscripts/download/tau_spatial_sound_events_2019.sh
Audio understanding, generation & events
TAU Urban Acoustic Scenes 2019
TAU Urban Acoustic Scenes 2019 Development dataset
Safe-first helper
Acoustic Scene Classification
Multi Device Audio Classification
Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_non_commercial
- Code license
- not_applicable
- License caution
- Zenodo record 2589280 lists license id "other-nc" without a more specific SPDX-style license. DCASE challenge and dataset terms should be checked before redistribution or commercial use.
- Download notes
- The helper downloads the Zenodo record JSON plus small doc/meta ZIPs by default. The 40-hour audio release is split across 21 ZIP files of roughly 1.3-1.8 GiB each, so audio download is an explicit opt-in.
Safe-first helperscripts/download/tau_asc_2019.sh
Audio understanding, generation & events
TAU Urban Acoustic Scenes 2020 Mobile
TAU Urban Acoustic Scenes 2020 Mobile: Development and Evaluation datasets
Safe-first helper
Acoustic Scene Classification
Device Robust Audio Classification
Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_non_commercial
- Code license
- not_applicable
- License caution
- Both Zenodo development and evaluation records list "Other (Non-Commercial)" without a more specific SPDX-style license. DCASE challenge terms should be checked before redistribution or commercial use.
- Download notes
- The helper downloads Zenodo record JSON plus small doc/meta ZIPs by default. Development audio is split across 16 ZIP files totaling about 27.4 GiB; evaluation audio is split across 8 ZIP files totaling about 13.1 GiB, so both are explicit opt-ins.
Safe-first helperscripts/download/tau_asc_2020_mobile.sh
Audio understanding, generation & events
TAU Urban Acoustic Scenes 2022 Mobile
TAU Urban Acoustic Scenes 2022 Mobile Development and 2025 Evaluation datasets
Safe-first helper
Acoustic Scene Classification
Device Robust Audio Classification
Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_non_commercial
- Code license
- not_applicable
- License caution
- Both official Zenodo records list "Other (Non-Commercial)" without a more specific SPDX-style license. Review the record and DCASE task terms before redistribution or commercial use.
- Download notes
- The helper downloads both Zenodo record JSON files plus the small doc/meta ZIPs by default. The 64-hour development release contains about 25.6 GiB across 16 audio ZIPs; the DCASE 2025 evaluation release contains about 19.2 GiB across 12 audio ZIPs, so each audio collection is an explicit opt-in. DCASE 2025 Task 1 reuses a restricted subset of the 2022 development data with a new split and adds the 2025 evaluation release.
Safe-first helperscripts/download/tau_asc_2022_mobile.sh
Audio understanding, generation & events
TAU-NIGENS Spatial Sound Events 2020
Safe-first helper
Sound Event Localization And Detection
Sound Source Localization
Acoustic Source Tracking
Sound Event Detection
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- custom_TAU_License
- License caution
- Zenodo lists the dataset as CC BY-NC 4.0. Source events come from the separately released NIGENS database, so preserve its provenance and review its active terms. The baseline README places most code under a custom TAU License and only its metrics folder under MIT.
- Download notes
- The public v1.2 release is the complete DCASE 2020 Task 3 development and evaluation corpus: 600 development and 200 evaluation sound scenes of one minute each, in 4-channel FOA and tetrahedral-microphone formats at 24 kHz. It uses real room impulse responses from 15 enclosures, static and moving sources from 14 classes, up to two overlapping events, and direction-of-arrival trajectories plus onset/offset labels. The helper downloads official pages, paper, record metadata, and the 17 KB dataset README by default; the approximately 14.0 GB audio archives remain on Zenodo, while the roughly 1.7 MB label archives are an explicit opt-in.
Safe-first helperscripts/download/tau_nigens_sse_2020.sh
Audio understanding, generation & events
TAU-NIGENS Spatial Sound Events 2021
Safe-first helper
Sound Event Localization And Detection
Sound Source Localization
Acoustic Source Tracking
Sound Event Detection
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- custom_TAU_License
- License caution
- Zenodo lists the dataset as CC BY-NC 4.0. Source events come from the separately released NIGENS database, so preserve its provenance and review its active terms. The baseline license allows experimental non-commercial use and prohibits commercial use.
- Download notes
- The public v1.1.0 release is the complete DCASE 2021 Task 3 development and evaluation corpus: 600 development and 200 evaluation sound scenes of one minute each, in 4-channel FOA and tetrahedral-microphone formats at 24 kHz. It adds directional non-target interferers, permits overlapping instances of the same target class, and supplies direction-of-arrival trajectories plus onset/offset labels for the development set. The helper downloads official pages, paper, record metadata, and the 23 KB dataset README by default; the approximately 14.2 GiB audio archives remain on Zenodo, while the roughly 1.8 MiB development labels are an explicit opt-in. The 200-file evaluation set intentionally has no public labels.
Safe-first helperscripts/download/tau_nigens_sse_2021.sh
Speech recognition
TEDx Spanish Corpus
TEDx Spanish Corpus: Audio and Transcripts in Spanish Taken from TEDx Talks
Safe-first helper
Spanish Asr
Spontaneous Speech Recognition
Speech Transcription
Access pathOpenSLR
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR lists CC BY-NC-ND 4.0. The corpus is derived from TEDx Talks, so downstream use should also respect TED/TEDx source terms.
- Download notes
- OpenSLR SLR67 hosts a single 2.3 GiB archive with Spanish speech and transcripts from TEDx Talks. The helper saves the OpenSLR page by default and downloads the archive only with TEDX_SPANISH_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/tedx_spanish.sh
Speaker, identity & emotion
TESS
Toronto Emotional Speech Set
Safe-first helper
Speech Emotion Recognition
Acted Emotional Speech
Auditory Emotion Perception
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The official Borealis record lists CC BY-NC 4.0. The corpus contains identifiable human voices and permits only non-commercial reuse under that license; preserve attribution and review voice-data ethics for downstream use.
- Download notes
- The owner-hosted University of Toronto Dataverse release contains 2,800 WAV stimuli: 200 target words spoken by two English-speaking actresses aged 26 and 64 in seven acted emotions. The helper saves official dataset metadata by default and requires TESS_DOWNLOAD_AUDIO=1 before downloading the complete ZIP.
Safe-first helperscripts/download/tess.sh
Speech generation
Text to Audio Human Preference Benchmark
Rapidata Text to Audio Human Preference Benchmark
Safe-first helper
Text To Speech Evaluation
Human Preference Evaluation
Speech Naturalness Evaluation
Speech Friendliness Evaluation
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_applicable
- License caution
- The official dataset card and Hugging Face API declare no license. The rows include model labels, aggregate scores, individual votes, and annotator demographic fields; public access does not imply permission to redistribute or reuse those records, audio references, or generated outputs.
- Download notes
- The public, ungated release contains 4,269 pairwise comparison rows and about 32,000 human responses judging generated voices for friendliness and naturalness. It stores audio references as strings rather than embedding audio. The helper downloads the dataset card and API metadata by default; the approximately 0.8 MB repository snapshot is opt-in.
Safe-first helperscripts/download/rapidata_tts_preference.sh
Speaker, identity & emotion
TFCL AFE
TFCL Paired Clean and Acoustic-Front-End-Processed Speech Dataset
Safe-first helper
Synthetic Speech Detection
Audio Deepfake Detection
Acoustic Front End Robustness
Paired Representation Consistency
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_hugging_face_card_uses_other
- Code license
- mit
- License caution
- The dataset card labels its license "other" but provides no actual dataset terms. The MIT license covers the GitHub code, not the derived audio or checkpoint. ASVspoof 2019, OpenSLR RIRs, MUSAN, DNS Challenge, AudioSet, and Freesound inputs retain their own terms; verify all component rights before reuse or redistribution.
- Download notes
- The helper downloads official dataset/repository documentation and Hub metadata by default. Set TFCL_AFE_DOWNLOAD_DATA=1 for the approximately 34.9 GB evaluation archive and 3.61 GB train/development archive, TFCL_AFE_DOWNLOAD_CHECKPOINT=1 for the approximately 1.27 GB checkpoint, or TFCL_AFE_CLONE_REPO=1 for the code. Clean ASVspoof 2019 LA data is obtained separately under its source terms.
Safe-first helperscripts/download/tfcl_afe.sh
Speech recognition
THCHS-30
THCHS-30: A Free Chinese Speech Corpus
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Noisy Speech Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_applicable
- License caution
- OpenSLR lists Apache License v2.0 and the resource description says the database is free to academic users. The original CSLT URL linked from OpenSLR returned 404 when checked, so use the current OpenSLR page and paper for access/provenance.
- Download notes
- OpenSLR SLR18 hosts a 6.4 GiB speech/transcript archive, a 1.9 GiB 0 dB noisy test archive, and a 24 MiB supplementary resource archive with lexicon/noise samples. The helper saves the OpenSLR page by default and only downloads selected archives through THCHS30_DOWNLOAD_PARTS.
Safe-first helperscripts/download/thchs_30.sh
Speech recognition
Thorsten-Voice
Thorsten-Voice Dataset 2021.02 (Neutral)
Safe-first helper
Word Level Timing
Speech Alignment
Automatic Speech Recognition
Speech Synthesis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC_BY_4_0_with_conflicting_CC0_project_signal
- Code license
- CC0_1_0
- License caution
- The exact Zenodo 5525342 record declares CC BY 4.0, while the official project repository has a CC0-1.0 LICENSE and describes all Thorsten-Voice datasets as free to use; the newer combined Hugging Face release also declares CC0. Apply the exact archive's CC BY 4.0 attribution terms unless the owner clarifies that CC0 supersedes them. The repository's CC0 signal does not silently replace the license attached to the cited Zenodo record.
- Download notes
- The helper downloads the official Zenodo record, project README, repository metadata, and license text by default. The approximately 2.74 GB version 3.0 archive requires explicit opt-in. The July 2026 controllable-verbatim-ASR paper cites this exact DOI and describes the 23-hour German neutral corpus as its read-speech word-boundary timing evaluation set; it reports mean absolute boundary error and F1 at 50, 100, and 200 ms collars.
Safe-first helperscripts/download/thorsten_voice.sh
Speaker, identity & emotion
TidyVoice
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
Manual or gated
Multilingual Speaker Verification
Cross Lingual Speaker Verification
Speaker Recognition
Language Mismatch Robustness
+1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC0-1.0_with_use_restrictions
- Code license
- Apache-2.0
- License caution
- Mozilla Data Collective labels TidyVoiceX_ASV CC0-1.0 but also states that it must only be used for speaker verification and forbids speaker identification or attempts to recover speaker identity. Treat those owner-stated usage rules and the current Common Voice terms as binding access conditions despite the permissive license label. Apache-2.0 covers the WeSpeaker baseline repository, not any separate model or derived artifact rights.
- Download notes
- The public Mozilla Data Collective release contains 321,711 utterances (457 hours) from 4,474 multilingual speakers across 40 languages, with training and development splits, pseudonymized speaker IDs, language metadata, and same-/cross-language target and non-target trial lists. The current archive is approximately 36.72 GB. Download requires a Data Collective account and API key, so the helper saves official public documentation and prints the owner-supported access path without accepting credentials or fetching audio. The January paper also describes the broader Tidy-M monolingual condition across 81 languages; this entry's reproducible download pointer is the released TidyVoiceX_ASV challenge package. AMECxSV section 4.1 evaluates a deterministic speaker-disjoint split derived from the 12-million-trial TidyVoiceX development protocol, not the challenge's hidden official evaluation set. The July language-normalization paper additionally reports the official 4-million-pair all-language and 1.28-million-pair unseen-language evaluation tracks. Its claim of a released scoring/submission pipeline has no linked public repository, so the index does not present that implementation as downloadable.
Safe-first helperscripts/download/tidyvoice.sh
Audio understanding, generation & events
TimeGround-1M
Safe-first helper
Temporal Audio Grounding
Temporal Audio Localization
Timestamped Audio Description
Timestamped Audio Summarization
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-3.0_on_hugging_face_card
- Code license
- not_specified
- License caution
- The official dataset card declares CC BY 3.0. Its recordings are selected from English YODAS2 YouTube-derived shards, so source-video rights, availability, attribution, and platform terms still require review. No separate license for generation or evaluation code is provided on the dataset card.
- Download notes
- The public, ungated English release contains separate train and test splits for temporal localization, temporal description, timed summarization, and recording-level nested annotations. The official card reports about 59,000 training and 4,200 test recordings totaling roughly 14,200 hours across duration buckets from under 10 minutes to 120 minutes. The GigaChat 3.1 Audio paper evaluates these generated tasks by duration bucket in section 4.1. The helper downloads the dataset card, repository API metadata, paper page, and model card by default; the Hugging Face API reports about 1.50 TB of repository storage, so the full snapshot requires explicit opt-in.
Safe-first helperscripts/download/timeground_1m.sh
Speech recognition
TIMIT
TIMIT Acoustic-Phonetic Continuous Speech Corpus
Manual or gated
Automatic Speech Recognition
Phone Recognition
Acoustic Phonetic Analysis
Speaker Dialect Coverage
Access pathLDC / licensed
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_ldc_license
- Code license
- not_applicable
- License caution
- LDC catalog pages list licensing instructions for Subscription/Standard Members and Non-Members, web download media, and fee visibility after login. Portions are copyright 1993 Trustees of the University of Pennsylvania; consult the current LDC agreement before use or redistribution.
- Download notes
- LDC distributes TIMIT by web download after login/licensing. The helper only prints official access steps because the corpus is paid/licensed and not publicly script-downloadable.
Safe-first helperscripts/download/timit.sh
Speech recognition
TORGO
TORGO Database of Acoustic and Articulatory Speech from Speakers with Dysarthria
Safe-first helper
Dysarthria Detection
Pathological Speech Recognition
Speech Intelligibility Assessment
Acoustic Articulatory Modeling
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_academic_nonprofit_only
- Code license
- not_applicable
- License caution
- The owner page says use is free for academic, non-profit purposes and requires citation of at least one listed TORGO paper. It supplies the data as-is and does not identify a standard open-data license or grant commercial use. The recordings contain identifiable voices and disability and health information, so ethical and privacy review remains necessary.
- Download notes
- The public University of Toronto release contains aligned 16 kHz acoustic recordings and measured 3D articulatory features from eight English speakers with cerebral palsy or amyotrophic lateral sclerosis and seven matched controls. Stimuli include non-words, isolated words, restricted sentences, and spontaneous descriptions. Four BZip2 archives are organized as female dysarthric (F), female control (FC), male dysarthric (M), and male control (MC); they total approximately 8.9 GiB compressed and 18 GB uncompressed. The helper downloads the official page, correction spreadsheet, and coil-location documentation by default. Archive downloads require explicit terms acknowledgment and a selected group list. The July 2026 voice-concept bottleneck paper evaluates only headMic recordings with leave-one-speaker-out cross-validation; its exact derived split is not separately released.
Safe-first helperscripts/download/torgo.sh
Audio understanding, generation & events
TREA
Temporal Reasoning Evaluation of Audio
Safe-first helper
Audio Question Answering
Temporal Audio Reasoning
Audio Event Ordering
Audio Event Counting
+2 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC0-1.0_repository_with_ESC-50_upstream_terms
- Code license
- CC0-1.0
- License caution
- The repository applies a CC0-1.0 LICENSE and GitHub detects CC0-1.0, but the paper states that every TREA audio file combines recordings from ESC-50, whose dataset is CC BY-NC 3.0 and whose ESC-10 subset clips are CC BY. Treat the restrictive upstream terms and clip attribution as surviving the derived release rather than assuming the repository-level CC0 waiver clears all source-audio rights.
- Download notes
- TREA is a public, ungated 600-item temporal-reasoning benchmark derived by combining ESC-50 clips. Its TREA-O, TREA-C, and TREA-D subsets each contain 200 ordering, counting, or duration questions. The July 2026 Audio-Zero paper evaluates both Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA alongside MMAU Test-mini and MMAR. The repository releases both four-option multiple-choice and open-text answer formats, audio, metadata, evaluation code, and uncertainty perturbation scripts. The helper downloads official documentation, repository metadata, license, paper page, and the lightweight CSV annotations by default; set TREA_CLONE_REPO=1 to clone the approximately 688 MiB GitHub repository and its audio.
Safe-first helperscripts/download/trea.sh
Speech generation
TTS Multilingual Test Set
MiniMaxAI TTS Multilingual Test Set
Safe-first helper
Multilingual Text To Speech
Zero Shot Voice Cloning
Cross Lingual Voice Cloning
Speech Intelligibility Evaluation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- The official Hugging Face dataset card lists CC BY-SA 4.0. Its 48 speaker prompts are selected from Mozilla Common Voice, whose data is CC0-1.0; retain benchmark attribution and share adaptations under the card's stated terms.
- Download notes
- The public, ungated Hugging Face repository contains 100 test sentences and two Common Voice-derived speaker prompts (one female and one male) for each of 24 languages. The helper downloads the official dataset card by default; the approximately 7.3 MB snapshot requires TTS_MULTILINGUAL_TEST_SET_DOWNLOAD_HF=1. Qwen3-TTS evaluates a 10-language subset for zero-shot multilingual and target-speaker generation, but the report does not identify the exact text rows used.
Safe-first helperscripts/download/tts_multilingual_test_set.sh
Audio understanding, generation & events
TUT Sound Events 2017
TUT Sound Events 2017: Sound Event Detection in Real-Life Audio
Safe-first helper
Sound Event Detection
Polyphonic Sound Event Detection
Temporal Audio Event Localization
Street Sound Event Detection
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_noncommercial
- Code license
- not_applicable
- License caution
- Zenodo labels both releases Other (Non-Commercial), and each documentation archive contains a controlling EULA. Review that packaged agreement before use or redistribution; the generic Zenodo label is not a permissive Creative Commons grant.
- Download notes
- The public version-2 development release contains 24 street recordings totaling 1:32:08 with verified strong annotations for six overlapping event classes and an official four-fold cross-validation setup. The public evaluation release contains eight recordings totaling 29:09 and now includes reference metadata. DCASE 2017 Task 3 ranks systems by one-second segment-based error rate. The helper downloads Zenodo record JSON, documentation, and small annotation archives by default; the approximately 1.55 GiB of 24-bit, 44.1 kHz audio requires explicit opt-in.
Safe-first helperscripts/download/tut_sound_events_2017.sh
Representation & general suites
UME-ERJ
UME English Speech Database Read by Japanese Students
Manual or gated
Second Language Pronunciation Assessment
Phone Scoring
Rhythm Scoring
Intonation Scoring
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_research_only_terms_with_application_review
- Code license
- not_applicable
- License caution
- The official NII record limits permission to research purposes and requires a usage pledge plus review; it does not publish a standard open-data license. Review the current agreement before use, do not redistribute the recordings, and account for identifiable learner voices and proficiency ratings in privacy and ethics planning.
- Download notes
- NII's Speech Resources Consortium distributes UME-ERJ only after an applicant submits a usage pledge and passes review. The official record describes 202 Japanese learners of English (100 male and 102 female), 16 kHz/16-bit mono WAV recordings, phone- and prosody-focused sentence and word sets, and ratings from native English teachers. The helper saves only the public official record and evaluation-paper metadata, then prints the manual application requirement.
Safe-first helperscripts/download/ume_erj.sh
Representation & general suites
UME-JRF
UME Japanese Speech Database Read by Foreign Students
Manual or gated
Second Language Pronunciation Assessment
Holistic Pronunciation Scoring
Rhythm Scoring
Intonation Scoring
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_research_only_terms_with_application_review
- Code license
- not_applicable
- License caution
- The official NII record limits permission to research purposes and requires a usage pledge plus review; it does not publish a standard open-data license. Review the current agreement before use, do not redistribute the recordings, and account for identifiable learner voices, native-language backgrounds, and pronunciation ratings in privacy and ethics planning.
- Download notes
- NII's Speech Resources Consortium distributes UME-JRF only after an applicant submits a usage pledge and passes review. The official record describes 141 intermediate-to-advanced learners from 26 native-language groups (72 male and 69 female), 16 kHz/16-bit mono WAV recordings, and ATR-balanced, difficult-sentence, prosody-sentence, and difficult-word materials. The helper saves only public provenance and evaluation-paper metadata, then prints the manual application requirement.
Safe-first helperscripts/download/ume_jrf.sh
Audio understanding, generation & events
UrBAN
UrBAN: Urban Beehive Acoustics and PheNotyping Dataset
Safe-first helper
Environmental Sound Classification
Beehive Acoustic Monitoring
Colony Strength Regression
Hive Health Monitoring
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_specified
- License caution
- The FRDR dataset record explicitly lists CC BY 4.0. The GitHub repository has no detected license, so the analysis notebooks and scripts should not be assumed to use the dataset license. The Scientific Data article itself is CC BY-NC-ND 4.0, distinct from the dataset terms.
- Download notes
- The public FRDR release contains longitudinal 2021-2022 raw 16 kHz beehive audio plus inspection, temperature, humidity, and weather metadata from a ten-hive Montréal rooftop apiary. The Scientific Data descriptor reports more than 3,000 hours, while the older FRDR record and repository README say more than 2,000 hours; the current FRDR landing page reports approximately 1.265 TB of files. Its benchmark protocols include random-split and hive-independent colony-strength regression. A 2026 follow-up evaluates modulation-tensorgram models on nine hives and emphasizes cross-hive generalization. The helper saves official landing pages, repository documentation, and API metadata only. Full data transfer remains a manual FRDR Globus workflow because of the corpus size and may require a Globus account and client.
Safe-first helperscripts/download/urban_beehive.sh
Audio understanding, generation & events
UrbanSound8K
UrbanSound8K: A Dataset and Taxonomy for Urban Sound Research
Safe-first helper
Urban Sound Classification
Environmental Sound Classification
Audio Tagging
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc
- Code license
- not_applicable
- License caution
- Zenodo lists CC BY-NC 4.0. The official Urban Sound site says UrbanSound/UrbanSound8K are free for non-commercial use under Creative Commons BY-NC 3.0; Freesound attributions are included in the dataset.
- Download notes
- The archive is about 6 GiB and contains 8732 WAV clips pre-sorted into 10 official folds. The helper downloads citation/license metadata by default and requires URBANSOUND8K_DOWNLOAD_AUDIO=1 for the full archive.
Safe-first helperscripts/download/urbansound8k.sh
Speech understanding & dialogue
URO-Bench-pro
URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
Safe-first helper
Spoken Dialogue Model Evaluation
Speech To Speech Evaluation
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- MIT
- License caution
- Qwen uses the pro track. HF card and GitHub repo list MIT.
Safe-first helperscripts/download/uro_bench_pro.sh
Representation & general suites
User-Intent Queries (UIQ)
User-Intent Queries benchmark from Omni-Embed-Audio
Safe-first helper
User Intent Audio Retrieval
Language Based Audio Retrieval
Query Reformulation Robustness
Exclusionary Query Understanding
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- CC BY 4.0 applies only to the released UIQ text queries. AudioCaps, Clotho, and MECAT audio is not redistributed and retains its original source terms; the top-level Omni-Embed-Audio code repository is MIT.
- Download notes
- The public, ungated release contains 13,053 text-query records over the AudioCaps test, Clotho evaluation, and MECAT pools: question, imperative, tagging, paraphrase, and exclusionary negative variants. The helper downloads the approximately 12 MiB of query JSONL files plus the benchmark README and license; it does not download source audio. Fusion Embedding section 6.3 independently reuses UIQ on the 1,045-clip Clotho pool and reports only the four positive query formulations.
Safe-first helperscripts/download/uiq.sh
Audiovisual & cross-modal
VABench
VABench: A Comprehensive Benchmark for Audio-Video Generation
Safe-first helper
Text To Audio Video Generation
Image To Audio Video Generation
Stereo Audio Generation
Audio Video Synchronization
+5 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_for_i2av_hugging_face_release
- Code license
- not_specified
- License caution
- The VABench_I2AV dataset card declares Apache 2.0. The GitHub repository has no license file or detected license, so public access does not establish reuse rights for its prompt/QA files or evaluation code. The large cache repository has no top-level dataset license and bundles third-party evaluators and weights; review every component's terms separately. The paper itself is CC BY 4.0.
- Download notes
- VABench is one benchmark family with 778 text-conditioned and 521 image-conditioned audio-video cases across seven content categories, plus a 116-prompt stereo-audio track. The GitHub release includes the complete 1,299-row prompt mapping, task JSON files with audio and visual QA, and the evaluation framework. The public Hugging Face image release contains a roughly 296 MB ZIP. The separate approximately 36.5 GB VABENCH_CACHE_DIR repository contains evaluator model caches rather than benchmark examples and is intentionally not downloaded by the helper. Safe defaults fetch only official documentation, repository metadata, the prompt mapping, and the image dataset card; toolkit cloning and the image ZIP are explicit opt-ins.
Safe-first helperscripts/download/vabench.sh
Speech generation
VCTK
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit
Safe-first helper
Text To Speech
Speech Synthesis
Voice Cloning
Multi Speaker Speech Synthesis
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- The official VCTK README and DataShare license_text identify Creative Commons Attribution 4.0 International. The newspaper text source was used with permission from Herald & Times Group.
- Download notes
- The official DataShare ZIP is about 10.94 GiB. The helper saves the official README and license text by default and requires VCTK_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/vctk.sh
Audiovisual & cross-modal
VGGSound
VGGSound: A Large-scale Audio-Visual Dataset
Safe-first helper
Audio Visual Event Classification
Audio Event Classification
Audio Tagging
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Official VGG page and repository license file list the dataset as CC BY 4.0 for commercial/research use, while copyright remains with original video owners. Re-check YouTube availability and upstream media terms before reconstructing clips.
- Download notes
- The official VGG page currently says the original dataset download links are no longer available from that website. The helper downloads the official CSV metadata, license, and optional pretrained model files only; it does not fetch or redistribute YouTube media.
Safe-first helperscripts/download/vggsound.sh
Audiovisual & cross-modal
Video-MME
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Safe-first helper
Audio Visual Question Answering
Long Video Understanding
Multimodal Reasoning
Audio Enabled Video Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_academic_research_only
- Code license
- not_specified
- License caution
- The official README prohibits commercial use and, without prior approval, distribution, publication, copying, dissemination, or modification of Video-MME in whole or in part. Video copyrights remain with their owners. The GitHub repository has no detected license; obtain approval and re-check source-video rights before reuse beyond the stated academic evaluation context.
- Download notes
- The public, ungated release contains 900 videos totaling 254 hours and 2,700 human-annotated question-answer pairs, with audio and subtitles available as evaluation modalities. The helper downloads only official documentation by default. The Hugging Face API reports about 389 GB of repository storage, so the media snapshot requires both VIDEO_MME_ACK_TERMS=1 and VIDEO_MME_DOWNLOAD_HF=1. Qwen3.5-Omni evaluates Video-MME with use_audio_in_video=True in section 5.1.4, Table 7.
Safe-first helperscripts/download/video_mme.sh
Audiovisual & cross-modal
video-SALMONN 2 Caption Benchmark
video-SALMONN 2 Human-Annotated Audio-Visual Caption Benchmark
Safe-first helper
Audio Visual Video Captioning
Detailed Video Captioning
Audio Visual Event Understanding
Caption Completeness Evaluation
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_card_label
- Code license
- Apache-2.0
- License caution
- The Hugging Face card labels the dataset Apache-2.0, and the official GitHub repository contains an Apache-2.0 LICENSE. The release does not document per-video provenance or underlying media licenses, so the card label must not be assumed to clear third-party video, audio, speech, music, likeness, or platform rights. Review source-media rights before redistribution or commercial use.
- Download notes
- The public, ungated test set contains 483 audio-bearing videos, each 30-60 seconds long, with a human-annotated detailed caption and manually refined visual, speech, and non-speech atomic events. The released evaluator uses an LLM to report missing-event, incorrect-event, hallucination, and total error rates. The helper downloads official documentation, API metadata, the approximately 3.5 MB annotation JSON, and evaluator by default. The current Hugging Face files total approximately 1.70 GB, so the 483 MP4 files require VIDEO_SALMONN2_DOWNLOAD_HF=1. ReMo evaluates this test set as video-SALMONN2 / video-SAL2 in section 5.1 of arXiv:2607.21179.
Safe-first helperscripts/download/video_salmonn2_caption.sh
Music
VocalSet
VocalSet: A Singing Voice Dataset
Safe-first helper
Singing Voice Analysis
Vocal Technique Classification
Vowel Classification
Singing Voice Synthesis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- The Zenodo record lists CC BY 4.0 and open access. Re-check subject-consent and attribution expectations before redistributing derivative voice data.
- Download notes
- Zenodo hosts a single VocalSet.zip archive of about 2.1 GB with 10.1 hours of monophonic professional singing from 20 singers, covering all five vowels across standard and extended vocal techniques. The helper saves the Zenodo record metadata by default and requires VOCALSET_DOWNLOAD_ARCHIVE=1 before downloading the full archive.
Safe-first helperscripts/download/vocalset.sh
Audio understanding, generation & events
VocalSound
VocalSound: A Dataset for Improving Human Vocal Sounds Recognition
Safe-first helper
Human Vocal Sound Classification
Vocalization Recognition
Audio Classification
Demographic Bias Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_specified
- License caution
- The official README includes a Creative Commons BY-SA 4.0 notice for the VocalSound dataset. GitHub API reports no repository-level license, so the code/baseline license is not specified; re-check before redistributing code or derived data.
- Download notes
- The official README describes 21,024 crowdsourced recordings from 3,365 subjects covering laughter, sighs, coughs, throat clearing, sneezes, and sniffs, with speaker metadata such as age, gender, native language, country, and health condition. The helper saves official documentation by default and requires VOCALSOUND_DOWNLOAD_ARCHIVE=1 before downloading the 1.7 GiB 16 kHz or 4.5 GiB 44.1 kHz ZIP.
Safe-first helperscripts/download/vocalsound.sh
Enhancement, separation & quality
VoiceBank-DEMAND
Noisy speech database for training speech enhancement algorithms and TTS models
Safe-first helper
Speech Enhancement
Speech Denoising
Noise Robust Tts
Clean Noisy Parallel Speech
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata lists Creative Commons Attribution 4.0 International Public License. The corpus derives clean speech from VCTK and noises from DEMAND plus speech-shaped/babble sources; re-check component/source terms before redistribution.
- Download notes
- The DataShare record exposes paired clean/noisy train and test ZIPs plus text/log files. The helper saves public metadata and license by default; text files and multi-GB audio archives are explicit opt-ins.
Safe-first helperscripts/download/voicebank_demand.sh
Speech understanding & dialogue
VoiceBench
VoiceBench: Benchmarking LLM-Based Voice Assistants
Safe-first helper
Voice Assistant Evaluation
Spoken Instruction Following
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0
- Code license
- Apache-2.0
- License caution
- HF dataset card and GitHub repo both list Apache-2.0.
Safe-first helperscripts/download/voicebench.sh
Speech recognition
VoiceCodeBench
VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
Safe-first helper
Asr
Structured Token Recovery
Entity Recovery
Workplace Speech Recognition
Access pathHugging Face
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The repository and dataset card declare MIT. The card says paid contributors consented to dataset use and release, but the audio contains identifiable voice characteristics and its stated intended-use guidance excludes speaker identification, biometric modeling, voice cloning, demographic profiling, and model training or post-training.
- Download notes
- The public, ungated test-only release contains 300 human-recorded English workplace-speech segments totaling 5.587 hours, with 85 anonymized speakers and 1,482 audited targets across 26 structured entity types. Its primary Canonical Token/Entity Match and Task Success Rate metrics test exact recovery of values such as email addresses, phone numbers, URLs, command-line flags, file paths, identifiers, dates, and measurements. The helper downloads official documentation, license, paper, API metadata, and the approximately 1.1 MB annotation JSONL by default. The complete Hugging Face repository is approximately 1.83 GiB and requires VOICECODEBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/voicecodebench.sh
Speaker, identity & emotion
VoiceMOS Challenge 2026
VoiceMOS Challenge 2026: Automatic Prediction of Human Ratings of Speech
Manual or gated
Mean Opinion Score Prediction
Speech Quality Assessment
Comparative Category Rating Prediction
Emotional Speech Naturalness Assessment
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_publicly_specified
- Code license
- Apache-2.0
- License caution
- The public challenge page and baseline README do not state dataset reuse or redistribution terms. Apache-2.0 covers the baseline repository only, not challenge audio, listener ratings, URGENT material, CodecMOS-Accent, or emotional-speech source data.
- Download notes
- The official site says training data were released to registered participants through a CodaBench page sent by email, with evaluation data scheduled for July 31, 2026. Track 1 covers 840 multilingual utterances in nine languages from six URGENT speech-enhancement systems; Track 2 covers emotional TTS and human speech; Track 3 uses 4,000 CodecMOS-Accent samples from 24 codec-resynthesis and TTS systems, 32 speakers, and ten accents. The helper saves public challenge and baseline documentation, then prints the registration path; it does not guess or expose the emailed CodaBench URL.
Safe-first helperscripts/download/voicemos_challenge_2026.sh
Speaker, identity & emotion
VoicePrivacy Challenge
VoicePrivacy Challenge 2024
Manual or gated
Voice Anonymization
Speaker Linkability Evaluation
Speaker Verification Privacy
Anonymized Speech Recognition
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- registered_challenge_with_mixed_component_terms
- Code license
- GPL-3.0
- License caution
- The 2024 repository's LICENSE is GPL-3.0. LibriSpeech is CC BY 4.0; IEMOCAP is separately request-controlled. The public challenge pages and repository do not state a standalone open-data license covering every model package, anonymised submission output, or derived evaluation artifact, so the code license must not be applied to them.
- Download notes
- The helper downloads the official 2024 README, GPL-3.0 license text, and GitHub repository metadata only. The public recipe defines LibriSpeech development/test enrolment and trial partitions for speaker-linkability privacy and ASR utility, plus IEMOCAP development and test data for emotion-recognition utility. Its official downloader requires a password issued after challenge registration, and IEMOCAP must be requested separately. No credentials or protected audio are accepted or downloaded by the helper.
Safe-first helperscripts/download/voiceprivacy_challenge.sh
Audiovisual & cross-modal
VoxBlink2
VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Safe-first helper
Speaker Verification
Open Set Speaker Identification
Speaker Recognition
Audio Visual Speaker Recognition
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- not_specified
- License caution
- The repository states that released annotation data is CC BY-NC-SA 4.0, but does not separately license its software. YouTube source-media rights, platform terms, privacy considerations, and local law remain separate and are not granted by the annotation license.
- Download notes
- The official release provides annotations, YouTube links, timestamps, speaker labels, ASR outputs, speaker metadata, and evaluation protocols rather than redistributing audio or video. The corpus describes approximately 10 million segments, more than 110,000 speakers, and 16,000 hours across more than 15 language families. The helper downloads official documentation and license text by default; the Google Drive resource bundle remains a manual download, and cloning evaluation/data-construction code is opt-in. Source media must be obtained separately and may be unavailable or removed.
Safe-first helperscripts/download/voxblink2.sh
Audiovisual & cross-modal
VoxCeleb
VoxCeleb speaker recognition datasets
Safe-first helper
Speaker Identification
Speaker Verification
Speaker Recognition
Audio Visual Speaker Recognition
Access pathOpenSLR
Upstream termsNot specified
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_for_original_media
- Code license
- not_applicable
- License caution
- Official VGG pages say provided VoxCeleb/VoxCeleb2 metadata is CC BY-SA 4.0 and the corpora consist of YouTube URLs with timestamps; original media rights and privacy terms remain with upstream owners. OpenSLR SLR49 lists its small metadata resource as not copyrighted.
- Download notes
- Official VGG pages currently say VoxCeleb1 and VoxCeleb2 audio, URL/timestamp, and identifying metadata files are no longer available from that website. The helper downloads small OpenSLR speaker-recognition recipe metadata and trial lists only; it does not fetch the original audio/video.
Safe-first helperscripts/download/voxceleb.sh
Audiovisual & cross-modal
VoxConverse
VoxConverse: A Large Scale Audio-Visual Diarisation Dataset
Safe-first helper
Speaker Diarization
Audio Visual Diarization
Overlapping Speech Diarization
Access pathOfficial / other
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- The official page and repository README say VoxConverse is available for research purposes under CC BY 4.0, while copyright remains with the original video owners. The GitHub repository does not expose a standalone license file through the API.
- Download notes
- The helper clones or updates the official annotation repository and saves the official page by default. The official page lists dev/test WAV ZIPs with MD5 checksums; audio downloads are explicit opt-ins because the dev ZIP is about 1.9 GiB and the test ZIP is also large.
Safe-first helperscripts/download/voxconverse.sh
Speaker, identity & emotion
VoxENES 2026
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Safe-first helper
Speech Spoofing Detection
Audio Deepfake Detection
Synthetic Speech Detection
Voice Conversion Detection
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0-with-upstream-terms
- Code license
- not_applicable
- License caution
- Kaggle declares CC BY 4.0 for the release. Bona fide speech derives from LibriSpeech and VoxPopuli, and synthetic samples incorporate source speech, speaker references, and outputs from multiple TTS/VC systems; review those upstream terms, voice-data rights, and model-output policies before redistribution or commercial use. The paper's CC BY 4.0 license applies to the paper, not by itself to every incorporated recording.
- Download notes
- The public Kaggle release contains 53,628 standardized 16 kHz mono WAV samples across English and Spanish, including 3,028 bona fide samples, 4,600 original synthetic samples from seven TTS and three voice-conversion systems, and 46,000 post-processed variants. The helper downloads Kaggle metadata by default; the approximately 23.3 GB dataset requires explicit opt-in and an authenticated Kaggle CLI.
Safe-first helperscripts/download/voxenes_2026.sh
Speech understanding & dialogue
VoxLingua107
VoxLingua107: a Dataset for Spoken Language Recognition
Safe-first helper
Spoken Language Identification
Language Recognition
Speech Classification
Access pathHugging Face
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The TalTechNLP Hugging Face dataset card lists cc-by-nc-4.0. The dataset is built from YouTube-derived speech segments, so source-media availability and platform terms still apply; the SpeechBrain recipe repository did not expose a detected license.
- Download notes
- The paper reports 6628 hours across 107 languages plus a 1609-utterance verified evaluation set. The helper downloads small Hugging Face metadata files by default and requires VOXLINGUA107_DOWNLOAD_HF=1 before attempting the larger mirrored dataset snapshot. The original TalTech host was not reliably reachable during the 2026-07-09 check, so verify upstream availability before large downloads.
Safe-first helperscripts/download/voxlingua107.sh
Speaker, identity & emotion
VoxParadox
VoxParadox: Adversarial Paralinguistic Speech QA under Linguistic-Acoustic Contradiction
Safe-first helper
Adversarial Paralinguistic Speech Qa
Acoustic Vs Lexical Evidence Grounding
Speech Emotion Recognition
Speaker Attribute And Identity Reasoning
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The custom USC Research License permits educational, research, and non-profit use with notice retention; commercial use requires a separate USC license. The card says audio was synthesized with ElevenLabs, GPT-4o, and Microsoft Azure and warns that commercial reuse is additionally subject to those vendors' terms. Do not treat public Hub access as an unrestricted or open-data grant.
- Download notes
- The public test-only release contains 2,000 verified English MCQs and 2,000 WAV files, with 200 items for each of ten tasks: age, gender, emotion, intonation, speaker identity, speaker counting, pitch, volume, speaking rate, and vocal range. Each transcript asserts a label that conflicts with the acoustic ground truth. The released scorer reports ground-truth accuracy and adversarial-label agreement; the code repository supplies Audio-Flamingo-3 and Qwen2-Audio runners, probing code, and PCLM/DPO model links. The helper downloads cards, terms, repository metadata, the JSONL/JSON manifests, and scorer by default. The Hugging Face API reports approximately 1.34 GB of used repository storage, so the full snapshot requires explicit opt-in; the larger code clone is a separate opt-in.
Safe-first helperscripts/download/voxparadox.sh
Speech recognition
VoxPopuli
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Safe-first helper
Multilingual Asr
Speech To Text Translation
Self Supervised Speech Representation Learning
Accented Speech Recognition
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc0-1.0
- Code license
- cc-by-nc-4.0
- License caution
- Official repo lists VoxPopuli data as CC0 and points users to the European Parliament legal notice for raw data; code and pretrained models are CC BY-NC 4.0.
- Download notes
- HF hosts converted Parquet shards and is about 673 GiB total; select a language/config and split before downloading.
Safe-first helperscripts/download/voxpopuli.sh
Speech understanding & dialogue
VoxSafeBench
VoxSafeBench: Not Just What Is Said, but Who, How, and Where
Safe-first helper
Speech Llm Safety Evaluation
Audio Conditioned Safety
Spoken Jailbreak Robustness
Agentic Action Safety
+6 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache_2.0_dataset_card_with_upstream_and_responsible_use_caveats
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares Apache-2.0 and the official code repository includes Apache-2.0. The paper says VoxSafeBench adapts and extends prior benchmarks, constructs new items, and uses some off-the-shelf data; its ethics statement identifies the intended use as auditing, red-teaming, and mitigation research and warns against re-identification or voice profiling. Do not assume the aggregate labels override upstream source rights or remove these safety and privacy obligations. The arXiv perpetual license covers the article, not third-party source material.
- Download notes
- The owner-linked, auto-gated Hugging Face release packages 17 splits across Safety Tier 1/2, Fairness Tier 1/2, and Privacy Tier 1/2 as raw JSONL plus audio. The paper defines 22 bilingual English/Chinese task families: Tier 1 compares matched text, clean speech, and diverse speech for content-centric risks, while Tier 2 holds benign transcripts fixed and tests whether speaker, paralinguistic, background, or overlap cues change safe, fair, or privacy-preserving behavior. The repository releases model runners, judge prompts, and evaluation code. Its README directs inferential-privacy evaluation to the separate HearSay family, which is not packaged as a VoxSafeBench split. The Hub API reports 25,357,365,873 bytes (about 23.6 GiB) of storage and 39,804 files, so the helper fetches only public paper, code, license, project, and Hub metadata by default; benchmark assets require accepted access terms, authentication, an acknowledgement flag, and an explicit download opt-in.
Safe-first helperscripts/download/voxsafebench.sh
Audiovisual & cross-modal
VSRo-200
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
Safe-first helper
Visual Speech Recognition
Audio Visual Speech Recognition
Sentence Level Lip Reading
Low Resource Romanian Speech
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting-noncommercial-terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0, while paper appendix A.1 says VSRo-200 is shared under CC BY-NC-SA 4.0. Both prohibit commercial use, but the share-alike discrepancy should be clarified before redistribution. The release contains metadata only and does not relicense the referenced YouTube audio or video. The official code repository exposes no license file, so do not assume its scripts or model artifacts inherit the dataset terms.
- Download notes
- The public, ungated Hugging Face release contains approximately 26.3 MB of metadata rather than podcast media: YouTube video IDs, segment timestamps, Romanian transcripts, and speaker-gender labels for 64,710 clips from 235 speakers. It provides separate human-annotated and automatically transcribed training/validation manifests, seen- and unseen-speaker tests, and four-category out-of-domain evaluation metadata. The official repository provides reconstruction, preprocessing, training, inference, audio-visual fusion, and LRRo transfer-evaluation code plus small sample face crops. The helper downloads the dataset card, API record, and five CSV manifests by default; cloning the toolkit is opt-in. Reconstructing clips requires accessing the referenced YouTube videos, which may disappear or carry separate platform and rights constraints.
Safe-first helperscripts/download/vsro_200.sh
Audio understanding, generation & events
WABAD
WABAD: A World Annotated Bird Acoustic Dataset for Passive Acoustic Monitoring
Safe-first helper
Bird Species Detection
Passive Acoustic Monitoring
Temporal Audio Event Localization
Time Frequency Event Localization
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_zenodo_metadata_treat_as_cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The Zenodo structured license field says CC BY 4.0, but the record's human-readable description explicitly says Creative Commons Attribution-NonCommercial 4.0. Treat the release as CC BY-NC 4.0 pending clarification from the maintainers; retain attribution and do not assume commercial-use permission from the structured field alone.
- Download notes
- The public, ungated release contains 5,047 minutes of passive-acoustic audio with 91,931 time-frequency-bounded vocalizations from 1,192 bird species, collected at 72 sites in 29 recording locations across 13 biomes. MetaPerch evaluates WABAD as an 84-hour multi-species detection benchmark in its results section. The helper downloads the Zenodo record, README, site metadata, pooled annotations, and species list by default; the 72 site archives total approximately 19.8 GiB and require explicit site-level opt-in.
Safe-first helperscripts/download/wabad.sh
Audio understanding, generation & events
WavCaps
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
Safe-first helper
Audio Captioning
Audio Language Retrieval
Audio Language Modeling
Zero Shot Audio Classification
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- academic_only
- Code license
- not_specified
- License caution
- The GitHub README and Hugging Face card say only academic uses are allowed for WavCaps audio. The HF metadata advertises CC BY 4.0, but the dataset card also points users to component source terms for FreeSound, BBC Sound Effects, SoundBible, and AudioSet; re-check those source licenses before redistribution or commercial use. Provided models are described as non-commercial research under a UK data copyright exemption.
- Download notes
- The Hugging Face repository exposes JSON metadata and split FLAC waveform ZIPs for FreeSound, BBC Sound Effects, SoundBible, and AudioSet SL. The full repository is hundreds of GiB, so the helper downloads README/JSON metadata by default and requires WAVCAPS_DOWNLOAD_ZIPS=1 plus WAVCAPS_ZIP_SOURCES for waveform archives.
Safe-first helperscripts/download/wavcaps.sh
Speaker, identity & emotion
WaveFake
WaveFake: A Data Set to Facilitate Audio Deepfake Detection
Safe-first helper
Audio Deepfake Detection
Synthetic Speech Detection
Vocoder Artifact Detection
Cross Generator Generalization
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-4.0
- Code license
- MIT
- License caution
- The Zenodo record and bundled dataset license declare CC BY-SA 4.0, while the GitHub toolkit is MIT. The archive contains generated audio only; LJSpeech and JSUT are not redistributed and retain their own source-corpus, speaker, and attribution terms. Review the licenses and usage conditions of the pretrained vocoders and TTS models when regenerating or extending the benchmark.
- Download notes
- WaveFake is one bilingual audio-deepfake family containing 104,885 generated 16-bit PCM WAV clips, approximately 175 hours in total. Its ten sample subsets cover six vocoder architectures and a complete TTS pipeline: eight English distributions based on LJSpeech and two Japanese distributions based on JSUT basic5000. The release does not redistribute either bona fide source corpus. The paper evaluates MFCC/LFCC GMM and RawNet2 baselines with EER, leave-one-generator-out, cross-language/speaker, novel-text, and simulated phone-recording protocols. The helper downloads the paper page, first-party documentation, Zenodo metadata, license, and datasheet by default. The 28,918,626,084-byte audio ZIP and toolkit clone are separate opt-ins.
Safe-first helperscripts/download/wavefake.sh
Speech recognition
WenetSpeech
Manual or gated
Mandarin Asr
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- non-commercial use under CC BY 4.0
- Code license
- Apache-2.0
- License caution
- Official site says WenetSpeech does not own audio copyright; original audio copyrights remain with owners.
Safe-first helperscripts/download/wenetspeech.sh
Enhancement, separation & quality
WHAM! / WHAMR!
WSJ0 Hipster Ambient Mixtures and WHAMR!: Noisy and Reverberant Single-Channel Speech Separation
Safe-first helper
Noisy Speech Separation
Speech Enhancement
Reverberant Speech Separation
Source Separation
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official WHAM page states the WHAM! and WHAM!48kHz noise datasets are CC BY-NC 4.0. Generated mixtures also depend on WSJ0/wsj0-2mix licensing, so redistribution or commercial use requires checking those upstream terms too.
- Download notes
- The helper downloads the official landing page and small WHAM!/WHAMR! generation script archives by default. WHAM! noise is 17 GiB compressed and WHAM!48kHz is 68.1 GiB compressed, so those archives are explicit opt-ins. Building full WHAM!/WHAMR! mixtures also requires separately licensed WSJ0/wsj0-2mix access.
Safe-first helperscripts/download/wham_whamr.sh
Speech recognition
Whisper-RIR-Mega
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
Safe-first helper
Automatic Speech Recognition
Reverberant Speech Recognition
Asr Robustness
Room Acoustics Robustness
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_CC-BY-NC-4.0_upstream_terms
- Code license
- not_specified_currently_unavailable
- License caution
- The benchmark dataset card declares CC BY 4.0 and identifies LibriSpeech as CC BY 4.0, but the current RIR-Mega v2 card declares CC BY-NC 4.0 for its RIR audio. Apply the stricter non-commercial upstream terms to the derived reverberant audio unless the owner clarifies otherwise. The benchmark card says its curation repository is MIT, but the linked repository was unavailable, so that code license could not be independently verified.
- Download notes
- The public, ungated release contains 2,000 English LibriSpeech test-clean utterances, each paired with a 16 kHz reverberant version made using one RIR-Mega room impulse response. Its deterministic, acoustically stratified split has 400 validation and 1,600 test pairs; evaluation reports clean/reverberant WER and CER plus the reverb penalty, with RT60 and DRR metadata when available. The helper saves the dataset card, API metadata, paper, and small leaderboard files by default. The Hugging Face API reports about 1.13 GB of repository storage, so the complete audio and Arrow snapshot requires WHISPER_RIRMEGA_DOWNLOAD_HF=1. The paper's cited GitHub code repository returned HTTP 404 when checked on 2026-07-22.
Safe-first helperscripts/download/whisper_rirmega.sh
Speech understanding & dialogue
WildSpeech-Bench
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
Safe-first helper
Speech To Speech Evaluation
Natural Speech Conversation
Access pathHugging Face
Upstream termsOpen / attribution signals
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC BY 4.0, except third-party datasets with their own terms
- Code license
- CC BY 4.0, except third-party datasets with their own terms
- License caution
- License.txt says users must comply with original licenses for third-party datasets.
Safe-first helperscripts/download/wildspeech_bench.sh
Audiovisual & cross-modal
WorldSense
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Safe-first helper
Audio Visual Question Answering
Omni Modal Video Understanding
Cross Modal Reasoning
Audio Visual Perception
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_cc_by_nc_sa_4_0_and_cc_by_4_0
- Code license
- not_specified
- License caution
- The WorldSense paper v3 Appendix G states CC BY-NC-SA 4.0, while the official repository README and Hugging Face card state CC BY 4.0. Apply the more restrictive CC BY-NC-SA 4.0 interpretation until the maintainers resolve the conflict. Videos are sourced primarily from FineVideo with selected MUSIC-AVQA material, so component-media terms and rights also require review. GitHub reports no detected repository license.
- Download notes
- The public, ungated release contains 1,662 synchronized audio-visual videos and 3,172 multiple-choice question-answer pairs across 26 tasks. The helper downloads official documentation and the approximately 4.3 MB QA JSON by default; the Hugging Face API reports approximately 18.1 GB of repository storage, so video and subtitle archives require WORLDSENSE_DOWNLOAD_HF=1. Qwen3.5-Omni reports WorldSense in section 5.1.4, Table 7.
Safe-first helperscripts/download/worldsense.sh
Speaker, identity & emotion
WSJ0-2mix / wsj0-mix
wsj0-mix: Single-channel multi-speaker speech separation mixtures from WSJ0
Safe-first helper
Speech Separation
Multi Speaker Speech Separation
Source Separation
Cocktail Party Speech Separation
Access pathLDC / licensed
Upstream termsMixed / custom — review
Paper citationsUnavailable
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ldc_restricted_derived
- Code license
- MERL script license not specified on the reachable page; pywsj0-mix is MIT.
- License caution
- The generated mixtures derive from the LDC CSR-I WSJ0 corpus, so access, use, and redistribution must follow the active LDC agreement. The MERL page provides scripts but does not publish the audio mixtures or a standalone data license.
- Download notes
- The helper downloads the official MERL page and generation scripts by default and can clone the MIT-licensed Python generator. It does not download WSJ0 audio; generation requires an already licensed local WSJ0 corpus from LDC and explicit WSJ0_2MIX_RUN_GENERATION=1. TF-MossFormer sections 3.1-3.3 use the standard 8 kHz two-speaker setup with 20,000 training, 5,000 validation, and 3,000 speaker-disjoint test mixtures and report SI-SDRi and SDRi; that paper adds no new mixture release.
Safe-first helperscripts/download/wsj0_2mix.sh
Speech recognition
WSYue-ASR-eval
WSYue-ASR-eval: Cantonese ASR Benchmark
Safe-first helper
Automatic Speech Recognition
Cantonese Speech Recognition
Code Switched Speech Recognition
Long Form Speech Recognition
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- Apache-2.0
- License caution
- The official Hugging Face card labels WSYue-ASR-eval CC BY-NC 4.0. The repository is Apache-2.0, but that code license must not be applied to the separately hosted benchmark audio. Review both sources before reuse, redistribution, or commercial evaluation.
- Download notes
- The public ungated Hugging Face release provides a 9.46-hour Short subset and a 1.97-hour Long subset, totaling about 1.06 GiB of audio, transcripts, and TextGrid annotations. The helper downloads only the dataset card and API metadata by default; set WSYUE_ASR_EVAL_DOWNLOAD_DATA=1 to fetch the released benchmark files. The companion 21,800-hour WenetSpeech-Yue corpus is training data and is not downloaded by this benchmark helper.
Safe-first helperscripts/download/wsyue_asr_eval.sh
Speech recognition
X-ARES
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Safe-first helper
Audio Representation Evaluation
Speech Representation Evaluation
Environmental Sound Representation Evaluation
Music Representation Evaluation
+11 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_component_terms
- Code license
- Apache-2.0
- License caution
- The repository code is Apache 2.0. Public task data are repackaged upstream corpora whose Zenodo records declare dataset-specific licenses (for example, the ESC-50 record declares CC BY-NC 3.0); review every selected record and original corpus terms before use or redistribution. The five private task datasets are not publicly released. No explicit paper license was found in the canonical arXiv or ISCA metadata.
- Download notes
- X-ARES is one benchmark protocol over component corpora, not a new collection of 22 independent datasets. The Interspeech 2025 paper evaluates frozen encoders with task-specific MLP and k-NN heads across speech, environmental-sound, and music domains. The current repository provides 22 public task implementations backed by 21 Zenodo records; its LibriSpeech ASR and male/female tasks share one record. Five additional vehicle or user-interaction sound tasks (finger snap, inside/outside car, key scratching, LiveEnv, and subway broadcast) are marked private and are skipped by the runner. SpeechOcean762 remains unchecked in the supported-task list and has no current task module. The helper downloads only official metadata, documentation, and license text by default, with an optional toolkit clone; it never launches the toolkit's automatic multi-dataset downloads.
Safe-first helperscripts/download/xares.sh
Speech recognition
XARES-LLM
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
Safe-first helper
Audio Encoder Evaluation
Large Audio Language Model Integration
Generative Audio Classification
Automatic Speech Recognition
+10 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_card_claim_with_mixed_component_terms
- Code license
- Apache-2.0
- License caution
- The repository code is Apache 2.0 and the xares_llm_data card declares Apache 2.0, but the snapshot repackages independently licensed source corpora such as ESC-50, FSD50K, VoxCeleb1, and Song Describer. The blanket card tag should not be assumed to override component licenses, consent conditions, source-media rights, or noncommercial restrictions. MECAT-Caption separately declares CC BY 3.0 while retaining upstream audio-source caveats. Hidden challenge assets have no public artifact license because they are not released.
- Download notes
- XARES-LLM is one Interspeech 2026 challenge and evaluation protocol, not 20 new dataset families. It freezes a submitted encoder and trains a projector plus LoRA-adapted SmolLM2-135M decoder. The current public toolkit implements Track A over 15 classification tasks and Track B over five understanding tasks; its ungated Hugging Face snapshot repackages 20 component corpora as WebDataset archives and reports about 128.7 GB of storage. MECAT-Caption is fetched from its separate public repository. The paper additionally reports hidden Track A and Track B challenge evaluations, but their exact shards and configs are not released. Its Track A table also lists a LibriSpeech gender task that is absent from the current repository configs and baseline table, so it is not claimed as reproducible public coverage. The helper downloads only official metadata, documentation, task configs, and license text by default; toolkit cloning and the full large dataset snapshots require separate explicit opt-ins.
Safe-first helperscripts/download/xares_llm.sh
Speech generation
ZeroSpeech 2019 TTS without T
The Zero Resource Speech Challenge 2019: TTS Without T
Safe-first helper
Acoustic Unit Discovery
Speaker Invariant Representation Learning
Discrete Speech Tokenization
Zero Resource Voice Conversion
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_unspecified_and_challenge_only_custom_terms
- Code license
- Not specified in the source record.
- License caution
- No standalone license is stated for the public English archives. The official surprise-data agreement limits use to the Zero Resource Speech Challenge, prohibits research and commercial uses outside the challenge, and prohibits redistribution while allowing cited examples in research reports. Public URLs do not relax these terms; review the current owner page and source-corpus conditions before use.
- Download notes
- The helper saves official challenge, paper, data, and results pages plus live archive headers by default. The 156 MB toy and 2.5 GB English archives require separate opt-ins. The 1.5 GB surprise-language ZIP is password protected and governed by challenge-only, no-redistribution terms, so the helper marks it manual_required rather than automating it.
Safe-first helperscripts/download/zerospeech_2019.sh
Representation & general suites
ZeroSpeech 2021 Spoken Language Modeling Benchmark
The Zero Resource Speech Benchmark 2021: Metrics and Baselines for Unsupervised Spoken Language Modeling
Safe-first helper
Acoustic Phonetic Discrimination
Spoken Word Nonword Discrimination
Spoken Grammaticality Judgment
Spoken Lexical Semantic Similarity
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- no_detected_license
- License caution
- Neither the archive endpoint nor the baseline repository states a standalone data license, and the repository has no detected license file. The suite includes Libri-light/LibriSpeech-derived recordings, synthetic stimuli produced with Google text-to-speech voices, and human similarity judgments compiled from thirteen source datasets. Public download access does not relicense those components; verify all upstream and service terms before reuse or redistribution.
- Download notes
- The official server exposes one approximately 30.61 GB ZIP archive. The helper saves the paper, baseline documentation, repository tree, and live archive headers by default; downloading the large archive and cloning the baseline repository are separate opt-ins.
Safe-first helperscripts/download/zerospeech_2021_slm.sh