Speech recognition
2nd MLC-SLM Challenge 2026
2nd Multilingual Conversational Speech Language Model Challenge 2026
Manual or gated
Multilingual Conversational Asr
Speaker Diarization
Speaker Attributed Asr
Acoustic Conversation Understanding
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_workshop_only_registration_agreement
- Code license
- not_specified
- License caution
- The official agreement limits the datasets to the 2026 MLC-SLM Workshop, prohibits redistribution and any other use, requires access controls, and requires return or destruction after termination. The challenge and baseline repositories have no detected license files, so their public visibility does not establish permission to reuse code. This second-edition release must not be conflated with the first MLC-SLM Eval annotation repository or its CC BY-SA 4.0 card label.
- Download notes
- The official challenge describes approximately 2,100 hours of two-speaker conversational training audio across 14 languages, plus approximately four development hours per language. Task 1 evaluates diarization and recognition with DER and time-constrained minimum-permutation WER/CER; Task 2 evaluates acoustic and semantic understanding through multilingual multiple-choice questions. The Task 1 system paper reports 150 development conversations across 21 language/accent categories and says evaluation references are not released. Access to training, development, and evaluation data requires challenge registration and acceptance of the data-use agreement. The helper downloads only public documentation, agreement, repository metadata, and paper metadata, then prints the manual registration path.
Safe-first helperscripts/download/mlc_slm_2nd_challenge.sh
Audio understanding, generation & events
ADQA-Bench
ADQA-Bench: Audio-Dependent Question Answering Evaluation Benchmark
Safe-first helper
Audio Question Answering
Audio Dependent Reasoning
Multiple Choice Question Answering
Shortcut Robustness
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_with_upstream_terms
- Code license
- not_applicable
- License caution
- The Hugging Face card declares Apache-2.0 for the release, but the dataset incorporates a portion of MMAU, MMAR, and MMSU plus newly annotated questions over other audio. Those component datasets and source recordings retain their own terms; review provenance and upstream media rights before redistribution or commercial use.
- Download notes
- The public, ungated DCASE 2026 Task 5 evaluation release contains 3,000 English multiple-choice questions with 3,000 WAV files spanning speech, music, and general sound understanding. Every item passed the organizers' four-stage Audio-Dependency Filtering process to reduce silent-audio and text-only shortcuts. The DCASE 2026 task-summary paper reports 1,607 development items, 14 participating teams, 36 ranked submissions, and two parameter-count tracks; its team results use the hidden 3,000-item evaluation set rather than the development set used for organizer baselines. The current release provides questions and choices without answers for challenge evaluation; the dataset card says answers will be released after the competition. The helper downloads official pages, the dataset card, API metadata, and the lightweight no-answer JSONL by default; the Hugging Face API reports approximately 2.94 GB of repository storage, so the audio snapshot requires explicit opt-in.
Safe-first helperscripts/download/adqa_bench.sh
Speech understanding & dialogue
ADReSS / ADReSSo
Alzheimer's Dementia Recognition through Spontaneous Speech Challenges
Manual or gated
Cognitive Impairment Detection
Alzheimers Dementia Classification
Mmse Score Regression
Cognitive Decline Prediction
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-3.0_with_password_protected_clinical_access
- Code license
- not_applicable
- License caution
- TalkBank says CC BY-NC-SA 3.0 governs its data unless otherwise indicated, prohibits incorporation into commercial products or large language models, and restricts DementiaBank membership to established researchers and clinicians or faculty-sponsored students. Password- protected data may not be shared with non-members or posted elsewhere. The recordings contain sensitive clinical and potentially identifiable speech; users must also follow the TalkBank Code of Ethics, NIH confidentiality protections, citation rules, and non-storage requirements for web processing.
- Download notes
- The official DementiaBank pages release age- and gender-balanced spontaneous-speech challenge sets after consortium approval. ADReSS 2020 provides enhanced full audio, normalized speech segments, transcripts, demographics, diagnosis labels, and MMSE scores for Alzheimer's-dementia classification and MMSE regression. ADReSSo 2021 is audio-only at evaluation time and adds longitudinal cognitive- decline prediction. The helper saves the public challenge pages, DementiaBank access rules, and the July 2026 cross-dataset evaluation paper, then prints the manual membership path; it never attempts to access password-protected participant recordings.
Safe-first helperscripts/download/adress_challenges.sh
Audio understanding, generation & events
AF-Reasoning-Eval
AF-Reasoning-Eval: Sound Reasoning Evaluation Benchmark
Safe-first helper
Audio Question Answering
Audio Reasoning
Audio Classification
Commonsense Reasoning
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC BY 4.0 metadata with mixed upstream audio terms
- Code license
- MIT
- License caution
- NVIDIA releases AF-Reasoning-Eval metadata under CC BY 4.0 and repository code under MIT. The 150 AQA items derive from Clotho-AQA, whose question-answer CSVs are MIT while Freesound audio retains per-file Creative Commons terms. The 7,227-item classification set derives from FSD50K, whose clips retain per-file CC0, CC-BY, CC-BY-NC, or CC Sampling+ terms in addition to the curated dataset's CC BY release.
- Download notes
- The helper downloads all four official JSON annotation files, totaling about 2.1 MB, plus the Sound-CoT README. It does not duplicate source audio. The AQA subset points to Clotho-AQA filenames and the classification subset points to FSD50K evaluation filenames; use those benchmarks' separate helpers and terms to obtain audio.
Safe-first helperscripts/download/af_reasoning_eval.sh
Music
AI-Generated Cover Song Diagnostics
A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
Safe-first helper
Cover Song Generation Evaluation
Music Quality Assessment
Music Error Diagnosis
Music Feature Analysis
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_for_released_tables_no_audio
- Code license
- MIT
- License caution
- The repository's MIT license covers its software and associated documentation, including the released score, manifest, and feature tables. It does not grant rights to the absent source songs or generated cover audio. No raw audio is publicly released, and users must supply locally authorized files to rerun audio feature extraction.
- Download notes
- The public release covers 30 generated covers from five source songs and six systems. It includes anonymized source-song identifiers, system/file mappings, expert scores for melody, harmony, key, style, and arrangement/production, nine extracted features, and the analysis pipeline. The helper downloads these lightweight tables and official documentation by default. The paper and repository explicitly withhold raw audio because of source-song copyright and commercial-API licensing constraints; cloning the small repository does not provide audio.
Safe-first helperscripts/download/ai_cover_song_diagnostics.sh
Audio understanding, generation & events
AIR-Bench
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Safe-first helper
Audio Question Answering
Audio Instruction Following
Audio Language Model Evaluation
Speech Understanding
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- Apache-2.0
- License caution
- The HF dataset card lists CC BY-NC 4.0 and enumerates component sources with mixed upstream licenses, including MusicCaps, Clotho, Fisher, SpokenWOZ, Common Voice, IEMOCAP, acoustic-scene datasets, MUSIC-AVQA, FMA, MTG-Jamendo, NSynth, SLURP, VoxCeleb, LibriSpeech, CoVoST 2, Fake-or-Real, and VocalSound. Re-check component terms before redistribution or commercial use.
- Download notes
- The helper downloads the official GitHub README and Hugging Face dataset card by default. The full HF audio snapshot is about 45.9 GB, so it requires AIR_BENCH_DOWNLOAD_HF=1; cloning the evaluation repo is also opt-in.
Safe-first helperscripts/download/air_bench.sh
Speech recognition
AISHELL-1
AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR33 lists Apache License v2.0 and also describes the data as free for academic use; re-check upstream terms before redistribution or commercial use.
- Download notes
- OpenSLR hosts a 15 GiB speech/transcript archive plus a small supplementary resource archive with lexicon and speaker info. The helper saves the OpenSLR page and supplementary archive by default; the large corpus archive is an explicit opt-in.
Safe-first helperscripts/download/aishell_1.sh
Speech generation
AISHELL-3
AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
Safe-first helper
Text To Speech
Speech Synthesis
Multi Speaker Speech Synthesis
Mandarin Speech Synthesis
+1 more
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_specified
- License caution
- OpenSLR SLR93 lists Apache License v2.0 for the resource. The aishelltech external URL redirected but returned HTTP 429 during the 2026-07-09 check, so OpenSLR was used as the primary access page.
- Download notes
- OpenSLR lists one 19 GiB speech/transcript archive. The helper saves the official OpenSLR page by default and requires AISHELL3_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/aishell_3.sh
Speech recognition
AISHELL-4
AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
Safe-first helper
Meeting Transcription
Multi Channel Asr
Speech Enhancement
Speech Separation
+2 more
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- Apache-2.0
- License caution
- OpenSLR SLR111 lists the corpus under CC BY-SA 4.0. The official baseline repository is Apache-2.0. The arXiv paper is CC BY 4.0, which is separate from the share-alike dataset terms.
- Download notes
- The public OpenSLR release contains 211 real Mandarin meeting sessions totaling 120 hours, recorded with an eight-channel circular microphone array in small, medium, and large rooms. It provides accurate transcripts and speaker activity for meetings with four to eight speakers. The helper saves the OpenSLR page and baseline documentation by default. The approximately 5.2 GB test archive and 46 GB of training archives are explicit opt-ins. VibeVoice-ASR-BitNet section 3.1 and Table 4 evaluate the AISHELL-4 test set with CER.
Safe-first helperscripts/download/aishell_4.sh
Speech recognition
AliMeeting
AliMeeting: A Free Mandarin Multi-channel Meeting Speech Corpus
Safe-first helper
Meeting Transcription
Multi Channel Asr
Multi Speaker Asr
Speaker Diarization
+1 more
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR119 lists AliMeeting under CC BY-SA 4.0. The corpus contains real Mandarin meetings with far-field microphone-array and near-field headset recordings; check challenge rules for benchmark submissions.
- Download notes
- The helper downloads small OpenSLR metadata by default. Corpus archives are large, about 73.24 GiB far-field train, 22.85 GiB near-field train, 3.42 GiB eval, and 8.90 GiB test, so they are explicit opt-ins.
Safe-first helperscripts/download/alimeeting.sh
Speech recognition
AMI
AMI Meeting Corpus
Safe-first helper
Meeting Speech Recognition
Multi Speaker Asr
Distant Speech Recognition
Meeting Understanding
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Official AMI pages say the corpus, signals, transcription, and some annotations are CC BY 4.0. OpenSLR SLR16 still lists an older modified CC BY-NC-SA v2.0 notice, so prefer the official AMI license page for current terms and re-check before redistribution.
- Download notes
- The helper downloads official annotation ZIPs by default. OpenSLR acoustic archives and the HF converted dataset are large, so audio downloads are explicit opt-ins.
Safe-first helperscripts/download/ami.sh
Representation & general suites
Androids Corpus
The Androids Corpus: A New Publicly Available Benchmark for Speech Based Depression Detection
Safe-first helper
Speech Based Depression Detection
Clinical Voice Assessment
Health Audio Classification
Paralinguistic Speech Analysis
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_academic_research_only_no_redistribution
- Code license
- not_specified
- License caution
- The official README limits the corpus to non-commercial academic research and prohibits redistribution, broadcasting, or making it publicly available anywhere else. No standalone code or data license file is present. The recordings expose depression/control labels plus age, gender, education, and identifiable voices, so users must also apply appropriate clinical-data, privacy, consent, and ethics review.
- Download notes
- The owner repository links a 3.69 GB archive containing 228 recordings from 118 native Italian speakers: 112 read-speech recordings and 116 spontaneous interview recordings, including 874 segmented interview clips, speaker metadata, turn timing, an openSMILE configuration, and official five-fold lists. The July 2026 voice-concept-bottleneck paper evaluates the 116-speaker interview subset with the official five-fold protocol. The helper saves owner documentation, repository metadata, and the primary Interspeech paper by default; archive download requires explicit acknowledgment of the restrictive terms.
Safe-first helperscripts/download/androids_corpus.sh
Speaker, identity & emotion
ASVspoof 2015
Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015) Database
Safe-first helper
Speaker Verification Anti Spoofing
Spoofed Speech Detection
Synthetic Speech Detection
Unseen Attack Generalization
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata declares Creative Commons Attribution 4.0 International. Retain attribution and review the packaged license and source-speech provenance before redistribution or commercial use.
- Download notes
- The public Edinburgh DataShare release contains genuine speech from 106 speakers and synthetic speech from ten known and unknown text-to-speech and voice-conversion attacks, partitioned into training, development, and evaluation sets. The helper downloads official metadata, README, evaluation plan, summary paper, file descriptions, and extraction instructions by default. The approximately 2.1 MB protocol archive is a separate opt-in, and the three-part approximately 24.1 GB WAV archive requires ASVSPOOF2015_DOWNLOAD_AUDIO=1. Section 4.1 of arXiv:2607.21127 uses attacks S3 and S10 for expert calibration.
Safe-first helperscripts/download/asvspoof_2015.sh
Speaker, identity & emotion
ASVspoof 2017 V2
The 2nd Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2017) Database, Version 2
Safe-first helper
Speaker Verification Anti Spoofing
Presentation Attack Detection
Spoofed Speech Detection
Replay Attack Detection
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata explicitly declares Creative Commons Attribution-NonCommercial 4.0 International. The database uses genuine and replayed RedDots speech; retain attribution and review the packaged files and upstream RedDots conditions before use or redistribution.
- Download notes
- The official Version 2 release contains 42 speakers and genuine and replayed RedDots speech recorded across 179 replay sessions in 61 unique room, replay-device, and recording-device configurations. Training, development, and evaluation archives total approximately 1.4 GiB. The helper downloads the DataShare metadata, README, change log, instructions, evaluation plan, and primary papers by default; the approximately 104 KiB protocol archive and speech archives require separate explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2017.sh
Speaker, identity & emotion
ASVspoof 2019
ASVspoof 2019: The 3rd Automatic Speaker Verification Spoofing and Countermeasures Challenge database
Safe-first helper
Speaker Verification Anti Spoofing
Presentation Attack Detection
Spoofed Speech Detection
Synthetic Speech Detection
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- odc-by-1.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata lists Open Data Commons Attribution License and ships the ODC Attribution license text. The license text itself cautions that database contents can have separate rights; ASVspoof 2019 is derived from VCTK, so re-check component terms before redistribution.
- Download notes
- The DataShare record exposes README, license, evaluation plan, paper PDF, and LA/PA archives. The helper downloads small documentation/license files by default; LA is about 7.6 GiB and PA is about 17.7 GiB, so archives are explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2019.sh
Speaker, identity & emotion
ASVspoof 2021
ASVspoof 2021: Logical Access, Physical Access, and Speech Deepfake Challenge databases
Safe-first helper
Speaker Verification Anti Spoofing
Presentation Attack Detection
Spoofed Speech Detection
Synthetic Speech Detection
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_specified
- License caution
- The official ASVspoof 2021 page says the databases are available under an Open Data Commons Attribution Licence. Zenodo currently lists ODC-BY for LA and PA, and ODC-ODbL for DF; re-check active Zenodo metadata and component-source terms before redistribution. The GitHub baseline repository did not expose a detected license via the GitHub API on 2026-07-10.
- Download notes
- The helper downloads the evaluation plan, LA/PA/DF keys and metadata, Zenodo record metadata, and file-map metadata by default. Evaluation speech archives are large: LA is about 7.8 GB, PA is split into seven parts totaling about 47 GB, and DF is split into four parts totaling about 34.5 GB, so speech archives are explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2021.sh
Speaker, identity & emotion
ASVspoof 5
ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
Safe-first helper
Speaker Verification Anti Spoofing
Spoofed Speech Detection
Synthetic Speech Detection
Deepfake Speech Detection
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- odc-by-1.0_database_and_cc-by-4.0_bona-fide_audio
- Code license
- not_specified
- License caution
- The packaged license applies ODC Attribution 1.0 to the database and CC BY 4.0 to bona fide data. ODC-BY explicitly does not grant every right in individual contents; preserve Multilingual LibriSpeech provenance and review privacy, personality, and generated-voice rights. The baseline repository has no detected license.
- Download notes
- The public release contains 182,357 training, 142,134 development, and 681,872 evaluation utterances at 16 kHz from crowdsourced speech by roughly 2,000 speakers. It covers more than 20 spoofing attacks, seven adversarial attacks, codec conditions, countermeasure evaluation, and spoofing-robust speaker verification. The helper downloads official metadata, README, license, evaluation plan, paper page, and baseline README by default. The approximately 19.7 MiB protocol archive is a separate opt-in; the complete approximately 142.3 GB audio release remains on Zenodo and is not downloaded by the helper.
Safe-first helperscripts/download/asvspoof_5.sh
Speech understanding & dialogue
Audio Agent Bench Suite
AudioAgentBench: Evaluating Multi-Turn Voice Agents on Real-World Tasks
Safe-first helper
End To End Speech Dialogue
Spoken Question Answering
Audio Agent Tool Use
Multi Turn State Tracking
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_CC-BY-4.0_and_MIT_metadata
- Code license
- not_specified
- License caution
- The suite-level card declares CC BY 4.0, while each of the six component dataset cards currently declares MIT. No separate license file or evaluation-code repository is linked. Treat this metadata conflict as unresolved and confirm the intended terms with Arcada Labs before redistribution or commercial reuse.
- Download notes
- The official public suite comprises six English, multi-turn domains: conference assistance (75 turns), laptop sales (31), grocery ordering (30), dental appointments (25), event planning (29), and personal assistance (31), for 221 scripted turns total. Each domain releases user audio, transcripts, reference responses, knowledge-base context, tool schemas, expected function calls, and scoring labels. The card says two consenting voice actors recorded the inputs. The helper saves the suite and component cards plus API metadata by default; downloading all six snapshots (about 209 MB of current repository storage) requires AUDIO_AGENT_BENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audio_agent_bench_suite.sh
Enhancement, separation & quality
Audio-Alpaca
Audio-Alpaca: A Preference Dataset for Aligning Text-to-Audio Models
Safe-first helper
Text To Audio Preference Modeling
Direct Preference Optimization
Text Audio Alignment
Temporal Event Alignment
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_apache-2.0_and_cc-by-nc-nd-4.0
- Code license
- cc-by-nc-nd-4.0
- License caution
- The Hugging Face dataset card declares Apache-2.0, while the linked official Tango repository applies CC BY-NC-ND 4.0 and does not explain whether that license excludes the dataset. Apply the more restrictive interpretation until the authors clarify scope. AudioCaps-derived captions, generated-audio/model terms, and any source-video rights also require separate review; neither license signal should be assumed to clear all upstream material.
- Download notes
- The public, ungated release contains 15,025 English prompt, chosen-audio, rejected-audio triplets across four construction strategies. Tango 2 generates candidate audio from AudioCaps training captions, perturbed prompts, and varied inference settings, then filters pairs with two CLAP models. Audio-Zero section 3.1 samples and filters 2,000 pairs for post-training but does not release its exact selection. The helper downloads official documentation and API metadata by default; the Hugging Face API reports approximately 9.71 GB of repository storage, so the audio snapshot requires AUDIO_ALPACA_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audio_alpaca.sh
Speech recognition
AudioBench
AudioBench: A Universal Benchmark for Audio Large Language Models
Safe-first helper
Audio Language Model Evaluation
Automatic Speech Recognition
Speech Translation
Spoken Question Answering
+4 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_upstream_terms
- Code license
- custom_noncommercial_unspecified_version
- License caution
- The paper introduces a suite of 8 tasks and 26 datasets, including 7 newly adapted or collected sets, while the maintained repository now supports more than 50 dataset configurations. The repository license file only says "Creative Commons NonCommercial" without a version, and its README states that each dataset remains under its respective license. Review every selected corpus and derived evaluation set before redistribution or commercial use.
- Download notes
- The helper downloads the official README, supported-dataset inventory, repository metadata, and license notice by default. Cloning the evaluation toolkit is opt-in because the repository is about 64 MB before Git history. It does not download the many upstream audio corpora, whose separate access paths and terms still apply.
Safe-first helperscripts/download/audiobench.sh
Enhancement, separation & quality
Audiobook Narration Appeal
Audio-Based Understanding of Audiobook Narration Appeal
Safe-first helper
Audiobook Appeal Prediction
Narration Quality Analysis
Paralinguistic Feature Analysis
Genre Conditioned Audio Modeling
+1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- Apache-2.0
- License caution
- The repository places its released CSV, supplementary material, and code under Apache-2.0. The CSV derives metadata and engagement counts from LibriVox and the Internet Archive; those services' terms and attribution requirements remain applicable. The license does not extend to separately hosted audiobook recordings or underlying texts. Verify per-item public-domain status in the intended jurisdiction and preserve source attribution. Proprietary Spotify engagement data is described only in aggregate and is not released.
- Download notes
- The public, ungated Spotify Research release contains one metadata row for each of 8,854 single-narrator English LibriVox audiobooks, covering 1,206 narrators and 65 genres. Fields include LibriVox and Internet Archive URLs, title grouping, narrator identifier, duration, genre, views, favorites, reviews, and days since publication. The paper uses time-normalized Internet Archive view rate as a noisy public proxy for narration appeal, evaluates global and genre-specific prediction, and ranks alternative narrations of the same text. The helper downloads the approximately 3.1 MB CSV, official documentation, license, paper, and supplementary material. Audiobook audio is not redistributed; users follow the released source URLs to public-domain LibriVox recordings. The paper's separate Spotify engagement analysis is proprietary and is not part of the public dataset.
Safe-first helperscripts/download/audiobook_narration_appeal.sh
Audio understanding, generation & events
AudioCaps
AudioCaps: Generating Captions for Audios in The Wild
Safe-first helper
Audio Captioning
Audio Language Modeling
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- academic_only
- Code license
- MIT
- License caution
- GitHub README says the code and dataset are free to use for academic purposes only. The repository has an MIT license, but the README adds the academic-use condition for repository material; re-check before redistribution or commercial use.
- Download notes
- Official CSVs contain captions, YouTube ids, and segment start times. Raw audio/video download requires the upstream form and is subject to AudioSet/YouTube availability and terms.
Safe-first helperscripts/download/audiocaps.sh
Audio understanding, generation & events
AudioCards / ASFx Eval
AudioCards: Structured Metadata Improves Audio Language Models for Sound Design
Safe-first helper
Structured Audio Captioning
Structured Metadata Generation
Text Audio Retrieval
Sound Effect Classification
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_annotations_with_separate_Adobe_audio_terms
- Code license
- not_specified
- License caution
- Zenodo declares CC BY 4.0 for the released AudioCard CSV. The underlying Adobe sound effects are not included there: Adobe describes them as royalty-free but states that downloading and using them is governed by the Adobe Audition and related-software EULA. The July 2026 paper is CC BY-NC-SA 4.0, but that paper license does not release or license its absent four-field augmentation or perturbation artifacts.
- Download notes
- The public, ungated Zenodo release contains 499 CSV rows with filenames and 13 structured semantic, caption, and UCS fields. The original paper and project page describe 500 manually screened AudioCards, but the released CSV currently has 499 data rows. Pair its filename column with the separately downloaded Adobe Audition Sound Effects library to reproduce ASFx Eval. The July 2026 evaluation-framework paper uses 499 clips. Its Table 1 selects ten released semantic fields and adds five computed acoustic targets (LUFS, pitch, onset, offset, and a frequency profile), but says that augmented dataset "will" be released; those five acoustic annotations are not part of the current Zenodo file. The helper safely downloads the record metadata, project page, papers, Adobe landing page, and approximately 323 KB annotation CSV; it does not download the Adobe audio archives.
Safe-first helperscripts/download/audiocards.sh
Audio understanding, generation & events
AudioGrounding
Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events
Safe-first helper
Text To Audio Grounding
Temporal Audio Grounding
Sound Event Localization
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_on_Zenodo_with_upstream_terms
- Code license
- MIT
- License caution
- Zenodo declares CC BY 4.0 for the release and the official repository is MIT. The recordings derive from AudioCaps and AudioSet YouTube clips, so retain record attribution and review source-video rights, availability, and platform terms before redistribution or commercial use.
- Download notes
- The public, ungated v2 release contains 3,994 training, 488 validation, and 492 test clips derived from AudioCaps/AudioSet, with caption phrases aligned to one or more onset-offset intervals. The July 2026 GigaChat Audio report evaluates the combined 980 validation/test samples as a short-clip temporal-grounding benchmark and reports mIoU. The helper downloads the official repository documentation, Zenodo metadata, and approximately 5.2 MB of JSON annotations by default; the approximately 2.33 GiB audio archive requires explicit opt-in.
Safe-first helperscripts/download/audiogrounding.sh
Speaker, identity & emotion
AudioMarkBench
AudioMarkBench: Benchmarking Robustness of Audio Watermarking
Safe-first helper
Audio Watermarking Robustness
Watermark Removal Detection
Watermark Forgery Detection
Adversarial Audio Robustness
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_upstream_terms_not_separately_specified
- Code license
- MPL-2.0
- License caution
- The repository's MPL-2.0 license covers its source code, but neither the README nor paper states a separate license for the released original, watermarked, or perturbed audio. AudioMarkData derives from Common Voice and the second corpus derives from CC BY 4.0 LibriSpeech; applicable Common Voice release terms, attribution requirements, speaker/privacy considerations, and rights in generated derivatives must be reviewed before reuse or redistribution.
- Download notes
- The public release evaluates AudioSeal/AudioSeal-B, Timbre, and WavMark against 12 no-box perturbation categories plus black-box and white-box adversarial attacks. AudioMarkData contains 20,000 five-second, 16 kHz Common Voice samples balanced for 25 languages, two reported biological-sex groups, and four age groups; the paper also samples 20,000 clips from LibriSpeech. The Drive folder releases original, watermarked, and perturbed audio, while the GitHub repository releases attack and evaluation code. The helper downloads official documentation, license text, repository metadata, and the paper by default; cloning the approximately 3.1 MB GitHub repository requires AUDIOMARKBENCH_CLONE_REPO=1. Google Drive audio remains a manual download so users can inspect its contents and upstream terms.
Safe-first helperscripts/download/audiomarkbench.sh
Speech understanding & dialogue
AudioMNIST
AudioMNIST: Exploring Explainable Artificial Intelligence for audio analysis on a simple benchmark
Safe-first helper
Spoken Digit Classification
Audio Classification
Speaker Metadata Analysis
Explainable Audio Ai
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT
- Code license
- MIT
- License caution
- The repository-level LICENSE is MIT and GitHub API reports MIT. Confirm whether downstream use of recorded voices raises consent/privacy obligations beyond the code/data license.
- Download notes
- The repository contains about 30,000 spoken-digit WAV files from 60 speakers plus speaker metadata and Caffe examples. Because the GitHub repository is large, the helper downloads README/LICENSE metadata by default and requires AUDIO_MNIST_DOWNLOAD_REPO=1 before cloning the full repository.
Safe-first helperscripts/download/audio_mnist.sh
Audio understanding, generation & events
AudioSet
Audio Set: An ontology and human-labeled dataset for audio events
Safe-first helper
Audio Event Classification
Audio Tagging
Sound Event Detection
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- AudioSet dataset annotations/features are CC BY 4.0; the ontology is CC BY-SA 4.0. Original YouTube media remains subject to upstream availability and terms.
- Download notes
- Official release provides segment CSVs and precomputed 128-dimensional audio features; it does not redistribute original YouTube audio.
Safe-first helperscripts/download/audioset.sh
Audio understanding, generation & events
AudioSetCaps
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Safe-first helper
Audio Captioning
Audio Text Retrieval
Audio Language Pretraining
Synthetic Audio Question Answering
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- academic_research_only
- Code license
- not_specified
- License caution
- The Hugging Face card metadata labels the release CC BY 4.0, but the same official card explicitly allows only academic and research use. Apply the stricter research-only statement pending clarification. The GitHub repository has no detected license, and AudioSet, YouTube-8M, and VGGSound source-media rights and platform terms still apply.
- Download notes
- The public, ungated release provides synthetic captions for 6,117,099 ten-second clips sourced from AudioSet, YouTube-8M, and VGGSound, plus 18,414,789 intermediate question-answer pairs. The Hugging Face repository contains caption/Q&A CSVs rather than the full source audio and currently reports about 20.2 GB of storage. The helper downloads only official documentation and repository metadata by default; the large CSV files require AUDIOSETCAPS_DOWNLOAD_METADATA=1. The maintainers warn that AudioCaps and VGGSound evaluation examples overlap the release and should be filtered before training.
Safe-first helperscripts/download/audiosetcaps.sh
Audiovisual & cross-modal
AV-SpeakerBench
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
Safe-first helper
Audio Visual Question Answering
Speaker Centric Audio Visual Reasoning
Audio Visual Temporal Grounding
Speech Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official project page, GitHub README, and Hugging Face card prose state CC BY-NC 4.0, while the Hugging Face card front matter incorrectly or inconsistently declares MIT. Use the more restrictive CC BY-NC 4.0 terms and re-check upstream source-video rights before redistribution or commercial use. GitHub reports no detected repository license.
- Download notes
- The public, ungated release contains 3,212 English multiple-choice questions plus aligned audio-only, visual-only, and audiovisual clips. The helper downloads official documentation by default; the Hugging Face API reports about 123 GB of repository storage, so the full snapshot requires AV_SPEAKERBENCH_DOWNLOAD_HF=1. Qwen3.5-Omni reports the benchmark in section 5.1.4, Table 7.
Safe-first helperscripts/download/av_speakerbench.sh
Audiovisual & cross-modal
AVA Active Speaker
AVA Active Speaker: An Audio-Visual Dataset for Active Speaker Detection
Safe-first helper
Active Speaker Detection
Audio Visual Speaker Localization
Speaking Activity Detection
Face Track Classification
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC BY 4.0
- Code license
- not_specified
- License caution
- Google's AVA download page states that all datasets listed there are CC BY 4.0. The GitHub repository has no detected license, so that statement should not be assumed to license repository code. Preserve attribution and review source-movie and hosting terms before redistributing media; video availability can change independently of the released annotations.
- Download notes
- The official v1.0 release associates visible face tracks with SPEAKING_AND_AUDIBLE, SPEAKING_BUT_NOT_AUDIBLE, or NOT_SPEAKING labels. Google reports 3.65 million labeled frames across approximately 39,000 face tracks; the current download page says dense labels cover 160 AVA movie clips that remained available on YouTube. The helper downloads official documentation and the 2.5 KB video-name manifest by default. Set AVA_ACTIVE_SPEAKER_DOWNLOAD_LABELS=1 for the approximately 23 MB train/validation annotation archives. It does not fetch the much larger source videos.
Safe-first helperscripts/download/ava_active_speaker.sh
Audiovisual & cross-modal
AVDC
AVDC: Audio-Visual Decoupled Captions
Safe-first helper
Audio Visual Decoupled Captioning
Audio Only Captioning
Visual Only Captioning
Joint Audio Visual Captioning
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The Hugging Face card has no license field, and neither the dataset repository nor the official GitHub repository exposes a license file. The arXiv publication license does not license the released annotations, generated captions or reasoning traces, code, or source videos. ShareGPT4Video, Vript, and each source platform or media owner retain applicable terms; obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated Hugging Face release provides avdc_caption.json with decoupled audio, visual, and joint captions and omni_qa.json with questions, answers, reasoning steps, and split labels. The paper describes 10,000 long-form videos, 10,000 caption-derived QA pairs, and an AVDC-test split for visible versus invisible sound-event evaluation. The current JSON blobs total approximately 134 MiB, so the helper downloads only official documentation and API metadata by default and requires AVDC_DOWNLOAD_HF=1 for the annotation snapshot. Source videos are referenced by video ID but are not redistributed in the Hugging Face release; users must obtain applicable ShareGPT4Video- and Vript-sourced media separately. The training/evaluation repository is an additional AVDC_CLONE_REPO=1 opt-in.
Safe-first helperscripts/download/avdc.sh
Audiovisual & cross-modal
AVE
Audio-Visual Event Localization in Unconstrained Videos
Safe-first helper
Audio Visual Event Localization
Audio Visual Event Classification
Cross Modal Localization
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository does not expose a detected GitHub license and the README does not state a standalone dataset license. AVE is built from unconstrained videos, so source-video copyright, platform terms, and redistribution rights should be checked before use.
- Download notes
- The helper downloads the project page and official README by default and can clone the code repository. The dataset, precomputed audio features, and visual features are linked from Google Drive; download those manually or with a user-selected Drive tool after reviewing terms.
Safe-first helperscripts/download/ave.sh
Audiovisual & cross-modal
AVE-Compass
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Safe-first helper
Instruction Based Audio Video Editing
Joint Audio Visual Editing
Speech Editing
Audio Only Editing
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for the benchmark release. The evaluation repository has no LICENSE file or GitHub-detected license. Source-video filenames identify a mixture that includes web-video-derived clips, so users should retain attribution and review uploader, platform, privacy, and underlying media rights before redistribution or derivative use.
- Download notes
- The public, ungated release contains 145 curated source videos, 196 human-verified English editing instructions, 196 checklist JSON files with 2,688 fine-grained items, and 28 editing operation types across joint, speech, video-only, and audio-only branches. The helper downloads official documentation and repository metadata by default. Set AVE_COMPASS_DOWNLOAD_METADATA=1 for the lightweight instruction, checklist, and Dataset Viewer metadata files, or AVE_COMPASS_DOWNLOAD_HF=1 for the complete approximately 442 MB Hugging Face snapshot. The official project currently provides a citation but no public paper URL.
Safe-first helperscripts/download/ave_compass.sh
Audiovisual & cross-modal
AVQA
AVQA: A Dataset for Audio-Visual Question Answering on Videos
Manual or gated
Audio Visual Question Answering
Multimodal Scene Understanding
Cross Modal Reasoning
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- noncommercial_or_permission_required
- Code license
- not_specified
- License caution
- The official page permits personal or classroom copying without fee only when it is not for profit or commercial advantage, and requires permission for broader copying, reposting, or redistribution. The repository has no license file or GitHub-detected license. Raw clips derive from VGGSound/YouTube, so source-media rights and availability also apply.
- Download notes
- The helper saves the official project page, repository README, and GitHub metadata, then prints the official OneDrive/Baidu download paths. It does not automate the combined archive or source-video retrieval. The public release provides QA annotations, a VGGSound-derived video manifest, raw-video access pointers, and large pre-extracted features.
Safe-first helperscripts/download/avqa.sh
Audiovisual & cross-modal
AVSBench
AVSBench: Audio-Visual Segmentation Benchmark
Manual or gated
Audio Visual Segmentation
Sounding Object Segmentation
Single Sound Source Segmentation
Multiple Sound Source Segmentation
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- Apache-2.0
- License caution
- The official project page licenses the AVSBench dataset published there under CC BY-NC 4.0, and the repository licenses the project under Apache-2.0. Source videos were collected from public YouTube material, so uploader rights, platform terms, and current video availability still apply; confirm whether the updated semantic release carries any additional terms during application.
- Download notes
- The benchmark covers semi-supervised Single Sound Source Segmentation (S4), fully supervised Multiple Sound Source Segmentation (MS3), and the later semantic-label AVSS task. The official project page publicly links the original AVSBench-object video-ID CSV and segmentation maps on Google Drive, while processed video/audio requires an email request. The repository directs users to the official application page for the updated object and semantic datasets. The helper saves lightweight official documentation, repository metadata, and license text, then prints these manual access paths; it does not automate Drive, email, or source-video retrieval.
Safe-first helperscripts/download/avsbench.sh
Audiovisual & cross-modal
AVSCapBench
AVSCapBench: Fine-Grained Audio-Visual Synergy Evaluation for Omni-Modal Video Captioning
Safe-first helper
Omni Modal Video Captioning
Audio Visual Captioning
Audio Event Captioning
Audio Visual Event Binding
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card and repository README declare CC BY-NC-SA 4.0, while the paper describes an academic-research-only restrictive release. The paper says clips come from YouTube, TikTok, and Video-MME, invokes fair use, and says only public URLs and timestamps are distributed, but the current Hugging Face repository lists 1,226 MP4 files. Treat the stricter academic/non-commercial interpretation as controlling and review source-platform, uploader, Video-MME, privacy, and copyright terms before downloading, redistributing, or publishing clips. The evaluation repository has no detected license.
- Download notes
- The public, ungated release contains 1,226 English video clips lasting 30 to 120 seconds, dense omni-modal captions, visual events, audio events separated into speech, music, and sound effects, and synergistic audio-visual events. The evaluation repository implements LLM-judged event recall. The helper downloads official documentation and repository metadata by default; set AVSCAPBENCH_DOWNLOAD_HF=1 for the approximately 19.6 GB Hugging Face snapshot.
Safe-first helperscripts/download/avscapbench.sh
Audiovisual & cross-modal
AVSD
Audio Visual Scene-Aware Dialog Dataset
Manual or gated
Audio Visual Dialogue
Video Question Answering
Multimodal Response Generation
Scene Understanding
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- unclear
- Code license
- MIT
- License caution
- The official repository is MIT-licensed, but it does not separately state that the MIT license covers the Google Drive dataset or underlying Charades videos. Treat dialog annotations and media rights conservatively, review any terms shown by the Drive/Charades access paths, and preserve source-video provenance before redistribution or commercial use.
- Download notes
- The CVPR paper introduces dialogs and final summaries for more than 11,000 Charades videos. The official DSTC7 repository reports 7,659 training, 1,787 validation, and 1,710 test dialogs and links the released challenge data through Google Drive. The helper saves the public repository documentation, license, and CVPR paper page, then prints the manual dataset and Charades media paths; it does not automate Google Drive or raw-video retrieval.
Safe-first helperscripts/download/avsd.sh
Audiovisual & cross-modal
AVUT
Audio-centric Video Understanding Benchmark without Text Shortcut
Safe-first helper
Audio Visual Question Answering
Audio Content Understanding
Audio Visual Alignment
Audio Event Localization
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the official repository nor Hugging Face card states a data or code license. The paper says AVUT contains only links to public YouTube videos and does not host or distribute video copies, while the current Hugging Face repository appears to include video files; users should review YouTube terms, source-video rights, and this discrepancy before downloading or redistributing media.
- Download notes
- The public, ungated release covers 2,662 English YouTube videos across 18 domains and 11,609 question-answer pairs in AV-Human and AV-Gemini. The helper downloads official documentation and four lightweight annotation JSON files by default; the Hugging Face API reports about 24.0 GB of repository storage, so the full snapshot requires AVUT_DOWNLOAD_HF=1. Qwen3.5-Omni reports AVUT in section 5.1.4, Table 7.
Safe-first helperscripts/download/avut.sh
Audiovisual & cross-modal
BAH
BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change
Manual or gated
Ambivalence Hesitancy Recognition
Multimodal Affect Recognition
Vocal Expression Analysis
Audio Visual Behavior Understanding
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- proprietary_research_only_eula
- Code license
- BSD-3-Clause
- License caution
- The ÉTS owner page expressly labels BAH proprietary and research-only. The current request process is limited to full-time faculty at an eligible university, higher-education institution, or equivalent organization; students and postdoctoral researchers cannot apply directly. The public repository's BSD-3-Clause license applies to code, not the gated recordings, transcripts, annotations, faces, or participant metadata. Review the signed EULA for storage, sharing, publication, retention, privacy, and downstream-use obligations.
- Download notes
- The release contains 1,427 videos totaling 10.60 hours from 300 participants across Canada, including 1.8 hours of annotated ambivalence/hesitancy moments. It provides raw videos, 16 kHz audio, timestamped transcripts, cropped and aligned faces, expert video- and frame-level labels and cues, participant metadata, and predefined participant-disjoint splits. Access is manual: an eligible full-time faculty member must submit the official form, list every team member, certify institutional eligibility, and sign the EULA. The helper saves only public owner, paper, repository, challenge, and request documentation and never downloads participant data.
Safe-first helperscripts/download/bah.sh
Audio understanding, generation & events
Big Bench Audio
Artificial Analysis Big Bench Audio
Safe-first helper
Spoken Question Answering
Speech Reasoning
Audio Question Answering
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares MIT and the Xiaomi MiMo evaluator is Apache-2.0. The 1,000 English recordings contain verbatim questions from four BIG-Bench Hard tasks and were synthesized with 23 OpenAI-, Azure-, and AWS-provided voices; review inherited task terms and provider-generated-audio conditions rather than assuming the card resolves every component right.
Safe-first helperscripts/download/big_bench_audio.sh
Enhancement, separation & quality
BVCC
BVCC: VoiceMOS Challenge 2022 Main-Track Dataset
Safe-first helper
Mean Opinion Score Prediction
Synthetic Speech Naturalness Assessment
Voice Conversion Quality Assessment
Text To Speech Quality Assessment
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_open_mixed_upstream_terms
- Code license
- not_separately_specified
- License caution
- Zenodo labels the release Other (Open), not a standard reusable data license. Its record explicitly prohibits redistribution of Blizzard Challenge samples and omits those files; Voice Conversion Challenge and ESPnet-TTS components retain their own terms. Treat ratings, metadata, audio, and scripts according to their component provenance, and do not infer broad redistribution or commercial rights from public access.
- Download notes
- The public VoiceMOS Challenge 2022 release provides unified MOS ratings and official train, development, and test splits for synthetic speech from past Voice Conversion and Blizzard Challenges plus ESPnet-TTS. The helper downloads official pages and Zenodo metadata by default. The approximately 273.4 MiB main-track archive, small out-of-domain package, and scoring package are separate opt-ins. Blizzard audio is intentionally absent from the archive; official scripts require users to obtain and preprocess that material under its original access terms. Section 4.2 of arXiv:2607.13477 constructs 60 BVCC test-set pairs with a human-MOS gap of at least 1.0 for its naturalness probe, but does not release the selected pair manifest.
Safe-first helperscripts/download/bvcc.sh
Speech generation
CapSpeech
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Safe-first helper
Style Captioned Text To Speech
Text To Speech With Sound Effects
Accent Captioned Text To Speech
Emotion Captioned Text To Speech
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_mixed_upstream_audio_terms
- Code license
- cc-by-nc-4.0
- License caution
- The dataset card, paper, and repository apply CC BY-NC 4.0 to CapSpeech resources. The release points to audio from Emilia, GigaSpeech, Common Voice, MLS, LibriTTS-R, VoxCeleb, EARS, Expresso, VCTK, VGGSound, FSDKaggle2018, ESC-50, and separate CapSpeech audio repositories; those recordings retain their own attribution, non-commercial, access, privacy, and media-rights constraints. Treat the CapSpeech license as covering its annotations and author contributions, not as overriding component-source terms.
- Download notes
- The public, ungated release contains more than 10 million machine-annotated and approximately 360,000 human-annotated English audio-caption records, with fixed pretraining and supervised-fine-tuning train, validation, and test splits across CapTTS, CapTTS-SE, AccCapTTS, EmoCapTTS, and AgentTTS. The main Hugging Face snapshot contains paths, transcripts, source labels, durations, and style captions rather than embedded source audio; the API reports approximately 4.31 GB compressed and 10.09 GB after processing. The helper downloads official documentation and API metadata by default, while the full metadata snapshot and code repository are separate opt-ins. The 2026 ProPS paper trains and evaluates prompt-conditioned speaker-profile distributions on CapSpeech's held-out splits.
Safe-first helperscripts/download/capspeech.sh
Audio understanding, generation & events
CASTELLA
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
Safe-first helper
Audio Moment Retrieval
Temporal Audio Grounding
Long Audio Retrieval
Audio Captioning
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_for_annotations_and_released_features
- Code license
- not_specified_for_audio_downloader
- License caution
- The annotation repository, its Hugging Face card, and the Zenodo feature record declare CC BY 4.0. That license covers the released annotations/features, not the underlying YouTube recordings. The separate raw-audio downloader repository has no detected license, and raw media remains governed by its owners and YouTube terms; verify availability and rights before downloading, redistribution, or commercial use.
- Download notes
- The public, ungated annotation release describes 1,862 real-world YouTube recordings split into 1,009 training, 213 validation, and 640 test items, with 3,925 human-written local captions and 11,308 temporal boundaries in English and Japanese. The Hugging Face mirror contains only the six lightweight annotation JSON files, not raw audio. The helper downloads those annotations plus official documentation and metadata by default. Precomputed MS-CLAP audio/text features are an approximately 2.78 GB Hugging Face opt-in (the Zenodo release is about 1.33 GB). Raw media must be reconstructed from YouTube IDs with the separate official downloader, subject to current availability and source-platform and recording rights; the helper only clones those tools when explicitly requested.
Safe-first helperscripts/download/castella.sh
Audiovisual & cross-modal
CH-SIMS
CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of Modality
Safe-first helper
Multimodal Sentiment Analysis
Audio Sentiment Analysis
Visual Sentiment Analysis
Text Sentiment Analysis
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The MMSA repository is MIT-licensed, but neither the paper nor the repository README expressly applies that license to the CH-SIMS video and annotation files in the shared Drive folders. Treat dataset terms and source-media rights as unspecified and verify permitted use before redistribution or commercial use.
- Download notes
- The paper introduces 2,281 Chinese in-the-wild video segments with multimodal sentiment labels and separate text, audio, and visual annotations. The official MMSA repository provides shared Baidu and Google Drive folders containing raw video, processed features, and labels for CH-SIMS alongside MOSI and MOSEI. The helper downloads official paper, README, license, and repository metadata only; dataset files remain a manual Drive download, and cloning the toolkit is opt-in.
Safe-first helperscripts/download/ch_sims.sh
Audiovisual & cross-modal
CH-SIMS v2
CH-SIMS v2.0: A Fine-grained Multi-label Chinese Multimodal Sentiment Analysis Dataset
Safe-first helper
Multimodal Sentiment Analysis
Audio Sentiment Analysis
Visual Sentiment Analysis
Text Sentiment Analysis
+2 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official project page, paper, and repository state no dataset or code license, and the repository has no LICENSE file. Treat the videos, annotations, features, and code as all-rights-reserved unless the authors provide terms; also review source-media, speaker, privacy, and platform rights before reuse or redistribution.
- Download notes
- The official release extends and re-annotates CH-SIMS with 4,402 supervised segments carrying multimodal and unimodal sentiment labels plus 10,161 unlabeled segments for semi-supervised evaluation. It provides raw videos, extracted features, IDs, splits, and labels through separate Google Drive and Baidu folders. The helper downloads only the official homepage, repository README/API metadata, and arXiv metadata; Drive data remains manual and the code clone is opt-in.
Safe-first helperscripts/download/ch_sims_v2.sh
Music
ChartGenEval
ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
Safe-first helper
Music Generation Evaluation
Rhythm Game Chart Generation
Chart Audio Alignment
Chart Structure Evaluation
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT_released_artifacts_corpus_not_released
- Code license
- Not specified in the source record.
- License caution
- The repository-level MIT license covers the released software and bundled artifacts. It does not grant rights to the absent chart/audio corpus; the README says potentially copyrighted community and commercial content is intentionally not redistributed. Users reproducing corpus-dependent experiments must supply a rights-compatible equivalent snapshot and review its chart, recording, and game-content terms.
- Download notes
- The public repository releases the NumPy-based evaluation toolkit, calibration artifact, corruption probes, plotting and reproduction scripts, and metric-only evaluation records. The paper validates seven output axes with nine held-out controlled-corruption tests across timing, audio response, structure, grammar, and human-reference gap measurements. The helper downloads official documentation and repository metadata by default; cloning the approximately 44 MB current file tree requires CHARTGENEVAL_CLONE_REPO=1. The underlying 3,880-chart calibration corpus and its audio are explicitly not distributed because they may contain copyrighted community and commercial material, so the released artifacts support verification of reported results but not full corpus-dependent reproduction.
Safe-first helperscripts/download/chartgeneval.sh
Speech recognition
CHILDES-Aligned
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling
Manual or gated
Asr
Child Speech Recognition
Long Form Speech Alignment
Forced Alignment
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0_with_talkbank_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC-SA 4.0 and its access agreement additionally limits use to non-commercial research, requires citation of the BEACON paper and every source CHILDES corpus used, incorporates the TalkBank Ground Rules, and prohibits audio redistribution. Access is manually reviewed. The paper's linked BEACON GitHub repository was not publicly reachable when checked, so no code license is claimed.
- Download notes
- The manually gated Hugging Face release contains a 413.3-hour general-purpose English child-speech configuration with corrected utterance timestamps and a quality-controlled 283-hour ASR-training configuration. The repository reports approximately 160.6 GB of storage. The helper prints the access steps by default and downloads a selected configuration only after the user has received access, authenticated with Hugging Face, and set CHILDES_ALIGNED_ACK_TERMS=1.
Safe-first helperscripts/download/childes_aligned.sh
Speech recognition
CHiME-6
CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings
Safe-first helper
Distant Speech Recognition
Multi Speaker Asr
Speaker Diarization
Speech Separation
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR150 lists CC BY-SA 4.0. CHiME says CHiME-6 is a corrected-alignment version of CHiME-5 and recommends CHiME-6 for new work.
- Download notes
- The helper downloads OpenSLR transcriptions, floorplans, and license by default. Audio archives are large, about 97 GiB train, 11 GiB dev, and 12 GiB eval, so they are explicit opt-ins.
Safe-first helperscripts/download/chime_6.sh
Speech recognition
CHiME-7 DASR
The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios
Manual or gated
Distant Automatic Speech Recognition
Speaker Attributed Automatic Speech Recognition
Speaker Diarization
Meeting Transcription
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- mixed_manual_agreements
- Code license
- Apache-2.0
- License caution
- The benchmark is a protocol over three separately controlled corpora, not a single uniformly licensed download. Retain the CHiME, DiPCo, and task-specific LDC/Mixer 6 terms independently. The official baseline is part of ESPnet, whose repository is Apache-2.0 licensed.
- Download notes
- CHiME-7 DASR evaluates one system across revised CHiME-6, DiPCo, and challenge-specific Mixer 6 Speech partitions, ranking submissions by macro-averaged diarization-attributed WER across the three scenarios. The official ESPnet recipe generates the task layout, but it can automatically obtain only DiPCo. CHiME-5/CHiME-6 must be obtained through the CHiME license path, and the task's Mixer 6 release requires a separate LDC evaluation agreement; the challenge warns that this Mixer 6 version differs from LDC2013S03. The helper saves only public task, data, paper, and baseline documentation before printing the manual access steps.
Safe-first helperscripts/download/chime_7_dasr.sh
Audio understanding, generation & events
Clotho
Clotho: An Audio Captioning Dataset
Safe-first helper
Audio Captioning
Language Based Audio Retrieval
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_specified
- License caution
- Zenodo lists rights as Other (Attribution). Audio clips keep their original Freesound licenses, mostly Creative Commons with attribution, recorded in metadata CSVs. Captions are under the Tampere University license, mainly non-commercial with attribution.
- Download notes
- Clotho v2.1 audio archives total about 7.1 GiB; the helper downloads captions/metadata by default and makes audio opt-in.
Safe-first helperscripts/download/clotho.sh
Audio understanding, generation & events
Clotho-Moment
Clotho-Moment: Simulated Long-Audio Dataset for Language-Based Audio Moment Retrieval
Safe-first helper
Language Based Audio Moment Retrieval
Temporal Audio Grounding
Long Audio Retrieval
Audio Text Retrieval
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_on_hugging_face_card_with_upstream_terms
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares Apache-2.0 and the Lighthouse repository includes an Apache-2.0 license. The generated audio incorporates Clotho/Freesound foreground clips and Walking Tours/YouTube background audio, so the card does not erase component recording licenses, attribution requirements, or source-platform terms; review packaged provenance before redistribution or commercial use.
- Download notes
- The public, ungated release contains 51,240 one-minute synthetic English recordings split into 37,930 training, 5,741 validation, and 7,569 test samples. Each sample pairs a text query with a temporal boundary for a Clotho foreground event overlaid at a random interval on Walking Tours background audio. DCASE 2026 Task 6 uses it as a development dataset and advertises a 16.1 GB download. The helper downloads official documentation, license, and repository metadata by default; the audio WebDataset snapshot requires explicit opt-in. Hugging Face currently reports approximately 213 GB of repository storage including history, so users should verify available disk space and select only needed splits or shards.
Safe-first helperscripts/download/clotho_moment.sh
Audio understanding, generation & events
ClothoAQA
Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
Safe-first helper
Audio Question Answering
Audio Language Understanding
Machine Listening
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_applicable
- License caution
- Zenodo lists rights as Other (Attribution). The train/validation/test question-answer CSVs are MIT licensed by Tampere University. Audio files keep per-file Freesound licenses, mostly Creative Commons with attribution, recorded in clotho_aqa_metadata.csv.
- Download notes
- The helper downloads the QA split CSVs, metadata, and license by default. The 3.1 GiB audio_files.zip archive is opt-in because it contains Clotho/Freesound-derived audio.
Safe-first helperscripts/download/clotho_aqa.sh
Enhancement, separation & quality
CMI-RewardBench / CMI-Pref
CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
Safe-first helper
Music Reward Model Evaluation
Music Preference Prediction
Music Quality Assessment
Text Music Alignment
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The CMI-Pref card and paper declare CC BY-NC-SA 4.0, and the evaluation repository contains Apache-2.0 code. The paper says some audio was generated through commercial APIs and describes a terms-aware release mechanism; generated-output and service terms may still apply. The composite CMI-RewardBench also incorporates PAM, MusicEval, and Music Arena, so their source licenses and music rights remain controlling for those subsets.
- Download notes
- The public, ungated CMI-Pref release contains 4,027 individual human preference votes over generated music, including a balanced 500-vote test split, 133.8 hours of English/Chinese material, and text, lyrics, reference-audio, musicality, alignment, confidence, and anonymized listener fields. CMI-RewardBench combines that test split with PAM, MusicEval, and Music Arena for music reward-model evaluation. The helper downloads the official cards, repository docs/configuration, the approximately 620 KB CMI-Pref test JSONL, and the approximately 4.8 MB composite test manifest by default. The Hugging Face API reports approximately 15.0 GB of repository storage, so all MP3 assets require explicit opt-in. A July 2026 full-song generation report uses CMI-Reward as one evaluator on its separate 500-example multilingual test set; that paper does not release those 500 evaluation inputs.
Safe-first helperscripts/download/cmi_rewardbench.sh
Audiovisual & cross-modal
CMU-MOSEI
CMU-MOSEI: CMU Multimodal Opinion Sentiment and Emotion Intensity
Safe-first helper
Multimodal Sentiment Analysis
Multimodal Emotion Recognition
Audio Sentiment Analysis
Speech Emotion Recognition
+2 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The CMU Multimodal SDK and MultiBench repositories are MIT-licensed, but neither repository expressly applies that license to MOSEI annotations, processed features, or source YouTube media. Treat dataset terms as unspecified and review creator, platform, privacy, and media rights before reuse, redistribution, or commercial use.
- Download notes
- The official SDK describes more than 65 hours of annotated YouTube monologue video from more than 1,000 speakers and 250 topics. Each sentence has a sentiment score and six non-exclusive emotion scores for happiness, sadness, anger, surprise, disgust, and fear. The SDK publishes labels plus processed acoustic, visual, and language computational sequences; MultiBench provides an additional word-aligned processed package through Google Drive. The helper saves official documentation, dataset definitions, and repository metadata only. Cloning either toolkit is opt-in, and the Drive package remains a manual download.
Safe-first helperscripts/download/cmu_mosei.sh
Audiovisual & cross-modal
CMU-MOSI
CMU-MOSI: Multimodal Opinion-level Sentiment Intensity Dataset
Safe-first helper
Multimodal Sentiment Analysis
Audio Sentiment Analysis
Subjectivity Analysis
Audio Visual Sentiment Analysis
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- Both official code repositories are MIT-licensed, but neither license expressly grants rights to the MOSI annotations, processed features, or source media. Raw YouTube videos are not redistributed. Treat dataset terms as unspecified and review creator, platform, privacy, and media rights before reuse or redistribution.
- Download notes
- The original paper introduces 2,199 opinion segments from 93 English YouTube review videos with sentiment-intensity, subjectivity, visual, and acoustic annotations. The official CMU Multimodal SDK publishes labels and anonymized processed acoustic, visual, and language computational sequences, but explicitly does not share raw videos because of YouTube creator privacy. MultiBench provides an additional word-aligned processed release through Google Drive. The helper saves official documentation and repository metadata only; cloning either toolkit is opt-in and the Drive package remains a manual download.
Safe-first helperscripts/download/cmu_mosi.sh
Speech generation
CN-NewsTTS Bench
CN-NewsTTS Bench: A Target-Level Automatic Benchmark for Raw-Input Chinese News TTS Pronunciation
Safe-first helper
Chinese Text To Speech Evaluation
Pronunciation Accuracy Evaluation
Text Normalization Evaluation
Target Level Error Analysis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- MIT
- License caution
- CC BY 4.0 covers benchmark data, fixed ASR transcripts, results, documentation, and metadata; MIT covers repository code. The Zenodo audio consists of generated outputs from seven commercial TTS providers and is published as an evaluation artifact. The maintainers warn that reuse may remain subject to each provider or API's terms, so do not treat those audio archives as unrestricted speech-training data without a separate rights review.
- Download notes
- The public v0.1 release contains 200 development records and 800 public-test records with 1,240 auto-evaluable pronunciation targets, fixed transcripts from a three-ASR ensemble, target-level scoring code, and results for seven TTS products. The helper downloads official documentation, licenses, Zenodo metadata, the two small JSONL benchmark splits, schema, scorer, and checksums by default. The approximately 1.58 MB core archive and 1.72 MB full-transcript archive are separate opt-ins. The approximately 425 MB development-audio and 1.74 GB public-test-audio archives require explicit provider-terms acknowledgement and opt-in.
Safe-first helperscripts/download/cn_news_tts_bench.sh
Speaker, identity & emotion
Codec-SUPERB
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
Safe-first helper
Neural Audio Codec Evaluation
Speech Reconstruction
Audio Reconstruction
Music Reconstruction
+3 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The Hugging Face card has no license field or provenance/rights statement. The repository README and badge call the project MIT, but the linked LICENSE file is absent and the GitHub API detects no license. Treat both data and code terms as unspecified until the maintainers publish authoritative license text, and verify the source field and upstream audio rights before reuse.
- Download notes
- The benchmark evaluates whether neural codecs preserve content, paralinguistics, speaker identity, and general audio information through downstream and signal-level metrics. The current official repository uses the public, ungated codec-superb-tiny release for regression runs: 6,000 rows split evenly across speech, audio, and music, with approximately 3.2 GB of downloads. The helper downloads official documentation by default; the dataset snapshot and repository clone are separate opt-ins.
Safe-first helperscripts/download/codec_superb.sh
Speech understanding & dialogue
CoDeTT
CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
Safe-first helper
Turn Taking
Full Duplex Dialogue
Context Aware Turn Decision
Spoken Dialogue Intent Classification
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_specified
- License caution
- The Hugging Face card declares Apache-2.0 for the dataset, but the GitHub repository has no LICENSE file or detected license. The paper says real samples come from Candor and MagicData-RAMC and synthetic speech uses references from KeSpeech and Emilia; verify those upstream corpus and voice-data terms before redistribution or commercial use.
- Download notes
- The public, ungated release contains more than 300 hours of English and Chinese multi-turn dialogue for four turn-taking actions and 14 fine-grained intent scenarios across system-speaking and system-idle states. It mixes synthetic material with real conversational samples derived from Candor and MagicData-RAMC. The helper downloads official documentation and API metadata by default. The Hugging Face API reports approximately 51.1 GB of repository storage, so the single CoDeTT.lz4 archive requires CODETT_DOWNLOAD_HF=1.
Safe-first helperscripts/download/codett.sh
Audiovisual & cross-modal
CoMind
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
Safe-first helper
Audio Visual Social Reasoning
Joint Attention Estimation
Socially Conditioned Object Interaction Anticipation
Collaborative Handover Prediction
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official project page declares CC BY-NC 4.0, although its displayed license and terms links are currently placeholder anchors rather than a separate terms document. Treat the data as non-commercial, preserve attribution, and review privacy, voice, face, gaze, biometric, and participant-consent implications before reuse. The paper states that participants consented to release of identifiable video, but that does not remove downstream ethical obligations. No separate license notice is embedded in the official Python downloader; the paper itself uses arXiv's perpetual non-exclusive license.
- Download notes
- The official public release covers 41 hours of unscripted cooking collaboration across 80 sessions, with two synchronized egocentric cameras, two exocentric views, audio and WhisperX transcripts, gaze, hand tracking, camera trajectories, scene/object scans, and annotations. Its three benchmarks accept audio or transcribed speech from a ten-second context window for joint-attention estimation, socially conditioned object-interaction anticipation, and collaborative handover prediction. The helper saves the official page, paper metadata, first-party downloader, and annotation manifest by default. The approximately 5.0 MiB annotation JSON files are opt-in; large recording components remain available through the saved official downloader and are never fetched automatically.
Safe-first helperscripts/download/comind.sh
Speech recognition
Common Voice
Mozilla Common Voice
Manual or gated
Asr
Access pathOfficial / other
Upstream termsOpen / attribution signals
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC0-1.0
- Code license
- MPL-2.0
- License caution
- Common Voice data is CC0-1.0; cv-dataset metadata repo is MPL-2.0.
- Download notes
- The helper requires a per-release, per-language URL generated by Mozilla Data Collective and never stores credentials or a private generated URL. Set COMMON_VOICE_FILENAME when the signed URL does not expose a useful archive name.
Safe-first helperscripts/download/common_voice.sh
Audio understanding, generation & events
Concerto Accompaniment Benchmark
Concerto Accompaniment Benchmark for Score-Free Piano Concerto Accompaniment
Safe-first helper
Music Audio Alignment
Automatic Accompaniment
Score Free Accompaniment Generation
Downbeat Alignment
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- MIT covers the public repository's code and released annotation/config files. It does not license the absent Music Minus One recordings or override the per-recording IMSLP terms listed in AudioDataSummary.csv, which include CC0, several CC variants, and public-domain status that may differ by jurisdiction. The repository does not state a license or public delivery path for the recorded solo-piano performances.
- Download notes
- The paper defines 150 alignment scenarios over four concerto movements, combining four recorded solo-piano performances, four commercial Music Minus One orchestra tracks, and eight IMSLP piano-orchestra mixes. The public repository provides the evaluation code, configuration tables, IMSLP source URLs, and measure-downbeat annotations. The helper saves those lightweight released artifacts by default and makes the repository clone opt-in. The commercial orchestra recordings are explicitly private and must be purchased separately. Although the paper calls the remaining data open source, the repository currently ignores audio and exposes no solo-piano recordings; do not infer a public audio download.
Safe-first helperscripts/download/concerto_accompaniment_benchmark.sh
Speech recognition
CoVoST 2
CoVoST 2: Massively Multilingual Speech-to-Text Translation Corpus
Safe-first helper
Speech To Text Translation
Speech Translation
Multilingual Asr
Machine Translation
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Meta/GitHub list CoVoST data as CC0; the current Hugging Face card lists CC BY-NC 4.0, and current Mozilla Data Collective packages for CoVoST 2/Common Voice segments may be CC BY-NC 4.0.
- Code license
- cc-by-nc-4.0
- License caution
- Treat packaged mirrors conservatively and re-check the active source before redistribution. The GitHub license table also says Tatoeba evaluation sentences are CC BY 2.0 FR and Tatoeba speech has per-row licenses.
- Download notes
- The helper downloads the official CoVoST 2 translation TSV archives and split-generation script. CoVoST 2 rows match Common Voice 4 validated.tsv entries, so users must obtain Common Voice audio separately under the applicable upstream terms.
Safe-first helperscripts/download/covost2.sh
Audiovisual & cross-modal
CREMA-D
CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset
Safe-first helper
Speech Emotion Recognition
Audio Visual Emotion Recognition
Acted Emotional Speech
Crowd Sourced Emotion Annotation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- odbl-1.0
- Code license
- not_specified
- License caution
- The official README/LICENSE say the database is under the Open Database License 1.0 and individual contents are under the Database Contents License 1.0. GitHub reports license as NOASSERTION, so keep the explicit upstream text as authority.
- Download notes
- The helper downloads small README/license/CSV metadata by default. Full audio and video live in Git LFS and require about 7.55 GiB for a complete clone; upstream asks repository users to fill out the access/community form.
Safe-first helperscripts/download/crema_d.sh
Speech generation
CV3-Eval
CV3-Eval: CosyVoice 3 in-the-wild zero-shot speech synthesis benchmark
Safe-first helper
Zero Shot Text To Speech
Multilingual Voice Cloning
Cross Lingual Voice Cloning
Emotion Cloning
+6 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0 repository license; mixed upstream media rights
- Code license
- Apache-2.0
- License caution
- The repository root applies Apache-2.0, but the README says reference speech comes from Common Voice, FLEURS, EmoBox, and web-crawled real-world audio. Treat the repository license as insufficient to clear every source recording, and verify component provenance and rights before redistribution or commercial use.
- Download notes
- The official repository includes objective multilingual, cross-lingual, and emotion-cloning subsets plus subjective expressive, continuation, and Chinese-accent subsets. Qwen3.5-Omni section 5.2.3 calls the public cross-lingual subset both CV3-Eval and the Cross-Lingual benchmark, and reports mixed error rate over 12 source-target directions among Chinese, English, Japanese, and Korean. The helper downloads the README and Apache-2.0 license by default; cloning the roughly 760 MiB repository, including evaluation audio and bundled scoring utilities/models, requires CV3_EVAL_CLONE_REPO=1.
Safe-first helperscripts/download/cv3_eval.sh
Audiovisual & cross-modal
Daily-Omni
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
Safe-first helper
Audio Visual Question Answering
Temporal Alignment
Cross Modal Reasoning
Audio Visual Event Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- GPL-3.0
- License caution
- The paper and Hugging Face card declare CC BY-NC-SA 4.0 for the benchmark, and GitHub reports GPL-3.0 for the repository. Videos are sampled from AudioSet, Video-MME, and FineVideo, so their upstream media rights and terms also require review.
- Download notes
- The public, ungated release contains 684 real-world videos and 1,197 English multiple-choice questions across six temporal audio-visual reasoning tasks. The helper downloads official documentation and qa.json by default; the Hugging Face API reports approximately 3.9 GB of storage, so Videos.tar requires DAILY_OMNI_DOWNLOAD_HF=1. Qwen3.5-Omni reports DailyOmni in section 5.1.4, Table 7.
Safe-first helperscripts/download/daily_omni.sh
Audiovisual & cross-modal
DAVE
DAVE: Diagnostic Benchmark for Audio Visual Evaluation
Safe-first helper
Audio Visual Alignment
Multimodal Synchronization
Sound Absence Detection
Sound Discrimination
+3 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT label with upstream dataset terms
- Code license
- MIT stated in README; no repository LICENSE file
- License caution
- The Hugging Face card labels DAVE as MIT and the repository README says everything is MIT, but the GitHub repository has no LICENSE file. DAVE is built on EPIC-KITCHENS and Ego4D, and its card says it inherits their risks; verify both upstream datasets' access and media terms before redistribution or commercial use.
- Download notes
- The public, ungated Hugging Face release has EPIC-KITCHENS- and Ego4D-derived splits with seven diagnostic task views. The helper downloads the official cards, loader, and approximately 9 MB of JSON annotations by default. Media archives are excluded because the Hugging Face API reports about 113.3 GB of repository storage; DAVE_DOWNLOAD_HF=1 explicitly opts into the full snapshot.
Safe-first helperscripts/download/dave.sh
Audio understanding, generation & events
DCASE 2024 Task 5
DCASE 2024 Task 5: Few-shot Bioacoustic Event Detection
Safe-first helper
Few Shot Bioacoustic Event Detection
Sound Event Detection
Animal Vocalization Detection
Five Shot Learning
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Both official Zenodo records declare Creative Commons Attribution 4.0 International. The benchmark combines multiple bioacoustic sources; retain the release attribution and review source-specific ethical or wildlife-recording constraints for downstream use.
- Download notes
- The official five-shot protocol provides the first five positive target events in each recording, then scores detection after the fifth event. The 2024 development release has 217 recordings: 174 training files covering 47 classes and 43 validation files covering seven classes. The official challenge reuses the 2023 evaluation release, which has 66 recordings across eight subsets. The helper downloads record metadata, class maps, and annotation-only archives by default. The current Zenodo files total approximately 20.4 GiB for development audio and 3.0 GiB for evaluation audio, so waveform archives require DCASE2024_TASK5_DOWNLOAD_AUDIO=1.
Safe-first helperscripts/download/dcase2024_task5.sh
Audio understanding, generation & events
DCASE 2024 Task 7 Sound Scene Synthesis
DCASE 2024 Task 7: Sound Scene Synthesis
Safe-first helper
Text To Audio Generation
Environmental Sound Scene Synthesis
Compositional Audio Generation
Audio Generation Quality Evaluation
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo declares CC BY 4.0 for the open-source dataset. The record says its audio is sourced from Freesound, so retain packaged source attribution and review per-clip provenance. The baseline repository has no LICENSE file and the task page states that it is mostly derived from the upstream AudioLDM repository; treat code terms as unspecified pending clarification.
- Download notes
- The public open-source release contains 310 manually composed four-second environmental sound scenes and corresponding structured text prompts. It uses only Freesound source audio and excludes the proprietary/private libraries present in the challenge reference data. The original protocol evaluates Fréchet Audio Distance with PANNs CNN14 Wavegram-Logmel embeddings plus listening tests for foreground fit, background fit, and audio quality. The challenge's 250 evaluation prompts and reference audios remain secret. The helper downloads official metadata and task documentation by default; the approximately 140 MiB public archive is an explicit opt-in.
Safe-first helperscripts/download/dcase2024_sound_scene_synthesis.sh
Enhancement, separation & quality
DCASE 2024 Task 9 LASS
DCASE 2024 Task 9: Language-Queried Audio Source Separation
Safe-first helper
Language Queried Audio Source Separation
Text Conditioned Audio Separation
Universal Sound Separation
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- All three official Zenodo records declare CC BY 4.0. Development audio remains in FSD50K and Clotho v2 under their own mixed or upstream terms, while validation/evaluation audio derives from Freesound; retain record attribution and inspect packaged per-clip provenance. The baseline repository has no LICENSE file or detected GitHub license and is largely derived from AudioSep, so its code terms are unspecified.
- Download notes
- The development release contains GPT-4-generated captions for the existing FSD50K development and evaluation clips; participants obtain the source FSD50K and Clotho v2 audio separately. The public validation release contains 3,000 synthetic mixtures built from 1,000 source clips with three captions per source. The evaluation release contains 3,000 additional synthetic mixtures plus 100 real overlapping Freesound clips annotated with two source queries each. Synthetic examples are scored with SDR; the real set uses listening tests for query relevance and overall quality. The helper downloads official task documentation, record metadata, and lightweight JSON/CSV annotations by default; approximately 1.14 GB of validation and evaluation audio is an explicit opt-in.
Safe-first helperscripts/download/dcase2024_lass.sh
Audio understanding, generation & events
DCASE 2025 Task 2 ASD
DCASE 2025 Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Safe-first helper
Anomalous Sound Detection
Machine Condition Monitoring
Domain Generalization
First Shot Learning
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- custom_dcase_challenge_license_v2.1
- License caution
- All three official Zenodo records declare CC BY-NC-SA 4.0, which prohibits commercial use and requires attribution and share-alike distribution. The evaluator repository includes a DCASE Challenge License v2.1 PDF rather than an SPDX-style open-source license; review it before reuse.
- Download notes
- The public first-shot protocol trains only on normal machine sounds, tests source/target-domain generalization, and uses different machine types for development versus final evaluation. The development record has seven machine types and approximately 2.36 GB of archives; the approximately 1.98 GB additional-training record and 358 MB evaluation record cover eight different types. Each evaluation section has 200 test clips, and the organizers have released labels and an evaluator. The helper downloads official task, record, and evaluator metadata by default; archives require DCASE2025_TASK2_DOWNLOAD_ARCHIVES=1 and an explicit part list.
Safe-first helperscripts/download/dcase2025_task2_asd.sh
Audio understanding, generation & events
DCASE 2025 Task 5 AudioQA
DCASE 2025 Task 5: Multi-Domain Audio Question Answering
Manual or gated
Audio Question Answering
Temporal Audio Reasoning
Bioacoustic Question Answering
Multiple Choice Question Answering
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- mit_on_hugging_face_card
- Code license
- unspecified
- License caution
- The Hugging Face card metadata declares MIT, but the benchmark incorporates audio from Watkins Marine Mammal Sound Database, AudioSet, Mira, and other sources named by the organizers. Upstream audio and source-platform terms may be narrower and still apply; confirm them before redistribution or commercial use. No separate code license was found for the release scripts.
- Download notes
- The official English multiple-choice benchmark combines Bioacoustics QA, Temporal Soundscapes QA, and Complex QA (MMAU). The challenge page reports approximately 8.1K training and 2.4K development question-answer pairs; the released repository also includes the evaluation set. Access is public but auto-approved gated: users must sign in and provide basic identity and affiliation fields. The helper saves public challenge, paper, and repository API metadata, then prints the manual acceptance and authenticated download steps; it never downloads audio automatically.
Safe-first helperscripts/download/dcase2025_audioqa.sh
Audio understanding, generation & events
DCASE 2026 Task 1 HAC
DCASE 2026 Task 1: Heterogeneous Audio Classification
Safe-first helper
Heterogeneous Audio Classification
Hierarchical Audio Classification
Multimodal Audio Classification
Domain Generalization
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- All three official Zenodo records declare CC BY 4.0. The releases include per-sound Freesound license and uploader provenance, so users must also honor each source recording's terms. The official baseline repository has no detected license or LICENSE file; its code terms are unspecified.
- Download notes
- The task predicts 23 second-level Broad Sound Taxonomy categories and scores macro-averaged hierarchical F-score. Development uses the curated BSD10k-v1.2 release (about 11,000 sounds and 35 hours) and the noisier crowd-sourced BSD35k-CS release (about 35,000 sounds and 150 hours), both with text metadata and Freesound provenance. The public evaluation archive contains audio and metadata but intentionally omits labels. The helper downloads official pages, Zenodo records, READMEs, and approximately 7 MB of development metadata by default. Roughly 200 MB of CLAP features and 47 GB of audio/evaluation archives require separate explicit opt-ins.
Safe-first helperscripts/download/dcase2026_task1_hac.sh
Audiovisual & cross-modal
DCASE2025 Task 3 Stereo SELD Dataset
DCASE2025 Task 3 Stereo Sound Event Localization and Detection Dataset
Safe-first helper
Sound Event Localization And Detection
Stereo Sound Source Localization
Audio Visual Sound Event Localization
Source Distance Estimation
+2 more
Access pathZenodo
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- MIT
- Code license
- mixed
- License caution
- Zenodo metadata and the included LICENSE identify the dataset as MIT. Sony's data-generator repository is MIT. The official baseline repository has no detected GitHub license or LICENSE file, so its code terms are not specified. Because the release is derived from STARSS23 recordings of people and rooms, review the original privacy/provenance context before sensitive visual use.
- Download notes
- The public Zenodo v1.1.0 release contains 30,000 labeled development clips (41.7 hours) and 10,000 unlabeled evaluation clips (13.9 hours), each five seconds long, with 24 kHz stereo audio and aligned perspective video. It is derived from STARSS23 by sampling and converting its FOA audio and 360-degree video, and adds folded azimuth, source-distance, and onscreen/offscreen labels. The helper downloads official pages, record metadata, README, license, paper page, and generator documentation by default; the approximately 15.2 MB label archive is opt-in, while the approximately 27.6 GB audio/video release remains on Zenodo.
Safe-first helperscripts/download/dcase2025_stereo_seld.sh
Audio understanding, generation & events
DESED
Domestic Environment Sound Event Detection Dataset
Safe-first helper
Sound Event Detection
Audio Tagging
Weakly Supervised Sound Event Detection
Synthetic Soundscape Generation
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- MIT
- License caution
- Zenodo records for DESED real and synthetic list CC BY 4.0. The GitHub README says the Python code is MIT and that component datasets include license files at their roots; source media comes from AudioSet/YouTube, Freesound, MUSAN, SINS, and related sources, so re-check component terms before redistribution.
- Download notes
- The helper downloads the official repo plus Zenodo record JSON and small metadata/JAMS files by default. Real and synthetic audio archives are multi-GB and require explicit opt-in flags.
Safe-first helperscripts/download/desed.sh
Speech generation
Designed Vocalizations Dataset
Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion
Safe-first helper
Non Human Voice Conversion
Designed Vocalization Generation
Timbre Transfer
Sound Design Reproduction
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0-with-mixed-upstream-terms
- Code license
- not_applicable
- License caution
- The dataset card applies CC BY 4.0 to the compilation and authors' original metadata and annotations. Raw and designed clips retain their applicable source terms: VCTK and HiFi-TTS are CC BY 4.0, while Freesound clips are individually CC0 1.0, CC BY 3.0, or CC BY 4.0. Preserve the release NOTICE and per-row license, attribution, creator, and source fields when redistributing. The project does not publish a separate evaluation-code repository.
- Download notes
- The public, ungated release contains 237,574 mono 44.1 kHz WAV clips embedded in Parquet: 5,654 raw training sources, 226,160 non-parallel designed training clips, 120 test sources, and 5,640 aligned test references. The test protocol crosses source timbres seen or unseen during training with 40 seen and seven unseen effect presets. The helper downloads official documentation, API metadata, licensing notices, preset metadata, and the approximately 532 KB test-pair manifest by default. The full Hugging Face repository is approximately 37.1 GB and requires DESIGNED_VOCALIZATIONS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/designed_vocalizations.sh
Audio understanding, generation & events
DHAuDS
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
Safe-first helper
Test Time Adaptation
Audio Classification Robustness
Dynamic Corruption Robustness
Speech Command Classification
+3 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_on_hugging_face_cards_with_upstream_terms
- Code license
- apache-2.0
- License caution
- All four Hugging Face cards declare Apache-2.0 and the code repository contains an Apache-2.0 LICENSE. The corrupted audio derives from Speech Commands V2, VocalSound, UrbanSound8K, ReefSet, QUT-NOISE, and DEMAND; those sources retain separate attribution, share-alike, non-commercial, or other terms. In particular, the Apache card labels should not be assumed to remove UrbanSound8K's non-commercial restriction or other upstream obligations.
- Download notes
- The public suite contains separately corrupted adaptation and evaluation sets derived from the held-out portions of Speech Commands V2, VocalSound, UrbanSound8K, and ReefSet. SC2-C, VS-C, and RS-C apply seven corruption categories at two severity levels; US8-C omits QUT-NOISE and DEMAND corruptions that overlap its target classes and uses four categories. The paper reports 908,196 derived samples in total and uses different random seeds for adaptation and evaluation corruptions. The helper downloads official documentation and repository metadata by default. The four Hugging Face repositories report about 50.0 GB of storage combined, so snapshots require DHAUDS_DOWNLOAD_HF=1 and an explicit DHAUDS_DATASETS selection.
Safe-first helperscripts/download/dhauds.sh
Speech recognition
Dialogs
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus for Dialog Assistants
Safe-first helper
Expressive Text To Speech
Conversational Text To Speech
Speech Emotion Classification
Automatic Speech Recognition
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- OpenRAIL
- Code license
- MIT
- License caution
- The dataset card links a custom OpenRAIL responsible-use license that permits use, modification, redistribution, and commercial use subject to its use-based restrictions. The paper and card state that performers gave written informed consent for public and commercial use. Review the complete LICENSE.md rather than treating OpenRAIL as an unrestricted permissive license. The linked VITS2 baseline repository is MIT.
- Download notes
- The public, ungated release contains 20.6 hours and 11,796 Russian utterances from face-to-face acted dialogues recorded in a studio by three professional performers. It provides transcripts, stress-marked text, speaker identifiers, and 12 style/emotion labels, with fixed 11,428/180/188 train, validation, and test splits. The paper evaluates the 188-item stratified test subset with six human-rated quality dimensions and trains a VITS2 expressive-TTS baseline. The helper downloads official documentation, API metadata, and the lightweight validation/test tables by default. The approximately 29.3 MB embedded- audio preview and 5.56 GB full Hugging Face snapshot are separate opt-ins.
Safe-first helperscripts/download/dialogs_ru.sh
Enhancement, separation & quality
Diamond Benchmark
Diamond Benchmark: 750 Real Degraded Speech Recordings for Evaluating Restoration Models
Safe-first helper
Speech Restoration
Speech Enhancement
Perceptual Speech Quality Evaluation
Content Preservation Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_unspecified
- Code license
- not_applicable
- License caution
- Hugging Face metadata labels the dataset "other", but the card provides no license text, source-corpus citation, consent statement, or redistribution terms. The manifest's `emolia_id` field appears to reference an upstream collection, but the card does not identify or license it. Treat the release as evaluation-only pending clarification and verify source-recording, speaker, transcript, and redistribution rights before reuse, especially for training or commercial purposes.
- Download notes
- The public, ungated release contains 750 English speech clips with real-world codec, bandwidth, noise, and clipping degradation, plus reference transcripts and speaker, duration, and sample-rate metadata. Its documented protocol combines DNSMOS-P.835 for perceptual quality with ASR character error rate for content preservation. The helper downloads the official card, API metadata, and approximately 197 KB manifest by default. The Hugging Face API reports approximately 340 MB of repository storage, so the audio snapshot is an explicit opt-in.
Safe-first helperscripts/download/diamond_benchmark.sh
Speaker, identity & emotion
DIHARD III
The Third DIHARD Speech Diarization Challenge
Manual or gated
Speaker Diarization
Speech Activity Detection
Overlapping Speech Diarization
Multisource Speech Diarization
Access pathLDC / licensed
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- ldc_user_agreement
- Code license
- not_applicable
- License caution
- LDC2022S12 and LDC2022S14 list the LDC User Agreement for Non-Members and are available through LDC membership/non-member access. Re-check active LDC terms and component source restrictions before use or redistribution.
- Download notes
- The LDC catalog records list web-download development and evaluation releases with user-agreement access. The helper only prints official access steps; it does not download LDC-controlled data.
Safe-first helperscripts/download/dihard_iii.sh
Enhancement, separation & quality
DNS Challenge
Deep Noise Suppression Challenge
Safe-first helper
Speech Enhancement
Speech Denoising
Dereverberation
Personalized Speech Enhancement
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The repository legal notice says documentation/content are CC BY 4.0 and code is MIT. DNS training data includes component sources such as AudioSet, Freesound, VCTK, VocalSet, and multilingual speech, so component/source-media terms should be re-checked before redistribution or commercial use.
- Download notes
- DNS5 development and blind test sets are multi-GB archives; the full training resources are hundreds of GB compressed and about 1 TB unpacked. The helper saves official README/license/downloader-script files by default and makes data archives explicit opt-ins.
Safe-first helperscripts/download/dns_challenge.sh
Representation & general suites
Doppelganger
Doppelganger: Sound Effects and Their Synthetic Twins
Safe-first helper
Synthetic Real Sound Effect Retrieval
Audio Instance Matching
Audio Representation Evaluation
Synthetic Audio Detection
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The Hugging Face card's MIT tag covers manifests, embeddings, and generation logs, not every audio asset. Stable Audio Open twins use the Stability AI Community License; ElevenLabs twins are redistributed under the author's ElevenLabs license; real audio retains FSD50K, UrbanSound8K, Freesound, or DCASE 2023 Task 7 source terms, and restricted real sources remain reference-by-ID only.
- Download notes
- The public, ungated release pairs 10,420 verified real sound-effect references across 34 Universal Category System events with Stable Audio Open synthetic twins, plus a controlled seven-class DCASE 2023 Task 7 corpus and text-only ElevenLabs controls. Real recordings are referenced by source ID and are not redistributed in bulk. The helper downloads official documentation and repository metadata by default; cloning the roughly 10 MB code/manifests repository or downloading the approximately 8.48 GB Hugging Face release requires separate opt-ins.
Safe-first helperscripts/download/doppelganger.sh
Speech understanding & dialogue
Dynamic-SUPERB
Dynamic-SUPERB: A dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
Safe-first helper
Spoken Language Model Evaluation
Instruction Following
Speech Understanding
Audio Understanding
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- not_specified
- License caution
- GitHub API reported no repository license on 2026-07-09. The benchmark aggregates tasks from many component datasets; use the official task metadata and each upstream corpus license before redistribution, training, or commercial use.
- Download notes
- The helper downloads the official README and leaderboard documentation by default and clones the benchmark repository only with DYNAMIC_SUPERB_CLONE_REPO=1. The benchmark is collaborative and spans many speech, music, and general sound tasks, so underlying task data should be checked through each component source before use.
Safe-first helperscripts/download/dynamic_superb.sh
Speech recognition
Earnings-21
Earnings-21: A Practical Benchmark for ASR in the Wild
Safe-first helper
Asr
Named Entity Recognition
Long Form Speech Recognition
Financial Speech Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0_text
- Code license
- not_specified
- License caution
- The dataset README shows a CC BY-SA 4.0 badge, but LICENSE.md expressly covers only transcripts and associated text files used for alignment. The repository has no detected top-level license, so confirm audio rights before redistribution or commercial use.
- Download notes
- Earnings-21 contains 44 English-language earnings calls totaling about 39 hours, plus a representative 10-hour Eval-10 subset. The helper downloads official documentation and lightweight file/speaker metadata by default; sparse checkout of the approximately 770 MB media tree, transcripts, RTTMs, and bias lists is opt-in.
Safe-first helperscripts/download/earnings_21.sh
Speech recognition
Earnings-22
Earnings-22: A Practical Benchmark for Accents in the Wild
Safe-first helper
Asr
Accented Speech Recognition
Long Form Speech Recognition
Financial Speech Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_specified
- License caution
- The earnings22 README shows a CC BY-SA 4.0 license badge, and LICENSE.md states that transcripts and associated text files are CC BY-SA 4.0. The top-level GitHub repository does not expose a detected repository license, and audio is stored through Git LFS, so re-check upstream terms before redistribution or commercial use.
- Download notes
- Earnings-22 contains 125 English-language earnings-call files totaling about 119 hours. The helper downloads README, license, and metadata by default; sparse checkout of transcripts/media is opt-in, and Git LFS audio pull is a second explicit opt-in.
Safe-first helperscripts/download/earnings_22.sh
Speaker, identity & emotion
EMO-SUPERB
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
Safe-first helper
Speech Emotion Recognition
Speech Representation Evaluation
Speaker Independent Cross Validation
Standardized Dataset Partitioning
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_component_terms
- Code license
- not_specified
- License caution
- The official repository has no LICENSE file and GitHub reports no detected license, so the evaluation code and released partition/label files have unspecified terms. Each of the six underlying speech corpora retains its own access agreement or license; a public benchmark repository does not make their audio freely redistributable.
- Download notes
- The public repository provides the evaluation implementation, corpus adapters, and standardized speaker-independent partitions for IEMOCAP, CREMA-D, MSP-IMPROV, and BIIC-NNIME. EMO-SUPERB evaluates six corpora in total, also including MSP-Podcast and BIIC-Podcast, with separate primary/secondary-emotion settings for some corpora. The helper downloads official documentation by default and makes the approximately 23 MB GitHub repository clone opt-in. It does not fetch corpus audio; users must obtain each component through its official EULA, form, or repository path.
Safe-first helperscripts/download/emo_superb.sh
Audiovisual & cross-modal
EmoPrefer
EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
Safe-first helper
Multimodal Emotion Preference Prediction
Audio Visual Emotion Understanding
Pairwise Emotion Description Evaluation
Multimodal Judge Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_research_terms
- Code license
- Apache-2.0_official_and_MIT_audit
- License caution
- The official EmoPrefer subdirectory includes Apache-2.0 but its README also calls the service a non-commercial research preview. The gated MER2025 card declares CC BY-NC 4.0 plus stricter academic-only, no-redistribution, and no-modification conditions that control the source media and annotations obtained there. The 2026 audit repository is MIT. Treat the narrower access terms as controlling where they conflict, and do not infer media rights from the public CSV release.
- Download notes
- The official repository publicly releases six small annotation tables: the original 574-pair EmoPrefer set, its V2 extension with 2,096 individual-annotator pairs, reverse-order variants for swap-consistency analysis, and variants exposing the two description-generator names for score calculation and shortcut auditing. The paired English audio/video comes from the separately gated MER2025 release and is not redistributed by the repository. The helper downloads the annotations, official documentation, licenses, and repository metadata only; users must accept the MER2025 access conditions themselves to obtain media. The 2026 audit paper and code provide reproducible content-blind, counterfactual, audio-visual judge, and ODIN-style diagnostics but intentionally distribute no private data, model weights, predictions, or checkpoints.
Safe-first helperscripts/download/emoprefer.sh
Speech generation
EmoV-DB
The Emotional Voices Database: Towards Controlling the Emotional Expressiveness in Voice Generation Systems
Safe-first helper
Emotional Speech Synthesis
Expressive Tts
Speech Emotion Recognition
Voice Conversion
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_non_commercial
- Code license
- not_specified
- License caution
- The EmoV-DB license permits non-commercial research, teaching, scientific publication, and personal experimentation, and asks users to contact the dataset owner for commercial use. GitHub API reports license NOASSERTION/Other.
- Download notes
- OpenSLR SLR115 hosts per-speaker/per-emotion archives. The helper downloads OpenSLR/GitHub docs and license by default; speech archives are opt-in with EMOV_DB_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/emov_db.sh
Audiovisual & cross-modal
EPIC-SOUNDS
EPIC-SOUNDS: A Large-Scale Dataset of Actions that Sound
Safe-first helper
Egocentric Audio Event Recognition
Sound Event Detection
Audio Visual Understanding
Action Sound Recognition
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The annotation README states all dataset files are published under Creative Commons Attribution-NonCommercial 4.0 International. Raw audio is derived from EPIC-KITCHENS-100 video recordings, so original dataset access terms and any HDF5 access approval should be checked before redistribution or commercial use.
- Download notes
- The helper downloads official docs and public annotation CSV files by default. Raw audio is not redistributed separately; the official README says to download EPIC-KITCHENS-100 videos and extract audio, or email the maintainers for access to an existing HDF5 file.
Safe-first helperscripts/download/epic_sounds.sh
Audio understanding, generation & events
ESC-50
ESC-50: Dataset for Environmental Sound Classification
Safe-first helper
Environmental Sound Classification
Audio Tagging
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-3.0
- Code license
- not_specified
- License caution
- ESC-10 subset clips are CC BY; ESC-50 as a whole is Creative Commons Attribution-NonCommercial. Per-clip Freesound attributions are in the repository LICENSE file.
Safe-first helperscripts/download/esc_50.sh
Audio understanding, generation & events
ESCUCHA
ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions
Safe-first helper
Spanish Speech Understanding
Audio Question Answering
Audio Reasoning
Multi Audio Comparison
+3 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The GitHub repository has no license file or detected GitHub license. The paper is CC BY 4.0 on arXiv, but that does not license the annotations, code, or source recordings. Review the rights and platform terms for each linked recording before downloading, redistribution, or commercial use.
- Download notes
- The public repository releases 1,000 Spanish questions as JSON and TSV, including 900 multiple-choice and 100 audio-instruction-following items. The helper downloads those approximately 2.2 MB of annotations plus the README and scorer by default; cloning the repository is opt-in. Audio is not redistributed: the release provides source URLs and a yt-dlp script for reconstructing up to 162.9 hours from public recordings, so availability can drift and source-platform terms apply.
Safe-first helperscripts/download/escucha.sh
Speech understanding & dialogue
Europarl-ST
Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates
Safe-first helper
Speech To Text Translation
Speech Translation
Multilingual Asr
Parliamentary Speech Translation
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The official README says the work carried out to construct Europarl-ST is released under CC BY-NC 4.0, while all rights of the underlying data belong to the European Union and respective copyright holders. Re-check EU/European Parliament reuse terms before redistribution or commercial use.
- Download notes
- The official page links release v1.1, which adds Romanian, Polish, and Dutch to German, English, Spanish, French, Italian, and Portuguese for 72 speech translation directions, plus a train-noisy set. The v1.1 archive is about 21 GB, so the helper downloads only the official page and README by default and requires EUROPARL_ST_DOWNLOAD_ARCHIVE=1 for the archive.
Safe-first helperscripts/download/europarl_st.sh
Speech recognition
Fisher English
Fisher English Training Speech and Transcripts
Manual or gated
Automatic Speech Recognition
Conversational Speech Recognition
Telephone Speech Recognition
Speech Transcription
Access pathLDC / licensed
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_ldc_license
- Code license
- not_applicable
- License caution
- The LDC catalog pages list the LDC User Agreement for Non-Members and availability for Subscription/Standard Members and Non-Members. Re-check the current LDC agreement before use or redistribution.
- Download notes
- Fisher English is distributed by LDC after login/licensing. The helper only prints official access steps because the speech and transcripts are paid/licensed catalog releases, not publicly script-downloadable archives.
Safe-first helperscripts/download/fisher_english.sh
Speech recognition
FLEURS
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Safe-first helper
Asr
Speech Translation
Language Identification
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- HF dataset card lists CC BY 4.0.
Safe-first helperscripts/download/fleurs.sh
Speech understanding & dialogue
Fluent Speech Commands
Fluent Speech Commands: A dataset for spoken language understanding research
Manual or gated
Spoken Language Understanding
Intent Classification
Slot Filling
Smart Home Voice Commands
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- Official Fluent.ai page says the dataset is released strictly for academic research only under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International, and not authorized for commercial use. No current public code repository was identified.
- Download notes
- The helper downloads the public dataset page and license PDF, then prints the manual Google Groups access path. The official page says the corpus contains 30,043 single-channel 16 kHz WAV utterances from 97 speakers with action, object, and location slot labels; do not commit granted links or downloaded audio.
Safe-first helperscripts/download/fluent_speech_commands.sh
Music
FMA
FMA: A Dataset For Music Analysis
Safe-first helper
Music Genre Classification
Music Auto Tagging
Music Recommendation
Artist Identification
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- Official GitHub README says metadata is CC BY 4.0, code is MIT, and audio is distributed under the license chosen by each artist because the dataset maintainers do not hold audio copyright. UCI lists the dataset record as CC BY 4.0, but per-track audio licenses should be checked before redistribution or commercial use.
- Download notes
- The helper downloads official README/license files by default. Metadata is 342 MiB; audio archives range from 7.2 GiB for fma_small to 879 GiB for fma_full, so metadata and audio are explicit opt-ins.
Safe-first helperscripts/download/fma.sh
Audio understanding, generation & events
ForestIR
ForestIR: Physics-Informed Forest Sound Simulation for Array-Based Bioacoustic Remote Sensing
Safe-first helper
Forest Impulse Response Simulation
Bioacoustic Sound Source Localization
Microphone Array Design
Spatial Audio Robustness
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- bundled_media_terms_not_separately_specified
- Code license
- MIT
- License caution
- The repository has an MIT LICENSE and GitHub detects MIT, but neither the paper nor README states separate licenses or complete provenance terms for bundled bird vocalizations and recorded environmental noise. Do not assume the software license clears those recordings; verify source-media attribution, wildlife-recording, and field-recording rights before redistribution or commercial use. The request-only processed validation data may carry additional terms.
- Download notes
- The public repository provides a physics-informed forest impulse-response simulator, command-line and Python interfaces, measured and synthetic tree/microphone/source geometry presets, example bird calls, environmental noise recordings, and manifest-producing array rendering. The helper downloads official documentation, license, repository metadata, and the paper by default; cloning the approximately 31 MB repository requires FORESTIR_CLONE_REPO=1. The paper says processed site-recorded data needed to reproduce its main validation analyses must be requested from the authors, so the public repository must not be represented as a complete release of those field measurements.
Safe-first helperscripts/download/forestir.sh
Audiovisual & cross-modal
Friend Bench
Friend Bench: Social relationship recognition from thin-slice dyadic interactions
Safe-first helper
Audio Visual Social Reasoning
Social Relationship Recognition
Familiar Stranger Classification
Multimodal Human Behavior Understanding
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The dataset card declares CC BY-NC 4.0 and says the benchmark inherits that license from Seamless Interaction. The release includes recorded human interactions, transcripts, relationship labels, and anonymized rater responses with bucketed demographic fields; users should retain attribution, limit use to non-commercial purposes, and review the source corpus's privacy, consent, biometric, and responsible-use terms.
- Download notes
- The public, ungated validation release contains 96 approximately 20-second dyadic clips, balanced across familiar/stranger and early/late conditions. Every item provides mixed audio, side-by-side video, a turn-level transcript, and labels for binary familiarity and six-way relationship type. It also releases anonymized human judgments under audio-, video-, and text-only conditions. The helper downloads the official card and API metadata by default. Set FRIEND_BENCH_DOWNLOAD_METADATA=1 for the two lightweight JSONL tables, or FRIEND_BENCH_DOWNLOAD_HF=1 for the complete approximately 433 MB snapshot. The card says a label-held-out test split is planned but not yet released, and currently provides no paper or citation.
Safe-first helperscripts/download/friend_bench.sh
Audio understanding, generation & events
FSD50K
FSD50K: An Open Dataset of Human-Labeled Sound Events
Safe-first helper
Sound Event Classification
Sound Event Tagging
Audio Tagging
Machine Listening
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- not_specified
- License caution
- Zenodo says individual Freesound clips keep per-clip CC0, CC-BY, CC-BY-NC, or CC Sampling+ licenses, with clip-level license mappings in the metadata JSON files. FSD50K as a curated dataset is additionally released under CC BY, but the maintainers warn that a single global license is not straightforward because items have different licenses.
- Download notes
- The helper downloads the small documentation, ground-truth, and metadata ZIPs by default. Audio is split across about 24.7 GiB of dev archives plus about 6.2 GiB of eval archives, so audio download is an explicit opt-in.
Safe-first helperscripts/download/fsd50k.sh
Audio understanding, generation & events
FSDKaggle2018
FSDKaggle2018: Freesound General-Purpose Audio Tagging Challenge Dataset
Safe-first helper
Audio Tagging
Sound Event Classification
General Purpose Audio Tagging
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- not_specified
- License caution
- The Zenodo record license id is other-at. The record states FSDKaggle2018 as a curated dataset is CC BY, while individual Freesound clips retain per-clip Creative Commons licenses listed in train_post_competition.csv and test_post_competition_scoring_clips.csv. Kaggle competition rules and Freesound source terms should also be checked for challenge use.
- Download notes
- The helper downloads the Zenodo record JSON plus small documentation and metadata ZIPs by default. The audio archives are about 4.6 GiB combined, so audio download is an explicit opt-in.
Safe-first helperscripts/download/fsdkaggle2018.sh
Audio understanding, generation & events
FSDKaggle2019
FSDKaggle2019: Freesound Audio Tagging 2019 / Audio Tagging with Noisy Labels and Minimal Supervision
Safe-first helper
Audio Tagging
Sound Event Classification
Noisy Label Learning
Weakly Labeled Audio Classification
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- MIT
- License caution
- Zenodo reports license id other-at. The record says FSDKaggle2019 as a curated dataset is CC BY, while individual Freesound clips keep per-clip CC0, CC-BY, CC-BY-NC, or CC Sampling+ licenses and Flickr/YFCC noisy-train clips keep CC-BY or CC BY-SA licenses recorded in metadata CSVs. Kaggle/DCASE challenge rules and source media terms should also be checked for challenge or redistribution use.
- Download notes
- The helper downloads the Zenodo record JSON plus small documentation and metadata ZIPs by default. Full audio is about 25 GiB and includes a split noisy-train archive, so audio download is an explicit opt-in with part selection.
Safe-first helperscripts/download/fsdkaggle2019.sh
Enhancement, separation & quality
FUSS
FUSS: Free Universal Sound Separation Dataset
Safe-first helper
Universal Sound Separation
Arbitrary Sound Separation
Reverberant Source Separation
Sound Event Separation
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- Apache-2.0
- License caution
- The official license document and Zenodo record release FUSS as a whole under CC BY 4.0. Its input audio clips are CC0 Freesound files selected using prerelease FSD50K labels; those labels are not distributed with FUSS. Google Research's sound-separation code repository is Apache-2.0.
- Download notes
- The public Zenodo release provides train, validation, and eval mixtures containing one to four arbitrary sound sources, with dry and reverberant references, simulated room responses, and CC0 source clips. The helper saves official documentation, repository metadata, and the small license archive by default. Data archives are about 1.9 to 8.9 GB each and require explicit selection with FUSS_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/fuss.sh
Audio understanding, generation & events
Geo-ATBench
Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
Safe-first helper
Multi Label Audio Tagging
Environmental Sound Classification
Audio Geospatial Context Fusion
Context Aware Audio Tagging
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_upstream_terms
- Code license
- MIT
- License caution
- The Zenodo record declares CC BY 4.0 and the official repository is MIT. The paper says audio comes from Freesound and a prior geotagged-audio dataset, while contextual metadata derives from OpenStreetMap; preserve source attribution and review per-clip audio licenses and OSM attribution/database terms before redistribution.
- Download notes
- The public, ungated release contains 3,854 ten-second mono WAV clips totaling 10.71 hours, 28 fine-grained sound-event labels, three coarse event groups, and POI-derived geospatial semantic context over 11 OpenStreetMap feature categories. The helper downloads official documentation and repository/Zenodo metadata by default; the approximately 850 MB dataset archive is opt-in.
Safe-first helperscripts/download/geo_atbench.sh
Speech recognition
Ghana Speech Eval
Ghana Speech Eval: ASR Evaluation for 10 Ghanaian Languages
Safe-first helper
Automatic Speech Recognition
Multilingual Speech Recognition
Low Resource Speech Recognition
African Language Speech Recognition
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_claim_with_unresolved_upstream_terms
- Code license
- not_applicable
- License caution
- The Ghana Speech Eval card declares CC BY 4.0 and says this follows the source dataset. However, the current AfriSpeech public-v1 card and API metadata do not state a repository-level license. Preserve attribution, review the original collection and speaker-consent terms, and confirm upstream reuse rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains 9,967 read-speech clips with verbatim transcripts across Ahanta, Dagaare, Dangme, Ewe, Fante, Frafra/Gurene, Ga, Nzema, Sehwi, and Twi. It samples up to 1,000 clips per language from AfriSpeech's public African speech corpus, merges the source splits into one evaluation split, and filters clips to 3-15 seconds. The card explicitly reserves the fixed set for ASR evaluation rather than training. The helper downloads official cards and API metadata by default; the approximately 594 MB compressed benchmark snapshot requires GHANA_SPEECH_EVAL_DOWNLOAD_HF=1.
Safe-first helperscripts/download/ghana_speech_eval.sh
Speech recognition
GigaSpeech
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Manual or gated
Asr
Large Scale Speech Recognition
Text To Speech
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- gated_non_commercial_research_educational
- Code license
- Apache-2.0
- License caution
- HF terms restrict the database to non-commercial research and educational purposes; the GitHub code repo is Apache-2.0.
- Download notes
- HF access is gated and the official repo also asks users to fill out the Google Form first; full HF dataset size is about 2.6 TB.
Safe-first helperscripts/download/gigaspeech.sh
Speech recognition
GigaSpeechBench
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Safe-first helper
Automatic Speech Recognition
Speech To Text Translation
Multilingual Speech Recognition
Dialectal Speech Recognition
+6 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the Hugging Face repository metadata nor the official GitHub repository declares a license, and the GitHub repository has no LICENSE file. Do not infer reuse or redistribution rights from public, ungated access; obtain clarification from SpeechColab before redistribution or commercial use.
- Download notes
- The paper defines a 680-hour in-the-wild benchmark with five modules covering 14 low-resource language/region sets, six Chinese dialects, six English accents, 12 Chinese and English terminology domains, and child/older-adult speech; it also reports Chinese and English translation references for 11 languages. The current public, ungated Hugging Face repository contains the Low-Resource-Languages, CH-EN-Dialects, and Vertical-Domain modules, including audio archives and JSON metadata, but no separate age-group module was visible when checked. Translation result files are public, while the exact release coverage of translation references should be verified per metadata file. The helper downloads official documentation and repository/API metadata by default; the Hugging Face API reports approximately 86.3 GB of repository storage, so the dataset snapshot requires GIGASPEECHBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/gigaspeechbench.sh
Speech recognition
Golos
Golos: Russian Dataset for Speech Research
Safe-first helper
Automatic Speech Recognition
Russian Speech Recognition
Far Field Speech Recognition
Language Modeling
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_golos_license
- Code license
- not_specified
- License caution
- OpenSLR points to SberDevices' English and Russian PDF license files rather than an SPDX-style open license. Re-check those PDFs before redistribution, commercial use, or training release claims.
- Download notes
- OpenSLR SLR114 mirrors an 18 GiB Opus archive with Russian speech and transcripts, a 71 MiB QuartzNet acoustic model, and 4.7 GiB KenLM language models. The helper saves the OpenSLR page, README, checksums, and license PDFs by default and requires explicit opt-ins before downloading large archives.
Safe-first helperscripts/download/golos.sh
Music
GTZAN
GTZAN Genre Collection
Safe-first helper
Music Genre Classification
Music Information Retrieval
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The reachable Hugging Face dataset card does not state a data license. Treat redistribution and commercial use as unclear until the Marsyas/original dataset terms are confirmed.
- Download notes
- The Hugging Face card describes 1,000 30-second mono WAV tracks across 10 genres and provides a reproducible datasets loader. The helper saves the HF dataset card by default and requires GTZAN_DOWNLOAD_HF=1 before downloading the audio snapshot.
Safe-first helperscripts/download/gtzan.sh
Speech recognition
HALAS
HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
Safe-first helper
Automatic Speech Recognition
Asr Hallucination Detection
Asr Error Analysis
Span Level Error Detection
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- unknown_with_cc-by-sa-4.0_upstream
- Code license
- not_specified
- License caution
- The Hugging Face card's machine-readable license field is "unknown," and neither the GitHub nor Hugging Face release includes a standalone license file. The card says HALAS derives from CC BY-SA 4.0 Earnings-22 and describes the authors' annotations as intended for the same license unless otherwise specified, but it also tells users to consult the repository for authoritative terms. Treat annotation and code rights as unspecified pending an explicit license; Earnings-22 source terms continue to apply to separately obtained audio.
- Download notes
- The public, ungated release contains human-reviewed predictions from seven ASR systems for 3,611 English Earnings-22 segments, with utterance labels and character-span annotations for hallucination, looping, and hallucinated looping. Meeting-disjoint stratified splits contain 2,866 train and 745 test segments; the test split has a 22.6% hallucination rate and excludes audio at or below one second or below three words. The release contains annotations, predictions, corrected references, and source identifiers rather than redistributed audio; audio must be obtained separately from Earnings-22. The helper downloads official documentation, prompts, and repository metadata by default. The approximately 0.86 MB test CSV is opt-in, while the full approximately 7.2 MB Hugging Face snapshot requires HALAS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/halas.sh
Representation & general suites
HEAR
HEAR: Holistic Evaluation of Audio Representations
Safe-first helper
Audio Representation Evaluation
Speech Classification
Environmental Sound Classification
Music Classification
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- Apache-2.0
- License caution
- Zenodo record lists CC BY 4.0 but explicitly says component datasets have different open licenses and to inspect each dataset's LICENSE.txt; eval kit repo is Apache-2.0.
- Download notes
- The helper downloads the Zenodo record metadata and LICENSE.txt by default. Individual task archives are large and opt-in; the Zenodo record notes that TFDS-derived crema-d, GTZAN genre, and GTZAN music/speech tasks in this release were retracted because of a preprocessing bug.
Safe-first helperscripts/download/hear.sh
Speech generation
Hi-Fi TTS
Hi-Fi Multi-Speaker English TTS Dataset
Safe-first helper
Text To Speech
Speech Synthesis
Multi Speaker Speech Synthesis
High Fidelity Speech Synthesis
+1 more
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR109 lists CC BY 4.0. The corpus is based on public-domain LibriVox audiobooks and Project Gutenberg texts, but downstream users should still preserve attribution and check packaged metadata.
- Download notes
- OpenSLR lists one 39-41 GiB speech/text archive. The helper saves the official OpenSLR page by default and requires HIFITTS_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/hifitts.sh
Speaker, identity & emotion
HI-MIA
HI-MIA: A Far-field Text-Dependent Speaker Verification Database and the Baselines
Safe-first helper
Text Dependent Speaker Verification
Far Field Speaker Verification
Wake Word Speaker Verification
Speaker Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR85 lists Apache License v2.0. The paper says HI-MIA contains recordings from 340 people in real-room far-field settings and is extracted from AISHELL-WakeUp-1.
- Download notes
- OpenSLR SLR85 hosts the AISHELL Speaker Verification Challenge 2019 far-field text-dependent speaker verification data. The helper downloads the official page and filename mapping by default; train/dev/test archives are multi-GB and require HIMIA_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/hi_mia.sh
Audiovisual & cross-modal
IEMOCAP
IEMOCAP: Interactive Emotional Dyadic Motion Capture Database
Manual or gated
Speech Emotion Recognition
Audio Visual Emotion Recognition
Multimodal Emotion Recognition
Dialogue Emotion Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_research_license
- Code license
- not_applicable
- License caution
- The official release page links a USC/SAIL data release form and Google request form. Treat access as manual/form-gated and re-check the signed release terms before redistribution, commercial use, or sharing derived copies.
- Download notes
- IEMOCAP is released by request after reading the license and submitting the official electronic release form; there is no public one-command archive URL.
Safe-first helperscripts/download/iemocap.sh
Speech understanding & dialogue
IFEval-Audio
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
Manual or gated
Audio Instruction Following
Speech Instruction Following
Structured Output Adherence
Semantic Correctness Evaluation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- apache-2.0_with_mixed_upstream_terms
- Code license
- cc_by_nc_unspecified_version
- License caution
- The Hugging Face card declares Apache-2.0, while the paper says clips are used as-is from Spoken SQuAD (CC BY-SA 4.0), TED-LIUM 3 (CC BY-NC-ND 3.0), MuChoMusic (CC BY-SA 4.0), WavCaps (academic use only), and AudioBench sources with inherited licenses. The AudioBench LICENSE says source code is Creative Commons NonCommercial without naming a version. Preserve the most restrictive source terms and verify per-clip provenance before redistribution or commercial use.
- Download notes
- The official release contains 280 English audio-instruction-answer triples spanning Content, Capitalization, Symbol, List Structure, Length, and Format requirements. It includes 240 speech triples and 40 music/environmental-sound triples, and reports Instruction Following Rate, Semantic Correctness Rate, and Overall Success Rate. The helper downloads public Hugging Face API metadata and official benchmark documentation by default; the approximately 45 MB compressed Hugging Face snapshot is opt-in and requires logging in and accepting the repository's access conditions.
Safe-first helperscripts/download/ifeval_audio.sh
Audiovisual & cross-modal
InCarEmo
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Manual or gated
Speech Emotion Recognition
Multimodal Emotion Recognition
Fatigue Detection
Driver Distraction Monitoring
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_academic_research
- Code license
- not_specified
- License caution
- The paper states that InCarEmo is licensed for non-commercial academic research and use, but neither the official repository nor linked Drive landing page provides a full license text. The repository contains no LICENSE file and no released code at the time checked. ArXiv's CC BY 4.0 license covers the paper, not the feature data.
- Download notes
- The paper describes synchronized Chinese RGB and infrared video, in-cabin audio, and dialogue text from 25 participants, with six emotion classes plus fatigue and distraction labels. It also defines an auxiliary English setting using translated text and synthesized English speech. Because of participant privacy, the official repository releases the dataset only as feature-level data through a public Google Drive folder, not as raw audio or video. The helper saves the public repository and paper metadata, then prints the manual Drive step; it does not download participant data.
Safe-first helperscripts/download/incaremo.sh
Speech recognition
IndicContextEval
IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Safe-first helper
Contextual Asr
Multilingual Asr
Named Entity Recognition
Context Utilisation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- The Hugging Face card and official repository README declare the benchmark CC BY 4.0. The GitHub repository has no LICENSE file and GitHub detects no repository license, so do not assume that its evaluation outputs or forthcoming evaluation code use the same terms without clarification.
- Download notes
- The public, ungated release contains 16,884 natural-speech utterances totaling 55.93 hours from 555 speakers across Bengali, Gujarati, Hindi, Malayalam, Marathi, Odia, Telugu, and Urdu. It covers 23 professional domains and supplies seven controlled prompt levels: no context, language, structured metadata, natural-language description, English- script entities, native-script entities, and incorrect-domain entities. The helper downloads official documentation, repository metadata, the small published results table, and prompt-taxonomy supplements by default. The Hugging Face card reports a 6.48 GB current download and its API reports approximately 19.64 GB of repository storage including history, so embedded audio requires INDIC_CONTEXT_EVAL_DOWNLOAD_HF=1.
Safe-first helperscripts/download/indic_context_eval.sh
Speech generation
InstructTTSEval
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
Safe-first helper
Controllable Text To Speech
Speech Instruction Following
Acoustic Parameter Control
Descriptive Style Control
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- not_specified
- License caution
- The Hugging Face dataset card lists MIT, while the GitHub repository has no detected license. The paper says reference audio was curated from NCSSD plus movies, TV dramas, variety shows, and other film/television sources; it also says the dataset is solely for academic and research use. Treat the card license as insufficient to clear third-party media rights, and verify source terms before redistribution or commercial use.
- Download notes
- The helper downloads the official GitHub README and Hugging Face dataset card by default. The public, ungated Hugging Face snapshot contains English and Chinese Parquet splits with embedded reference audio and uses about 1.8 GB, so data download requires INSTRUCT_TTS_EVAL_DOWNLOAD_HF=1; cloning the evaluation repository is separately opt-in.
Safe-first helperscripts/download/instruct_tts_eval.sh
Audiovisual & cross-modal
InterPet4D
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
Safe-first helper
Audio Conditioned Motion Generation
Human Pet Interaction Modeling
Multimodal Motion Generation
Animal Behavior Understanding
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_noncommercial_terms
- Code license
- not_applicable
- License caution
- The current public Hugging Face card declares CC BY-NC 4.0 and says participants consented to data release. The paper's ethical statement instead says users must agree to a data-use agreement prohibiting redistribution, surveillance, and biometric identification, but no such agreement or click-through gate is exposed on the current ungated dataset page. Apply the stricter paper restrictions pending clarification, do not redistribute the corpus, and do not infer that arXiv's CC BY 4.0 paper license governs the released participant recordings.
- Download notes
- The public, ungated v1 release contains 227 approximately 17-20-second egocentric clips from 113 interaction sessions involving 13 dogs and about 23 human participants. It provides synchronized 48 kHz stereo MP3 audio, MERT embeddings, and human-hand, human-body, dog-skeleton, and SMAL motion parameters. The helper downloads the official dataset card and API metadata by default; the Hugging Face API reports about 10.7 GB of repository storage, so the full snapshot requires INTERPET4D_DOWNLOAD_HF=1. The paper also describes 12-view and egocentric RGB video, but the current Hugging Face file inventory releases audio and motion artifacts rather than those raw videos.
Safe-first helperscripts/download/interpet4d.sh
Speech recognition
JASMIN-CGN
JASMIN-CGN: Extension of the Spoken Dutch Corpus with Speech of Elderly People, Children and Non-Natives
Manual or gated
Automatic Speech Recognition
Child Speech Recognition
Accented Speech Recognition
Non Native Speech Recognition
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_signed_license_noncommercial
- Code license
- not_applicable
- License caution
- The official non-commercial product is free but requires a signed license and account login; the owner page directs commercial users to a separate commercial product. Review the current agreement before use or redistribution.
- Download notes
- The Dutch Language Institute distributes the approximately 115-hour Dutch/Flemish speech corpus after account login and a signed license agreement. Its official page says the corpus contains read speech and human-machine dialogues from adolescents, non-native speakers, and seniors, with WAV audio plus TXT and TextGrid annotations. The helper prints the official access steps only; it does not bypass login or download corpus audio. The 2026 evaluation paper selects 120 human-machine-interaction test utterances from children, older adults, and Flemish speakers and also reports ASR results on the full corresponding test sets.
Safe-first helperscripts/download/jasmin_cgn.sh
Speech recognition
KeSpeech
KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects
Manual or gated
Mandarin Asr
Dialect Asr
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_no_adaptations_no_distribution
- Code license
- not_specified
- License caution
- Downloading data means accepting the custom license in dataset_license.md.
Safe-first helperscripts/download/kespeech.sh
Speaker, identity & emotion
KVoiceBench / KOpenAudioBench / KMMAU
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
Safe-first helper
Korean Spoken Question Answering
Speech Instruction Following
Speech Reasoning
Speech Safety Evaluation
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_Apache-2.0_and_CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The KVoiceBench and KOpenAudioBench cards declare Apache-2.0; KMMAU declares CC BY-NC-SA 4.0. These benchmarks adapt prior benchmark content or source speech corpora, so the declared repository licenses may not replace upstream terms. Review VoiceBench/OpenAudioBench provenance and the KSS, KMSAV, and Seoul Corpus conditions before redistribution or commercial use. Raon-Eval code is Apache-2.0.
- Download notes
- The three public, ungated Hugging Face releases contain 12,345 Korean test samples: 7,306 KVoiceBench spoken-QA items transferred from VoiceBench, 2,835 KOpenAudioBench items transferred from OpenAudioBench, and 2,204 KMMAU audio-understanding items derived from KSS, KMSAV, and Seoul Corpus. The helper downloads official cards, repository metadata, and the paper page by default. Full snapshots require KOREAN_SPEECHLM_BENCHMARKS_DOWNLOAD_HF=1 because Hugging Face reports approximately 4.64 GB, 608 MB, and 4.49 GB of repository storage respectively.
Safe-first helperscripts/download/korean_speechlm_benchmarks.sh
Speech recognition
L2-ARCTIC
L2-ARCTIC: A Non-native English Speech Corpus
Manual or gated
Asr
Accented Speech Recognition
Mispronunciation Detection
Accent Conversion
+2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- Official homepage states the corpus is released under CC BY-NC 4.0; usage outside that license requires contacting the TAMU dataset owner. The current release covers 24 non-native English speakers plus suitcase-corpus material, while the Interspeech 2018 paper describes the initial v1.0 release.
- Download notes
- Official access requires reviewing the CC BY-NC 4.0 license terms and submitting the download form with name, email, and affiliation. The project sends a generated Google Drive link by email, so the helper prints manual access steps rather than storing or using private generated URLs.
Safe-first helperscripts/download/l2_arctic.sh
Enhancement, separation & quality
L3DAS21
L3DAS21 Challenge: Machine Learning for 3D Audio Signal Processing
Safe-first helper
Three Dimensional Speech Enhancement
Speech Enhancement
Sound Event Localization And Detection
Sound Source Localization
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- DataCite lists CC BY 4.0 for the Zenodo dataset. The generated mixtures incorporate LibriSpeech speech and FSD50K sound events, so preserve LibriSpeech attribution and review FSD50K's per-clip Creative Commons terms, including non-commercial clips. The official GitHub repository exposes no LICENSE file or GitHub-detected license, so clarify code terms before reuse or redistribution.
- Download notes
- The public Zenodo V1 release is a 65-hour corpus of multi-source, multi-perspective B-format Ambisonics audio generated from impulse responses measured at 252 positions with two first-order Ambisonics microphones. Task 1 contains more than 30,000 spatial speech mixtures for 3D speech enhancement, with clean monaural speech targets. Task 2 contains 900 one-minute soundscapes for sound-event localization and detection, with up to three simultaneous events and 100 ms activity, class, and Cartesian-location targets. Both tasks have one- and two-microphone tracks. The helper downloads official pages, repository metadata, DataCite metadata, and the paper by default; cloning the code and downloading selected large archives are separate opt-ins.
Safe-first helperscripts/download/l3das21.sh
Enhancement, separation & quality
L3DAS22
L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
Safe-first helper
Three Dimensional Speech Enhancement
Speech Enhancement
Sound Event Localization And Detection
Sound Source Localization
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- Kaggle's official API lists CC BY 4.0 for L3DAS22. The generated mixtures incorporate LibriSpeech speech and FSD50K sound events; preserve LibriSpeech attribution and check FSD50K's per-clip Creative Commons licenses, including non-commercial clips. The official GitHub repository exposes no LICENSE file or GitHub-detected license, so clarify code terms before redistribution.
- Download notes
- The public Kaggle release contains 105,757,713,362 bytes of multi-source, multi-perspective B-format Ambisonics audio. Task 1 has more than 40,000 spatial speech mixtures totaling nearly 90 hours at 16 kHz, with up to three overlapping background noises and clean monaural speech targets. Task 2 has 900 30-second soundscapes totaling 7.5 hours at 32 kHz, with 14 office-relevant event classes, up to three overlaps, and 100 ms activity and Cartesian-location targets. Both tasks provide one- and two-microphone tracks using one or two first-order Ambisonics microphones. The helper downloads official pages, repository metadata, the paper, documentation, and Kaggle API metadata by default; cloning the code and downloading the 105.8 GB dataset are separate opt-ins. Full data download requires the Kaggle CLI and account credentials.
Safe-first helperscripts/download/l3das22.sh
Audio understanding, generation & events
LAT-Bench
LAT-Bench: Long-form Audio Temporal Awareness Benchmark
Safe-first helper
Long Form Audio Understanding
Temporal Audio Reasoning
Dense Audio Captioning
Temporal Audio Grounding
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_claim_with_upstream_media_terms
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for LAT-Bench, while the GitHub repository is Apache-2.0. The release references externally hosted source audio rather than relicensing or redistributing it; review each recording's rights and platform terms before retrieval, redistribution, or commercial use.
- Download notes
- The public, ungated release provides Chinese and English metadata plus task JSONL files for a human-verified 40-hour benchmark, with 25 hours of Chinese and 15 hours of English audio up to 30 minutes long. Dense Audio Captioning, Temporal Audio Grounding, and Targeted Audio Captioning annotations are included. Audio is not bundled; the metadata records source URLs, so availability can drift and source-platform terms apply. The helper downloads the approximately 2.6 MB annotations, official documentation, and repository metadata by default; cloning the evaluation repository is opt-in.
Safe-first helperscripts/download/lat_bench.sh
Speech recognition
Libri-Light
Libri-Light: A Benchmark for ASR with Limited or No Supervision
Safe-first helper
Automatic Speech Recognition
Self Supervised Speech Representation
Semi Supervised Asr
Unsupervised Speech Representation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The GitHub repository says Libri-Light code is MIT, but the reviewed README and data-preparation docs do not state a standalone data license. The paper describes the data as derived from open-source LibriVox audiobooks, so re-check per-source audiobook terms and attribution requirements before redistribution.
- Download notes
- The helper downloads official README/license/data-preparation/evaluation docs by default. Limited-supervision finetuning data is about 0.6 GiB, ABX item data is small, and unlabeled audio ranges from 35 GiB small to 321 GiB medium and 3.05 TiB large, so all data archives require explicit opt-in.
Safe-first helperscripts/download/libri_light.sh
Speech recognition
LibriCSS
LibriCSS: Continuous Speech Separation Dataset
Safe-first helper
Continuous Speech Separation
Overlapped Speech Recognition
Multi Channel Asr
Speaker Diarization
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_upstream_basis_data_license_not_separately_stated
- Code license
- MIT
- License caution
- The repository LICENSE applies MIT to the software and documentation and identifies the source LibriSpeech corpus as CC BY 4.0, but it does not state a separate license for the replayed LibriCSS recordings. Preserve LibriSpeech attribution and confirm recording-level rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains ten approximately one-hour sessions made from LibriSpeech utterances replayed through eight loudspeakers and captured with a seven-channel circular microphone array in an office meeting room. Each session has six ten-minute mini-sessions spanning 0%-40% overlap, including separate short-gap and long-gap 0% conditions. The official repository provides preparation, Kaldi ASR, and WER-scoring tools for utterance-wise and continuous-input evaluation. The helper downloads official documentation and repository metadata by default; the direct Google Drive archive is 6,407,297,637 bytes (about 5.97 GiB) and requires LIBRICSS_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/libricss.sh
Enhancement, separation & quality
LibriMix
LibriMix: An Open-Source Dataset for Generalizable Speech Separation
Safe-first helper
Speech Separation
Speech Enhancement
Noisy Speech Separation
Source Separation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_derived
- Code license
- MIT
- License caution
- The LibriMix repository license is MIT for code/scripts. Generated data is derived from LibriSpeech, which OpenSLR lists as CC BY 4.0, plus WHAM noise; re-check WHAM terms and cite all upstream components before redistribution or commercial use.
- Download notes
- The helper clones the official generator/metadata repository by default. Running generation downloads LibriSpeech clean subsets and WHAM noise, then creates many mixtures; the upstream README estimates about 430 GiB for Libri2Mix plus 332 GiB for Libri3Mix, with additional source storage, so generation requires an explicit opt-in and an external storage directory.
Safe-first helperscripts/download/librimix.sh
Speech recognition
LibriSpeech
LibriSpeech ASR corpus
Safe-first helper
Asr
Speech Reconstruction
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR12 and HF dataset card list CC BY 4.0.
- Download notes
- The helper downloads the OpenSLR landing page and checksums by default; corpus archives require LIBRISPEECH_DOWNLOAD_ARCHIVES=1. Qwen3-TTS section 4.1.2 evaluates tokenizer reconstruction on all 2,620 utterances in test-clean with PESQ, STOI, UTMOS, and WavLM-based speaker similarity. Qwen-Audio-VAE sections 4.1-4.3 use LibriSpeech for speech reconstruction and ablation evaluation.
Safe-first helperscripts/download/librispeech.sh
Speech generation
LibriTTS
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Safe-first helper
Text To Speech
Speech Synthesis
Voice Cloning
Multi Speaker Speech Synthesis
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR60 lists CC BY 4.0. LibriTTS is derived from LibriSpeech, which in turn derives from LibriVox audio and Project Gutenberg text.
- Download notes
- OpenSLR lists seven archives totaling tens of GiB; the helper saves the official OpenSLR page by default and requires LIBRITTS_DOWNLOAD_ARCHIVES=1 before downloading archives.
Safe-first helperscripts/download/libritts.sh
Speech recognition
LibriWASN
LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices
Safe-first helper
Meeting Transcription
Continuous Speech Separation
Speaker Diarization
Multi Channel Asr
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- Zenodo declares the LibriWASN release CC BY 4.0 and includes the license text; the official tools repository is MIT. The recordings replay LibriSpeech/LibriCSS material, so preserve that provenance and review upstream speech-data terms when redistributing derived data.
- Download notes
- The public release contains 20 hours of meeting-like recordings in two rooms with approximately 200 ms and 800 ms reverberation times. Five smartphones and four microphone arrays provide 29 asynchronous audio channels, with 0%-40% overlap conditions and ground-truth diarization. The same LibriSpeech sentences and speakers as LibriCSS were replayed. The helper downloads official metadata, README/license files, paper page, and repository documentation by default. Twelve per-room and per-overlap ZIP archives total about 55.8 GB and require explicit opt-in; the official repository's full downloader also fetches LibriCSS as a transcription/reference dependency.
Safe-first helperscripts/download/libriwasn.sh
Speech recognition
Live Gurbani Captioning Benchmark v1
Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan
Safe-first helper
Closed Vocabulary Singing Captioning
Live Audio Tracking
Lyrics Alignment
Sung Scripture Identification
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_annotations
- Code license
- MIT
- License caution
- The repository LICENSE applies CC BY 4.0 to ground-truth annotations in test/ and the committed baselines, and MIT to the software and documentation. The four underlying Kirtan recordings are referenced by YouTube ID rather than bundled; the repository license does not grant rights to those recordings, and platform terms, uploader rights, performer rights, and local cultural or legal considerations remain separate.
- Download notes
- The public v1 repository contains ground-truth timelines for four hand-reviewed Sikh Kirtan recordings, each evaluated from 0%, 33%, and 66% start offsets, giving 12 cases and approximately 57 minutes of scored audio. It also releases a standard-library Python scorer, visualization and annotation tools, and empty, shifted, and perfect baselines. The primary metric is one-second frame accuracy with a one-second boundary collar. The helper downloads official documentation and repository metadata by default; cloning the small repository with annotations and evaluation code is opt-in. Audio is not redistributed: the official README identifies four YouTube video IDs and provides yt-dlp/ffmpeg preparation commands, so availability and source-media rights must be checked separately.
Safe-first helperscripts/download/live_gurbani_captioning_v1.sh
Speech generation
LJSpeech
The LJ Speech Dataset
Safe-first helper
Text To Speech
Speech Synthesis
Single Speaker Speech Synthesis
Asr
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- public_domain
- Code license
- not_applicable
- License caution
- Official page says text, audio, and annotations are public domain; the HF mirror lists unlicense.
- Download notes
- The official archive is about 2.6 GiB. The helper saves the official dataset page by default and requires LJSPEECH_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/ljspeech.sh
Audiovisual & cross-modal
LLP
Look, Listen, and Parse: Audio-Visual Video Parsing Dataset
Safe-first helper
Audio Visual Video Parsing
Temporal Event Localization
Audio Event Detection
Visual Event Detection
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- unclear
- License caution
- The official repository does not contain a license file and GitHub reports no detected license. Its README says GPLv3 but links to a GPLv3 license in an unrelated repository, so neither that statement nor the linked file establishes clear terms for LLP annotations or code. Obtain clarification before redistribution or commercial use; source videos also retain their original rights and YouTube terms.
- Download notes
- The official repository releases lightweight weak-label train/validation/test CSVs and dense audio and visual event annotations for validation and test. The full CSV contains 11,849 ten-second YouTube segment references; the published split files contain 10,000 train, 649 validation, and 1,200 test rows. The helper downloads documentation and all annotation CSVs by default. Extracted VGGish, ResNet-152, and R(2+1)D features remain a manual Google Drive download, while raw videos must be reconstructed from referenced YouTube segments subject to availability and platform terms.
Safe-first helperscripts/download/llp.sh
Audio understanding, generation & events
LOCATA
LOCATA: IEEE-AASP Challenge on Acoustic Source Localization and Tracking
Safe-first helper
Sound Source Localization
Acoustic Source Tracking
Multi Source Localization
Moving Source Localization
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ODC-By-1.0
- Code license
- not_specified
- License caution
- Zenodo identifies the final dataset release as Open Data Commons Attribution, and the official corpus page says both LOCATA and its VCTK speech material use Open Data Commons terms. Neither official MATLAB repository exposes a LICENSE file or a GitHub-detected license, so clarify software terms before redistribution.
- Download notes
- The open final release contains corrected development and evaluation datasets with close-talking speech, distant multichannel recordings from four microphone-array configurations, and OptiTrack ground truth for sources and microphones. Its six tasks span static and moving, single- and multi-source scenarios with static or moving arrays. The dev.zip and eval.zip archives total about 19.3 GB, so the helper downloads only official pages, documentation, Zenodo metadata, and tool metadata by default.
Safe-first helperscripts/download/locata.sh
Audiovisual & cross-modal
LRRo
LRRo: A Lip Reading Data Set for the Under-resourced Romanian Language
Safe-first helper
Visual Speech Recognition
Lip Reading
Isolated Word Recognition
Low Resource Romanian Speech
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- The Zenodo record declares CC BY 4.0. Wild LRRo was collected from Internet videos and the release contains identifiable speaker imagery, so attribution, privacy, likeness, and any source-media rights should still be reviewed for redistribution, biometric use, or deployment. The documentation-only GitHub repository has no detected license.
- Download notes
- The public, ungated Zenodo release contains Wild LRRo, with more than 20 hours, over 35 speakers, 1,100 word instances, and a 21-word vocabulary, plus Lab LRRo, with more than 5 hours, 19 speakers, 6,400 word instances, and a 48-word vocabulary. Both isolated-word collections provide train, validation, and test subsets. VSRo-200 section 4.6 evaluates transfer on both LRRo subsets. The helper saves the official Zenodo record and repository documentation by default; the 291,504,956-byte archive is opt-in with LRRO_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/lrro.sh
Audiovisual & cross-modal
LRS2-BBC
The Oxford-BBC Lip Reading Sentences 2 Dataset
Manual or gated
Audio Visual Speech Recognition
Visual Speech Recognition
Lip Reading
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- non-commercial_academic_access_agreement
- Code license
- not_applicable
- License caution
- BBC R&D restricts LRS2 to non-commercial research by universities, reputable academic institutions, and relevant public organizations; companies and independent researchers are not eligible. The signed agreement controls use, and the BBC asks researchers to obtain approval before publishing sample images. Review the current agreement and BBC broadcast-content rights before use.
- Download notes
- The official VGG page describes 144,482 pre-train/train/validation/test utterances from BBC television, with date-disjoint validation and test broadcasts, and links a password-protected 50 GB package plus public split file links. Access requires submitting the BBC LRS2 Data Sharing Agreement from an official academic address; approved researchers receive a countersigned agreement and password. The helper saves the official pages, agreement, and paper metadata only and does not request credentials or download the corpus.
Safe-first helperscripts/download/lrs2.sh
Audiovisual & cross-modal
LRS3-TED
LRS3-TED: A Large-Scale Dataset for Visual Speech Recognition
Manual or gated
Audio Visual Speech Recognition
Visual Speech Recognition
Lip Reading
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- unspecified
- Code license
- not_applicable
- License caution
- No LRS3 data license is stated on the currently accessible official landing page or paper record. TED/TEDx source-video rights remain with their owners; obtain the official terms before use and do not infer reuse rights from third-party mirrors.
- Download notes
- The paper introduces more than 400 hours of aligned face tracks, audio, subtitles, and word boundaries from TED and TEDx for visual and audio-visual speech recognition. The official VGG landing page still identifies LRS3 and links its dataset page, but that linked page returned HTTP 404 when checked on 2026-07-22. The helper preserves the live official landing page and paper metadata and reports the unavailable official download route; it does not substitute an unofficial mirror.
Safe-first helperscripts/download/lrs3.sh
Audiovisual & cross-modal
LVOmniBench
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Manual or gated
Audio Visual Question Answering
Long Video Understanding
Cross Modal Reasoning
Temporal Localization
+2 more
Access pathHugging Face
Upstream termsNot specified
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The paper says every source video carries a Creative Commons license, but neither the official dataset card nor repository states a benchmark-level license or records each video's exact Creative Commons variant in public documentation. The GitHub repository has no LICENSE file or detected license. Review the access form, per-video rights, attribution requirements, and YouTube terms before reuse or redistribution; the paper's CC BY 4.0 license covers the article, not automatically the dataset or evaluation code.
- Download notes
- The gated Hugging Face release contains 275 English Creative Commons-licensed YouTube videos totaling 140 hours, with durations from 10 to 90 minutes, plus 1,014 manually authored four-option question-answer pairs. Questions span nine categories and require joint reasoning over speech, music, or sound with visual evidence. The helper downloads public repository documentation and API metadata by default. The Hugging Face API reports approximately 187.4 GB of repository storage, so the complete snapshot requires approval, authentication, and LVOMNIBENCH_DOWNLOAD_HF=1. OmniReasoner evaluates LVOmniBench in its main audio-visual table and reports duration-wise tool-use behavior on the benchmark.
Safe-first helperscripts/download/lvomnibench.sh
Enhancement, separation & quality
Lyra-SA
Lyra Lab Singing Assessment Dataset
Manual or gated
Singing Quality Assessment
Full Song Singing Assessment
Singing Score Prediction
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- not_applicable
- License caution
- The official page states CC BY-NC 4.0 for non-commercial use, requires attribution to the source page and notice, reserves copyright to Tencent Music Entertainment Group, and requires separate permission for commercial use. The recordings are WeSing user performances of copyrighted songs; rely on the official authorization and application terms rather than inferring broader music or performance rights from the Creative Commons label.
- Download notes
- The official Tencent Music Lyra Lab page describes 1,000 complete mobile-karaoke recordings: 100 user covers for each of 10 Chinese songs, with no repeated singer. The release includes 44.1 kHz 16-bit mono WAV audio, listener-provided overall singing scores, rough singer labels, timed lyrics, and reference MIDI. Access is application-based: users must complete the official form and accept its terms, after which Lyra Lab says it emails a download link within three days. The helper saves official documentation and the recent SongSQA paper, then prints the manual application path; it never guesses or bypasses an emailed archive URL.
Safe-first helperscripts/download/lyra_sa.sh
Audio understanding, generation & events
MACS
MACS: Multi-Annotator Captioned Soundscapes
Safe-first helper
Audio Captioning
Audio Tagging
Multi Annotator Audio Labeling
Acoustic Scene Captioning
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other-nc
- Code license
- not_applicable
- License caution
- Zenodo lists MACS as Other (Non-Commercial), and LICENSE.txt grants experimental non-commercial use with attribution to Tampere University/Machine Listening Group. Audio files come from TAU Urban Acoustic Scenes 2019, whose Zenodo record also lists Other (Non-Commercial).
- Download notes
- The helper downloads the MACS annotations, competence scores, license, and TAU Urban Acoustic Scenes 2019 docs/metadata by default. The upstream TAU 2019 audio shards are large, so audio download is an explicit opt-in.
Safe-first helperscripts/download/macs.sh
Music
MADB
MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
Safe-first helper
Music Aesthetic Assessment
Music Quality Assessment
Music Score Regression
Multimodal Music Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_upstream_audio_rights
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 and research-only use, but also warns that audio may remain subject to original copyright restrictions. No license file is present in the official GitHub repository. Apply the non-commercial dataset terms and independently review MuChin, generator-service/output, and unidentified online-source rights before redistributing or using audio.
- Download notes
- The paper describes 9,999 tracks rated by 30 trained annotators, with about ten ratings per track across ten perceptual dimensions and an overall score plus comments and tags. The public, ungated Hugging Face release includes all annotations and 1,730 Suno/Levo audio tracks; its card points to the separate MuChin repository for another 4,400 tracks and says remaining tracks came from diverse online sources. The helper downloads official documentation and repository metadata by default, makes the approximately 69 MB annotation tables opt-in, and requires MADB_DOWNLOAD_HF=1 for the approximately 18.6 GB Hugging Face snapshot.
Safe-first helperscripts/download/madb.sh
Representation & general suites
MAEB
MAEB: Massive Audio Embedding Benchmark
Safe-first helper
Audio Embedding Evaluation
Audio Text Retrieval
Audio Classification
Audio Clustering
+6 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_component_terms
- Code license
- Apache-2.0
- License caution
- Apache-2.0 covers the MTEB software and benchmark registry, not the underlying MAEB task datasets. The 30 tasks draw on multiple speech, music, and environmental-audio sources with their own licenses, access controls, attribution requirements, and media rights. Review every selected task's metadata and upstream dataset terms before downloading, redistributing, or using it commercially.
- Download notes
- The public MTEB registry defines the 30-task MAEB beta suite across speech, music, environmental sound, and audio-text evaluation in more than 100 languages. It includes retrieval, classification, clustering, pair classification, reranking, multilabel classification, and zero-shot classification tasks. The helper downloads official documentation, repository metadata, the Apache-2.0 license, and the lightweight benchmark registry by default; cloning the approximately 55 MB MTEB source repository is opt-in. It does not fetch component datasets, which MTEB tasks acquire separately and which can be large or restricted.
Safe-first helperscripts/download/maeb.sh
Music
MAESTRO
MAESTRO: MIDI and Audio Edited for Synchronous TRacks and Organization
Safe-first helper
Automatic Music Transcription
Piano Transcription
Music Synthesis
Symbolic Music Modeling
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_applicable
- License caution
- Official Magenta page says the dataset is made available by Google LLC under Creative Commons Attribution Non-Commercial Share-Alike 4.0.
- Download notes
- The helper downloads v3.0.0 CSV/JSON metadata by default. The MIDI-only archive is about 56 MiB and the full WAV+MIDI archive is about 101 GiB, so both are explicit opt-ins.
Safe-first helperscripts/download/maestro.sh
Audio understanding, generation & events
MAESTRO Real
MAESTRO Real: Multi-Annotator Estimated Strong Labels
Safe-first helper
Sound Event Detection
Soft Label Sound Event Detection
Multi Annotator Label Aggregation
Long Form Sound Event Detection
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_noncommercial
- Code license
- not_applicable
- License caution
- The packaged Tampere University license permits copying and use only for experimental and non-commercial purposes with the full copyright notice and source acknowledgment. It explicitly prohibits commercial use, including selling or distributing results or content achieved through use of the dataset.
- Download notes
- The public development release contains 49 real-life recordings from five acoustic scenes with crowdsourced soft strong labels for 17 classes. Recordings are three to five minutes long and come from subsets of TUT Sound Events 2016 and 2017. DCASE 2024 Task 4 combines MAESTRO Real with DESED and evaluates 11 MAESTRO classes; its separate 26-file MAESTRO evaluation set is not included in the public development archive. The helper downloads official metadata, README, license, and the sub-megabyte annotation archive by default; the approximately 2.43 GiB audio archive requires explicit opt-in. The official Zenodo description reports 189 minutes 52 seconds total while its packaged README reports 97 minutes 4 seconds, so verify duration against the downloaded files rather than assuming either figure.
Safe-first helperscripts/download/maestro_real.sh
Speech recognition
MAGICDATA Mandarin Chinese Read Speech Corpus
Safe-first helper
Automatic Speech Recognition
Speaker Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR lists CC BY-NC-ND 4.0 and says the corpus is freely published for non-commercial or academic use. Re-check the current official page before redistribution or commercial use.
- Download notes
- The helper downloads the OpenSLR page and small metadata archive by default. Speech archives are large, including about 52 GiB train, 1.0 GiB dev, and 2.2 GiB test, so archive download is an explicit opt-in.
Safe-first helperscripts/download/magicdata_mandarin.sh
Music
MagnaTagATune
MagnaTagATune: A Music Annotation Benchmark from the TagATune Game
Safe-first helper
Music Auto Tagging
Music Annotation
Music Similarity
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-3.0
- Code license
- GPL-3.0
- License caution
- The original TagATune details page says the data is CC BY-NC-SA 3.0 except scripts released under GPL v3, enabling non-commercial research redistribution. Audio clips are Magnatune excerpts with artist/album URLs for purchase or commercial licensing.
- Download notes
- City University MIRG hosts metadata, annotations, comparisons, Echo Nest features, and three 1 GiB MP3 split archives. The helper downloads CSV metadata by default and makes features/audio explicit opt-ins.
Safe-first helperscripts/download/magnatagatune.sh
Speaker, identity & emotion
MCR-Bench
MCR-Bench: Modal Conflict Resolution Benchmark for Large Audio-Language Models
Safe-first helper
Cross Modal Conflict Resolution
Audio Text Robustness
Audio Question Answering
Speech Emotion Recognition
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_upstream_terms
- Code license
- apache-2.0
- License caution
- The repository contains an Apache-2.0 LICENSE, although its README badge says MIT. Neither statement clearly relicenses the released benchmark archive or embedded source audio. ClothoAQA/Clotho, MELD's copyrighted Friends clips, and VocalSound retain their own terms, so verify item-level provenance and upstream permissions before redistribution or commercial use.
- Download notes
- The public Google Drive release contains approximately 3,000 English samples across audio question answering, speech emotion recognition, and vocal-sound classification. Each audio item is paired with faithful, adversarial, irrelevant, and neutral text conditions to measure whether an audio-language model follows contradictory text instead of audio evidence. The benchmark derives its task audio from ClothoAQA, MELD, and VocalSound. The helper downloads official documentation and repository metadata by default and prints the manual Drive path; cloning the documentation-only repository is opt-in.
Safe-first helperscripts/download/mcr_bench.sh
Enhancement, separation & quality
MedleyDB
MedleyDB: A Multitrack Dataset for Annotation-Intensive MIR Research
Safe-first helper
Music Information Retrieval
Melody Extraction
Music Source Separation
Instrument Recognition
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- MIT
- License caution
- The official downloads page says MedleyDB is free for non-commercial research use only and identifies the dataset as Creative Commons Attribution-NonCommercial-ShareAlike 4.0. It also asks users not to republish the dataset in full or in part without consent, even though redistribution is technically allowed under the license. The GitHub tooling repository is MIT.
- Download notes
- The helper saves official pages and repository license/README by default. It checks the Zenodo request records only with MEDLEYDB_CHECK_ZENODO=1, downloads the public sample archive only with MEDLEYDB_DOWNLOAD_SAMPLE=1, and clones the annotation/metadata/tooling repo only with MEDLEYDB_CLONE_REPO=1. Full MedleyDB and MedleyDB 2.0 audio require requesting access through the official Zenodo records.
Safe-first helperscripts/download/medleydb.sh
Audiovisual & cross-modal
MELD
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
Safe-first helper
Speech Emotion Recognition
Multimodal Emotion Recognition
Dialogue Emotion Recognition
Sentiment Analysis
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- gpl-3.0
- Code license
- GPL-3.0
- License caution
- GitHub and the Hugging Face dataset card list GPL-3.0. MELD clips are derived from the Friends TV series, so downstream users should re-check media rights and any fair-use/research assumptions before redistribution or commercial use.
- Download notes
- The helper downloads official project, README, license, and dataset-card metadata by default. Raw audio/video tarballs and extracted feature/model tarballs are opt-in because they are larger and include TV-derived media clips.
Safe-first helperscripts/download/meld.sh
Audiovisual & cross-modal
MER2023
MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
Manual or gated
Multimodal Emotion Recognition
Discrete Emotion Classification
Dimensional Emotion Recognition
Modality Robustness
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_additional_gated_terms
- Code license
- not_specified
- License caution
- The current Hugging Face card declares CC BY-NC 4.0 and limits access to academic research. Its gate also prohibits handing the database or derived labeling files to third parties and prohibits modification without written consent. The challenge paper describes a separate EULA with academic-only, no-editing, and no-upload conditions. The baseline repository has no LICENSE file. Clips were collected from movies and television, so underlying media, performer, privacy, and platform rights remain separate.
- Download notes
- The request-gated release extends CHEAVD and provides 3,373 labeled Train&Val clips, 411 MER-MULTI test clips, 412 corrupted-modality MER-NOISE test clips, and MER-SEMI with 834 labeled plus 73,148 unlabeled clips. MER-MULTI evaluates joint discrete-emotion and valence prediction, MER-NOISE tests robustness to noisy audio and blurred video, and MER-SEMI evaluates semi-supervised discrete emotion recognition. The helper saves public paper, repository, and Hugging Face API metadata only. The current Hugging Face repository is about 140 GB, password-protected, and requires approval, so benchmark files are left as a manual download.
Safe-first helperscripts/download/mer2023.sh
Audiovisual & cross-modal
MER2024
MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
Manual or gated
Multimodal Emotion Recognition
Discrete Emotion Classification
Semi Supervised Emotion Recognition
Modality Robustness
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_additional_gated_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0 and non-commercial use. Its gate prohibits transfer of the database or derived labeling files to third parties and modification without written consent. The challenge page additionally limits the dataset to academic research and prohibits uploading samples. The MER2024 README displays an Apache-2.0 badge but links to a license under the later MER2025 directory; the repository root and MER2024 directory contain no applicable LICENSE file, so the MER2024 code license is recorded as unspecified. Source-video, performer, privacy, and platform rights remain separate.
- Download notes
- This request-gated extension of MER2023 provides 5,030 labeled Train&Val clips and 115,595 unlabeled clips. MER-SEMI evaluates 1,169 annotated clips from the unlabeled pool, MER-NOISE evaluates 1,170 clips with additive audio noise and image blur, and MER-OV evaluates free-form emotion labels. Table 1 of the paper reports 322 MER-OV samples, while the surrounding prose says 332; the index preserves that primary-source discrepancy. The current Hugging Face tree is approximately 218.4 GB and requires approval. The helper saves only public paper, project, repository, and Hugging Face API metadata.
Safe-first helperscripts/download/mer2024.sh
Speech recognition
MInDS-14
MInDS-14: Multilingual and Cross-Lingual Intent Detection from Spoken Data
Safe-first helper
Spoken Language Understanding
Intent Classification
Automatic Speech Recognition
Multilingual Speech Understanding
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Hugging Face dataset card lists CC BY 4.0 and exposes 14 spoken e-banking intents across 14 language varieties. No separate code license was identified for the dataset card.
- Download notes
- The helper downloads the Hugging Face dataset card by default. Dataset snapshots include audio and are opt-in; choose one locale such as en-US or all with MINDS14_CONFIG.
Safe-first helperscripts/download/minds14.sh
Speech generation
Ming-Freeform-Audio-Edit
Ming-Freeform-Audio-Edit: Free-form instruction-based speech editing benchmark
Safe-first helper
Instruction Based Speech Editing
Semantic Speech Editing
Speech Content Deletion
Speech Content Insertion
+6 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_with_upstream_terms
- Code license
- not_specified
- License caution
- The Hugging Face dataset card declares Apache-2.0, but the released audio is derived from multiple upstream corpora whose terms still apply. Seed-TTS Eval states no data license, LibriTTS is CC BY 4.0, and GigaSpeech uses its own agreement/access conditions. The evaluation repository has no license file or detected GitHub license; review each source corpus and clarify annotation/code rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains Chinese and English source speech, natural-language editing instructions, and metadata for semantic deletion, insertion, and substitution plus acoustic emotion, dialect, speed, pitch, and volume changes. The paper's section 6.3 constructs the semantic set from 896 Chinese and 655 English Seed-TTS test samples and reports separate basic and full instruction versions; its acoustic sets also use Seed-TTS audio. The current dataset card additionally names LibriTTS and GigaSpeech as source corpora. The helper downloads official documentation and repository metadata by default; the Hugging Face API reports approximately 1.07 GB of repository storage, so audio and annotations require explicit opt-in.
Safe-first helperscripts/download/ming_freeform_audio_edit.sh
Speech recognition
MIR-1K vocal
MIR-1K
Safe-first helper
Singing Voice Transcription
Singing Voice Separation
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_on_official_page
- Code license
- not_applicable
- License caution
- Official MIR Lab page has direct downloads but no visible license statement; its MIR-1K.rar URL returned 404 on 2026-07-09. Figshare mirror lists CC BY 4.0.
Safe-first helperscripts/download/mir_1k_vocal.sh
Speech recognition
MLC-SLM Eval
Multilingual Conversational Speech Language Model Challenge Eval Ground Truth
Safe-first helper
Multilingual Conversational Asr
Speaker Diarization
Speaker Attributed Asr
Long Form Speech Recognition
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0_annotation_card_label
- Code license
- not_specified
- License caution
- The Hugging Face card labels the released ground-truth repository CC BY-SA 4.0. It does not publish a separate license file or establish terms for the absent evaluation recordings. The official baseline repository has no detected license, so do not assume the annotation label licenses its code. Obtain the audio and its terms from the challenge owners before attempting full benchmark reproduction.
- Download notes
- The public, ungated Hugging Face release contains approximately 6.55 MB of oracle segmentation, speaker labels, and transcriptions for the challenge's 32-hour Eval-1 and 32-hour Eval-2 sets. It covers English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese; English additionally spans five accent groups. The repository currently contains 224 annotation files and no audio. Challenge participants previously received evaluation recordings, but the official paper, dataset card, and baseline repository provide no current public audio URL. The helper downloads official documentation and repository metadata by default; the lightweight annotation snapshot requires MLC_SLM_EVAL_DOWNLOAD_HF=1 and does not include audio. VibeVoice-ASR evaluates the MLC-Challenge set, while VibeVoice-ASR-BitNet section 3.1 and Table 4 report six MLC language subsets.
Safe-first helperscripts/download/mlc_slm_eval.sh
Speech recognition
MLS
MLS: A Large-Scale Multilingual Dataset for Speech Research
Safe-first helper
Multilingual Asr
Asr
Language Modeling
Limited Supervision Asr
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR94 lists CC BY 4.0. MLS is derived from LibriVox audiobooks and provides public-file-hosted archives plus MD5 checksums.
- Download notes
- OpenSLR links original FLAC and compressed OPUS archives for English, German, Dutch, French, Spanish, Italian, Portuguese, and Polish; archives range from about 1.6 GiB to multiple TiB, so the helper downloads only the page and checksums by default and requires MLS_DOWNLOAD_ARCHIVES=1 for audio.
Safe-first helperscripts/download/mls.sh
Enhancement, separation & quality
MMAE
MMAE: A Massive Multitask Audio Editing Benchmark
Safe-first helper
Instruction Based Audio Editing
Multi Round Audio Editing
Multi Hop Audio Editing
Speech Editing
+4 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The Hugging Face card does not declare a dataset license or identify licenses for the source audio. The official GitHub repository has an MIT LICENSE for its code, but that must not be assumed to license the benchmark audio. Review source-media rights and obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated release contains 2,000 high-fidelity input samples across sound, speech, music, and mixtures, organized by six complexity levels, two granularities, and eight operation types. Its 17,741 rubric criteria evaluate instruction following and context consistency. The helper downloads official documentation and repository metadata by default; cloning the evaluation repository is opt-in, and the Hugging Face API reports approximately 4.43 GB of repository storage, so the audio snapshot requires MMAE_DOWNLOAD_HF=1.
Safe-first helperscripts/download/mmae.sh
Audio understanding, generation & events
MMAR
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Safe-first helper
Audio Question Answering
Audio Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- GitHub repo did not expose a detected license; HF dataset card lists cc-by-nc-4.0.
Safe-first helperscripts/download/mmar.sh
Audio understanding, generation & events
MMAU
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Safe-first helper
Audio Question Answering
Audio Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- HF cards: MMAU-test-mini is cc-by-nc-4.0; MMAU-test is mit
- Code license
- Apache-2.0
- License caution
- The two HF dataset cards list different licenses; re-check before redistribution.
Safe-first helperscripts/download/mmau.sh
Audio understanding, generation & events
MMAU-Pro
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Safe-first helper
Audio Question Answering
Audio Reasoning
Long Audio Understanding
Multi Audio Reasoning
+5 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_upstream_terms
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0, but the paper says almost all audio was sourced from in-the-wild recordings and the spatial subset reuses EasyCom. Review source-media and EasyCom terms before redistribution or commercial use. The official GitHub repository has no license file or detected license, so evaluator code terms are unspecified.
- Download notes
- The public, ungated release contains 5,305 expert-authored multiple-choice and open-ended QA instances spanning 49 skills across speech, environmental sound, music, and their mixtures. It includes multiple-audio, spatial, instruction-following, and up-to-10-minute long-form cases. The helper downloads official documentation, repository metadata, the evaluator, and Hugging Face metadata by default; the Hugging Face API reports approximately 47.5 GB of repository storage, so the audio and test snapshot requires MMAU_PRO_DOWNLOAD_HF=1.
Safe-first helperscripts/download/mmau_pro.sh
Speech generation
MMGenre
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Safe-first helper
Singing Voice Synthesis
Genre Conditioned Singing Voice Synthesis
Singing Voice Genre Alignment
Score Conditioned Singing Voice Synthesis
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_generated_source_caveat
- Code license
- cc-by-4.0
- License caution
- The dataset card, dataset LICENSE, and repository LICENSE declare CC BY 4.0. The benchmark was derived from music generated with Suno V4.5 and then source-separated and automatically aligned; the release license does not independently resolve any terms or rights attached to the generation service or generated source content, so review those conditions for downstream use.
- Download notes
- The public, ungated Hugging Face release contains 3,152 aligned Chinese singing-voice and symbolic-score segments from 148 generated songs, totaling about 4.36 hours across 10 major genres and 26 subgenres. The paper body says 27 subgenres, but the dataset card identifies this as a counting error and treats the released 26-subgenre taxonomy as authoritative. The helper downloads official documentation and API metadata by default, can clone the lightweight score/code repository, and requires explicit opt-in for the approximately 5.54 GB Hugging Face snapshot.
Safe-first helperscripts/download/mmgenre.sh
Audiovisual & cross-modal
MMOU
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
Safe-first helper
Audio Visual Question Answering
Long Video Understanding
Cross Modal Reasoning
Temporal Grounding
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0 metadata label; non-commercial research restriction in paper
- Code license
- not_applicable
- License caution
- Both Hugging Face cards declare Apache-2.0, but the MMOU paper says the dataset is released solely for non-commercial research. Apply the stricter non-commercial restriction pending clarification. Videos were collected from public web platforms including YouTube, so source copyright, platform terms, availability, and any per-video rights also apply.
- Download notes
- The public, ungated NVIDIA release contains 20,000 English questions over 11,877 long-form web videos. The 5,000-item Test Mini split includes labels for local evaluation; answers for the main 15,000-item split are withheld for evaluator submission. The helper downloads official cards and API metadata by default, makes the approximately 48 MB question files opt-in, and requires a separate opt-in for the approximately 322.8 GB community-hosted video snapshot. Audio-Visual Flamingo evaluates MMOU as an omni-modal benchmark.
Safe-first helperscripts/download/mmou.sh
Speech understanding & dialogue
MMSU
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Safe-first helper
Spoken Language Understanding
Speech Reasoning
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- not_specified
- License caution
- HF dataset card lists MIT; GitHub code repo did not expose a detected license.
Safe-first helperscripts/download/mmsu.sh
Enhancement, separation & quality
MoisesDB
MoisesDB: A Dataset for Source Separation Beyond 4-Stems
Safe-first helper
Music Source Separation
Multi Stem Source Separation
Instrument Source Separation
Singing Voice Separation
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- cc-by-nc-sa-4.0
- License caution
- The official repository applies CC BY-NC-SA 4.0 to MoisesDB and its packaged loader/evaluation materials, and the Music AI page limits the dataset to non-commercial research use. Attribution and ShareAlike obligations apply; commercial use is not permitted by this release.
- Download notes
- The official Music AI research page provides the dataset through its browser download flow. The release contains 240 songs by 47 artists across 12 high-level genres, totaling 14 hours, 24 minutes, and 46 seconds, with mixtures and a hierarchical stem/source taxonomy extending beyond the common four-stem setup. The helper saves the official page plus repository README and license by default; it does not fetch the large audio archive, and cloning the loader/evaluation repository is opt-in.
Safe-first helperscripts/download/moisesdb.sh
Speech generation
MS-SNSD
Microsoft Scalable Noisy Speech Dataset
Safe-first helper
Speech Enhancement
Speech Denoising
Noisy Speech Synthesis
Subjective Speech Quality Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The README says Microsoft provides the datasets as-is under the original terms received. It lists PTDB-TUG under ODbL 1.0, Edinburgh/VoiceBank material under its DataShare license, selected Freesound noise as CC0, and DEMAND as CC BY-SA 3.0. Re-check component terms before redistribution or commercial use.
- Download notes
- The official GitHub repository contains clean speech, noise, noisy test, and clean test directories plus scripts for generating noisy speech at configurable SNRs. The helper saves README/license/generator files by default; cloning the large repository is an explicit opt-in.
Safe-first helperscripts/download/ms_snsd.sh
Speaker, identity & emotion
MSP-Podcast
MSP-Podcast: A Large Naturalistic Speech Emotional Dataset
Manual or gated
Speech Emotion Recognition
Dimensional Emotion Recognition
Categorical Emotion Recognition
Speaker Independent Emotion Recognition
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_academic_license
- Code license
- not_applicable
- License caution
- The owner page currently calls the release an Academic License and requires an institution-signed FDP data-transfer agreement. Although the page says source podcasts were chosen under permissive licenses, the signed corpus agreement controls access and reuse; review it directly before commercial use, redistribution, or sharing copies.
- Download notes
- Version 2.0 contains 264,705 naturalistic podcast speaking turns totaling 409 hours, with speaker-independent train, development, and three test partitions. It provides categorical emotion and activation, dominance, and valence labels; Test3 releases audio but withholds labels, speaker information, transcripts, and alignments for web-based evaluation. Access is free but institution/form-gated: an authorized institutional representative must sign the official academic agreement and send it to the corpus owner. There is no public archive URL.
Safe-first helperscripts/download/msp_podcast.sh
Speech understanding & dialogue
MSU-Bench
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Safe-first helper
Multi Speaker Conversation Understanding
Speaker Identification
Speaker Attribute Recognition
Speaker Verification
+4 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_upstream_terms
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY-NC 4.0 and limits the release to non-commercial academic research because audio derives from third-party film/TV, telephone, meeting, and podcast sources. The repository LICENSE applies MIT only to code and explicitly gives the benchmark data separate non-commercial academic-research terms; review each source corpus and media right before redistribution. The paper reports approximately 731 hours of source corpora, not 731 released hours.
- Download notes
- The public, ungated release contains 2,847 English and Chinese four-choice QA items over 241 multi-speaker audio clips, including 2,223 human-verified items across 16 tasks. The helper downloads official documentation, repository metadata, and the approximately 5.8 MB test JSONL by default; the Hugging Face API reports approximately 2.5 GB of repository storage, so the audio and annotations snapshot requires MSU_BENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/msu_bench.sh
Speech understanding & dialogue
MSWC
Multilingual Spoken Words Corpus
Safe-first helper
Keyword Spotting
Spoken Term Search
Multilingual Speech Classification
Forced Alignment
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- MLCommons and the HF card list CC BY 4.0. MSWC is derived from crowd-sourced sentence-level audio, including Common Voice, so preserve source attribution and re-check the active source terms for downstream redistribution.
- Download notes
- The official MLCommons page and HF card describe 50 languages, more than 340,000 keywords, 23.4 million 1-second examples, and over 6,000 hours. The helper saves official docs and the HF card by default; direct per-language audio, splits, and alignments are opt-in with MSWC_DOWNLOAD_ARCHIVES=1 because high-resource language audio archives can be many GiB.
Safe-first helperscripts/download/mswc.sh
Speech recognition
mTEDx
The Multilingual TEDx Corpus for Speech Recognition and Translation
Safe-first helper
Multilingual Asr
Speech To Text Translation
Speech Translation
Sentence Level Alignment
Access pathOpenSLR
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- OpenSLR SLR100 lists CC BY-NC-ND 4.0. The corpus is derived from TEDx Talks; respect TED/TEDx source terms as well as the corpus license.
- Download notes
- OpenSLR SLR100 hosts ASR-only language archives, speech-translation language-pair archives, IWSLT 2021 test sets, and a small French talk gender annotation CSV. Archives are multi-GiB, so the helper downloads only the OpenSLR page and small CSV by default and requires MTEDX_DOWNLOAD_ARCHIVES=1 for archive downloads.
Safe-first helperscripts/download/mtedx.sh
Music
MTG-Jamendo
MTG-Jamendo Dataset for Automatic Music Tagging
Safe-first helper
Music Auto Tagging
Music Genre Classification
Musical Instrument Recognition
Music Mood Theme Recognition
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- Apache-2.0
- License caution
- The GitHub README says repository metadata is CC BY-NC-SA 4.0, code is Apache-2.0, audio files keep individual Creative Commons licenses listed in audio_licenses.txt, and the dataset is made available solely for non-commercial research and academic use unless Jamendo grants separate authorization.
- Download notes
- The helper clones or updates the official metadata/scripts repository by default and saves Zenodo record metadata. The upstream downloader can fetch very large archives: raw_30s audio is about 508 GiB, raw_30s audio-low is about 156 GiB, and autotagging_moodtheme audio-low is about 46 GiB, so media downloads require MTG_JAMENDO_DOWNLOAD_MEDIA=1.
Safe-first helperscripts/download/mtg_jamendo.sh
Speaker, identity & emotion
MUGEN
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
Safe-first helper
Multi Audio Understanding
Audio Grounding
Speech Understanding
Speaker And Demographic Understanding
+5 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_with_mixed_upstream_terms
- Code license
- MIT
- License caution
- The 35 Hugging Face dataset cards do not declare a data license. The repository's MIT license identifies its subject as software and associated documentation, so it should not be assumed to relicense embedded audio. The paper describes public corpora, specialized academic corpora, Mozilla Data Collective speech, and synthesized speech as sources; review each task's upstream corpus and generation terms before redistribution or commercial use.
- Download notes
- The public, ungated release provides 35 separate Hugging Face task repositories with 1,750 five-way audio-grounding problems and 9,250 audio clips across seven dimensions. Ten tasks add a reference clip, producing six concurrent audio inputs. The helper downloads official documentation, repository metadata, the license, and the Hugging Face collection inventory by default; set MUGEN_DOWNLOAD_TASK to one of the documented task names to explicitly download that task. The current cards report about 10.7 GB of compressed files across all task repositories and also expose reduced-candidate splits for input-scaling analysis.
Safe-first helperscripts/download/mugen.sh
Music
MulTTiPop
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
Safe-first helper
Automatic Music Transcription
Multitrack Music Transcription
Audio Midi Alignment
Music Information Retrieval
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_for_released_midi_and_metadata
- Code license
- not_applicable
- License caution
- The dataset card declares CC BY 4.0 and says its aligned MIDI is adapted from the CC BY 4.0 Lakh MIDI Dataset. Source audio is not licensed or redistributed; users are instructed to obtain only the referenced segments and use MulTTiPop for evaluation rather than training. YouTube availability, platform terms, and commercial-song rights remain applicable.
- Download notes
- The public, ungated release contains aligned multitrack MIDI and metadata for 572 commercial-pop segments (about 3.5 hours), divided into artist-disjoint development and test splits. It does not redistribute source audio; metadata provides YouTube video identifiers and segment timestamps. The helper downloads the dataset card, API metadata, and lightweight split manifests by default, while the full MIDI/metadata snapshot requires MULTTIPOP_DOWNLOAD_HF=1.
Safe-first helperscripts/download/multtipop.sh
Speaker, identity & emotion
MUSAN
MUSAN: A Music, Speech, and Noise Corpus
Safe-first helper
Voice Activity Detection
Music Speech Discrimination
Speech Music Noise Classification
Speaker Recognition Augmentation
Access pathOpenSLR
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR17 lists Attribution 4.0 International (CC BY 4.0). The paper describes MUSAN as music, speech, and noise recordings released under a flexible Creative Commons license.
- Download notes
- OpenSLR lists the corpus archive as 11 GiB. The helper downloads the OpenSLR landing page by default and requires MUSAN_DOWNLOAD_ARCHIVE=1 for the large archive.
Safe-first helperscripts/download/musan.sh
Enhancement, separation & quality
MUSDB18
MUSDB18: A Corpus for Music Separation
Safe-first helper
Music Source Separation
Singing Voice Separation
Audio Source Separation
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other-non-commercial
- Code license
- MIT
- License caution
- Zenodo lists "Other (Non-Commercial)" and the official pages state educational/academic use only, with commercial use requiring copyright-holder permission. Track sources include DSD100/Mixing Secrets, MedleyDB CC BY-NC-SA 4.0, Native Instruments stems, and Easton Ellises/Heise CC BY-NC-SA 3.0 material.
- Download notes
- The helper saves the official SigSep and Zenodo pages by default. The 4.7 GiB compressed STEMS archive and 22.7 GiB uncompressed HQ archive require explicit terms acknowledgement and opt-in.
Safe-first helperscripts/download/musdb18.sh
Audiovisual & cross-modal
MUSIC-AVQA
MUSIC-AVQA: Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Safe-first helper
Audio Visual Question Answering
Multimodal Scene Understanding
Spatiotemporal Reasoning
Music Understanding
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- conflicting: GitHub API/LICENSE file reports MIT, while the README License section mentions GPLv3
- License caution
- No standalone dataset license was found on the official page or README on 2026-07-10. Raw musical-performance videos and extracted features should be treated as upstream media with terms to verify before redistribution or commercial use.
- Download notes
- The helper downloads the official project page, README/LICENSE, and public JSON QA annotations by default. Raw videos and large feature files are hosted through Google Drive and Baidu Drive links on the project page/README; they are not downloaded automatically.
Safe-first helperscripts/download/music_avqa.sh
Audio understanding, generation & events
MusICA-MetaBench
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music
Safe-first helper
Music Perception Question Answering
Audio Music Understanding
Symbolic Music Understanding
Sheet Music Understanding
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_CC-BY-4.0_and_CC-BY-SA-4.0_by_source
- Code license
- GPL-3.0-or-later
- License caution
- ChoraleBricks-derived benchmark instances are CC BY 4.0, while ChoralSynth-derived instances are CC BY-SA 4.0. Original code, question templates, ontology, and configurations are GPL-3.0-or-later. Model outputs and inference logs are provided without a separate license. Source recordings and scores are downloaded separately and retain the upstream dataset terms; select the license by each benchmark file's documented source rather than treating the repository as uniformly licensed.
- Download notes
- The public repository releases the evaluation and benchmark-generation pipeline, question templates, ontology, configurations, inference logs, and pre-generated benchmark instances. The paper's main ChoraleBricks instance and its ChoralSynth validation instance each contain 300 five-option questions balanced across audio, symbolic, and sheet-image modalities; 20% use "none of the other options" as the correct answer. Items test pitch, rhythm, and harmony perception. The helper downloads official documentation, licenses, configs, and both sub-500 KB TSV instances by default. Cloning the approximately 49 MB GitHub repository is opt-in. Source audio and scores are not stored in that repository and must be obtained from ChoraleBricks or ChoralSynth under their respective terms.
Safe-first helperscripts/download/musica_metabench.sh
Audio understanding, generation & events
MusicCaps
MusicCaps: A Dataset of Music Captions
Safe-first helper
Music Captioning
Text To Music Evaluation
Music Understanding
Audio Language Modeling
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- The Hugging Face card and Kaggle-converted metadata list CC BY-SA 4.0 for the annotation CSV. The referenced media are 10-second clips from AudioSet/YouTube, so original media copyright, platform terms, and availability still apply.
- Download notes
- The public release is a CSV of 5,521 music-text pairs with YouTube IDs, segment timestamps, AudioSet labels, aspect lists, and musician-written captions. Raw audio is not mirrored by the dataset and must be reconstructed from YouTube/AudioSet subject to upstream availability and terms.
Safe-first helperscripts/download/musiccaps.sh
Music
MusicNet
MusicNet: A Dataset for Music Transcription and Multi-label Classification
Safe-first helper
Music Transcription
Multi Label Music Classification
Note Onset Labeling
Musical Instrument Recognition
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo lists CC BY 4.0 for the MusicNet release. The record says audio recordings are Creative Commons licensed and Public Domain performances from the Isabella Stewart Gardner Museum, the European Archive Foundation, and Musopen, with per-recording provenance described in the metadata.
- Download notes
- The helper saves the Zenodo record JSON and 44 KiB metadata CSV by default. Reference MIDI files are a small opt-in download, while the full audio/label archive is about 10.3 GiB and requires MUSICNET_DOWNLOAD_AUDIO=1.
Safe-first helperscripts/download/musicnet.sh
Enhancement, separation & quality
NISQA
NISQA Speech Quality Corpus
Safe-first helper
Speech Quality Assessment
Mean Opinion Score Prediction
Non Intrusive Speech Quality
Multidimensional Speech Quality
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- MIT
- License caution
- The GitHub README says the corpus is provided under the original terms of the source speech and noise samples, generally non-commercial research with some subsets permitting commercial use; individual README/license files inside the archive should be checked. The Zenodo record reports license id other-at, and model weights are CC BY-NC-SA 4.0.
- Download notes
- The helper downloads the official README, corpus wiki markdown, model-weight license, and Zenodo record JSON by default. The full NISQA_Corpus.zip archive is about 15.9 GB and requires NISQA_DOWNLOAD_CORPUS=1.
Safe-first helperscripts/download/nisqa.sh
Speech recognition
NOTSOFAR-1
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
Safe-first helper
Distant Automatic Speech Recognition
Speaker Attributed Automatic Speech Recognition
Speaker Diarization
Continuous Speech Separation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- Microsoft applies CC BY 4.0 to the released data and MIT to the baseline repository. The repository notes that challenge-only Dev-set-2 is not part of the current open release and is restricted to publications about systems developed during the challenge; use the current public subsets for new work.
- Download notes
- The public English recorded-meeting release currently documents 237 meetings averaging six minutes across 30 conference rooms, with 4-8 attendees and 35 speakers. It provides train, dev, an 80-meeting eval-small set matching the challenge evaluation set, and a 129-meeting eval-full set, with ground truth available for both evaluation releases. The original challenge and paper describe roughly 280 meetings; the post-challenge open release removes Dev-set-2 for legal and quality reasons and adds eval-full, so benchmark versions must be reported. Single-channel and known-geometry seven-channel tracks use speaker-attributed tcpWER for ranking. The helper downloads only official documentation and license files by default; use the repository's versioned download utilities for selected audio subsets. The Hugging Face repository has very large historical storage and must not be snapshotted wholesale.
Safe-first helperscripts/download/notsofar_1.sh
Music
NSynth
NSynth: Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
Safe-first helper
Audio Synthesis
Music Synthesis
Musical Instrument Modeling
Timbre Modeling
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- Apache-2.0
- License caution
- Official Magenta dataset page lists the dataset under CC BY 4.0; Magenta code is Apache-2.0.
- Download notes
- The full TFDS download is about 73 GiB. The helper saves the official dataset page by default and requires NSYNTH_DOWNLOAD_ARCHIVES=1 before downloading split archives.
Safe-first helperscripts/download/nsynth.sh
Audiovisual & cross-modal
Omni-Cloze
Omni-Cloze: Omni Detailed Captioning Benchmark
Safe-first helper
Detailed Audio Captioning
Detailed Visual Captioning
Detailed Audio Visual Captioning
Cloze Question Answering
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- Neither the official Hugging Face card nor the GitHub repository states a data or code license, and GitHub reports no detected license. The release includes video files without documented source-media provenance or reuse terms; obtain clarification and review media rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains 2,320 audio-visual files across nine main domains and 47 subcategories, with approximately 70,000 fine-grained cloze blanks. The helper downloads official documentation and evaluation scripts by default; the 25.1 MB JSONL metadata and Hugging Face snapshot are separate opt-ins. The current repository files total about 6.1 GB, while the Hugging Face API reports about 11.3 GB of repository storage including history. Qwen3.5-Omni evaluates detailed audio-visual captioning on Omni-Cloze in section 5.1.4, Table 7.
Safe-first helperscripts/download/omni_cloze.sh
Audiovisual & cross-modal
OmniBench
OmniBench: Towards the Future of Universal Omni-Language Models
Safe-first helper
Omni Modal Question Answering
Audio Visual Question Answering
Tri Modal Reasoning
Speech Understanding
+2 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official dataset card has no license field, and the official repository has no LICENSE file or GitHub-detected license. The project page's footer links CC BY-SA 4.0 but does not clearly state that it covers the benchmark data, code, or component media. Obtain clarification and review source-image/audio rights before redistribution or commercial use.
- Download notes
- The public, ungated release contains 1,142 four-choice questions that jointly pair an image, audio, and text prompt across seven task types. Audio spans speech, sound events, and music. The helper downloads official documentation and repository metadata by default; the Hugging Face card reports approximately 1.26 GB of downloads, so the media-bearing Parquet snapshot requires OMNIBENCH_DOWNLOAD_HF=1. OPOD evaluates OmniBench as its omni-modal benchmark and reports accuracy in the Experiments benchmark block of arXiv:2607.20918.
Safe-first helperscripts/download/omnibench.sh
Audiovisual & cross-modal
OmniGAIA
OmniGAIA: Towards Native Omni-Modal AI Agents
Safe-first helper
Omni Modal Question Answering
Audio Visual Reasoning
Multi Hop Reasoning
Tool Use
+2 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0
- Code license
- MIT
- License caution
- The Hugging Face card declares Apache-2.0 and the GitHub repository declares MIT. The benchmark curates media from FineVideo, LongVideoBench, LongVideo-Reason, COCO 2017, and other Hugging Face sources; verify component media rights and attribution requirements before redistribution or commercial use.
- Download notes
- The public, ungated release contains 360 English test tasks with audio, image, and video inputs across nine domains. The helper downloads official documentation and the lightweight test metadata JSON by default; the Hugging Face API reports about 9.9 GB of repository storage, so the full media snapshot requires OMNIGAIA_DOWNLOAD_HF=1. Qwen3.5-Omni reports OmniGAIA without a thinking prompt or answer-tag formatting in section 5.1.4, Table 7.
Safe-first helperscripts/download/omnigaia.sh
Audiovisual & cross-modal
OmniRetriever-Bench
OmniRetriever-Bench: 12-Direction Audio-Video-Text Retrieval Benchmark
Safe-first helper
Audio Text Retrieval
Audio Video Retrieval
Audio Video Text Retrieval
Cross Modal Retrieval
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_research_use
- Code license
- not_specified
- License caution
- The Hugging Face card labels the annotations Apache-2.0 but adds a biometric-identification, profiling, and surveillance prohibition, while the paper calls the release a custom research-use license. Treat the additional restriction as nonstandard and controlling pending a standalone license file. Underlying TikTok media remains owned by its uploaders and governed by platform terms. The evaluator repository has no detected license.
- Download notes
- The public, ungated release contains a 923,982-byte CSV with 3,782 held-out English-captioned audio-video-text triples and evaluates six single-modal plus six dual-modal retrieval directions. The helper downloads the official dataset card, CSV annotations, evaluator README, and paper page. Media is not redistributed; each row contains a TikTok source URL and clip interval, so availability depends on the source platform and users must obtain media themselves under applicable platform and uploader terms.
Safe-first helperscripts/download/omniretriever_bench.sh
Audiovisual & cross-modal
OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Manual or gated
Audio Visual Question Answering
Audio Visual Reasoning
Long Video Understanding
Temporal Reasoning
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- conflicting_CC-BY-NC-SA-4.0_and_CC-BY-NC-ND-4.0
- Code license
- not_specified
- License caution
- The official GitHub README says CC BY-NC-SA 4.0, while the Hugging Face card metadata declares CC BY-NC-ND 4.0 and the access form additionally requires non-commercial research use and no redistribution without permission. Apply the stricter gated terms pending clarification. The authors explicitly do not own the raw-video copyrights, so source-media rights remain separate; the GitHub repository has no detected license for its evaluation code.
- Download notes
- The gated Hugging Face release contains 628 English videos lasting from several seconds to 30 minutes and 1,000 manually verified question-answer pairs. Every question requires complementary audio and visual evidence and includes atomic step-by-step reasoning annotations; the paper reports 762 speech, 147 sound, and 91 music questions across 13 reasoning types. The helper downloads public repository documentation and API metadata by default. The Hugging Face API reports approximately 114 GB of repository storage, so the complete snapshot requires approval through the dataset questionnaire, authentication, and OMNIVIDEOBENCH_DOWNLOAD_HF=1. OmniReasoner evaluates OmniVideoBench in its main audio-visual benchmark table and reports duration- and audio-type-specific results.
Safe-first helperscripts/download/omnivideobench.sh
Speech recognition
Opencpop-test
Opencpop
Manual or gated
Singing Voice Transcription
Mandarin Singing Voice
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- License page is spelled /liscense/ on the official site.
Safe-first helperscripts/download/opencpop_test.sh
Music
OpenMIC-2018
OpenMIC-2018: An Open Dataset for Multiple Instrument Recognition
Safe-first helper
Musical Instrument Recognition
Music Auto Tagging
Multi Label Audio Classification
Music Information Retrieval
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo says Spotify AB releases the dataset under CC BY 4.0 and includes full license terms in the archive. The included metadata contains licenses for each audio recording, so check per-track metadata before redistribution or commercial use.
- Download notes
- The Zenodo archive is about 2.6 GiB and contains 10-second OGG clips, VGGish features, crowd-sourced labels, metadata with per-recording licenses, and train/test partitions. The helper saves the Zenodo record JSON and official README by default and requires OPENMIC_DOWNLOAD_ARCHIVE=1 before downloading the full archive.
Safe-first helperscripts/download/openmic_2018.sh
Speech generation
OpenSTBench
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
Safe-first helper
Speech To Text Translation Evaluation
Speech To Speech Translation Evaluation
Streaming Speech Translation Evaluation
Translation Quality Evaluation
+6 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other
- Code license
- MIT_with_CC-BY-SA-4.0_adapted_components
- License caution
- The paired-set card marks the dataset license as "other" and says users must comply with the original data and synthesis-component terms. LibriTTS is CC BY 4.0, but neither the repository's MIT license nor the paper's CC BY-NC-SA 4.0 license establishes standalone reuse terms for all translated metadata or Qwen3-TTS-generated reference audio. OpenSTBench's original code is MIT, while adapted SimulEval latency components are CC BY-SA 4.0. MSLT, RAVDESS, MCAE-SPPS, NonverbalTTS, and SynParaSpeech retain their own access and license terms.
- Download notes
- OpenSTBench provides a public evaluation package for offline and streaming speech-to-text and speech-to-speech translation. The paper's section 4.2 evaluates separate public source datasets for translation, speech quality, emotion, paralinguistics, temporal consistency, and latency, and constructs a 300-sample, 35-speaker LibriTTS-based paired set for speaker preservation. That paired set is public and ungated on Hugging Face and includes original and prompt LibriTTS audio, translated text, and Qwen3-TTS-synthesized reference speech. The helper downloads the paper, repository documentation, license notices, dataset card, and Hugging Face API metadata by default. Cloning the evaluation toolkit is opt-in, and downloading the approximately 511 MiB paired-set snapshot requires OPENSTBENCH_DOWNLOAD_PAIRED_SET=1. The helper does not fetch the other component datasets.
Safe-first helperscripts/download/openstbench.sh
Audiovisual & cross-modal
OV-MERD
OV-MERD: Open-Vocabulary Multimodal Emotion Recognition Dataset
Manual or gated
Open Vocabulary Multimodal Emotion Recognition
Audio Visual Emotion Understanding
Acoustic Emotion Cue Reasoning
Free Form Multi Label Emotion Prediction
+1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0_with_additional_gated_terms
- Code license
- Apache-2.0_with_noncommercial_notice
- License caution
- The paper and Hugging Face card identify OV-MERD as CC BY-NC 4.0. MER2025's gated terms further limit use to academic research and non-commercial purposes, prohibit distribution of the dataset or derivative annotation/label files to third parties, and prohibit modification without prior written consent. The OV-MER repository includes Apache-2.0 code terms but also calls the service a non-commercial research preview. Source clips derive from MER2023 movie and television media, so underlying media rights remain separate.
- Download notes
- OV-MERD extends a consented subset of MER2023 movie and television clips with human-checked acoustic and visual clues, merged multimodal descriptions, and open-vocabulary emotion labels. The paper reports 236 emotion categories, one to nine labels per sample (most have two to four), and clips that are mostly one to four seconds long. The current official repository points to the gated MER2025 release, which bundles OV-MERD label and description tables with audio, video, subtitles, and face features for the broader challenge corpus. The Hugging Face API reports approximately 442 GB of repository storage. The helper downloads only public official documentation, repository metadata, the paper landing page, and the Hugging Face API response; users must request access and download any benchmark files manually.
Safe-first helperscripts/download/ov_merd.sh
Speech recognition
Pansori-TEDxKR
Pansori-TEDxKR: Korean Speech Corpus Generated from Korean Language TEDx Talks
Safe-first helper
Korean Asr
Speech Transcription
Subtitle Aligned Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_specified
- License caution
- OpenSLR lists Creative Commons BY-NC-ND 4.0. The corpus is derived from TEDx talks, so downstream use should also respect TED/TEDx source terms and original media rights.
- Download notes
- OpenSLR SLR58 hosts about 3 hours of Korean TEDx speech from 41 speakers with subtitle-boundary segmentation and manual alignment checks. The helper downloads the OpenSLR page, about page, info, and checksum by default; the 174 MB corpus archive is opt-in.
Safe-first helperscripts/download/pansori_tedxkr.sh
Speaker, identity & emotion
ParaPairAudioBench
ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge
Safe-first helper
Paralinguistic Pairwise Judgment
Audio Language Model Judge Evaluation
Speaking Style Comparison
Speech Rate Comparison
+4 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_non_commercial_and_gated
- Code license
- not_specified
- License caution
- The benchmark repository has no license file or detected GitHub license. Its README lists SVC as non-commercial academic research, EARS and Expresso as CC BY-NC 4.0, and LibriTTS as CC BY 4.0. Treat the pair annotations and builder code as rights-unspecified, obtain SVC approval for the affected age/gender rows, and preserve each source corpus's terms.
- Download notes
- The official repository releases pairwise JSON annotations and swapped-order variants for speech rate, emphasis, and style, plus source-pair metadata and builders for age and gender. The paper reports 5,175 pairs across five criteria, including tie and same-/cross-transcript conditions. The helper downloads official documentation and repository metadata by default; cloning the approximately 6 MB repository is opt-in and does not fetch underlying audio. Age and part of gender require manually approved SVC access; the remaining source audio comes from EARS, Expresso, and LibriTTS under their own download terms.
Safe-first helperscripts/download/parapair_audio_bench.sh
Speaker, identity & emotion
PartialEdit
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
Safe-first helper
Partial Speech Deepfake Detection
Partial Speech Deepfake Localization
Neural Speech Editing Detection
Codec Artifact Analysis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_upstream_terms_and_partial_release
- Code license
- not_applicable
- License caution
- Zenodo declares CC BY 4.0 for the released record. The audio is derived from VCTK, whose official release is also CC BY 4.0, but users should preserve both provenances and review neural-editor output terms. The license does not make the withheld Audiobox-derived E3/E4 subsets public or grant rights to reconstruct them.
- Download notes
- The official Zenodo release contains the E1 (VoiceCraft), E1-Codec, E2 (SSR-Speech), and E2-Codec subsets derived from VCTK, plus the E1/E2 protocol CSV and modified-text metadata. The four audio archives total approximately 21.9 GB. The project and Zenodo description state that E3 (Audiobox-Speech) and E4 (Audiobox) cannot be released under Audiobox's license. The helper downloads official pages and Zenodo record metadata by default; the approximately 7.7 MB protocol/text metadata and large audio archives are separate opt-ins. SALMONN-2 section IV-E and Table VIII evaluate temporal spoof localization on PartialEdit using mean intersection over union.
Safe-first helperscripts/download/partialedit.sh
Speech recognition
PazaBench
PazaBench: A Benchmark for Automatic Speech Recognition on Low Resource Languages
Safe-first helper
Automatic Speech Recognition
Low Resource Asr
Multilingual Asr
Asr Efficiency
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The Hugging Face Space card declares MIT, although its visible tree has no standalone LICENSE file. That declaration does not relicense the 11 source dataset groups: their displayed terms range from CC0 and CC BY to CC BY-SA, CC BY-NC-SA, gated CC BY-NC, Apache-2.0, and mixed terms. Review each upstream dataset card and access agreement before use, redistribution, or commercial deployment.
- Download notes
- The public Microsoft Research Africa leaderboard currently reports WER, CER, and inverse real-time factor across 61 African languages and 53 ASR/language models. It evaluates 16 kHz mono speech from 11 named public or community dataset groups, including African Next Voices, ALFFA, FLEURS, Common Voice 23.0, WAXAL, Naija Voices, and TWB Voice. The public Space exposes the leaderboard, submission interface, dataset inventory, and implementation, but its repository does not contain a standalone unified audio snapshot, frozen item manifest, or result CSV; obtain evaluation audio from each named upstream provider under that provider's access terms. The helper saves official documentation, Space metadata, implementation metadata, and the dataset inventory only; it does not download source corpora or model weights.
Safe-first helperscripts/download/pazabench.sh
Speech generation
PodEval
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
Safe-first helper
Podcast Generation Evaluation
Long Form Audio Generation Evaluation
Dialogue Naturalness Evaluation
Speech Quality Evaluation
+3 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_for_real_pod_manifest_and_linked_audio
- Code license
- MIT
- License caution
- The repository's MIT license covers the software and accompanying documentation, but no separate data license is stated for the Real-Pod manifest. The linked podcast recordings remain hosted by third parties and retain creator, publisher, platform, voice, music, and other media rights. The maintainers direct users to follow legal and ethical rules and use the reference data for research and education; public links do not grant redistribution or commercial-use rights.
- Download notes
- The public framework evaluates podcast generation across text, speech, and audio using objective metrics, LLM-based judging, and structured listening tests. Its Real-Pod reference manifest contains 51 topics across 17 categories, with one publicly accessible Apple Podcasts episode link per topic. The repository does not redistribute those recordings. The helper downloads official documentation, the small JSON manifest, license, repository metadata, and arXiv metadata by default; cloning the approximately 12 MB MIT toolkit is opt-in and still does not download podcast audio.
Safe-first helperscripts/download/podeval.sh
Speech recognition
Primewords Chinese Corpus Set 1
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR lists Attribution-NonCommercial-NoDerivatives 4.0 International and describes the corpus as free for academic use. Re-check upstream terms before redistribution or commercial use.
- Download notes
- OpenSLR hosts a 9.0 GiB Mandarin speech/transcript archive recorded from 296 native Chinese speakers. The helper saves the OpenSLR page by default and makes the full archive an explicit opt-in.
Safe-first helperscripts/download/primewords_chinese.sh
Enhancement, separation & quality
PVQD
Perceptual Voice Qualities Database
Safe-first helper
Clinical Voice Quality Assessment
Pathological Voice Assessment
Perceptual Voice Rating
Cape V Prediction
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- Mendeley Data v4 and its DataCite DOI record declare CC BY 4.0. Preserve attribution and modification notices. The recordings contain human voices and clinical voice-quality information, and the release includes demographics; lawful and ethical handling of identifiable and health-related data remains necessary even though access is ungated.
- Download notes
- The public, ungated Mendeley Data v4 release contains 296 mono 44.1 kHz, 16-bit WAV recordings of sustained /a/ and /i/ vowels and six CAPE-V sentences, plus 13 XLSX files with demographics and experienced clinicians' CAPE-V and GRBAS ratings. The complete 310-file release is approximately 514.5 MiB. The helper downloads the official dataset page, DataCite DOI record, and live Mendeley file manifest by default. PVQD_DOWNLOAD_ANNOTATIONS=1 fetches the approximately 0.6 MiB PDF/XLSX documentation and labels; PVQD_DOWNLOAD_ALL=1 explicitly downloads the complete release including identifiable clinical voice recordings. The July 2026 voice-concept bottleneck paper uses an 80:20 speaker split and derives per-utterance segments with voice activity detection; those split and segment artifacts are not part of PVQD itself.
Safe-first helperscripts/download/pvqd.sh
Audiovisual & cross-modal
QIVD
Qualcomm Interactive Video Dataset: Can Vision-Language Models Answer Face to Face Questions in the Real-World?
Manual or gated
Situated Audio Visual Question Answering
Real Time Audio Visual Understanding
Spoken Query Understanding
When To Answer Prediction
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_publicly_specified_account_terms_apply
- Code license
- not_applicable
- License caution
- No public dataset license text was found on the official landing page or in the paper on 2026-07-21. The paper says crowd contributors signed consent permitting research and commercial use of their video and audio, but that consent statement is not itself a downstream dataset license. Review and retain any terms shown in Qualcomm's account/download flow before use or redistribution.
- Download notes
- The official Qualcomm release page describes 2,900 short English video files across 13 semantic categories. Each clip contains raw audio with a spoken question, an annotated transcription, a text answer, and a timestamp indicating when enough context is available to answer. The helper saves the public landing page, then prints the manual Qualcomm account/download path because the release flow is JavaScript-driven and may present account-specific terms. Qwen3.5-Omni calls it Qualcomm IVD and evaluates audio-query interaction in section 5.1.4, Table 7.
Safe-first helperscripts/download/qivd.sh
Audiovisual & cross-modal
RAVDESS
The Ryerson Audio-Visual Database of Emotional Speech and Song
Safe-first helper
Speech Emotion Recognition
Audio Visual Emotion Recognition
Acted Emotional Speech
Emotional Song Recognition
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-sa-4.0
- Code license
- not_applicable
- License caution
- Zenodo lists CC BY-NC-SA 4.0 for the dataset and says commercial licenses are available separately. The linked PLOS ONE paper itself is CC BY, but that article license is not the dataset license.
- Download notes
- The helper downloads the Zenodo metadata JSON by default. The audio-only speech and song ZIPs are about 215 MB and 198 MB respectively; video archives are much larger and are not downloaded by default.
Safe-first helperscripts/download/ravdess.sh
Speaker, identity & emotion
REAL-TSE Challenge
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Manual or gated
Target Speaker Extraction
Online Target Speaker Extraction
Offline Target Speaker Extraction
Real Conversational Speech Separation
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- access_restricted_terms_not_publicly_specified
- Code license
- MIT
- License caution
- The official evaluation repository is MIT, but that code license does not license the challenge audio. The public challenge page restricts DEV/EVAL use to validation or final evaluation, forbids training and fine-tuning, and says access was limited to registered teams; it does not state a standalone dataset license. DEV and EVAL-1 derive from AISHELL-4, AliMeeting, AMI, DiPCo, and CHiME-6, whose upstream terms also remain applicable.
- Download notes
- The challenge reports 6,991 Mandarin/English mixture-enrollment trials over 2,309 real conversational mixtures and 11.3 hours of mixture audio. DEV has 1,991 REAL-T-derived pairs; EVAL-1 and EVAL-2 contain 2,000 seen and 3,000 unseen pairs without public references. The public helper saves official pages and repository metadata only. Organizers distributed password-protected data by email exclusively to registered teams, registration closed on May 31, 2026, and no public dataset URL is currently provided.
Safe-first helperscripts/download/real_tse.sh
Audio understanding, generation & events
RealDESED
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Safe-first helper
Sound Event Detection
Domestic Sound Event Detection
Temporal Audio Event Localization
Multi Annotator Label Aggregation
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc0_or_cc-by_per_audio_file_and_cc-by-4.0_annotations
- Code license
- MIT
- License caution
- The Zenodo record is open and lists CC BY 4.0 at record level, but its description and official repository clarify that each audio recording and corresponding metadata row uses the per-file license recorded in metadata.csv, either CC0 or CC BY; remaining metadata and annotations are CC BY 4.0. Preserve creator attribution for CC BY recordings and consult metadata.csv before redistribution. The baseline repository is MIT.
- Download notes
- The public release contains 5,710 real-home recordings (37.85 hours) from 652 participants, with 64,430 temporal annotations across 15 domestic event classes. Multiple annotators label each recording, validation and test annotations receive additional review, and metadata covers recording devices, placement, environments, and scene descriptions. The helper saves official Zenodo/GitHub metadata and documentation by default; the train, validation, and test archives total approximately 8.74 GB and require explicit opt-in.
Safe-first helperscripts/download/realdesed.sh
Speaker, identity & emotion
RealMAN
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
Safe-first helper
Multichannel Speech Enhancement
Speech Source Localization
Dynamic Speaker Localization
Variable Array Generalization
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- unspecified
- License caution
- The official repository README declares the dataset CC BY 4.0, but the repository has no detected standalone license file and GitHub reports no repository license. Treat CC BY 4.0 as covering the dataset release, preserve attribution and notices, and verify terms separately for the baseline code and any downstream derived artifacts.
- Download notes
- The public, ungated release contains 83.7 hours of 32-channel speech recorded in 32 scenes and 144.5 hours of background noise recorded in 31 scenes, with direct-path speech, transcriptions, source locations, and scene and speaker metadata. The repository lists approximately 531.4 GB of training data, 27.5 GB of validation mixtures, 39.3 GB of test mixtures, 158 GB of raw validation/test recordings, and 129 MB of dataset information; the Hugging Face API currently reports about 812.0 GB of repository storage. The helper downloads only official documentation and API metadata by default, while the complete snapshot requires explicit opt-in. A July 2026 geometry-aware enhancement paper evaluates RealMAN after resampling to 8 kHz and fixed four-second test segments, and separately tests array generalization on CHiME-4.
Safe-first helperscripts/download/realman.sh
Speech understanding & dialogue
RealSI
RealSI: Open Benchmark for Simultaneous Interpretation in Real-world Scenarios
Safe-first helper
Simultaneous Speech To Text Translation
Simultaneous Speech To Speech Translation
Long Form Speech Translation
Streaming Latency Evaluation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- CC-BY-4.0_repository_license
- License caution
- The repository declares the dataset CC BY 4.0, but its README also says the authors do not own the copyright in the source videos and describes the release as annotations plus public video links for educational and informational use. The current repository additionally contains WAV derivatives. Treat CC BY 4.0 as covering author-created annotations, and independently review source-video copyright, platform terms, and the README disclaimer before using or redistributing audio.
- Download notes
- The official public repository contains timestamped Chinese-English and English-Chinese transcripts and translations for 20 natural, approximately 3-8 minute recordings across ten domains. The release totals 95 minutes 29 seconds and 778 utterance segments (431 En-to-Zh and 347 Zh-to-En), and its current tree also includes 20 WAV files. SimulS2ST-Omni section 4.1 and appendix D.3 reuse RealSI for sentence-level and long-form streaming S2TT/S2ST evaluation. The helper downloads official documentation, repository metadata, and all 20 lightweight JSON annotation files by default. The repository's WAV payload is about 351 MiB, so cloning the repository requires REALSI_CLONE_REPO=1.
Safe-first helperscripts/download/realsi.sh
Music
RUBATO
RUBATO: A Multi-Version Benchmark for Robust Music Transcription and Analysis
Safe-first helper
Automatic Music Transcription
Beat Tracking
Downbeat And Measure Tracking
Local Key Estimation
+2 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed_creative_commons
- Code license
- not_separately_specified
- License caution
- Zenodo labels the deposition CC BY 3.0, but metadata_versions.csv assigns per-recording terms that include CC0, CC BY, CC BY-SA, CC BY-ND, CC BY-NC, CC BY-NC-SA, CC BY-NC-ND, ambiguous "CC"/"CC0?", and EEF. Treat the per-recording field as controlling and review it before redistribution or commercial use; scripts inside the archive have no separate license statement.
- Download notes
- The open Zenodo v0.3 release contains 566 versions of 15 musical works (about 42.9 hours), including 22.05 kHz mono audio, aligned score MIDI/MuseScore/PDF/images, performance video, note/beat/measure/local-key/structure annotations, and audio-to-score warping paths. The helper downloads the 83 KB version metadata and Zenodo API record by default; the approximately 6.26 GB archive requires RUBATO_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/rubato.sh
Audio understanding, generation & events
RUL-MuchoMusic
RUL-MuchoMusic from RUListening / Are You Really Listening?
Safe-first helper
Music Question Answering
Perceptual Music Understanding
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- RUL repo/HF card: mit; upstream MuChoMusic dataset: CC BY-SA 4.0
- Code license
- MIT
- License caution
- RUL-MuchoMusic derives from MuChoMusic; check upstream audio/source terms too.
Safe-first helperscripts/download/rul_muchomusic.sh
Speech recognition
S-DiverSe
S-DiverSe: Spanish Diverse Speech
Safe-first helper
Automatic Speech Recognition
Pathological Speech Recognition
Spanish Speech Recognition
Asr Robustness
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository has no license file or detected GitHub license. The paper is CC BY 4.0 on arXiv, but that does not license the annotations, reconstruction code, or linked source recordings. The manifest includes health-condition metadata and potentially identifiable speech/transcripts; review consent, privacy, research ethics, source rights, and platform terms before use or redistribution.
- Download notes
- The public repository releases a TSV manifest for 444 manually transcribed Spanish segments totaling 3.2 hours from 22 speakers with amyotrophic lateral sclerosis, Parkinson's disease, or post-stroke effects. Metadata includes speaker ID, sex, condition, intelligibility, source URL, timestamp, and duration. Audio is not redistributed; the repository provides a yt-dlp/ffmpeg reconstruction script for public video sources, whose availability and platform terms can change. The paper evaluates corpus-level and condition-specific WER after lowercasing, punctuation removal, digit expansion, and retention of filled pauses. The helper downloads annotations, documentation, and reconstruction code only; cloning the repository is opt-in.
Safe-first helperscripts/download/s_diverse.sh
Speaker, identity & emotion
SALMon
SALMon: A Suite for Acoustic Language Model Evaluation
Safe-first helper
Acoustic Consistency Evaluation
Acoustic Semantic Alignment
Speaker Consistency
Speaker Gender Consistency
+5 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official repository and Hugging Face card license the SALMon dataset under CC BY-NC 4.0 because some source datasets are non-commercial. The release derives speech or acoustics from Expresso, VCTK, LJSpeech, FSD50K, and EchoThief and also includes Azure TTS output, so preserve component provenance and review upstream terms. The evaluation repository has no standalone license file and GitHub reports no detected code license.
- Download notes
- The public, ungated release contains 1,600 positive/negative audio pairs across eight configurations for acoustic consistency and acoustic-semantic alignment. It covers speaker identity, speaker gender, sentiment, background sound, and room impulse response, using a likelihood-ranking protocol for speech language models. The helper saves official documentation and repository metadata by default; set SALMON_DOWNLOAD_HF=1 for the approximately 562 MB Hugging Face snapshot. Google Drive provides the same benchmark as raw WAV files.
Safe-first helperscripts/download/salmon.sh
Representation & general suites
SEABAD
SEABAD: Southeast Asian Bird Activity Detection
Manual or gated
Bird Activity Detection
Binary Bird Presence Detection
Passive Acoustic Monitoring
Tropical Bioacoustic Detection
+1 more
Access pathZenodo
Upstream termsNot specified
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- MIT_claimed_in_readme_no_license_file
- License caution
- Zenodo's structured record declares CC BY 4.0 for the compilation, while the official repository says individual positive clips retain their Xeno-Canto licenses, including CC BY, CC BY-SA, CC BY-NC, and CC BY-NC-SA, and negative clips retain each source dataset's terms. Use the included provenance metadata and comply per recording; do not infer that the record-level license removes noncommercial, share-alike, or attribution requirements. The repository README calls the curation code MIT, but the repository currently has no LICENSE file and GitHub detects no license, so that code statement should be clarified before reuse.
- Download notes
- Zenodo v1.0.0 releases 50,000 balanced three-second, 16 kHz mono WAV clips: 25,000 bird-present clips spanning 1,677 Southeast Asian species and 25,000 bird-absent clips. Fixed stratified train, validation, and test splits contain 40,000, 5,000, and 5,000 clips. Positive recordings derive from Xeno-Canto; negatives derive from BirdVox-DCASE-20k, Freefield1010, Warblr, FSC-22, ESC-50, and DataSEC. The helper downloads official Zenodo, paper, and repository metadata by default. The single approximately 3.87 GiB mybad.zip archive requires explicit source-terms acknowledgment and an audio opt-in. DrongoNet section 7.1 evaluates the held-out 5,000-clip test split across five seeds and reports AUC, accuracy, recall, and F1.
Safe-first helperscripts/download/seabad.sh
Speech generation
Seed-TTS Eval
Seed-TTS objective zero-shot speech generation evaluation set
Safe-first helper
Zero Shot Text To Speech
Voice Cloning
Zero Shot Voice Conversion
Speech Intelligibility Evaluation
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The official repository has no LICENSE file and GitHub reports no detected license. The objective set selects 1,000 English samples from Common Voice and 2,000 Mandarin samples from DiDiSpeech-2, so verify both component-source terms before redistribution or commercial use. Public access does not imply an open license.
- Download notes
- The helper downloads the official README and lightweight evaluation scripts by default; cloning the evaluation repository is opt-in. The public objective EN/ZH test set is linked through Google Drive and must be downloaded manually. The Seed-TTS paper states that the 100-sample-per-language subjective set is not released because of copyright restrictions.
Safe-first helperscripts/download/seed_tts_eval.sh
Speech generation
SILMA Open-source Arabic TTS Benchmark
Open-source Arabic TTS Benchmark
Safe-first helper
Speech Synthesis
Text To Speech
Arabic Speech Synthesis
Dialectal Speech Synthesis
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_declared_at_space_level
- Code license
- Apache-2.0_declared_at_space_level
- License caution
- The Space metadata declares Apache-2.0, but the repository contains no separate license file and does not document the provenance or reuse terms of its Arabic prompts. Generated clips may also remain subject to the licenses and acceptable-use terms of the evaluated TTS models. Treat the Space-level declaration as insufficient to resolve all prompt, voice, and model-output rights before redistribution or commercial use.
- Download notes
- SILMA's public, ungated Hugging Face Space provides fixed prompts and generated model audio for direct listening comparisons in Modern Standard Arabic, Egyptian Arabic, and Saudi Arabic. The current release has 10 MSA prompts across four systems, five Egyptian prompts across five systems, and five Saudi prompts across three systems. SILMA says this release deliberately prioritizes auditory assessment because WER, CER, speaker similarity, and UTMOS do not fully capture Arabic speech nuances; it does not publish aggregate human ratings or a formal automatic scoring protocol. The helper downloads the official README, application source, three prompt CSVs, and Space API metadata by default. Cloning the approximately 29.6 MB Space, including generated evaluation audio, requires SILMA_ARABIC_TTS_CLONE_SPACE=1.
Safe-first helperscripts/download/silma_open_source_arabic_tts.sh
Enhancement, separation & quality
SingMOS-Pro
SingMOS-Pro: A Comprehensive Benchmark for Singing Quality Assessment
Safe-first helper
Singing Quality Assessment
Singing Mos Prediction
Singing Voice Synthesis Evaluation
Singing Voice Conversion Evaluation
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_upstream_source_terms
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY 4.0 and the official predictor repository is MIT. SingMOS-Pro includes ground truth and outputs from singing synthesis, conversion, resynthesis, and song-generation systems built from 12 source datasets. The paper and card do not provide a per-file license inventory, so retain source-corpus, performer, composition, model-output, and service terms rather than assuming the card clears all embedded audio rights.
- Download notes
- The public, ungated release contains 7,981 Chinese and Japanese singing clips totaling 11.15 hours, generated by 41 models across 12 source datasets. At least five experienced annotators rated every clip for overall MOS, and 4,155 clips additionally have lyrics and melody scores. The helper downloads official documentation, API metadata, split definitions, and system metadata by default. The approximately 11.6 MB sample/rating annotations require a separate opt-in, while the Hugging Face API reports approximately 2.83 GB of repository storage for the full audio snapshot.
Safe-first helperscripts/download/singmos_pro.sh
Enhancement, separation & quality
Slakh2100
Slakh2100: The Synthesized Lakh Dataset
Safe-first helper
Music Source Separation
Multi Instrument Automatic Transcription
Music Information Retrieval
Synthetic Multitrack Music
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- MIT
- License caution
- The official Slakh page states that Slakh2100 and Flakh2100 are licensed under Creative Commons Attribution 4.0 International. The slakh-utils repository is MIT. Slakh is synthesized from Lakh MIDI Dataset v0.1, so keep source MIDI attribution/provenance in downstream use.
- Download notes
- The helper downloads the official Slakh page and utility README/LICENSE by default. Zenodo currently hosts the full Slakh2100 record plus a tiny prototyping subset; the helper can save the Zenodo landing pages with SLAKH_CHECK_ZENODO=1, but archive file selection should be made from the live Zenodo records because the full corpus is large.
Safe-first helperscripts/download/slakh2100.sh
Speech recognition
SLUE
SLUE: Spoken Language Understanding Evaluation
Safe-first helper
Spoken Language Understanding
Automatic Speech Recognition
Named Entity Recognition
Named Entity Localization
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- SLUE-VoxPopuli is CC0; SLUE-VoxCeleb is CC BY 4.0; HF dataset metadata also advertises cc0-1.0 and cc-by-4.0 tags.
- Code license
- MIT
- License caution
- SLUE redistributes curated subsets of VoxPopuli and VoxCeleb plus task annotations. The VoxCeleb license notice says original and cropped video copyrights remain with the original owners.
- Download notes
- The helper downloads official toolkit docs and component license files by default. Hugging Face dataset snapshots are opt-in because they contain audio-derived benchmark data; use SLUE_DATASETS to choose slue, slue-phase-2, or both.
Safe-first helperscripts/download/slue.sh
Speech understanding & dialogue
SLURP
SLURP: A Spoken Language Understanding Resource Package
Safe-first helper
Spoken Language Understanding
Intent Classification
Slot Filling
Semantic Entity Labeling
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Textual annotations are CC BY 4.0; Zenodo-hosted audio is non-commercial (Zenodo license id other-nc; GitHub README states CC BY-NC 4.0).
- Code license
- not_specified
- License caution
- GitHub README separates textual data and audio licensing and says a less-strict audio license may be available by contacting the dataset owner. No standalone code license file was found in the repo via raw GitHub on 2026-07-09.
- Download notes
- The helper clones or updates the official annotation/code repository by default and downloads Zenodo LICENSE.txt. Audio archives are about 3.9 GiB real plus 2.8 GiB synthetic, so audio download is an explicit opt-in.
Safe-first helperscripts/download/slurp.sh
Speech recognition
SmartGlasses Challenge 2026
SLT 2026 SmartGlasses Challenge: Egocentric Speech Interaction on AI Glasses
Manual or gated
Time Stamped Speaker Attributed Asr
Meeting Transcription
Speaker Diarization
Spoken Language Understanding
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- access_restricted_terms_not_publicly_specified
- Code license
- not_specified
- License caution
- The challenge page does not publish a standalone dataset license and reserves organizer control over the participation terms. The public evaluation repository has no LICENSE file and GitHub reports no detected license, so its availability must not be treated as permission to redistribute the toolkit or corpus. Obtain permission from the organizers before reusing data or code beyond the challenge.
- Download notes
- The official challenge covers dyadic conversations and multi-party meetings recorded with a four-channel microphone array on smart glasses. Across train, development, and test, Track 1 reports 518 sessions and 44.95 hours, while Track 2 reports 196 sessions and 62.03 hours. Each track evaluates time-stamped speaker-attributed ASR with tcpWER and multiple-choice spoken-language understanding; public reference answers are limited to development data. The helper saves the official challenge page, public evaluation-toolkit documentation and metadata, and the July 2026 system paper. Corpus access required registration and agreement to challenge rules, download links were emailed to participating teams, registration closed in June 2026, and no current public corpus URL is provided.
Safe-first helperscripts/download/smartglasses_challenge_2026.sh
Enhancement, separation & quality
Song Describer Dataset
The Song Describer Dataset: A Corpus of Audio Captions for Music-and-Language Evaluation
Safe-first helper
Music Captioning
Text To Music Generation Evaluation
Music Text Retrieval
Music Codec Reconstruction
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-SA-4.0
- Code license
- MIT
- License caution
- Zenodo and the repository declare CC BY-SA 4.0 for the dataset and MIT for code. The audio originates from MTG-Jamendo and retains per-track Creative Commons licenses recorded in audio_licenses.txt; preserve attribution and apply each track's terms in addition to the dataset license.
- Download notes
- The public, ungated release contains 706 approximately two-minute MTG-Jamendo tracks with 1,106 crowdsourced English captions. Its human-validated evaluation subset contains 546 tracks and 746 captions. The helper downloads the official annotations, per-track audio-license list, metadata, dataset documentation, and Zenodo record by default; the approximately 3.09 GiB audio archive requires explicit opt-in. Qwen-Music section 4.2.2 evaluates codec reconstruction on all 546 validated tracks. Qwen-Audio-VAE sections 4.1-4.2 also use the dataset for music reconstruction evaluation.
Safe-first helperscripts/download/song_describer.sh
Enhancement, separation & quality
SongEval
SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
Safe-first helper
Song Aesthetics Assessment
Music Quality Prediction
Full Song Generation Evaluation
Human Preference Modeling
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC-SA 4.0 and the GitHub toolkit includes Apache-2.0. The paper says the audio includes outputs from five open and commercial song generators plus real and deliberately poor examples; generated-output, service, and any underlying music rights may still apply, and the release does not provide per-item provenance in metadata.jsonl. Review those rights before redistributing audio or relying on the card license alone.
- Download notes
- The public, ungated release contains 2,399 complete English and Chinese songs (about 140 hours) spanning nine mainstream genres. Sixteen musically trained annotators rated coherence, memorability, vocal breathing and phrasing naturalness, structural clarity, and overall musicality on five-point scales. The helper downloads official cards, API metadata, the approximately 1.27 MB rating JSONL, and toolkit documentation by default. The Hugging Face API reports approximately 16.1 GB of repository storage, so fetching all MP3 files requires explicit opt-in.
Safe-first helperscripts/download/songeval.sh
Audio understanding, generation & events
SongFormBench
Safe-first helper
Music Structure Analysis
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- cc-by-4.0
- License caution
- Dataset card lists CC BY 4.0. Audio reconstruction notes reference HarmonixSet and BigVGAN resources.
Safe-first helperscripts/download/songformbench.sh
Audiovisual & cross-modal
Sonic Seasoning
Sonic Seasoning: a Multi-Source Perceptual Dataset of Taste-Evoking Sounds
Safe-first helper
Taste From Audio Regression
Taste Conditioned Music Retrieval
Music Representation Evaluation
Crossmodal Audio Perception
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- Apache-2.0
- License caution
- The Hugging Face card declares CC BY-NC 4.0 for the compilation, ratings, splits, and redistributed clips, and the repository code is Apache-2.0. Audio provenance varies: the release combines pre-existing music, MusicGen outputs, and stimuli from 13 prior studies. The card limits redistributed clips to non-commercial research and instructs users to cite the originating studies; review underlying music, publication-stimulus, performer, and generated-output rights in addition to the compilation license.
- Download notes
- The public, ungated Hugging Face release contains 377 uniformly encoded WAV clips with normalized sweet, bitter, salty, sour, and spicy ratings; some subsets also provide temperature and emotion annotations. Its fixed split column contains 269 train, 68 validation, and 40 test items. The paper evaluates ten frozen audio encoders with a shared multi-task regression protocol and uses a 309-item pool for taste-conditioned retrieval. The helper downloads the official card, repository docs, API metadata, and approximately 34 KB ratings/path Parquet file by default. The Hugging Face API reports approximately 797 MB of repository storage, so the audio snapshot and the approximately 642 KB code repository are separate opt-ins. The repository README still calls the training dataset private, which conflicts with the current public, ungated Hugging Face release; the helper follows the live dataset state.
Safe-first helperscripts/download/sonic_seasoning.sh
Audio understanding, generation & events
SONYC-UST-V2
SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context
Safe-first helper
Audio Tagging
Urban Sound Tagging
Multilabel Sound Classification
Spatiotemporal Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Zenodo v2.3 lists CC BY 4.0 and the record README says the SONYC-UST dataset is offered under Creative Commons Attribution 4.0 International. Challenge rules also restrict private external data for reproducible task submissions.
- Download notes
- The helper downloads Zenodo record JSON plus README, annotations, taxonomy, and unpack script by default. Audio is split across 19 archive shards totaling about 12.8 GiB, so audio download is an explicit opt-in.
Safe-first helperscripts/download/sonyc_ust_v2.sh
Audio understanding, generation & events
Soroll-IA
Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring
Safe-first helper
Multi Label Audio Tagging
Weakly Supervised Sound Event Classification
Industrial Sound Monitoring
Real World Environmental Audio Classification
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- Kaggle and the official benchmark README declare CC BY-NC 4.0, which prohibits commercial use and requires attribution. The benchmark repository has no LICENSE file or GitHub-detected license, so its code terms are unspecified.
- Download notes
- The public Kaggle release contains 7,396 FLAC clips (approximately 22 hours and 2.17 GB) recorded by two fixed sensing nodes in the Port of Valencia, with 26 weakly labeled industrial sound classes. It provides two ground-truth variants: a permissive non-cross-validated annotation set and a conservative set requiring agreement from at least two-thirds of annotators, plus five-fold assignments. The helper downloads official metadata, paper, and benchmark documentation by default; the audio and annotations require explicit opt-in and an authenticated Kaggle CLI.
Safe-first helperscripts/download/soroll_ia.sh
Audio understanding, generation & events
Spatial LibriSpeech
Spatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning
Safe-first helper
Sound Source Localization
Source Distance Estimation
Room Acoustics Estimation
Direct To Reverberant Ratio Estimation
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- Apple's copyrights in the dataset are CC BY 4.0. Its license preserves upstream terms for LibriSpeech, Microsoft DNS noise, AudioSet, and CC0 Freesound material and says Apple makes no representations about those upstream rights; review component provenance before redistribution.
- Download notes
- The public Apple release provides more than 650 hours of 16 kHz first-order ambisonic speech and optional distractor noise, synthesized from LibriSpeech across more than 200,000 acoustic conditions and 8,000 synthetic rooms. Labels cover 3D source position and speaking direction, room geometry, C50, DRR, EDT, T20, and T30. The helper downloads official documentation by default; the approximately 365 MiB metadata Parquet file and individual FLAC samples are explicit opt-ins. The README still describes raw 19-channel audio as forthcoming and directs users to contact Apple rather than exposing a public download.
Safe-first helperscripts/download/spatial_librispeech.sh
Speech understanding & dialogue
SPEARBench
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Safe-first helper
Streaming Speech To Speech Evaluation
Conversational Naturalness Evaluation
Response Latency Evaluation
Interruption Evaluation
+4 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_derived_from_seamless_interaction
- Code license
- MIT
- License caution
- The GitHub repository's MIT license covers the website and helper code. Neither the paper nor project page states a separate license for the downloadable benchmark audio, which is extracted from Seamless Interaction; public access does not establish redistribution or commercial-use rights, so verify the source corpus and package terms before reuse.
- Download notes
- The public project page links a SharePoint package containing 5,419 selected English question-answer dialogues from the Seamless Interaction development and test sets (37.33 hours including contexts, questions, and human answers). The helper downloads only official documentation, leaderboard metadata, an example submission CSV, and inference instructions; obtain the audio package manually from the project page and keep its directory structure intact.
Safe-first helperscripts/download/spearbench.sh
Speech understanding & dialogue
Speech Commands
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Safe-first helper
Keyword Spotting
Limited Vocabulary Speech Recognition
Audio Classification
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Google Research blog and Hugging Face dataset card list Creative Commons BY 4.0; HF card also asks users not to try to identify speakers.
- Download notes
- TensorFlow Datasets reports v0.02 as about 2.37 GiB download / 8.17 GiB extracted; avoid accidental full downloads in automated checks.
Safe-first helperscripts/download/speech_commands.sh
Speech generation
SpeechEditBench
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
Safe-first helper
Instruction Guided Speech Editing
Content Editing
Speaker Editing
Emotion Editing
+5 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0_with_upstream_terms
- Code license
- Apache-2.0
- License caution
- The repository and Hugging Face card release contributor-authored code, documentation, metadata, and benchmark assets under Apache-2.0, while requiring compliance with source-corpus terms. The paper's appendix identifies mixed upstream conditions: CC BY sources, Apache-2.0 sources, CC BY-NC and CC BY-NC-SA sources, StoryTTS research-only restrictions, MagicData-RAMC custom terms, and the IEMOCAP access agreement. Apply those component restrictions to affected rows and audio rather than treating the aggregate label as overriding them.
- Download notes
- The public, ungated v1.1 release contains 4,700 English and Chinese source-instruction pairs and 5,400 audio files across seven atomic editing tasks and a compositional split. Evaluation separately measures target success, lexical-content preservation, and joint success. The helper downloads official documentation, release metadata, and the eight sample JSONL files by default. The Hugging Face API reports approximately 3.75 GB of repository storage, so audio requires SPEECH_EDIT_BENCH_DOWNLOAD_HF=1; cloning the roughly 3.2 MB evaluation repository is a separate opt-in.
Safe-first helperscripts/download/speech_edit_bench.sh
Speech understanding & dialogue
SpeechEQ
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models
Safe-first helper
Spoken Dialogue Emotional Intelligence
Paralinguistic Reasoning
Multi Turn Speech Understanding
Acoustic Multiple Choice
+1 more
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_specified
- License caution
- The Hugging Face card has no license field, and neither the dataset repository nor the official code repository exposes a license file. The paper itself is CC BY-NC-SA 4.0, but that publication license must not be assumed to license the released benchmark audio, annotations, or code. Obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated English release contains 2,265 six-turn dialogues totaling 42 hours 23 minutes across 15 EQ-i 2.0 subscales. Evaluation selects between high- and low-EQ acoustic renditions of identical text at turns four and six, testing pitch, energy, rate, pauses, and sustained conversational context. The helper downloads official documentation and repository metadata by default. The five Parquet shards contain embedded audio and total approximately 2.45 GB, so the full Hugging Face snapshot requires SPEECHEQ_DOWNLOAD_HF=1.
Safe-first helperscripts/download/speecheq.sh
Speech understanding & dialogue
SpeechRole
SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
Safe-first helper
Speech Role Playing
Speech Dialogue Evaluation
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- not_specified
- License caution
- HF cards list MIT; GitHub repo has no separate detected license.
Safe-first helperscripts/download/speechrole.sh
Audiovisual & cross-modal
SpEmoC
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
Manual or gated
Speech Emotion Recognition
Audio Visual Emotion Recognition
Multimodal Emotion Recognition
Cross Dataset Emotion Generalization
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_non_commercial_academic_eula
- Code license
- not_specified
- License caution
- The official agreement limits use to academic research, education, scientific publication, and other non-commercial research; prohibits redistribution and sharing download links; and leaves copyright in the source movie and television clips with their respective owners. The public benchmark repository has no LICENSE file or detected GitHub license, so code and public split/metadata terms remain unspecified.
- Download notes
- The benchmark contains 30,000 refined clips curated from 306,544 raw speaking segments across 3,100 English-language movies and television series, with aligned audio, visual, and text modalities and a near-balanced seven-emotion label distribution. The official project says the full dataset is available only after a requestor and faculty advisor or principal investigator sign the access agreement and submit it from an institutional email address. The helper saves public project, paper, repository, and agreement metadata, then prints the manual application steps; it never downloads restricted media.
Safe-first helperscripts/download/spemoc.sh
Speech recognition
SPGISpeech
SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition
Manual or gated
Asr
Fully Formatted Transcription
Financial Speech Recognition
Access pathHugging Face
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- gated_academic_research_internal_use
- Code license
- not_specified
- License caution
- HF terms say the content is for academic research purposes and internal use only, prohibit redistribution without prior written consent, and include additional restrictions on creating competing databases/products and identifying individuals.
- Download notes
- Hugging Face access requires logging in and accepting Kensho terms. The HF card lists split sizes from 11 GiB for dev/test to 530 GiB for the L training subset; the helper refuses to download until SPGISPEECH_ACK_TERMS=1 is set.
Safe-first helperscripts/download/spgispeech.sh
Enhancement, separation & quality
SpInt
SpInt: A Spanish Speech Intelligibility Dataset
Safe-first helper
Speech Intelligibility Assessment
Objective Intelligibility Metric Evaluation
Speech Enhancement Evaluation
Spanish Speech
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_for_released_SpInt_artifacts
- Code license
- Not specified in the source record.
- License caution
- Zenodo declares CC BY 4.0 for the released SpInt package. The record explicitly withholds the original clean Spanish Matrix Test recordings to respect their licensing conditions, so the release is not a standalone audio corpus and CC BY 4.0 must not be extended to those absent recordings.
- Download notes
- The public Zenodo v1.0 release provides behavioral intelligibility labels for 5,148 processed Spanish utterances, per-stimulus and listener-response metadata, complex speech-enhancement masks, noise signals, and a reconstruction script. The clean Spanish Matrix Test recordings are deliberately excluded because of their separate license; users must obtain that corpus independently to reconstruct the stimuli. The helper downloads the official record, README, reconstruction script, and approximately 2.7 MB JSON metadata by default. The approximately 807 MiB noise and 3.08 GiB mask archives require explicit opt-in.
Safe-first helperscripts/download/spint.sh
Speaker, identity & emotion
SpoofCeleb
SpoofCeleb: Speech Deepfake Detection and SASV in the Wild
Manual or gated
Speech Deepfake Detection
Synthetic Speech Detection
Spoofing Robust Speaker Verification
In The Wild Anti Spoofing
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_source_media_rights
- Code license
- not_applicable
- License caution
- The official project and Hugging Face tag state CC BY 4.0, while the project explicitly says copyright in the human speech files remains with the original video owners. Access is granted only after a request and agreement to Hugging Face terms. Treat the Creative Commons label as covering the released compilation and author contributions, not as clearance of every underlying video, voice, likeness, privacy, or generated-speech right.
- Download notes
- The gated author-owned Hugging Face release contains more than 2.5 million bona fide and synthetic utterances from 1,251 VoxCeleb1 speakers. Its 23 TTS attacks and speaker-disjoint train, validation, and evaluation partitions support both speech deepfake detection and spoofing-robust speaker verification protocols. The helper downloads the official project page, paper pages, and Hugging Face API metadata by default. The API reports approximately 268.3 GB of repository storage, so the snapshot requires author approval, local Hugging Face authentication, explicit acceptance of the terms, and both SPOOFCELEB_ACK_TERMS=1 and SPOOFCELEB_DOWNLOAD_HF=1. Section 4.1 of arXiv:2607.21127 evaluates balanced TTS attacks from all SpoofCeleb splits but does not publish its clipped row selection.
Safe-first helperscripts/download/spoofceleb.sh
Audio understanding, generation & events
SpurAudio
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification
Safe-first helper
Few Shot Audio Classification
Environmental Sound Classification
Shortcut Learning Evaluation
Background Shift Robustness
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0_with_mixed_upstream_terms
- Code license
- MIT
- License caution
- The Hugging Face card declares CC BY 4.0 for the released mixtures, but the benchmark derives audio from five upstream datasets with separate terms. In particular, ESC-50 and UrbanSound8K include non-commercial restrictions; review all component licenses before redistribution or commercial use. The evaluation repository is MIT.
- Download notes
- The public, ungated Hugging Face release contains 16,381 WAV files in train, validation, and test splits and reports approximately 7.69 GB of repository storage. It mixes foreground events from ESC-50, UrbanSound8K, VocalSound, WILD DESED, and USM with unrelated background textures to measure IID-versus-OOD shortcut reliance in 1-shot and 5-shot classification. The helper downloads official documentation and repository metadata by default; the audio snapshot is opt-in.
Safe-first helperscripts/download/spuraudio.sh
Speech recognition
ST-CMDS
ST-CMDS-20170001_1: Free ST Chinese Mandarin Corpus
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR SLR38 lists Creative Commons BY-NC-ND 4.0 and asks users to cite the data as "ST-CMDS-20170001_1, Free ST Chinese Mandarin Corpus." Re-check upstream terms before redistribution or commercial use.
- Download notes
- OpenSLR hosts an 8.2 GiB archive with cellphone-recorded Mandarin speech, transcriptions, and metadata from 855 speakers and 102,600 utterances. The helper saves the OpenSLR page by default and requires ST_CMDS_DOWNLOAD_ARCHIVE=1 for the large archive.
Safe-first helperscripts/download/st_cmds.sh
Audio understanding, generation & events
STAR-Bench
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
Safe-first helper
Foundational Acoustic Perception
Pitch Perception
Loudness Perception
Duration Perception
+6 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- conflicting_MIT_and_Apache-2.0_signals
- License caution
- The Hugging Face card declares the dataset CC BY-NC 4.0 and labels use as research-only. Preserve the terms and provenance of Clotho, FSD50K, STARSS23, and internet-sourced audio. The repository's LICENSE file is MIT, but its README badge says Apache-2.0 and describes both data and code as research-only; clarify the intended software terms before redistribution or commercial use.
- Download notes
- The public, ungated v1.0 release contains 2,353 English multiple-choice questions: 951 for foundational perception, 900 for temporal reasoning, and 502 for spatial reasoning. Its current metadata revises the v0.5 questions reported in the paper. Foundational audio is synthesized; temporal tasks draw on Clotho and FSD50K, while spatial tasks use STARSS23 and in-the-wild audio. The helper downloads official documentation, repository metadata, and the approximately 2 MB of JSON question metadata by default. The 2.74 GB audio archive and evaluation repository are separate opt-ins.
Safe-first helperscripts/download/star_bench.sh
Audiovisual & cross-modal
StoryAD-QA
StoryAD-QA: Narrative-Comprehension Evaluation for Long-Form Audio Description
Safe-first helper
Long Form Audio Description Evaluation
Narrative Comprehension
Multiple Choice Question Answering
Context Conditioned Reasoning
+1 more
Access pathOfficial / other
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- not_specified_pending_finalization
- License caution
- The repository's LICENSE is a placeholder that says release terms will be finalized and merely recommends MIT or Apache-2.0 for code and CC BY-NC 4.0 or another author-approved license for annotations. Those recommendations are not grants. The release contains no movie video, audio, frames, subtitles, or scripts; clip identifiers derive from ten MAD-Eval movies in LSMDC, and users must obtain lawful access to the underlying copyrighted media separately.
- Download notes
- The official ECCV 2026 repository releases 2,572 manually verified, five-option question-answer pairs across two tracks: 1,609 segment-only questions over 30-, 60-, 120-, and 240-second windows, and 963 context-conditioned questions using 30, 60, or 90 seconds of preceding context plus a 30-second target clip. The repository includes full annotation CSVs, question-only files, answer keys with rationales, prompts, and a local accuracy scorer. The paper reports 2,574 retained questions, while the repository README says its public release contains 2,572 after validation and cleanup; this entry uses the released-file total. The helper saves official documentation, license notice, summary, evaluator, and repository metadata by default; cloning the approximately 8 MB repository is opt-in.
Safe-first helperscripts/download/storyad_qa.sh
Speech recognition
SUPERB
SUPERB: Speech Processing Universal PERformance Benchmark
Safe-first helper
Speech Representation Evaluation
Phoneme Recognition
Automatic Speech Recognition
Keyword Spotting
+11 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mixed
- Code license
- Apache-2.0
- License caution
- S3PRL is mostly Apache-2.0, with repository notes indicating some Facebook-authored files are CC BY-NC. SUPERB tasks use multiple external datasets, so each component corpus must be downloaded and licensed through its own official source.
- Download notes
- SUPERB is a benchmark suite over multiple upstream corpora. The helper downloads official documentation/license files by default and only clones the S3PRL toolkit with SUPERB_CLONE_TOOLKIT=1; underlying corpora such as LibriSpeech, Speech Commands, VoxCeleb, and IEMOCAP keep their own access paths and licenses.
Safe-first helperscripts/download/superb.sh
Representation & general suites
Surge Pitch Dataset
Pitch Audio Dataset (Surge synthesizer)
Safe-first helper
Musical Pitch Classification
Musical Pitch Ranking
Synthesizer Preset Classification
Audio Representation Evaluation
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_applicable
- License caution
- Zenodo declares CC BY 4.0 for the released dataset. The paper's publication license and the Surge synthesizer and preset licenses are separate; retain the dataset citation and review synthesizer/preset terms if regenerating or redistributing modified renders.
- Download notes
- The public, ungated Zenodo release contains 3.4 hours of four-second sounds generated from 2,084 human-authored Surge presets. Each preset is rendered at MIDI pitches 21 through 108 with velocity 64, a three-second note-on duration, and RMS normalization. The helper saves official record metadata and the paper by default; the approximately 7.58 GB tar archive requires explicit opt-in. NABEATs section 4.1 uses this release for downstream pitch classification under clean and constructed noisy conditions.
Safe-first helperscripts/download/surge_pitch.sh
Speech recognition
Switchboard
Switchboard-1 Release 2 conversational telephone speech corpus
Manual or gated
Automatic Speech Recognition
Conversational Speech Recognition
Telephone Speech Recognition
Speaker Recognition
Access pathLDC / licensed
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_ldc_license
- Code license
- not_applicable
- License caution
- LDC catalog pages list membership/licensing terms and web-download access. Re-check the current LDC agreement before use or redistribution; do not treat benchmark recipes or transcripts as granting rights to the underlying audio.
- Download notes
- Switchboard-1 Release 2 and the 2000 HUB5 English Evaluation Speech set are distributed by LDC after login/licensing. The helper only prints official access steps because the corpus and standard evaluation audio are not publicly script-downloadable.
Safe-first helperscripts/download/switchboard.sh
Audiovisual & cross-modal
SyncBench
SyncBench: Causal-Semantic Audio-Visual Synchronization Evaluation for Generative Models
Safe-first helper
Audio Visual Synchronization Evaluation
Audio Visual Generation Evaluation
Video To Audio Evaluation
Causal Semantic Alignment
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- MIT
- License caution
- The official code repository is MIT, but the Hugging Face dataset has no card, license tag, or license file. The paper is CC BY 4.0, which does not license the released generated videos or prompts. Model-output provider terms and any rights in prompt/source content may also apply; obtain clarification before redistribution or commercial use.
- Download notes
- The public, ungated Hugging Face repository contains 1,185 generated MP4s across six model directories plus two small evaluator-score JSON files, and reports approximately 12.9 GB of repository storage. The paper's section 4.4 defines SyncBench as 185 curated prompts across five audio-visual domains, while the current release has up to 200 numbered clips per model and does not include a dataset card or prompt manifest. The helper downloads official documentation, repository metadata, Hugging Face metadata, and the two lightweight score files by default; all videos require SYNCBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/syncbench.sh
Speaker, identity & emotion
SynSFX
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
Manual or gated
Non Speech Audio Deepfake Detection
Synthetic Sound Effect Detection
Unseen Generator Robustness
Cross Domain Audio Forensics
+1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- academic_research_only
- Code license
- not_released
- License caution
- The official release page labels SynSFX "Academic research only" but does not publish a full standalone dataset license or redistribution terms. The authentic partition incorporates AudioCaps, Clotho, ESC-50, TACoS, and WavCaps material, so their source-media and per-clip terms remain applicable. The arXiv paper uses arXiv's perpetual non-exclusive publication license, which does not license the dataset. No official evaluation-code release was linked when checked.
- Download notes
- The official release page provides a direct private-storage download route for the academic-research-only corpus. The paper reports 43,374 clips totaling 178 hours: 16,922 authentic clips from AudioCaps, Clotho, ESC-50, TACoS, and WavCaps, plus 26,452 clips synthesized by seven text-to-audio systems. It also defines a 1,890-prompt controlled subset shared across all seven generators and train, validation, in-domain test, and unseen-generator test protocols. The official page rounds the duration to approximately 180 hours and inconsistently lists 26,460 synthetic clips, despite retaining the 43,374 total; this entry uses the internally consistent paper counts. The helper saves the official page and paper by default. Because the uncompressed-WAV archive is large and the publisher states research-only access, downloading it requires both SYNSFX_ACK_RESEARCH_ONLY=1 and SYNSFX_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/synsfx.sh
Speech recognition
Tadabur
Tadabur: A Large-Scale Quran Audio Dataset
Safe-first helper
Quranic Speech Recognition
Arabic Speech Recognition
Reciter Identification
Word Level Alignment
+2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0_with_source_and_cultural_caveats
- Code license
- not_specified
- License caution
- The Hugging Face card declares CC BY-NC 4.0, limits use to research and education, and adds ethical and cultural expectations for respectful Qur'anic use. Audio was collected from public Qur'anic repositories and archives, but the release does not provide per-recording source-license provenance; confirm source rights before redistribution. The linked repository has no standalone LICENSE file, so no code license is claimed.
- Download notes
- The public, ungated Hugging Face release contains more than 365,000 verse-level Arabic recitation examples totaling over 1,400 hours from more than 600 reciters, with simple and Uthmani text plus automatically derived word timestamps. It exposes one training split rather than a fixed held-out evaluation split. The helper downloads official documentation and repository metadata by default; the Hugging Face API reports approximately 1.94 TB of repository storage, so the audio-bearing snapshot requires TADABUR_DOWNLOAD_HF=1.
Safe-first helperscripts/download/tadabur.sh
Audio understanding, generation & events
TAU Spatial Sound Events 2019
TAU Spatial Sound Events 2019: Ambisonic and Microphone Array
Safe-first helper
Sound Event Localization And Detection
Sound Source Localization
Sound Event Detection
Spatial Audio Understanding
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_TAU_noncommercial
- Code license
- custom_TAU_noncommercial
- License caution
- Both Zenodo records use a custom Tampere University license permitting experimental non-commercial use with attribution and prohibiting commercial use. The source events derive from the DCASE 2016 Task 2 isolated-event dataset, so preserve that provenance. The baseline repository applies closely matching custom experimental/non-commercial terms to its code.
- Download notes
- The public DCASE 2019 Task 3 release has 400 development and 100 evaluation scenes of one minute each at 48 kHz, in matching 4-channel first-order Ambisonic and tetrahedral-microphone formats. Scenes use stationary sources from 11 classes, real impulse responses measured at 504 azimuth-elevation-distance combinations across five indoor locations, natural ambient noise, and zero or up to two overlapping events. Version 2 includes temporal and azimuth/elevation labels for both development and evaluation audio. The helper downloads official pages, papers, record metadata, READMEs, and license files by default; the approximately 10.1 GB audio release remains on Zenodo, while the roughly 490 KB label archives are an explicit opt-in.
Safe-first helperscripts/download/tau_spatial_sound_events_2019.sh
Audio understanding, generation & events
TAU Urban Acoustic Scenes 2019
TAU Urban Acoustic Scenes 2019 Development dataset
Safe-first helper
Acoustic Scene Classification
Multi Device Audio Classification
Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_non_commercial
- Code license
- not_applicable
- License caution
- Zenodo record 2589280 lists license id "other-nc" without a more specific SPDX-style license. DCASE challenge and dataset terms should be checked before redistribution or commercial use.
- Download notes
- The helper downloads the Zenodo record JSON plus small doc/meta ZIPs by default. The 40-hour audio release is split across 21 ZIP files of roughly 1.3-1.8 GiB each, so audio download is an explicit opt-in.
Safe-first helperscripts/download/tau_asc_2019.sh
Audio understanding, generation & events
TAU Urban Acoustic Scenes 2020 Mobile
TAU Urban Acoustic Scenes 2020 Mobile: Development and Evaluation datasets
Safe-first helper
Acoustic Scene Classification
Device Robust Audio Classification
Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_non_commercial
- Code license
- not_applicable
- License caution
- Both Zenodo development and evaluation records list "Other (Non-Commercial)" without a more specific SPDX-style license. DCASE challenge terms should be checked before redistribution or commercial use.
- Download notes
- The helper downloads Zenodo record JSON plus small doc/meta ZIPs by default. Development audio is split across 16 ZIP files totaling about 27.4 GiB; evaluation audio is split across 8 ZIP files totaling about 13.1 GiB, so both are explicit opt-ins.
Safe-first helperscripts/download/tau_asc_2020_mobile.sh
Audio understanding, generation & events
TAU Urban Acoustic Scenes 2022 Mobile
TAU Urban Acoustic Scenes 2022 Mobile Development and 2025 Evaluation datasets
Safe-first helper
Acoustic Scene Classification
Device Robust Audio Classification
Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- other_non_commercial
- Code license
- not_applicable
- License caution
- Both official Zenodo records list "Other (Non-Commercial)" without a more specific SPDX-style license. Review the record and DCASE task terms before redistribution or commercial use.
- Download notes
- The helper downloads both Zenodo record JSON files plus the small doc/meta ZIPs by default. The 64-hour development release contains about 25.6 GiB across 16 audio ZIPs; the DCASE 2025 evaluation release contains about 19.2 GiB across 12 audio ZIPs, so each audio collection is an explicit opt-in. DCASE 2025 Task 1 reuses a restricted subset of the 2022 development data with a new split and adds the 2025 evaluation release.
Safe-first helperscripts/download/tau_asc_2022_mobile.sh
Audio understanding, generation & events
TAU-NIGENS Spatial Sound Events 2020
Safe-first helper
Sound Event Localization And Detection
Sound Source Localization
Acoustic Source Tracking
Sound Event Detection
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- custom_TAU_License
- License caution
- Zenodo lists the dataset as CC BY-NC 4.0. Source events come from the separately released NIGENS database, so preserve its provenance and review its active terms. The baseline README places most code under a custom TAU License and only its metrics folder under MIT.
- Download notes
- The public v1.2 release is the complete DCASE 2020 Task 3 development and evaluation corpus: 600 development and 200 evaluation sound scenes of one minute each, in 4-channel FOA and tetrahedral-microphone formats at 24 kHz. It uses real room impulse responses from 15 enclosures, static and moving sources from 14 classes, up to two overlapping events, and direction-of-arrival trajectories plus onset/offset labels. The helper downloads official pages, paper, record metadata, and the 17 KB dataset README by default; the approximately 14.0 GB audio archives remain on Zenodo, while the roughly 1.7 MB label archives are an explicit opt-in.
Safe-first helperscripts/download/tau_nigens_sse_2020.sh
Audio understanding, generation & events
TAU-NIGENS Spatial Sound Events 2021
Safe-first helper
Sound Event Localization And Detection
Sound Source Localization
Acoustic Source Tracking
Sound Event Detection
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-4.0
- Code license
- custom_TAU_License
- License caution
- Zenodo lists the dataset as CC BY-NC 4.0. Source events come from the separately released NIGENS database, so preserve its provenance and review its active terms. The baseline license allows experimental non-commercial use and prohibits commercial use.
- Download notes
- The public v1.1.0 release is the complete DCASE 2021 Task 3 development and evaluation corpus: 600 development and 200 evaluation sound scenes of one minute each, in 4-channel FOA and tetrahedral-microphone formats at 24 kHz. It adds directional non-target interferers, permits overlapping instances of the same target class, and supplies direction-of-arrival trajectories plus onset/offset labels for the development set. The helper downloads official pages, paper, record metadata, and the 23 KB dataset README by default; the approximately 14.2 GiB audio archives remain on Zenodo, while the roughly 1.8 MiB development labels are an explicit opt-in. The 200-file evaluation set intentionally has no public labels.
Safe-first helperscripts/download/tau_nigens_sse_2021.sh
Speech recognition
TEDx Spanish Corpus
TEDx Spanish Corpus: Audio and Transcripts in Spanish Taken from TEDx Talks
Safe-first helper
Spanish Asr
Spontaneous Speech Recognition
Speech Transcription
Access pathOpenSLR
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-nd-4.0
- Code license
- not_applicable
- License caution
- OpenSLR lists CC BY-NC-ND 4.0. The corpus is derived from TEDx Talks, so downstream use should also respect TED/TEDx source terms.
- Download notes
- OpenSLR SLR67 hosts a single 2.3 GiB archive with Spanish speech and transcripts from TEDx Talks. The helper saves the OpenSLR page by default and downloads the archive only with TEDX_SPANISH_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/tedx_spanish.sh
Speaker, identity & emotion
TESS
Toronto Emotional Speech Set
Safe-first helper
Speech Emotion Recognition
Acted Emotional Speech
Auditory Emotion Perception
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The official Borealis record lists CC BY-NC 4.0. The corpus contains identifiable human voices and permits only non-commercial reuse under that license; preserve attribution and review voice-data ethics for downstream use.
- Download notes
- The owner-hosted University of Toronto Dataverse release contains 2,800 WAV stimuli: 200 target words spoken by two English-speaking actresses aged 26 and 64 in seven acted emotions. The helper saves official dataset metadata by default and requires TESS_DOWNLOAD_AUDIO=1 before downloading the complete ZIP.
Safe-first helperscripts/download/tess.sh
Speech generation
Text to Audio Human Preference Benchmark
Rapidata Text to Audio Human Preference Benchmark
Safe-first helper
Text To Speech Evaluation
Human Preference Evaluation
Speech Naturalness Evaluation
Speech Friendliness Evaluation
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified
- Code license
- not_applicable
- License caution
- The official dataset card and Hugging Face API declare no license. The rows include model labels, aggregate scores, individual votes, and annotator demographic fields; public access does not imply permission to redistribute or reuse those records, audio references, or generated outputs.
- Download notes
- The public, ungated release contains 4,269 pairwise comparison rows and about 32,000 human responses judging generated voices for friendliness and naturalness. It stores audio references as strings rather than embedding audio. The helper downloads the dataset card and API metadata by default; the approximately 0.8 MB repository snapshot is opt-in.
Safe-first helperscripts/download/rapidata_tts_preference.sh
Speech recognition
THCHS-30
THCHS-30: A Free Chinese Speech Corpus
Safe-first helper
Automatic Speech Recognition
Mandarin Speech Recognition
Noisy Speech Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Apache-2.0
- Code license
- not_applicable
- License caution
- OpenSLR lists Apache License v2.0 and the resource description says the database is free to academic users. The original CSLT URL linked from OpenSLR returned 404 when checked, so use the current OpenSLR page and paper for access/provenance.
- Download notes
- OpenSLR SLR18 hosts a 6.4 GiB speech/transcript archive, a 1.9 GiB 0 dB noisy test archive, and a 24 MiB supplementary resource archive with lexicon/noise samples. The helper saves the OpenSLR page by default and only downloads selected archives through THCHS30_DOWNLOAD_PARTS.
Safe-first helperscripts/download/thchs_30.sh
Speaker, identity & emotion
TidyVoice
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
Manual or gated
Multilingual Speaker Verification
Cross Lingual Speaker Verification
Speaker Recognition
Language Mismatch Robustness
+1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- CC0-1.0_with_use_restrictions
- Code license
- Apache-2.0
- License caution
- Mozilla Data Collective labels TidyVoiceX_ASV CC0-1.0 but also states that it must only be used for speaker verification and forbids speaker identification or attempts to recover speaker identity. Treat those owner-stated usage rules and the current Common Voice terms as binding access conditions despite the permissive license label. Apache-2.0 covers the WeSpeaker baseline repository, not any separate model or derived artifact rights.
- Download notes
- The public Mozilla Data Collective release contains 321,711 utterances (457 hours) from 4,474 multilingual speakers across 40 languages, with training and development splits, pseudonymized speaker IDs, language metadata, and same-/cross-language target and non-target trial lists. The current archive is approximately 36.72 GB. Download requires a Data Collective account and API key, so the helper saves official public documentation and prints the owner-supported access path without accepting credentials or fetching audio. The January paper also describes the broader Tidy-M monolingual condition across 81 languages; this entry's reproducible download pointer is the released TidyVoiceX_ASV challenge package. AMECxSV section 4.1 evaluates a deterministic speaker-disjoint split derived from the 12-million-trial TidyVoiceX development protocol, not the challenge's hidden official evaluation set.
Safe-first helperscripts/download/tidyvoice.sh
Audio understanding, generation & events
TimeGround-1M
Safe-first helper
Temporal Audio Grounding
Temporal Audio Localization
Timestamped Audio Description
Timestamped Audio Summarization
+1 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-3.0_on_hugging_face_card
- Code license
- not_specified
- License caution
- The official dataset card declares CC BY 3.0. Its recordings are selected from English YODAS2 YouTube-derived shards, so source-video rights, availability, attribution, and platform terms still require review. No separate license for generation or evaluation code is provided on the dataset card.
- Download notes
- The public, ungated English release contains separate train and test splits for temporal localization, temporal description, timed summarization, and recording-level nested annotations. The official card reports about 59,000 training and 4,200 test recordings totaling roughly 14,200 hours across duration buckets from under 10 minutes to 120 minutes. The GigaChat 3.1 Audio paper evaluates these generated tasks by duration bucket in section 4.1. The helper downloads the dataset card, repository API metadata, paper page, and model card by default; the Hugging Face API reports about 1.50 TB of repository storage, so the full snapshot requires explicit opt-in.
Safe-first helperscripts/download/timeground_1m.sh
Speech recognition
TIMIT
TIMIT Acoustic-Phonetic Continuous Speech Corpus
Manual or gated
Automatic Speech Recognition
Phone Recognition
Acoustic Phonetic Analysis
Speaker Dialect Coverage
Access pathLDC / licensed
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- custom_ldc_license
- Code license
- not_applicable
- License caution
- LDC catalog pages list licensing instructions for Subscription/Standard Members and Non-Members, web download media, and fee visibility after login. Portions are copyright 1993 Trustees of the University of Pennsylvania; consult the current LDC agreement before use or redistribution.
- Download notes
- LDC distributes TIMIT by web download after login/licensing. The helper only prints official access steps because the corpus is paid/licensed and not publicly script-downloadable.
Safe-first helperscripts/download/timit.sh
Speech recognition
TORGO
TORGO Database of Acoustic and Articulatory Speech from Speakers with Dysarthria
Safe-first helper
Dysarthria Detection
Pathological Speech Recognition
Speech Intelligibility Assessment
Acoustic Articulatory Modeling
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_academic_nonprofit_only
- Code license
- not_applicable
- License caution
- The owner page says use is free for academic, non-profit purposes and requires citation of at least one listed TORGO paper. It supplies the data as-is and does not identify a standard open-data license or grant commercial use. The recordings contain identifiable voices and disability and health information, so ethical and privacy review remains necessary.
- Download notes
- The public University of Toronto release contains aligned 16 kHz acoustic recordings and measured 3D articulatory features from eight English speakers with cerebral palsy or amyotrophic lateral sclerosis and seven matched controls. Stimuli include non-words, isolated words, restricted sentences, and spontaneous descriptions. Four BZip2 archives are organized as female dysarthric (F), female control (FC), male dysarthric (M), and male control (MC); they total approximately 8.9 GiB compressed and 18 GB uncompressed. The helper downloads the official page, correction spreadsheet, and coil-location documentation by default. Archive downloads require explicit terms acknowledgment and a selected group list. The July 2026 voice-concept bottleneck paper evaluates only headMic recordings with leave-one-speaker-out cross-validation; its exact derived split is not separately released.
Safe-first helperscripts/download/torgo.sh
Audio understanding, generation & events
TREA
Temporal Reasoning Evaluation of Audio
Safe-first helper
Audio Question Answering
Temporal Audio Reasoning
Audio Event Ordering
Audio Event Counting
+2 more
Access pathOfficial / other
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC0-1.0_repository_with_ESC-50_upstream_terms
- Code license
- CC0-1.0
- License caution
- The repository applies a CC0-1.0 LICENSE and GitHub detects CC0-1.0, but the paper states that every TREA audio file combines recordings from ESC-50, whose dataset is CC BY-NC 3.0 and whose ESC-10 subset clips are CC BY. Treat the restrictive upstream terms and clip attribution as surviving the derived release rather than assuming the repository-level CC0 waiver clears all source-audio rights.
- Download notes
- TREA is a public, ungated 600-item temporal-reasoning benchmark derived by combining ESC-50 clips. Its TREA-O, TREA-C, and TREA-D subsets each contain 200 ordering, counting, or duration questions. The July 2026 Audio-Zero paper evaluates both Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA alongside MMAU Test-mini and MMAR. The repository releases both four-option multiple-choice and open-text answer formats, audio, metadata, evaluation code, and uncertainty perturbation scripts. The helper downloads official documentation, repository metadata, license, paper page, and the lightweight CSV annotations by default; set TREA_CLONE_REPO=1 to clone the approximately 688 MiB GitHub repository and its audio.
Safe-first helperscripts/download/trea.sh
Speech generation
TTS Multilingual Test Set
MiniMaxAI TTS Multilingual Test Set
Safe-first helper
Multilingual Text To Speech
Zero Shot Voice Cloning
Cross Lingual Voice Cloning
Speech Intelligibility Evaluation
+2 more
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_applicable
- License caution
- The official Hugging Face dataset card lists CC BY-SA 4.0. Its 48 speaker prompts are selected from Mozilla Common Voice, whose data is CC0-1.0; retain benchmark attribution and share adaptations under the card's stated terms.
- Download notes
- The public, ungated Hugging Face repository contains 100 test sentences and two Common Voice-derived speaker prompts (one female and one male) for each of 24 languages. The helper downloads the official dataset card by default; the approximately 7.3 MB snapshot requires TTS_MULTILINGUAL_TEST_SET_DOWNLOAD_HF=1. Qwen3-TTS evaluates a 10-language subset for zero-shot multilingual and target-speaker generation, but the report does not identify the exact text rows used.
Safe-first helperscripts/download/tts_multilingual_test_set.sh
Audio understanding, generation & events
TUT Sound Events 2017
TUT Sound Events 2017: Sound Event Detection in Real-Life Audio
Safe-first helper
Sound Event Detection
Polyphonic Sound Event Detection
Temporal Audio Event Localization
Street Sound Event Detection
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_noncommercial
- Code license
- not_applicable
- License caution
- Zenodo labels both releases Other (Non-Commercial), and each documentation archive contains a controlling EULA. Review that packaged agreement before use or redistribution; the generic Zenodo label is not a permissive Creative Commons grant.
- Download notes
- The public version-2 development release contains 24 street recordings totaling 1:32:08 with verified strong annotations for six overlapping event classes and an official four-fold cross-validation setup. The public evaluation release contains eight recordings totaling 29:09 and now includes reference metadata. DCASE 2017 Task 3 ranks systems by one-second segment-based error rate. The helper downloads Zenodo record JSON, documentation, and small annotation archives by default; the approximately 1.55 GiB of 24-bit, 44.1 kHz audio requires explicit opt-in.
Safe-first helperscripts/download/tut_sound_events_2017.sh
Audio understanding, generation & events
UrBAN
UrBAN: Urban Beehive Acoustics and PheNotyping Dataset
Safe-first helper
Environmental Sound Classification
Beehive Acoustic Monitoring
Colony Strength Regression
Hive Health Monitoring
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- not_specified
- License caution
- The FRDR dataset record explicitly lists CC BY 4.0. The GitHub repository has no detected license, so the analysis notebooks and scripts should not be assumed to use the dataset license. The Scientific Data article itself is CC BY-NC-ND 4.0, distinct from the dataset terms.
- Download notes
- The public FRDR release contains longitudinal 2021-2022 raw 16 kHz beehive audio plus inspection, temperature, humidity, and weather metadata from a ten-hive Montréal rooftop apiary. The Scientific Data descriptor reports more than 3,000 hours, while the older FRDR record and repository README say more than 2,000 hours; the current FRDR landing page reports approximately 1.265 TB of files. Its benchmark protocols include random-split and hive-independent colony-strength regression. A 2026 follow-up evaluates modulation-tensorgram models on nine hives and emphasizes cross-hive generalization. The helper saves official landing pages, repository documentation, and API metadata only. Full data transfer remains a manual FRDR Globus workflow because of the corpus size and may require a Globus account and client.
Safe-first helperscripts/download/urban_beehive.sh
Audio understanding, generation & events
UrbanSound8K
UrbanSound8K: A Dataset and Taxonomy for Urban Sound Research
Safe-first helper
Urban Sound Classification
Environmental Sound Classification
Audio Tagging
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc
- Code license
- not_applicable
- License caution
- Zenodo lists CC BY-NC 4.0. The official Urban Sound site says UrbanSound/UrbanSound8K are free for non-commercial use under Creative Commons BY-NC 3.0; Freesound attributions are included in the dataset.
- Download notes
- The archive is about 6 GiB and contains 8732 WAV clips pre-sorted into 10 official folds. The helper downloads citation/license metadata by default and requires URBANSOUND8K_DOWNLOAD_AUDIO=1 for the full archive.
Safe-first helperscripts/download/urbansound8k.sh
Speech understanding & dialogue
URO-Bench-pro
URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
Safe-first helper
Spoken Dialogue Model Evaluation
Speech To Speech Evaluation
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- mit
- Code license
- MIT
- License caution
- Qwen uses the pro track. HF card and GitHub repo list MIT.
Safe-first helperscripts/download/uro_bench_pro.sh
Representation & general suites
User-Intent Queries (UIQ)
User-Intent Queries benchmark from Omni-Embed-Audio
Safe-first helper
User Intent Audio Retrieval
Language Based Audio Retrieval
Query Reformulation Robustness
Exclusionary Query Understanding
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0
- Code license
- MIT
- License caution
- CC BY 4.0 applies only to the released UIQ text queries. AudioCaps, Clotho, and MECAT audio is not redistributed and retains its original source terms; the top-level Omni-Embed-Audio code repository is MIT.
- Download notes
- The public, ungated release contains 13,053 text-query records over the AudioCaps test, Clotho evaluation, and MECAT pools: question, imperative, tagging, paraphrase, and exclusionary negative variants. The helper downloads the approximately 12 MiB of query JSONL files plus the benchmark README and license; it does not download source audio. Fusion Embedding section 6.3 independently reuses UIQ on the 1,045-clip Clotho pool and reports only the four positive query formulations.
Safe-first helperscripts/download/uiq.sh
Speech generation
VCTK
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit
Safe-first helper
Text To Speech
Speech Synthesis
Voice Cloning
Multi Speaker Speech Synthesis
+1 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- The official VCTK README and DataShare license_text identify Creative Commons Attribution 4.0 International. The newspaper text source was used with permission from Herald & Times Group.
- Download notes
- The official DataShare ZIP is about 10.94 GiB. The helper saves the official README and license text by default and requires VCTK_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/vctk.sh
Audiovisual & cross-modal
VGGSound
VGGSound: A Large-scale Audio-Visual Dataset
Safe-first helper
Audio Visual Event Classification
Audio Event Classification
Audio Tagging
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- Official VGG page and repository license file list the dataset as CC BY 4.0 for commercial/research use, while copyright remains with original video owners. Re-check YouTube availability and upstream media terms before reconstructing clips.
- Download notes
- The official VGG page currently says the original dataset download links are no longer available from that website. The helper downloads the official CSV metadata, license, and optional pretrained model files only; it does not fetch or redistribute YouTube media.
Safe-first helperscripts/download/vggsound.sh
Audiovisual & cross-modal
Video-MME
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Safe-first helper
Audio Visual Question Answering
Long Video Understanding
Multimodal Reasoning
Audio Enabled Video Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- custom_academic_research_only
- Code license
- not_specified
- License caution
- The official README prohibits commercial use and, without prior approval, distribution, publication, copying, dissemination, or modification of Video-MME in whole or in part. Video copyrights remain with their owners. The GitHub repository has no detected license; obtain approval and re-check source-video rights before reuse beyond the stated academic evaluation context.
- Download notes
- The public, ungated release contains 900 videos totaling 254 hours and 2,700 human-annotated question-answer pairs, with audio and subtitles available as evaluation modalities. The helper downloads only official documentation by default. The Hugging Face API reports about 389 GB of repository storage, so the media snapshot requires both VIDEO_MME_ACK_TERMS=1 and VIDEO_MME_DOWNLOAD_HF=1. Qwen3.5-Omni evaluates Video-MME with use_audio_in_video=True in section 5.1.4, Table 7.
Safe-first helperscripts/download/video_mme.sh
Audiovisual & cross-modal
video-SALMONN 2 Caption Benchmark
video-SALMONN 2 Human-Annotated Audio-Visual Caption Benchmark
Safe-first helper
Audio Visual Video Captioning
Detailed Video Captioning
Audio Visual Event Understanding
Caption Completeness Evaluation
+1 more
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0_card_label
- Code license
- Apache-2.0
- License caution
- The Hugging Face card labels the dataset Apache-2.0, and the official GitHub repository contains an Apache-2.0 LICENSE. The release does not document per-video provenance or underlying media licenses, so the card label must not be assumed to clear third-party video, audio, speech, music, likeness, or platform rights. Review source-media rights before redistribution or commercial use.
- Download notes
- The public, ungated test set contains 483 audio-bearing videos, each 30-60 seconds long, with a human-annotated detailed caption and manually refined visual, speech, and non-speech atomic events. The released evaluator uses an LLM to report missing-event, incorrect-event, hallucination, and total error rates. The helper downloads official documentation, API metadata, the approximately 3.5 MB annotation JSON, and evaluator by default. The current Hugging Face files total approximately 1.70 GB, so the 483 MP4 files require VIDEO_SALMONN2_DOWNLOAD_HF=1. ReMo evaluates this test set as video-SALMONN2 / video-SAL2 in section 5.1 of arXiv:2607.21179.
Safe-first helperscripts/download/video_salmonn2_caption.sh
Music
VocalSet
VocalSet: A Singing Voice Dataset
Safe-first helper
Singing Voice Analysis
Vocal Technique Classification
Vowel Classification
Singing Voice Synthesis
+1 more
Access pathZenodo
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- The Zenodo record lists CC BY 4.0 and open access. Re-check subject-consent and attribution expectations before redistributing derivative voice data.
- Download notes
- Zenodo hosts a single VocalSet.zip archive of about 2.1 GB with 10.1 hours of monophonic professional singing from 20 singers, covering all five vowels across standard and extended vocal techniques. The helper saves the Zenodo record metadata by default and requires VOCALSET_DOWNLOAD_ARCHIVE=1 before downloading the full archive.
Safe-first helperscripts/download/vocalset.sh
Audio understanding, generation & events
VocalSound
VocalSound: A Dataset for Improving Human Vocal Sounds Recognition
Safe-first helper
Human Vocal Sound Classification
Vocalization Recognition
Audio Classification
Demographic Bias Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-sa-4.0
- Code license
- not_specified
- License caution
- The official README includes a Creative Commons BY-SA 4.0 notice for the VocalSound dataset. GitHub API reports no repository-level license, so the code/baseline license is not specified; re-check before redistributing code or derived data.
- Download notes
- The official README describes 21,024 crowdsourced recordings from 3,365 subjects covering laughter, sighs, coughs, throat clearing, sneezes, and sniffs, with speaker metadata such as age, gender, native language, country, and health condition. The helper saves official documentation by default and requires VOCALSOUND_DOWNLOAD_ARCHIVE=1 before downloading the 1.7 GiB 16 kHz or 4.5 GiB 44.1 kHz ZIP.
Safe-first helperscripts/download/vocalsound.sh
Enhancement, separation & quality
VoiceBank-DEMAND
Noisy speech database for training speech enhancement algorithms and TTS models
Safe-first helper
Speech Enhancement
Speech Denoising
Noise Robust Tts
Clean Noisy Parallel Speech
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_applicable
- License caution
- Edinburgh DataShare metadata lists Creative Commons Attribution 4.0 International Public License. The corpus derives clean speech from VCTK and noises from DEMAND plus speech-shaped/babble sources; re-check component/source terms before redistribution.
- Download notes
- The DataShare record exposes paired clean/noisy train and test ZIPs plus text/log files. The helper saves public metadata and license by default; text files and multi-GB audio archives are explicit opt-ins.
Safe-first helperscripts/download/voicebank_demand.sh
Speech understanding & dialogue
VoiceBench
VoiceBench: Benchmarking LLM-Based Voice Assistants
Safe-first helper
Voice Assistant Evaluation
Spoken Instruction Following
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- apache-2.0
- Code license
- Apache-2.0
- License caution
- HF dataset card and GitHub repo both list Apache-2.0.
Safe-first helperscripts/download/voicebench.sh
Speech recognition
VoiceCodeBench
VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
Safe-first helper
Asr
Structured Token Recovery
Entity Recovery
Workplace Speech Recognition
Access pathHugging Face
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- Not specified in the source record.
- Code license
- Not specified in the source record.
- License caution
- The repository and dataset card declare MIT. The card says paid contributors consented to dataset use and release, but the audio contains identifiable voice characteristics and its stated intended-use guidance excludes speaker identification, biometric modeling, voice cloning, demographic profiling, and model training or post-training.
- Download notes
- The public, ungated test-only release contains 300 human-recorded English workplace-speech segments totaling 5.587 hours, with 85 anonymized speakers and 1,482 audited targets across 26 structured entity types. Its primary Canonical Token/Entity Match and Task Success Rate metrics test exact recovery of values such as email addresses, phone numbers, URLs, command-line flags, file paths, identifiers, dates, and measurements. The helper downloads official documentation, license, paper, API metadata, and the approximately 1.1 MB annotation JSONL by default. The complete Hugging Face repository is approximately 1.83 GiB and requires VOICECODEBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/voicecodebench.sh
Speaker, identity & emotion
VoiceMOS Challenge 2026
VoiceMOS Challenge 2026: Automatic Prediction of Human Ratings of Speech
Manual or gated
Mean Opinion Score Prediction
Speech Quality Assessment
Comparative Category Rating Prediction
Emotional Speech Naturalness Assessment
+3 more
Access pathOfficial / other
Upstream termsMixed / custom — review
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- not_publicly_specified
- Code license
- Apache-2.0
- License caution
- The public challenge page and baseline README do not state dataset reuse or redistribution terms. Apache-2.0 covers the baseline repository only, not challenge audio, listener ratings, URGENT material, CodecMOS-Accent, or emotional-speech source data.
- Download notes
- The official site says training data were released to registered participants through a CodaBench page sent by email, with evaluation data scheduled for July 31, 2026. Track 1 covers 840 multilingual utterances in nine languages from six URGENT speech-enhancement systems; Track 2 covers emotional TTS and human speech; Track 3 uses 4,000 CodecMOS-Accent samples from 24 codec-resynthesis and TTS systems, 32 speakers, and ten accents. The helper saves public challenge and baseline documentation, then prints the registration path; it does not guess or expose the emailed CodaBench URL.
Safe-first helperscripts/download/voicemos_challenge_2026.sh
Audiovisual & cross-modal
VoxBlink2
VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Safe-first helper
Speaker Verification
Open Set Speaker Identification
Speaker Recognition
Audio Visual Speaker Recognition
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-NC-SA-4.0
- Code license
- not_specified
- License caution
- The repository states that released annotation data is CC BY-NC-SA 4.0, but does not separately license its software. YouTube source-media rights, platform terms, privacy considerations, and local law remain separate and are not granted by the annotation license.
- Download notes
- The official release provides annotations, YouTube links, timestamps, speaker labels, ASR outputs, speaker metadata, and evaluation protocols rather than redistributing audio or video. The corpus describes approximately 10 million segments, more than 110,000 speakers, and 16,000 hours across more than 15 language families. The helper downloads official documentation and license text by default; the Google Drive resource bundle remains a manual download, and cloning evaluation/data-construction code is opt-in. Source media must be obtained separately and may be unavailable or removed.
Safe-first helperscripts/download/voxblink2.sh
Audiovisual & cross-modal
VoxCeleb
VoxCeleb speaker recognition datasets
Safe-first helper
Speaker Identification
Speaker Verification
Speaker Recognition
Audio Visual Speaker Recognition
Access pathOpenSLR
Upstream termsNot specified
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- not_specified_for_original_media
- Code license
- not_applicable
- License caution
- Official VGG pages say provided VoxCeleb/VoxCeleb2 metadata is CC BY-SA 4.0 and the corpora consist of YouTube URLs with timestamps; original media rights and privacy terms remain with upstream owners. OpenSLR SLR49 lists its small metadata resource as not copyrighted.
- Download notes
- Official VGG pages currently say VoxCeleb1 and VoxCeleb2 audio, URL/timestamp, and identifying metadata files are no longer available from that website. The helper downloads small OpenSLR speaker-recognition recipe metadata and trial lists only; it does not fetch the original audio/video.
Safe-first helperscripts/download/voxceleb.sh
Audiovisual & cross-modal
VoxConverse
VoxConverse: A Large Scale Audio-Visual Diarisation Dataset
Safe-first helper
Speaker Diarization
Audio Visual Diarization
Overlapping Speech Diarization
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0
- Code license
- not_specified
- License caution
- The official page and repository README say VoxConverse is available for research purposes under CC BY 4.0, while copyright remains with the original video owners. The GitHub repository does not expose a standalone license file through the API.
- Download notes
- The helper clones or updates the official annotation repository and saves the official page by default. The official page lists dev/test WAV ZIPs with MD5 checksums; audio downloads are explicit opt-ins because the dev ZIP is about 1.9 GiB and the test ZIP is also large.
Safe-first helperscripts/download/voxconverse.sh
Speaker, identity & emotion
VoxENES 2026
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Safe-first helper
Speech Spoofing Detection
Audio Deepfake Detection
Synthetic Speech Detection
Voice Conversion Detection
+2 more
Access pathOfficial / other
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-4.0-with-upstream-terms
- Code license
- not_applicable
- License caution
- Kaggle declares CC BY 4.0 for the release. Bona fide speech derives from LibriSpeech and VoxPopuli, and synthetic samples incorporate source speech, speaker references, and outputs from multiple TTS/VC systems; review those upstream terms, voice-data rights, and model-output policies before redistribution or commercial use. The paper's CC BY 4.0 license applies to the paper, not by itself to every incorporated recording.
- Download notes
- The public Kaggle release contains 53,628 standardized 16 kHz mono WAV samples across English and Spanish, including 3,028 bona fide samples, 4,600 original synthetic samples from seven TTS and three voice-conversion systems, and 46,000 post-processed variants. The helper downloads Kaggle metadata by default; the approximately 23.3 GB dataset requires explicit opt-in and an authenticated Kaggle CLI.
Safe-first helperscripts/download/voxenes_2026.sh
Speech understanding & dialogue
VoxLingua107
VoxLingua107: a Dataset for Spoken Language Recognition
Safe-first helper
Spoken Language Identification
Language Recognition
Speech Classification
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The TalTechNLP Hugging Face dataset card lists cc-by-nc-4.0. The dataset is built from YouTube-derived speech segments, so source-media availability and platform terms still apply; the SpeechBrain recipe repository did not expose a detected license.
- Download notes
- The paper reports 6628 hours across 107 languages plus a 1609-utterance verified evaluation set. The helper downloads small Hugging Face metadata files by default and requires VOXLINGUA107_DOWNLOAD_HF=1 before attempting the larger mirrored dataset snapshot. The original TalTech host was not reliably reachable during the 2026-07-09 check, so verify upstream availability before large downloads.
Safe-first helperscripts/download/voxlingua107.sh
Speech recognition
VoxPopuli
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Safe-first helper
Multilingual Asr
Speech To Text Translation
Self Supervised Speech Representation Learning
Accented Speech Recognition
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc0-1.0
- Code license
- cc-by-nc-4.0
- License caution
- Official repo lists VoxPopuli data as CC0 and points users to the European Parliament legal notice for raw data; code and pretrained models are CC BY-NC 4.0.
- Download notes
- HF hosts converted Parquet shards and is about 673 GiB total; select a language/config and split before downloading.
Safe-first helperscripts/download/voxpopuli.sh
Audio understanding, generation & events
WABAD
WABAD: A World Annotated Bird Acoustic Dataset for Passive Acoustic Monitoring
Safe-first helper
Bird Species Detection
Passive Acoustic Monitoring
Temporal Audio Event Localization
Time Frequency Event Localization
+1 more
Access pathZenodo
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_zenodo_metadata_treat_as_cc-by-nc-4.0
- Code license
- not_applicable
- License caution
- The Zenodo structured license field says CC BY 4.0, but the record's human-readable description explicitly says Creative Commons Attribution-NonCommercial 4.0. Treat the release as CC BY-NC 4.0 pending clarification from the maintainers; retain attribution and do not assume commercial-use permission from the structured field alone.
- Download notes
- The public, ungated release contains 5,047 minutes of passive-acoustic audio with 91,931 time-frequency-bounded vocalizations from 1,192 bird species, collected at 72 sites in 29 recording locations across 13 biomes. MetaPerch evaluates WABAD as an 84-hour multi-species detection benchmark in its results section. The helper downloads the Zenodo record, README, site metadata, pooled annotations, and species list by default; the 72 site archives total approximately 19.8 GiB and require explicit site-level opt-in.
Safe-first helperscripts/download/wabad.sh
Audio understanding, generation & events
WavCaps
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
Safe-first helper
Audio Captioning
Audio Language Retrieval
Audio Language Modeling
Zero Shot Audio Classification
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- academic_only
- Code license
- not_specified
- License caution
- The GitHub README and Hugging Face card say only academic uses are allowed for WavCaps audio. The HF metadata advertises CC BY 4.0, but the dataset card also points users to component source terms for FreeSound, BBC Sound Effects, SoundBible, and AudioSet; re-check those source licenses before redistribution or commercial use. Provided models are described as non-commercial research under a UK data copyright exemption.
- Download notes
- The Hugging Face repository exposes JSON metadata and split FLAC waveform ZIPs for FreeSound, BBC Sound Effects, SoundBible, and AudioSet SL. The full repository is hundreds of GiB, so the helper downloads README/JSON metadata by default and requires WAVCAPS_DOWNLOAD_ZIPS=1 plus WAVCAPS_ZIP_SOURCES for waveform archives.
Safe-first helperscripts/download/wavcaps.sh
Speech recognition
WenetSpeech
Manual or gated
Mandarin Asr
Access pathOfficial / other
Upstream termsNon-commercial / research terms
Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.
Access, terms & download helper
- Data license / terms
- non-commercial use under CC BY 4.0
- Code license
- Apache-2.0
- License caution
- Official site says WenetSpeech does not own audio copyright; original audio copyrights remain with owners.
Safe-first helperscripts/download/wenetspeech.sh
Enhancement, separation & quality
WHAM! / WHAMR!
WSJ0 Hipster Ambient Mixtures and WHAMR!: Noisy and Reverberant Single-Channel Speech Separation
Safe-first helper
Noisy Speech Separation
Speech Enhancement
Reverberant Speech Separation
Source Separation
Access pathOfficial / other
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- cc-by-nc-4.0
- Code license
- not_specified
- License caution
- The official WHAM page states the WHAM! and WHAM!48kHz noise datasets are CC BY-NC 4.0. Generated mixtures also depend on WSJ0/wsj0-2mix licensing, so redistribution or commercial use requires checking those upstream terms too.
- Download notes
- The helper downloads the official landing page and small WHAM!/WHAMR! generation script archives by default. WHAM! noise is 17 GiB compressed and WHAM!48kHz is 68.1 GiB compressed, so those archives are explicit opt-ins. Building full WHAM!/WHAMR! mixtures also requires separately licensed WSJ0/wsj0-2mix access.
Safe-first helperscripts/download/wham_whamr.sh
Speech recognition
Whisper-RIR-Mega
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
Safe-first helper
Automatic Speech Recognition
Reverberant Speech Recognition
Asr Robustness
Room Acoustics Robustness
Access pathHugging Face
Upstream termsNon-commercial / research terms
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC-BY-4.0_with_CC-BY-NC-4.0_upstream_terms
- Code license
- not_specified_currently_unavailable
- License caution
- The benchmark dataset card declares CC BY 4.0 and identifies LibriSpeech as CC BY 4.0, but the current RIR-Mega v2 card declares CC BY-NC 4.0 for its RIR audio. Apply the stricter non-commercial upstream terms to the derived reverberant audio unless the owner clarifies otherwise. The benchmark card says its curation repository is MIT, but the linked repository was unavailable, so that code license could not be independently verified.
- Download notes
- The public, ungated release contains 2,000 English LibriSpeech test-clean utterances, each paired with a 16 kHz reverberant version made using one RIR-Mega room impulse response. Its deterministic, acoustically stratified split has 400 validation and 1,600 test pairs; evaluation reports clean/reverberant WER and CER plus the reverb penalty, with RT60 and DRR metadata when available. The helper saves the dataset card, API metadata, paper, and small leaderboard files by default. The Hugging Face API reports about 1.13 GB of repository storage, so the complete audio and Arrow snapshot requires WHISPER_RIRMEGA_DOWNLOAD_HF=1. The paper's cited GitHub code repository returned HTTP 404 when checked on 2026-07-22.
Safe-first helperscripts/download/whisper_rirmega.sh
Speech understanding & dialogue
WildSpeech-Bench
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
Safe-first helper
Speech To Speech Evaluation
Natural Speech Conversation
Access pathHugging Face
Upstream termsOpen / attribution signals
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- CC BY 4.0, except third-party datasets with their own terms
- Code license
- CC BY 4.0, except third-party datasets with their own terms
- License caution
- License.txt says users must comply with original licenses for third-party datasets.
Safe-first helperscripts/download/wildspeech_bench.sh
Audiovisual & cross-modal
WorldSense
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Safe-first helper
Audio Visual Question Answering
Omni Modal Video Understanding
Cross Modal Reasoning
Audio Visual Perception
Access pathHugging Face
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- conflicting_cc_by_nc_sa_4_0_and_cc_by_4_0
- Code license
- not_specified
- License caution
- The WorldSense paper v3 Appendix G states CC BY-NC-SA 4.0, while the official repository README and Hugging Face card state CC BY 4.0. Apply the more restrictive CC BY-NC-SA 4.0 interpretation until the maintainers resolve the conflict. Videos are sourced primarily from FineVideo with selected MUSIC-AVQA material, so component-media terms and rights also require review. GitHub reports no detected repository license.
- Download notes
- The public, ungated release contains 1,662 synchronized audio-visual videos and 3,172 multiple-choice question-answer pairs across 26 tasks. The helper downloads official documentation and the approximately 4.3 MB QA JSON by default; the Hugging Face API reports approximately 18.1 GB of repository storage, so video and subtitle archives require WORLDSENSE_DOWNLOAD_HF=1. Qwen3.5-Omni reports WorldSense in section 5.1.4, Table 7.
Safe-first helperscripts/download/worldsense.sh
Speaker, identity & emotion
WSJ0-2mix / wsj0-mix
wsj0-mix: Single-channel multi-speaker speech separation mixtures from WSJ0
Safe-first helper
Speech Separation
Multi Speaker Speech Separation
Source Separation
Cocktail Party Speech Separation
Access pathLDC / licensed
Upstream termsMixed / custom — review
The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.
Access, terms & download helper
- Data license / terms
- ldc_restricted_derived
- Code license
- MERL script license not specified on the reachable page; pywsj0-mix is MIT.
- License caution
- The generated mixtures derive from the LDC CSR-I WSJ0 corpus, so access, use, and redistribution must follow the active LDC agreement. The MERL page provides scripts but does not publish the audio mixtures or a standalone data license.
- Download notes
- The helper downloads the official MERL page and generation scripts by default and can clone the MIT-licensed Python generator. It does not download WSJ0 audio; generation requires an already licensed local WSJ0 corpus from LDC and explicit WSJ0_2MIX_RUN_GENERATION=1. TF-MossFormer sections 3.1-3.3 use the standard 8 kHz two-speaker setup with 20,000 training, 5,000 validation, and 3,000 speaker-disjoint test mixtures and report SI-SDRi and SDRi; that paper adds no new mixture release.
Safe-first helperscripts/download/wsj0_2mix.sh