Living benchmark directory

Find the right benchmark for speech, audio, and music

Search 288 benchmark families by task, access path, and upstream terms. Start from official sources, use the safe-first helpers, and adapt the evaluation workflow to your model.

Benchmarks288
Download helpers288
Manual or gated43
Last checked

A starting point—not a ranking or rights grant

This index does not mirror datasets or endorse one benchmark over another. Upstream licenses and access terms control; re-check them before training, redistribution, commercial use, or publishing results. Helpers default to documentation or metadata when large or restricted files require an explicit opt-in.

Start with the task

What are you evaluating?

Choose a broad family to narrow the catalog. Detailed task tags remain searchable inside every benchmark record.

Benchmark catalog

Browse all 288 benchmarks

Search names, tasks, source notes, or license terms. Filters change discovery only—the list is alphabetical, not ranked.

Showing 288 of 288 benchmarks.

Speech recognition

2nd MLC-SLM Challenge 2026

2nd Multilingual Conversational Speech Language Model Challenge 2026

Manual or gated
Multilingual Conversational Asr Speaker Diarization Speaker Attributed Asr Acoustic Conversation Understanding +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_workshop_only_registration_agreement
Code license
not_specified
License caution
The official agreement limits the datasets to the 2026 MLC-SLM Workshop, prohibits redistribution and any other use, requires access controls, and requires return or destruction after termination. The challenge and baseline repositories have no detected license files, so their public visibility does not establish permission to reuse code. This second-edition release must not be conflated with the first MLC-SLM Eval annotation repository or its CC BY-SA 4.0 card label.
Download notes
The official challenge describes approximately 2,100 hours of two-speaker conversational training audio across 14 languages, plus approximately four development hours per language. Task 1 evaluates diarization and recognition with DER and time-constrained minimum-permutation WER/CER; Task 2 evaluates acoustic and semantic understanding through multilingual multiple-choice questions. The Task 1 system paper reports 150 development conversations across 21 language/accent categories and says evaluation references are not released. Access to training, development, and evaluation data requires challenge registration and acceptance of the data-use agreement. The helper downloads only public documentation, agreement, repository metadata, and paper metadata, then prints the manual registration path.
Safe-first helperscripts/download/mlc_slm_2nd_challenge.sh
View helper
Audio understanding, generation & events

ADQA-Bench

ADQA-Bench: Audio-Dependent Question Answering Evaluation Benchmark

Safe-first helper
Audio Question Answering Audio Dependent Reasoning Multiple Choice Question Answering Shortcut Robustness
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0_with_upstream_terms
Code license
not_applicable
License caution
The Hugging Face card declares Apache-2.0 for the release, but the dataset incorporates a portion of MMAU, MMAR, and MMSU plus newly annotated questions over other audio. Those component datasets and source recordings retain their own terms; review provenance and upstream media rights before redistribution or commercial use.
Download notes
The public, ungated DCASE 2026 Task 5 evaluation release contains 3,000 English multiple-choice questions with 3,000 WAV files spanning speech, music, and general sound understanding. Every item passed the organizers' four-stage Audio-Dependency Filtering process to reduce silent-audio and text-only shortcuts. The DCASE 2026 task-summary paper reports 1,607 development items, 14 participating teams, 36 ranked submissions, and two parameter-count tracks; its team results use the hidden 3,000-item evaluation set rather than the development set used for organizer baselines. The current release provides questions and choices without answers for challenge evaluation; the dataset card says answers will be released after the competition. The helper downloads official pages, the dataset card, API metadata, and the lightweight no-answer JSONL by default; the Hugging Face API reports approximately 2.94 GB of repository storage, so the audio snapshot requires explicit opt-in.
Safe-first helperscripts/download/adqa_bench.sh
View helper
Speech understanding & dialogue

ADReSS / ADReSSo

Alzheimer's Dementia Recognition through Spontaneous Speech Challenges

Manual or gated
Cognitive Impairment Detection Alzheimers Dementia Classification Mmse Score Regression Cognitive Decline Prediction +2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC-BY-NC-SA-3.0_with_password_protected_clinical_access
Code license
not_applicable
License caution
TalkBank says CC BY-NC-SA 3.0 governs its data unless otherwise indicated, prohibits incorporation into commercial products or large language models, and restricts DementiaBank membership to established researchers and clinicians or faculty-sponsored students. Password- protected data may not be shared with non-members or posted elsewhere. The recordings contain sensitive clinical and potentially identifiable speech; users must also follow the TalkBank Code of Ethics, NIH confidentiality protections, citation rules, and non-storage requirements for web processing.
Download notes
The official DementiaBank pages release age- and gender-balanced spontaneous-speech challenge sets after consortium approval. ADReSS 2020 provides enhanced full audio, normalized speech segments, transcripts, demographics, diagnosis labels, and MMSE scores for Alzheimer's-dementia classification and MMSE regression. ADReSSo 2021 is audio-only at evaluation time and adds longitudinal cognitive- decline prediction. The helper saves the public challenge pages, DementiaBank access rules, and the July 2026 cross-dataset evaluation paper, then prints the manual membership path; it never attempts to access password-protected participant recordings.
Safe-first helperscripts/download/adress_challenges.sh
View helper
Audio understanding, generation & events

AF-Reasoning-Eval

AF-Reasoning-Eval: Sound Reasoning Evaluation Benchmark

Safe-first helper
Audio Question Answering Audio Reasoning Audio Classification Commonsense Reasoning
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC BY 4.0 metadata with mixed upstream audio terms
Code license
MIT
License caution
NVIDIA releases AF-Reasoning-Eval metadata under CC BY 4.0 and repository code under MIT. The 150 AQA items derive from Clotho-AQA, whose question-answer CSVs are MIT while Freesound audio retains per-file Creative Commons terms. The 7,227-item classification set derives from FSD50K, whose clips retain per-file CC0, CC-BY, CC-BY-NC, or CC Sampling+ terms in addition to the curated dataset's CC BY release.
Download notes
The helper downloads all four official JSON annotation files, totaling about 2.1 MB, plus the Sound-CoT README. It does not duplicate source audio. The AQA subset points to Clotho-AQA filenames and the classification subset points to FSD50K evaluation filenames; use those benchmarks' separate helpers and terms to obtain audio.
Safe-first helperscripts/download/af_reasoning_eval.sh
View helper
Music

AI-Generated Cover Song Diagnostics

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features

Safe-first helper
Cover Song Generation Evaluation Music Quality Assessment Music Error Diagnosis Music Feature Analysis
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT_for_released_tables_no_audio
Code license
MIT
License caution
The repository's MIT license covers its software and associated documentation, including the released score, manifest, and feature tables. It does not grant rights to the absent source songs or generated cover audio. No raw audio is publicly released, and users must supply locally authorized files to rerun audio feature extraction.
Download notes
The public release covers 30 generated covers from five source songs and six systems. It includes anonymized source-song identifiers, system/file mappings, expert scores for melody, harmony, key, style, and arrangement/production, nine extracted features, and the analysis pipeline. The helper downloads these lightweight tables and official documentation by default. The paper and repository explicitly withhold raw audio because of source-song copyright and commercial-API licensing constraints; cloning the small repository does not provide audio.
Safe-first helperscripts/download/ai_cover_song_diagnostics.sh
View helper
Audio understanding, generation & events

AIR-Bench

AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Safe-first helper
Audio Question Answering Audio Instruction Following Audio Language Model Evaluation Speech Understanding +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
Apache-2.0
License caution
The HF dataset card lists CC BY-NC 4.0 and enumerates component sources with mixed upstream licenses, including MusicCaps, Clotho, Fisher, SpokenWOZ, Common Voice, IEMOCAP, acoustic-scene datasets, MUSIC-AVQA, FMA, MTG-Jamendo, NSynth, SLURP, VoxCeleb, LibriSpeech, CoVoST 2, Fake-or-Real, and VocalSound. Re-check component terms before redistribution or commercial use.
Download notes
The helper downloads the official GitHub README and Hugging Face dataset card by default. The full HF audio snapshot is about 45.9 GB, so it requires AIR_BENCH_DOWNLOAD_HF=1; cloning the evaluation repo is also opt-in.
Safe-first helperscripts/download/air_bench.sh
View helper
Speech recognition

AISHELL-1

AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline

Safe-first helper
Automatic Speech Recognition Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0
Code license
not_applicable
License caution
OpenSLR SLR33 lists Apache License v2.0 and also describes the data as free for academic use; re-check upstream terms before redistribution or commercial use.
Download notes
OpenSLR hosts a 15 GiB speech/transcript archive plus a small supplementary resource archive with lexicon and speaker info. The helper saves the OpenSLR page and supplementary archive by default; the large corpus archive is an explicit opt-in.
Safe-first helperscripts/download/aishell_1.sh
View helper
Speech generation

AISHELL-3

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Safe-first helper
Text To Speech Speech Synthesis Multi Speaker Speech Synthesis Mandarin Speech Synthesis +1 more
Access pathOpenSLR
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0
Code license
not_specified
License caution
OpenSLR SLR93 lists Apache License v2.0 for the resource. The aishelltech external URL redirected but returned HTTP 429 during the 2026-07-09 check, so OpenSLR was used as the primary access page.
Download notes
OpenSLR lists one 19 GiB speech/transcript archive. The helper saves the official OpenSLR page by default and requires AISHELL3_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/aishell_3.sh
View helper
Speech recognition

AISHELL-4

AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Safe-first helper
Meeting Transcription Multi Channel Asr Speech Enhancement Speech Separation +2 more
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
Apache-2.0
License caution
OpenSLR SLR111 lists the corpus under CC BY-SA 4.0. The official baseline repository is Apache-2.0. The arXiv paper is CC BY 4.0, which is separate from the share-alike dataset terms.
Download notes
The public OpenSLR release contains 211 real Mandarin meeting sessions totaling 120 hours, recorded with an eight-channel circular microphone array in small, medium, and large rooms. It provides accurate transcripts and speaker activity for meetings with four to eight speakers. The helper saves the OpenSLR page and baseline documentation by default. The approximately 5.2 GB test archive and 46 GB of training archives are explicit opt-ins. VibeVoice-ASR-BitNet section 3.1 and Table 4 evaluate the AISHELL-4 test set with CER.
Safe-first helperscripts/download/aishell_4.sh
View helper
Speech recognition

AliMeeting

AliMeeting: A Free Mandarin Multi-channel Meeting Speech Corpus

Safe-first helper
Meeting Transcription Multi Channel Asr Multi Speaker Asr Speaker Diarization +1 more
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
not_applicable
License caution
OpenSLR SLR119 lists AliMeeting under CC BY-SA 4.0. The corpus contains real Mandarin meetings with far-field microphone-array and near-field headset recordings; check challenge rules for benchmark submissions.
Download notes
The helper downloads small OpenSLR metadata by default. Corpus archives are large, about 73.24 GiB far-field train, 22.85 GiB near-field train, 3.42 GiB eval, and 8.90 GiB test, so they are explicit opt-ins.
Safe-first helperscripts/download/alimeeting.sh
View helper
Speech recognition

AMI

AMI Meeting Corpus

Safe-first helper
Meeting Speech Recognition Multi Speaker Asr Distant Speech Recognition Meeting Understanding
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
Official AMI pages say the corpus, signals, transcription, and some annotations are CC BY 4.0. OpenSLR SLR16 still lists an older modified CC BY-NC-SA v2.0 notice, so prefer the official AMI license page for current terms and re-check before redistribution.
Download notes
The helper downloads official annotation ZIPs by default. OpenSLR acoustic archives and the HF converted dataset are large, so audio downloads are explicit opt-ins.
Safe-first helperscripts/download/ami.sh
View helper
Representation & general suites

Androids Corpus

The Androids Corpus: A New Publicly Available Benchmark for Speech Based Depression Detection

Safe-first helper
Speech Based Depression Detection Clinical Voice Assessment Health Audio Classification Paralinguistic Speech Analysis
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_academic_research_only_no_redistribution
Code license
not_specified
License caution
The official README limits the corpus to non-commercial academic research and prohibits redistribution, broadcasting, or making it publicly available anywhere else. No standalone code or data license file is present. The recordings expose depression/control labels plus age, gender, education, and identifiable voices, so users must also apply appropriate clinical-data, privacy, consent, and ethics review.
Download notes
The owner repository links a 3.69 GB archive containing 228 recordings from 118 native Italian speakers: 112 read-speech recordings and 116 spontaneous interview recordings, including 874 segmented interview clips, speaker metadata, turn timing, an openSMILE configuration, and official five-fold lists. The July 2026 voice-concept-bottleneck paper evaluates the 116-speaker interview subset with the official five-fold protocol. The helper saves owner documentation, repository metadata, and the primary Interspeech paper by default; archive download requires explicit acknowledgment of the restrictive terms.
Safe-first helperscripts/download/androids_corpus.sh
View helper
Speaker, identity & emotion

ASVspoof 2015

Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015) Database

Safe-first helper
Speaker Verification Anti Spoofing Spoofed Speech Detection Synthetic Speech Detection Unseen Attack Generalization
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
Edinburgh DataShare metadata declares Creative Commons Attribution 4.0 International. Retain attribution and review the packaged license and source-speech provenance before redistribution or commercial use.
Download notes
The public Edinburgh DataShare release contains genuine speech from 106 speakers and synthetic speech from ten known and unknown text-to-speech and voice-conversion attacks, partitioned into training, development, and evaluation sets. The helper downloads official metadata, README, evaluation plan, summary paper, file descriptions, and extraction instructions by default. The approximately 2.1 MB protocol archive is a separate opt-in, and the three-part approximately 24.1 GB WAV archive requires ASVSPOOF2015_DOWNLOAD_AUDIO=1. Section 4.1 of arXiv:2607.21127 uses attacks S3 and S10 for expert calibration.
Safe-first helperscripts/download/asvspoof_2015.sh
View helper
Speaker, identity & emotion

ASVspoof 2017 V2

The 2nd Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2017) Database, Version 2

Safe-first helper
Speaker Verification Anti Spoofing Presentation Attack Detection Spoofed Speech Detection Replay Attack Detection +1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_applicable
License caution
Edinburgh DataShare metadata explicitly declares Creative Commons Attribution-NonCommercial 4.0 International. The database uses genuine and replayed RedDots speech; retain attribution and review the packaged files and upstream RedDots conditions before use or redistribution.
Download notes
The official Version 2 release contains 42 speakers and genuine and replayed RedDots speech recorded across 179 replay sessions in 61 unique room, replay-device, and recording-device configurations. Training, development, and evaluation archives total approximately 1.4 GiB. The helper downloads the DataShare metadata, README, change log, instructions, evaluation plan, and primary papers by default; the approximately 104 KiB protocol archive and speech archives require separate explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2017.sh
View helper
Speaker, identity & emotion

ASVspoof 2019

ASVspoof 2019: The 3rd Automatic Speaker Verification Spoofing and Countermeasures Challenge database

Safe-first helper
Speaker Verification Anti Spoofing Presentation Attack Detection Spoofed Speech Detection Synthetic Speech Detection +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
odc-by-1.0
Code license
not_applicable
License caution
Edinburgh DataShare metadata lists Open Data Commons Attribution License and ships the ODC Attribution license text. The license text itself cautions that database contents can have separate rights; ASVspoof 2019 is derived from VCTK, so re-check component terms before redistribution.
Download notes
The DataShare record exposes README, license, evaluation plan, paper PDF, and LA/PA archives. The helper downloads small documentation/license files by default; LA is about 7.6 GiB and PA is about 17.7 GiB, so archives are explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2019.sh
View helper
Speaker, identity & emotion

ASVspoof 2021

ASVspoof 2021: Logical Access, Physical Access, and Speech Deepfake Challenge databases

Safe-first helper
Speaker Verification Anti Spoofing Presentation Attack Detection Spoofed Speech Detection Synthetic Speech Detection +2 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
not_specified
License caution
The official ASVspoof 2021 page says the databases are available under an Open Data Commons Attribution Licence. Zenodo currently lists ODC-BY for LA and PA, and ODC-ODbL for DF; re-check active Zenodo metadata and component-source terms before redistribution. The GitHub baseline repository did not expose a detected license via the GitHub API on 2026-07-10.
Download notes
The helper downloads the evaluation plan, LA/PA/DF keys and metadata, Zenodo record metadata, and file-map metadata by default. Evaluation speech archives are large: LA is about 7.8 GB, PA is split into seven parts totaling about 47 GB, and DF is split into four parts totaling about 34.5 GB, so speech archives are explicit opt-ins.
Safe-first helperscripts/download/asvspoof_2021.sh
View helper
Speaker, identity & emotion

ASVspoof 5

ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale

Safe-first helper
Speaker Verification Anti Spoofing Spoofed Speech Detection Synthetic Speech Detection Deepfake Speech Detection +2 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
odc-by-1.0_database_and_cc-by-4.0_bona-fide_audio
Code license
not_specified
License caution
The packaged license applies ODC Attribution 1.0 to the database and CC BY 4.0 to bona fide data. ODC-BY explicitly does not grant every right in individual contents; preserve Multilingual LibriSpeech provenance and review privacy, personality, and generated-voice rights. The baseline repository has no detected license.
Download notes
The public release contains 182,357 training, 142,134 development, and 681,872 evaluation utterances at 16 kHz from crowdsourced speech by roughly 2,000 speakers. It covers more than 20 spoofing attacks, seven adversarial attacks, codec conditions, countermeasure evaluation, and spoofing-robust speaker verification. The helper downloads official metadata, README, license, evaluation plan, paper page, and baseline README by default. The approximately 19.7 MiB protocol archive is a separate opt-in; the complete approximately 142.3 GB audio release remains on Zenodo and is not downloaded by the helper.
Safe-first helperscripts/download/asvspoof_5.sh
View helper
Speech understanding & dialogue

Audio Agent Bench Suite

AudioAgentBench: Evaluating Multi-Turn Voice Agents on Real-World Tasks

Safe-first helper
End To End Speech Dialogue Spoken Question Answering Audio Agent Tool Use Multi Turn State Tracking +2 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_CC-BY-4.0_and_MIT_metadata
Code license
not_specified
License caution
The suite-level card declares CC BY 4.0, while each of the six component dataset cards currently declares MIT. No separate license file or evaluation-code repository is linked. Treat this metadata conflict as unresolved and confirm the intended terms with Arcada Labs before redistribution or commercial reuse.
Download notes
The official public suite comprises six English, multi-turn domains: conference assistance (75 turns), laptop sales (31), grocery ordering (30), dental appointments (25), event planning (29), and personal assistance (31), for 221 scripted turns total. Each domain releases user audio, transcripts, reference responses, knowledge-base context, tool schemas, expected function calls, and scoring labels. The card says two consenting voice actors recorded the inputs. The helper saves the suite and component cards plus API metadata by default; downloading all six snapshots (about 209 MB of current repository storage) requires AUDIO_AGENT_BENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audio_agent_bench_suite.sh
View helper
Enhancement, separation & quality

Audio-Alpaca

Audio-Alpaca: A Preference Dataset for Aligning Text-to-Audio Models

Safe-first helper
Text To Audio Preference Modeling Direct Preference Optimization Text Audio Alignment Temporal Event Alignment +1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
conflicting_apache-2.0_and_cc-by-nc-nd-4.0
Code license
cc-by-nc-nd-4.0
License caution
The Hugging Face dataset card declares Apache-2.0, while the linked official Tango repository applies CC BY-NC-ND 4.0 and does not explain whether that license excludes the dataset. Apply the more restrictive interpretation until the authors clarify scope. AudioCaps-derived captions, generated-audio/model terms, and any source-video rights also require separate review; neither license signal should be assumed to clear all upstream material.
Download notes
The public, ungated release contains 15,025 English prompt, chosen-audio, rejected-audio triplets across four construction strategies. Tango 2 generates candidate audio from AudioCaps training captions, perturbed prompts, and varied inference settings, then filters pairs with two CLAP models. Audio-Zero section 3.1 samples and filters 2,000 pairs for post-training but does not release its exact selection. The helper downloads official documentation and API metadata by default; the Hugging Face API reports approximately 9.71 GB of repository storage, so the audio snapshot requires AUDIO_ALPACA_DOWNLOAD_HF=1.
Safe-first helperscripts/download/audio_alpaca.sh
View helper
Speech recognition

AudioBench

AudioBench: A Universal Benchmark for Audio Large Language Models

Safe-first helper
Audio Language Model Evaluation Automatic Speech Recognition Speech Translation Spoken Question Answering +4 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_upstream_terms
Code license
custom_noncommercial_unspecified_version
License caution
The paper introduces a suite of 8 tasks and 26 datasets, including 7 newly adapted or collected sets, while the maintained repository now supports more than 50 dataset configurations. The repository license file only says "Creative Commons NonCommercial" without a version, and its README states that each dataset remains under its respective license. Review every selected corpus and derived evaluation set before redistribution or commercial use.
Download notes
The helper downloads the official README, supported-dataset inventory, repository metadata, and license notice by default. Cloning the evaluation toolkit is opt-in because the repository is about 64 MB before Git history. It does not download the many upstream audio corpora, whose separate access paths and terms still apply.
Safe-first helperscripts/download/audiobench.sh
View helper
Enhancement, separation & quality

Audiobook Narration Appeal

Audio-Based Understanding of Audiobook Narration Appeal

Safe-first helper
Audiobook Appeal Prediction Narration Quality Analysis Paralinguistic Feature Analysis Genre Conditioned Audio Modeling +1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0
Code license
Apache-2.0
License caution
The repository places its released CSV, supplementary material, and code under Apache-2.0. The CSV derives metadata and engagement counts from LibriVox and the Internet Archive; those services' terms and attribution requirements remain applicable. The license does not extend to separately hosted audiobook recordings or underlying texts. Verify per-item public-domain status in the intended jurisdiction and preserve source attribution. Proprietary Spotify engagement data is described only in aggregate and is not released.
Download notes
The public, ungated Spotify Research release contains one metadata row for each of 8,854 single-narrator English LibriVox audiobooks, covering 1,206 narrators and 65 genres. Fields include LibriVox and Internet Archive URLs, title grouping, narrator identifier, duration, genre, views, favorites, reviews, and days since publication. The paper uses time-normalized Internet Archive view rate as a noisy public proxy for narration appeal, evaluates global and genre-specific prediction, and ranks alternative narrations of the same text. The helper downloads the approximately 3.1 MB CSV, official documentation, license, paper, and supplementary material. Audiobook audio is not redistributed; users follow the released source URLs to public-domain LibriVox recordings. The paper's separate Spotify engagement analysis is proprietary and is not part of the public dataset.
Safe-first helperscripts/download/audiobook_narration_appeal.sh
View helper
Audio understanding, generation & events

AudioCaps

AudioCaps: Generating Captions for Audios in The Wild

Safe-first helper
Audio Captioning Audio Language Modeling
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
academic_only
Code license
MIT
License caution
GitHub README says the code and dataset are free to use for academic purposes only. The repository has an MIT license, but the README adds the academic-use condition for repository material; re-check before redistribution or commercial use.
Download notes
Official CSVs contain captions, YouTube ids, and segment start times. Raw audio/video download requires the upstream form and is subject to AudioSet/YouTube availability and terms.
Safe-first helperscripts/download/audiocaps.sh
View helper
Audio understanding, generation & events

AudioCards / ASFx Eval

AudioCards: Structured Metadata Improves Audio Language Models for Sound Design

Safe-first helper
Structured Audio Captioning Structured Metadata Generation Text Audio Retrieval Sound Effect Classification +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_annotations_with_separate_Adobe_audio_terms
Code license
not_specified
License caution
Zenodo declares CC BY 4.0 for the released AudioCard CSV. The underlying Adobe sound effects are not included there: Adobe describes them as royalty-free but states that downloading and using them is governed by the Adobe Audition and related-software EULA. The July 2026 paper is CC BY-NC-SA 4.0, but that paper license does not release or license its absent four-field augmentation or perturbation artifacts.
Download notes
The public, ungated Zenodo release contains 499 CSV rows with filenames and 13 structured semantic, caption, and UCS fields. The original paper and project page describe 500 manually screened AudioCards, but the released CSV currently has 499 data rows. Pair its filename column with the separately downloaded Adobe Audition Sound Effects library to reproduce ASFx Eval. The July 2026 evaluation-framework paper uses 499 clips. Its Table 1 selects ten released semantic fields and adds five computed acoustic targets (LUFS, pitch, onset, offset, and a frequency profile), but says that augmented dataset "will" be released; those five acoustic annotations are not part of the current Zenodo file. The helper safely downloads the record metadata, project page, papers, Adobe landing page, and approximately 323 KB annotation CSV; it does not download the Adobe audio archives.
Safe-first helperscripts/download/audiocards.sh
View helper
Audio understanding, generation & events

AudioGrounding

Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events

Safe-first helper
Text To Audio Grounding Temporal Audio Grounding Sound Event Localization
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_on_Zenodo_with_upstream_terms
Code license
MIT
License caution
Zenodo declares CC BY 4.0 for the release and the official repository is MIT. The recordings derive from AudioCaps and AudioSet YouTube clips, so retain record attribution and review source-video rights, availability, and platform terms before redistribution or commercial use.
Download notes
The public, ungated v2 release contains 3,994 training, 488 validation, and 492 test clips derived from AudioCaps/AudioSet, with caption phrases aligned to one or more onset-offset intervals. The July 2026 GigaChat Audio report evaluates the combined 980 validation/test samples as a short-clip temporal-grounding benchmark and reports mIoU. The helper downloads the official repository documentation, Zenodo metadata, and approximately 5.2 MB of JSON annotations by default; the approximately 2.33 GiB audio archive requires explicit opt-in.
Safe-first helperscripts/download/audiogrounding.sh
View helper
Speaker, identity & emotion

AudioMarkBench

AudioMarkBench: Benchmarking Robustness of Audio Watermarking

Safe-first helper
Audio Watermarking Robustness Watermark Removal Detection Watermark Forgery Detection Adversarial Audio Robustness +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_upstream_terms_not_separately_specified
Code license
MPL-2.0
License caution
The repository's MPL-2.0 license covers its source code, but neither the README nor paper states a separate license for the released original, watermarked, or perturbed audio. AudioMarkData derives from Common Voice and the second corpus derives from CC BY 4.0 LibriSpeech; applicable Common Voice release terms, attribution requirements, speaker/privacy considerations, and rights in generated derivatives must be reviewed before reuse or redistribution.
Download notes
The public release evaluates AudioSeal/AudioSeal-B, Timbre, and WavMark against 12 no-box perturbation categories plus black-box and white-box adversarial attacks. AudioMarkData contains 20,000 five-second, 16 kHz Common Voice samples balanced for 25 languages, two reported biological-sex groups, and four age groups; the paper also samples 20,000 clips from LibriSpeech. The Drive folder releases original, watermarked, and perturbed audio, while the GitHub repository releases attack and evaluation code. The helper downloads official documentation, license text, repository metadata, and the paper by default; cloning the approximately 3.1 MB GitHub repository requires AUDIOMARKBENCH_CLONE_REPO=1. Google Drive audio remains a manual download so users can inspect its contents and upstream terms.
Safe-first helperscripts/download/audiomarkbench.sh
View helper
Speech understanding & dialogue

AudioMNIST

AudioMNIST: Exploring Explainable Artificial Intelligence for audio analysis on a simple benchmark

Safe-first helper
Spoken Digit Classification Audio Classification Speaker Metadata Analysis Explainable Audio Ai
Access pathOfficial / other
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT
Code license
MIT
License caution
The repository-level LICENSE is MIT and GitHub API reports MIT. Confirm whether downstream use of recorded voices raises consent/privacy obligations beyond the code/data license.
Download notes
The repository contains about 30,000 spoken-digit WAV files from 60 speakers plus speaker metadata and Caffe examples. Because the GitHub repository is large, the helper downloads README/LICENSE metadata by default and requires AUDIO_MNIST_DOWNLOAD_REPO=1 before cloning the full repository.
Safe-first helperscripts/download/audio_mnist.sh
View helper
Audio understanding, generation & events

AudioSet

Audio Set: An ontology and human-labeled dataset for audio events

Safe-first helper
Audio Event Classification Audio Tagging Sound Event Detection
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
AudioSet dataset annotations/features are CC BY 4.0; the ontology is CC BY-SA 4.0. Original YouTube media remains subject to upstream availability and terms.
Download notes
Official release provides segment CSVs and precomputed 128-dimensional audio features; it does not redistribute original YouTube audio.
Safe-first helperscripts/download/audioset.sh
View helper
Audio understanding, generation & events

AudioSetCaps

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Safe-first helper
Audio Captioning Audio Text Retrieval Audio Language Pretraining Synthetic Audio Question Answering
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
academic_research_only
Code license
not_specified
License caution
The Hugging Face card metadata labels the release CC BY 4.0, but the same official card explicitly allows only academic and research use. Apply the stricter research-only statement pending clarification. The GitHub repository has no detected license, and AudioSet, YouTube-8M, and VGGSound source-media rights and platform terms still apply.
Download notes
The public, ungated release provides synthetic captions for 6,117,099 ten-second clips sourced from AudioSet, YouTube-8M, and VGGSound, plus 18,414,789 intermediate question-answer pairs. The Hugging Face repository contains caption/Q&A CSVs rather than the full source audio and currently reports about 20.2 GB of storage. The helper downloads only official documentation and repository metadata by default; the large CSV files require AUDIOSETCAPS_DOWNLOAD_METADATA=1. The maintainers warn that AudioCaps and VGGSound evaluation examples overlap the release and should be filtered before training.
Safe-first helperscripts/download/audiosetcaps.sh
View helper
Audiovisual & cross-modal

AV-SpeakerBench

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

Safe-first helper
Audio Visual Question Answering Speaker Centric Audio Visual Reasoning Audio Visual Temporal Grounding Speech Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The official project page, GitHub README, and Hugging Face card prose state CC BY-NC 4.0, while the Hugging Face card front matter incorrectly or inconsistently declares MIT. Use the more restrictive CC BY-NC 4.0 terms and re-check upstream source-video rights before redistribution or commercial use. GitHub reports no detected repository license.
Download notes
The public, ungated release contains 3,212 English multiple-choice questions plus aligned audio-only, visual-only, and audiovisual clips. The helper downloads official documentation by default; the Hugging Face API reports about 123 GB of repository storage, so the full snapshot requires AV_SPEAKERBENCH_DOWNLOAD_HF=1. Qwen3.5-Omni reports the benchmark in section 5.1.4, Table 7.
Safe-first helperscripts/download/av_speakerbench.sh
View helper
Audiovisual & cross-modal

AVA Active Speaker

AVA Active Speaker: An Audio-Visual Dataset for Active Speaker Detection

Safe-first helper
Active Speaker Detection Audio Visual Speaker Localization Speaking Activity Detection Face Track Classification
Access pathOfficial / other
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC BY 4.0
Code license
not_specified
License caution
Google's AVA download page states that all datasets listed there are CC BY 4.0. The GitHub repository has no detected license, so that statement should not be assumed to license repository code. Preserve attribution and review source-movie and hosting terms before redistributing media; video availability can change independently of the released annotations.
Download notes
The official v1.0 release associates visible face tracks with SPEAKING_AND_AUDIBLE, SPEAKING_BUT_NOT_AUDIBLE, or NOT_SPEAKING labels. Google reports 3.65 million labeled frames across approximately 39,000 face tracks; the current download page says dense labels cover 160 AVA movie clips that remained available on YouTube. The helper downloads official documentation and the 2.5 KB video-name manifest by default. Set AVA_ACTIVE_SPEAKER_DOWNLOAD_LABELS=1 for the approximately 23 MB train/validation annotation archives. It does not fetch the much larger source videos.
Safe-first helperscripts/download/ava_active_speaker.sh
View helper
Audiovisual & cross-modal

AVDC

AVDC: Audio-Visual Decoupled Captions

Safe-first helper
Audio Visual Decoupled Captioning Audio Only Captioning Visual Only Captioning Joint Audio Visual Captioning +3 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The Hugging Face card has no license field, and neither the dataset repository nor the official GitHub repository exposes a license file. The arXiv publication license does not license the released annotations, generated captions or reasoning traces, code, or source videos. ShareGPT4Video, Vript, and each source platform or media owner retain applicable terms; obtain clarification before redistribution or commercial use.
Download notes
The public, ungated Hugging Face release provides avdc_caption.json with decoupled audio, visual, and joint captions and omni_qa.json with questions, answers, reasoning steps, and split labels. The paper describes 10,000 long-form videos, 10,000 caption-derived QA pairs, and an AVDC-test split for visible versus invisible sound-event evaluation. The current JSON blobs total approximately 134 MiB, so the helper downloads only official documentation and API metadata by default and requires AVDC_DOWNLOAD_HF=1 for the annotation snapshot. Source videos are referenced by video ID but are not redistributed in the Hugging Face release; users must obtain applicable ShareGPT4Video- and Vript-sourced media separately. The training/evaluation repository is an additional AVDC_CLONE_REPO=1 opt-in.
Safe-first helperscripts/download/avdc.sh
View helper
Audiovisual & cross-modal

AVE

Audio-Visual Event Localization in Unconstrained Videos

Safe-first helper
Audio Visual Event Localization Audio Visual Event Classification Cross Modal Localization
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The official repository does not expose a detected GitHub license and the README does not state a standalone dataset license. AVE is built from unconstrained videos, so source-video copyright, platform terms, and redistribution rights should be checked before use.
Download notes
The helper downloads the project page and official README by default and can clone the code repository. The dataset, precomputed audio features, and visual features are linked from Google Drive; download those manually or with a user-selected Drive tool after reviewing terms.
Safe-first helperscripts/download/ave.sh
View helper
Audiovisual & cross-modal

AVE-Compass

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Safe-first helper
Instruction Based Audio Video Editing Joint Audio Visual Editing Speech Editing Audio Only Editing +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The Hugging Face card declares CC BY-NC 4.0 for the benchmark release. The evaluation repository has no LICENSE file or GitHub-detected license. Source-video filenames identify a mixture that includes web-video-derived clips, so users should retain attribution and review uploader, platform, privacy, and underlying media rights before redistribution or derivative use.
Download notes
The public, ungated release contains 145 curated source videos, 196 human-verified English editing instructions, 196 checklist JSON files with 2,688 fine-grained items, and 28 editing operation types across joint, speech, video-only, and audio-only branches. The helper downloads official documentation and repository metadata by default. Set AVE_COMPASS_DOWNLOAD_METADATA=1 for the lightweight instruction, checklist, and Dataset Viewer metadata files, or AVE_COMPASS_DOWNLOAD_HF=1 for the complete approximately 442 MB Hugging Face snapshot. The official project currently provides a citation but no public paper URL.
Safe-first helperscripts/download/ave_compass.sh
View helper
Audiovisual & cross-modal

AVQA

AVQA: A Dataset for Audio-Visual Question Answering on Videos

Manual or gated
Audio Visual Question Answering Multimodal Scene Understanding Cross Modal Reasoning
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
noncommercial_or_permission_required
Code license
not_specified
License caution
The official page permits personal or classroom copying without fee only when it is not for profit or commercial advantage, and requires permission for broader copying, reposting, or redistribution. The repository has no license file or GitHub-detected license. Raw clips derive from VGGSound/YouTube, so source-media rights and availability also apply.
Download notes
The helper saves the official project page, repository README, and GitHub metadata, then prints the official OneDrive/Baidu download paths. It does not automate the combined archive or source-video retrieval. The public release provides QA annotations, a VGGSound-derived video manifest, raw-video access pointers, and large pre-extracted features.
Safe-first helperscripts/download/avqa.sh
View helper
Audiovisual & cross-modal

AVSBench

AVSBench: Audio-Visual Segmentation Benchmark

Manual or gated
Audio Visual Segmentation Sounding Object Segmentation Single Sound Source Segmentation Multiple Sound Source Segmentation +1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
Apache-2.0
License caution
The official project page licenses the AVSBench dataset published there under CC BY-NC 4.0, and the repository licenses the project under Apache-2.0. Source videos were collected from public YouTube material, so uploader rights, platform terms, and current video availability still apply; confirm whether the updated semantic release carries any additional terms during application.
Download notes
The benchmark covers semi-supervised Single Sound Source Segmentation (S4), fully supervised Multiple Sound Source Segmentation (MS3), and the later semantic-label AVSS task. The official project page publicly links the original AVSBench-object video-ID CSV and segmentation maps on Google Drive, while processed video/audio requires an email request. The repository directs users to the official application page for the updated object and semantic datasets. The helper saves lightweight official documentation, repository metadata, and license text, then prints these manual access paths; it does not automate Drive, email, or source-video retrieval.
Safe-first helperscripts/download/avsbench.sh
View helper
Audiovisual & cross-modal

AVSCapBench

AVSCapBench: Fine-Grained Audio-Visual Synergy Evaluation for Omni-Modal Video Captioning

Safe-first helper
Omni Modal Video Captioning Audio Visual Captioning Audio Event Captioning Audio Visual Event Binding +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
not_specified
License caution
The Hugging Face card and repository README declare CC BY-NC-SA 4.0, while the paper describes an academic-research-only restrictive release. The paper says clips come from YouTube, TikTok, and Video-MME, invokes fair use, and says only public URLs and timestamps are distributed, but the current Hugging Face repository lists 1,226 MP4 files. Treat the stricter academic/non-commercial interpretation as controlling and review source-platform, uploader, Video-MME, privacy, and copyright terms before downloading, redistributing, or publishing clips. The evaluation repository has no detected license.
Download notes
The public, ungated release contains 1,226 English video clips lasting 30 to 120 seconds, dense omni-modal captions, visual events, audio events separated into speech, music, and sound effects, and synergistic audio-visual events. The evaluation repository implements LLM-judged event recall. The helper downloads official documentation and repository metadata by default; set AVSCAPBENCH_DOWNLOAD_HF=1 for the approximately 19.6 GB Hugging Face snapshot.
Safe-first helperscripts/download/avscapbench.sh
View helper
Audiovisual & cross-modal

AVSD

Audio Visual Scene-Aware Dialog Dataset

Manual or gated
Audio Visual Dialogue Video Question Answering Multimodal Response Generation Scene Understanding
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
unclear
Code license
MIT
License caution
The official repository is MIT-licensed, but it does not separately state that the MIT license covers the Google Drive dataset or underlying Charades videos. Treat dialog annotations and media rights conservatively, review any terms shown by the Drive/Charades access paths, and preserve source-video provenance before redistribution or commercial use.
Download notes
The CVPR paper introduces dialogs and final summaries for more than 11,000 Charades videos. The official DSTC7 repository reports 7,659 training, 1,787 validation, and 1,710 test dialogs and links the released challenge data through Google Drive. The helper saves the public repository documentation, license, and CVPR paper page, then prints the manual dataset and Charades media paths; it does not automate Google Drive or raw-video retrieval.
Safe-first helperscripts/download/avsd.sh
View helper
Audiovisual & cross-modal

AVUT

Audio-centric Video Understanding Benchmark without Text Shortcut

Safe-first helper
Audio Visual Question Answering Audio Content Understanding Audio Visual Alignment Audio Event Localization +1 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
Neither the official repository nor Hugging Face card states a data or code license. The paper says AVUT contains only links to public YouTube videos and does not host or distribute video copies, while the current Hugging Face repository appears to include video files; users should review YouTube terms, source-video rights, and this discrepancy before downloading or redistributing media.
Download notes
The public, ungated release covers 2,662 English YouTube videos across 18 domains and 11,609 question-answer pairs in AV-Human and AV-Gemini. The helper downloads official documentation and four lightweight annotation JSON files by default; the Hugging Face API reports about 24.0 GB of repository storage, so the full snapshot requires AVUT_DOWNLOAD_HF=1. Qwen3.5-Omni reports AVUT in section 5.1.4, Table 7.
Safe-first helperscripts/download/avut.sh
View helper
Audiovisual & cross-modal

BAH

BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change

Manual or gated
Ambivalence Hesitancy Recognition Multimodal Affect Recognition Vocal Expression Analysis Audio Visual Behavior Understanding +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
proprietary_research_only_eula
Code license
BSD-3-Clause
License caution
The ÉTS owner page expressly labels BAH proprietary and research-only. The current request process is limited to full-time faculty at an eligible university, higher-education institution, or equivalent organization; students and postdoctoral researchers cannot apply directly. The public repository's BSD-3-Clause license applies to code, not the gated recordings, transcripts, annotations, faces, or participant metadata. Review the signed EULA for storage, sharing, publication, retention, privacy, and downstream-use obligations.
Download notes
The release contains 1,427 videos totaling 10.60 hours from 300 participants across Canada, including 1.8 hours of annotated ambivalence/hesitancy moments. It provides raw videos, 16 kHz audio, timestamped transcripts, cropped and aligned faces, expert video- and frame-level labels and cues, participant metadata, and predefined participant-disjoint splits. Access is manual: an eligible full-time faculty member must submit the official form, list every team member, certify institutional eligibility, and sign the EULA. The helper saves only public owner, paper, repository, challenge, and request documentation and never downloads participant data.
Safe-first helperscripts/download/bah.sh
View helper
Audio understanding, generation & events

Big Bench Audio

Artificial Analysis Big Bench Audio

Safe-first helper
Spoken Question Answering Speech Reasoning Audio Question Answering
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mit
Code license
Apache-2.0
License caution
The Hugging Face card declares MIT and the Xiaomi MiMo evaluator is Apache-2.0. The 1,000 English recordings contain verbatim questions from four BIG-Bench Hard tasks and were synthesized with 23 OpenAI-, Azure-, and AWS-provided voices; review inherited task terms and provider-generated-audio conditions rather than assuming the card resolves every component right.
Safe-first helperscripts/download/big_bench_audio.sh
View helper
Enhancement, separation & quality

BVCC

BVCC: VoiceMOS Challenge 2022 Main-Track Dataset

Safe-first helper
Mean Opinion Score Prediction Synthetic Speech Naturalness Assessment Voice Conversion Quality Assessment Text To Speech Quality Assessment +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other_open_mixed_upstream_terms
Code license
not_separately_specified
License caution
Zenodo labels the release Other (Open), not a standard reusable data license. Its record explicitly prohibits redistribution of Blizzard Challenge samples and omits those files; Voice Conversion Challenge and ESPnet-TTS components retain their own terms. Treat ratings, metadata, audio, and scripts according to their component provenance, and do not infer broad redistribution or commercial rights from public access.
Download notes
The public VoiceMOS Challenge 2022 release provides unified MOS ratings and official train, development, and test splits for synthetic speech from past Voice Conversion and Blizzard Challenges plus ESPnet-TTS. The helper downloads official pages and Zenodo metadata by default. The approximately 273.4 MiB main-track archive, small out-of-domain package, and scoring package are separate opt-ins. Blizzard audio is intentionally absent from the archive; official scripts require users to obtain and preprocess that material under its original access terms. Section 4.2 of arXiv:2607.13477 constructs 60 BVCC test-set pairs with a human-MOS gap of at least 1.0 for its naturalness probe, but does not release the selected pair manifest.
Safe-first helperscripts/download/bvcc.sh
View helper
Speech generation

CapSpeech

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

Safe-first helper
Style Captioned Text To Speech Text To Speech With Sound Effects Accent Captioned Text To Speech Emotion Captioned Text To Speech +2 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0_with_mixed_upstream_audio_terms
Code license
cc-by-nc-4.0
License caution
The dataset card, paper, and repository apply CC BY-NC 4.0 to CapSpeech resources. The release points to audio from Emilia, GigaSpeech, Common Voice, MLS, LibriTTS-R, VoxCeleb, EARS, Expresso, VCTK, VGGSound, FSDKaggle2018, ESC-50, and separate CapSpeech audio repositories; those recordings retain their own attribution, non-commercial, access, privacy, and media-rights constraints. Treat the CapSpeech license as covering its annotations and author contributions, not as overriding component-source terms.
Download notes
The public, ungated release contains more than 10 million machine-annotated and approximately 360,000 human-annotated English audio-caption records, with fixed pretraining and supervised-fine-tuning train, validation, and test splits across CapTTS, CapTTS-SE, AccCapTTS, EmoCapTTS, and AgentTTS. The main Hugging Face snapshot contains paths, transcripts, source labels, durations, and style captions rather than embedded source audio; the API reports approximately 4.31 GB compressed and 10.09 GB after processing. The helper downloads official documentation and API metadata by default, while the full metadata snapshot and code repository are separate opt-ins. The 2026 ProPS paper trains and evaluates prompt-conditioned speaker-profile distributions on CapSpeech's held-out splits.
Safe-first helperscripts/download/capspeech.sh
View helper
Audio understanding, generation & events

CASTELLA

CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

Safe-first helper
Audio Moment Retrieval Temporal Audio Grounding Long Audio Retrieval Audio Captioning
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_for_annotations_and_released_features
Code license
not_specified_for_audio_downloader
License caution
The annotation repository, its Hugging Face card, and the Zenodo feature record declare CC BY 4.0. That license covers the released annotations/features, not the underlying YouTube recordings. The separate raw-audio downloader repository has no detected license, and raw media remains governed by its owners and YouTube terms; verify availability and rights before downloading, redistribution, or commercial use.
Download notes
The public, ungated annotation release describes 1,862 real-world YouTube recordings split into 1,009 training, 213 validation, and 640 test items, with 3,925 human-written local captions and 11,308 temporal boundaries in English and Japanese. The Hugging Face mirror contains only the six lightweight annotation JSON files, not raw audio. The helper downloads those annotations plus official documentation and metadata by default. Precomputed MS-CLAP audio/text features are an approximately 2.78 GB Hugging Face opt-in (the Zenodo release is about 1.33 GB). Raw media must be reconstructed from YouTube IDs with the separate official downloader, subject to current availability and source-platform and recording rights; the helper only clones those tools when explicitly requested.
Safe-first helperscripts/download/castella.sh
View helper
Audiovisual & cross-modal

CH-SIMS

CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of Modality

Safe-first helper
Multimodal Sentiment Analysis Audio Sentiment Analysis Visual Sentiment Analysis Text Sentiment Analysis +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
MIT
License caution
The MMSA repository is MIT-licensed, but neither the paper nor the repository README expressly applies that license to the CH-SIMS video and annotation files in the shared Drive folders. Treat dataset terms and source-media rights as unspecified and verify permitted use before redistribution or commercial use.
Download notes
The paper introduces 2,281 Chinese in-the-wild video segments with multimodal sentiment labels and separate text, audio, and visual annotations. The official MMSA repository provides shared Baidu and Google Drive folders containing raw video, processed features, and labels for CH-SIMS alongside MOSI and MOSEI. The helper downloads official paper, README, license, and repository metadata only; dataset files remain a manual Drive download, and cloning the toolkit is opt-in.
Safe-first helperscripts/download/ch_sims.sh
View helper
Audiovisual & cross-modal

CH-SIMS v2

CH-SIMS v2.0: A Fine-grained Multi-label Chinese Multimodal Sentiment Analysis Dataset

Safe-first helper
Multimodal Sentiment Analysis Audio Sentiment Analysis Visual Sentiment Analysis Text Sentiment Analysis +2 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The official project page, paper, and repository state no dataset or code license, and the repository has no LICENSE file. Treat the videos, annotations, features, and code as all-rights-reserved unless the authors provide terms; also review source-media, speaker, privacy, and platform rights before reuse or redistribution.
Download notes
The official release extends and re-annotates CH-SIMS with 4,402 supervised segments carrying multimodal and unimodal sentiment labels plus 10,161 unlabeled segments for semi-supervised evaluation. It provides raw videos, extracted features, IDs, splits, and labels through separate Google Drive and Baidu folders. The helper downloads only the official homepage, repository README/API metadata, and arXiv metadata; Drive data remains manual and the code clone is opt-in.
Safe-first helperscripts/download/ch_sims_v2.sh
View helper
Music

ChartGenEval

ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

Safe-first helper
Music Generation Evaluation Rhythm Game Chart Generation Chart Audio Alignment Chart Structure Evaluation +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT_released_artifacts_corpus_not_released
Code license
Not specified in the source record.
License caution
The repository-level MIT license covers the released software and bundled artifacts. It does not grant rights to the absent chart/audio corpus; the README says potentially copyrighted community and commercial content is intentionally not redistributed. Users reproducing corpus-dependent experiments must supply a rights-compatible equivalent snapshot and review its chart, recording, and game-content terms.
Download notes
The public repository releases the NumPy-based evaluation toolkit, calibration artifact, corruption probes, plotting and reproduction scripts, and metric-only evaluation records. The paper validates seven output axes with nine held-out controlled-corruption tests across timing, audio response, structure, grammar, and human-reference gap measurements. The helper downloads official documentation and repository metadata by default; cloning the approximately 44 MB current file tree requires CHARTGENEVAL_CLONE_REPO=1. The underlying 3,880-chart calibration corpus and its audio are explicitly not distributed because they may contain copyrighted community and commercial material, so the released artifacts support verification of reported results but not full corpus-dependent reproduction.
Safe-first helperscripts/download/chartgeneval.sh
View helper
Speech recognition

CHILDES-Aligned

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

Manual or gated
Asr Child Speech Recognition Long Form Speech Alignment Forced Alignment +1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0_with_talkbank_terms
Code license
not_specified
License caution
The Hugging Face card declares CC BY-NC-SA 4.0 and its access agreement additionally limits use to non-commercial research, requires citation of the BEACON paper and every source CHILDES corpus used, incorporates the TalkBank Ground Rules, and prohibits audio redistribution. Access is manually reviewed. The paper's linked BEACON GitHub repository was not publicly reachable when checked, so no code license is claimed.
Download notes
The manually gated Hugging Face release contains a 413.3-hour general-purpose English child-speech configuration with corrected utterance timestamps and a quality-controlled 283-hour ASR-training configuration. The repository reports approximately 160.6 GB of storage. The helper prints the access steps by default and downloads a selected configuration only after the user has received access, authenticated with Hugging Face, and set CHILDES_ALIGNED_ACK_TERMS=1.
Safe-first helperscripts/download/childes_aligned.sh
View helper
Speech recognition

CHiME-6

CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings

Safe-first helper
Distant Speech Recognition Multi Speaker Asr Speaker Diarization Speech Separation
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
not_applicable
License caution
OpenSLR SLR150 lists CC BY-SA 4.0. CHiME says CHiME-6 is a corrected-alignment version of CHiME-5 and recommends CHiME-6 for new work.
Download notes
The helper downloads OpenSLR transcriptions, floorplans, and license by default. Audio archives are large, about 97 GiB train, 11 GiB dev, and 12 GiB eval, so they are explicit opt-ins.
Safe-first helperscripts/download/chime_6.sh
View helper
Speech recognition

CHiME-7 DASR

The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios

Manual or gated
Distant Automatic Speech Recognition Speaker Attributed Automatic Speech Recognition Speaker Diarization Meeting Transcription +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
mixed_manual_agreements
Code license
Apache-2.0
License caution
The benchmark is a protocol over three separately controlled corpora, not a single uniformly licensed download. Retain the CHiME, DiPCo, and task-specific LDC/Mixer 6 terms independently. The official baseline is part of ESPnet, whose repository is Apache-2.0 licensed.
Download notes
CHiME-7 DASR evaluates one system across revised CHiME-6, DiPCo, and challenge-specific Mixer 6 Speech partitions, ranking submissions by macro-averaged diarization-attributed WER across the three scenarios. The official ESPnet recipe generates the task layout, but it can automatically obtain only DiPCo. CHiME-5/CHiME-6 must be obtained through the CHiME license path, and the task's Mixer 6 release requires a separate LDC evaluation agreement; the challenge warns that this Mixer 6 version differs from LDC2013S03. The helper saves only public task, data, paper, and baseline documentation before printing the manual access steps.
Safe-first helperscripts/download/chime_7_dasr.sh
View helper
Audio understanding, generation & events

Clotho

Clotho: An Audio Captioning Dataset

Safe-first helper
Audio Captioning Language Based Audio Retrieval
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
not_specified
License caution
Zenodo lists rights as Other (Attribution). Audio clips keep their original Freesound licenses, mostly Creative Commons with attribution, recorded in metadata CSVs. Captions are under the Tampere University license, mainly non-commercial with attribution.
Download notes
Clotho v2.1 audio archives total about 7.1 GiB; the helper downloads captions/metadata by default and makes audio opt-in.
Safe-first helperscripts/download/clotho.sh
View helper
Audio understanding, generation & events

Clotho-Moment

Clotho-Moment: Simulated Long-Audio Dataset for Language-Based Audio Moment Retrieval

Safe-first helper
Language Based Audio Moment Retrieval Temporal Audio Grounding Long Audio Retrieval Audio Text Retrieval
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0_on_hugging_face_card_with_upstream_terms
Code license
Apache-2.0
License caution
The Hugging Face card declares Apache-2.0 and the Lighthouse repository includes an Apache-2.0 license. The generated audio incorporates Clotho/Freesound foreground clips and Walking Tours/YouTube background audio, so the card does not erase component recording licenses, attribution requirements, or source-platform terms; review packaged provenance before redistribution or commercial use.
Download notes
The public, ungated release contains 51,240 one-minute synthetic English recordings split into 37,930 training, 5,741 validation, and 7,569 test samples. Each sample pairs a text query with a temporal boundary for a Clotho foreground event overlaid at a random interval on Walking Tours background audio. DCASE 2026 Task 6 uses it as a development dataset and advertises a 16.1 GB download. The helper downloads official documentation, license, and repository metadata by default; the audio WebDataset snapshot requires explicit opt-in. Hugging Face currently reports approximately 213 GB of repository storage including history, so users should verify available disk space and select only needed splits or shards.
Safe-first helperscripts/download/clotho_moment.sh
View helper
Audio understanding, generation & events

ClothoAQA

Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering

Safe-first helper
Audio Question Answering Audio Language Understanding Machine Listening
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
not_applicable
License caution
Zenodo lists rights as Other (Attribution). The train/validation/test question-answer CSVs are MIT licensed by Tampere University. Audio files keep per-file Freesound licenses, mostly Creative Commons with attribution, recorded in clotho_aqa_metadata.csv.
Download notes
The helper downloads the QA split CSVs, metadata, and license by default. The 3.1 GiB audio_files.zip archive is opt-in because it contains Clotho/Freesound-derived audio.
Safe-first helperscripts/download/clotho_aqa.sh
View helper
Enhancement, separation & quality

CMI-RewardBench / CMI-Pref

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

Safe-first helper
Music Reward Model Evaluation Music Preference Prediction Music Quality Assessment Text Music Alignment +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-SA-4.0
Code license
Apache-2.0
License caution
The CMI-Pref card and paper declare CC BY-NC-SA 4.0, and the evaluation repository contains Apache-2.0 code. The paper says some audio was generated through commercial APIs and describes a terms-aware release mechanism; generated-output and service terms may still apply. The composite CMI-RewardBench also incorporates PAM, MusicEval, and Music Arena, so their source licenses and music rights remain controlling for those subsets.
Download notes
The public, ungated CMI-Pref release contains 4,027 individual human preference votes over generated music, including a balanced 500-vote test split, 133.8 hours of English/Chinese material, and text, lyrics, reference-audio, musicality, alignment, confidence, and anonymized listener fields. CMI-RewardBench combines that test split with PAM, MusicEval, and Music Arena for music reward-model evaluation. The helper downloads the official cards, repository docs/configuration, the approximately 620 KB CMI-Pref test JSONL, and the approximately 4.8 MB composite test manifest by default. The Hugging Face API reports approximately 15.0 GB of repository storage, so all MP3 assets require explicit opt-in. A July 2026 full-song generation report uses CMI-Reward as one evaluator on its separate 500-example multilingual test set; that paper does not release those 500 evaluation inputs.
Safe-first helperscripts/download/cmi_rewardbench.sh
View helper
Audiovisual & cross-modal

CMU-MOSEI

CMU-MOSEI: CMU Multimodal Opinion Sentiment and Emotion Intensity

Safe-first helper
Multimodal Sentiment Analysis Multimodal Emotion Recognition Audio Sentiment Analysis Speech Emotion Recognition +2 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
MIT
License caution
The CMU Multimodal SDK and MultiBench repositories are MIT-licensed, but neither repository expressly applies that license to MOSEI annotations, processed features, or source YouTube media. Treat dataset terms as unspecified and review creator, platform, privacy, and media rights before reuse, redistribution, or commercial use.
Download notes
The official SDK describes more than 65 hours of annotated YouTube monologue video from more than 1,000 speakers and 250 topics. Each sentence has a sentiment score and six non-exclusive emotion scores for happiness, sadness, anger, surprise, disgust, and fear. The SDK publishes labels plus processed acoustic, visual, and language computational sequences; MultiBench provides an additional word-aligned processed package through Google Drive. The helper saves official documentation, dataset definitions, and repository metadata only. Cloning either toolkit is opt-in, and the Drive package remains a manual download.
Safe-first helperscripts/download/cmu_mosei.sh
View helper
Audiovisual & cross-modal

CMU-MOSI

CMU-MOSI: Multimodal Opinion-level Sentiment Intensity Dataset

Safe-first helper
Multimodal Sentiment Analysis Audio Sentiment Analysis Subjectivity Analysis Audio Visual Sentiment Analysis +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
MIT
License caution
Both official code repositories are MIT-licensed, but neither license expressly grants rights to the MOSI annotations, processed features, or source media. Raw YouTube videos are not redistributed. Treat dataset terms as unspecified and review creator, platform, privacy, and media rights before reuse or redistribution.
Download notes
The original paper introduces 2,199 opinion segments from 93 English YouTube review videos with sentiment-intensity, subjectivity, visual, and acoustic annotations. The official CMU Multimodal SDK publishes labels and anonymized processed acoustic, visual, and language computational sequences, but explicitly does not share raw videos because of YouTube creator privacy. MultiBench provides an additional word-aligned processed release through Google Drive. The helper saves official documentation and repository metadata only; cloning either toolkit is opt-in and the Drive package remains a manual download.
Safe-first helperscripts/download/cmu_mosi.sh
View helper
Speech generation

CN-NewsTTS Bench

CN-NewsTTS Bench: A Target-Level Automatic Benchmark for Raw-Input Chinese News TTS Pronunciation

Safe-first helper
Chinese Text To Speech Evaluation Pronunciation Accuracy Evaluation Text Normalization Evaluation Target Level Error Analysis +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
MIT
License caution
CC BY 4.0 covers benchmark data, fixed ASR transcripts, results, documentation, and metadata; MIT covers repository code. The Zenodo audio consists of generated outputs from seven commercial TTS providers and is published as an evaluation artifact. The maintainers warn that reuse may remain subject to each provider or API's terms, so do not treat those audio archives as unrestricted speech-training data without a separate rights review.
Download notes
The public v0.1 release contains 200 development records and 800 public-test records with 1,240 auto-evaluable pronunciation targets, fixed transcripts from a three-ASR ensemble, target-level scoring code, and results for seven TTS products. The helper downloads official documentation, licenses, Zenodo metadata, the two small JSONL benchmark splits, schema, scorer, and checksums by default. The approximately 1.58 MB core archive and 1.72 MB full-transcript archive are separate opt-ins. The approximately 425 MB development-audio and 1.74 GB public-test-audio archives require explicit provider-terms acknowledgement and opt-in.
Safe-first helperscripts/download/cn_news_tts_bench.sh
View helper
Speaker, identity & emotion

Codec-SUPERB

Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Safe-first helper
Neural Audio Codec Evaluation Speech Reconstruction Audio Reconstruction Music Reconstruction +3 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The Hugging Face card has no license field or provenance/rights statement. The repository README and badge call the project MIT, but the linked LICENSE file is absent and the GitHub API detects no license. Treat both data and code terms as unspecified until the maintainers publish authoritative license text, and verify the source field and upstream audio rights before reuse.
Download notes
The benchmark evaluates whether neural codecs preserve content, paralinguistics, speaker identity, and general audio information through downstream and signal-level metrics. The current official repository uses the public, ungated codec-superb-tiny release for regression runs: 6,000 rows split evenly across speech, audio, and music, with approximately 3.2 GB of downloads. The helper downloads official documentation by default; the dataset snapshot and repository clone are separate opt-ins.
Safe-first helperscripts/download/codec_superb.sh
View helper
Speech understanding & dialogue

CoDeTT

CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation

Safe-first helper
Turn Taking Full Duplex Dialogue Context Aware Turn Decision Spoken Dialogue Intent Classification
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0
Code license
not_specified
License caution
The Hugging Face card declares Apache-2.0 for the dataset, but the GitHub repository has no LICENSE file or detected license. The paper says real samples come from Candor and MagicData-RAMC and synthetic speech uses references from KeSpeech and Emilia; verify those upstream corpus and voice-data terms before redistribution or commercial use.
Download notes
The public, ungated release contains more than 300 hours of English and Chinese multi-turn dialogue for four turn-taking actions and 14 fine-grained intent scenarios across system-speaking and system-idle states. It mixes synthetic material with real conversational samples derived from Candor and MagicData-RAMC. The helper downloads official documentation and API metadata by default. The Hugging Face API reports approximately 51.1 GB of repository storage, so the single CoDeTT.lz4 archive requires CODETT_DOWNLOAD_HF=1.
Safe-first helperscripts/download/codett.sh
View helper
Audiovisual & cross-modal

CoMind

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

Safe-first helper
Audio Visual Social Reasoning Joint Attention Estimation Socially Conditioned Object Interaction Anticipation Collaborative Handover Prediction +2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The official project page declares CC BY-NC 4.0, although its displayed license and terms links are currently placeholder anchors rather than a separate terms document. Treat the data as non-commercial, preserve attribution, and review privacy, voice, face, gaze, biometric, and participant-consent implications before reuse. The paper states that participants consented to release of identifiable video, but that does not remove downstream ethical obligations. No separate license notice is embedded in the official Python downloader; the paper itself uses arXiv's perpetual non-exclusive license.
Download notes
The official public release covers 41 hours of unscripted cooking collaboration across 80 sessions, with two synchronized egocentric cameras, two exocentric views, audio and WhisperX transcripts, gaze, hand tracking, camera trajectories, scene/object scans, and annotations. Its three benchmarks accept audio or transcribed speech from a ten-second context window for joint-attention estimation, socially conditioned object-interaction anticipation, and collaborative handover prediction. The helper saves the official page, paper metadata, first-party downloader, and annotation manifest by default. The approximately 5.0 MiB annotation JSON files are opt-in; large recording components remain available through the saved official downloader and are never fetched automatically.
Safe-first helperscripts/download/comind.sh
View helper
Speech recognition

Common Voice

Mozilla Common Voice

Manual or gated
Asr
Access pathOfficial / other
Upstream termsOpen / attribution signals

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC0-1.0
Code license
MPL-2.0
License caution
Common Voice data is CC0-1.0; cv-dataset metadata repo is MPL-2.0.
Download notes
The helper requires a per-release, per-language URL generated by Mozilla Data Collective and never stores credentials or a private generated URL. Set COMMON_VOICE_FILENAME when the signed URL does not expose a useful archive name.
Safe-first helperscripts/download/common_voice.sh
View helper
Audio understanding, generation & events

Concerto Accompaniment Benchmark

Concerto Accompaniment Benchmark for Score-Free Piano Concerto Accompaniment

Safe-first helper
Music Audio Alignment Automatic Accompaniment Score Free Accompaniment Generation Downbeat Alignment +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Not specified in the source record.
Code license
Not specified in the source record.
License caution
MIT covers the public repository's code and released annotation/config files. It does not license the absent Music Minus One recordings or override the per-recording IMSLP terms listed in AudioDataSummary.csv, which include CC0, several CC variants, and public-domain status that may differ by jurisdiction. The repository does not state a license or public delivery path for the recorded solo-piano performances.
Download notes
The paper defines 150 alignment scenarios over four concerto movements, combining four recorded solo-piano performances, four commercial Music Minus One orchestra tracks, and eight IMSLP piano-orchestra mixes. The public repository provides the evaluation code, configuration tables, IMSLP source URLs, and measure-downbeat annotations. The helper saves those lightweight released artifacts by default and makes the repository clone opt-in. The commercial orchestra recordings are explicitly private and must be purchased separately. Although the paper calls the remaining data open source, the repository currently ignores audio and exposes no solo-piano recordings; do not infer a public audio download.
Safe-first helperscripts/download/concerto_accompaniment_benchmark.sh
View helper
Speech recognition

CoVoST 2

CoVoST 2: Massively Multilingual Speech-to-Text Translation Corpus

Safe-first helper
Speech To Text Translation Speech Translation Multilingual Asr Machine Translation
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Meta/GitHub list CoVoST data as CC0; the current Hugging Face card lists CC BY-NC 4.0, and current Mozilla Data Collective packages for CoVoST 2/Common Voice segments may be CC BY-NC 4.0.
Code license
cc-by-nc-4.0
License caution
Treat packaged mirrors conservatively and re-check the active source before redistribution. The GitHub license table also says Tatoeba evaluation sentences are CC BY 2.0 FR and Tatoeba speech has per-row licenses.
Download notes
The helper downloads the official CoVoST 2 translation TSV archives and split-generation script. CoVoST 2 rows match Common Voice 4 validated.tsv entries, so users must obtain Common Voice audio separately under the applicable upstream terms.
Safe-first helperscripts/download/covost2.sh
View helper
Audiovisual & cross-modal

CREMA-D

CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset

Safe-first helper
Speech Emotion Recognition Audio Visual Emotion Recognition Acted Emotional Speech Crowd Sourced Emotion Annotation
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
odbl-1.0
Code license
not_specified
License caution
The official README/LICENSE say the database is under the Open Database License 1.0 and individual contents are under the Database Contents License 1.0. GitHub reports license as NOASSERTION, so keep the explicit upstream text as authority.
Download notes
The helper downloads small README/license/CSV metadata by default. Full audio and video live in Git LFS and require about 7.55 GiB for a complete clone; upstream asks repository users to fill out the access/community form.
Safe-first helperscripts/download/crema_d.sh
View helper
Speech generation

CV3-Eval

CV3-Eval: CosyVoice 3 in-the-wild zero-shot speech synthesis benchmark

Safe-first helper
Zero Shot Text To Speech Multilingual Voice Cloning Cross Lingual Voice Cloning Emotion Cloning +6 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0 repository license; mixed upstream media rights
Code license
Apache-2.0
License caution
The repository root applies Apache-2.0, but the README says reference speech comes from Common Voice, FLEURS, EmoBox, and web-crawled real-world audio. Treat the repository license as insufficient to clear every source recording, and verify component provenance and rights before redistribution or commercial use.
Download notes
The official repository includes objective multilingual, cross-lingual, and emotion-cloning subsets plus subjective expressive, continuation, and Chinese-accent subsets. Qwen3.5-Omni section 5.2.3 calls the public cross-lingual subset both CV3-Eval and the Cross-Lingual benchmark, and reports mixed error rate over 12 source-target directions among Chinese, English, Japanese, and Korean. The helper downloads the README and Apache-2.0 license by default; cloning the roughly 760 MiB repository, including evaluation audio and bundled scoring utilities/models, requires CV3_EVAL_CLONE_REPO=1.
Safe-first helperscripts/download/cv3_eval.sh
View helper
Audiovisual & cross-modal

Daily-Omni

Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities

Safe-first helper
Audio Visual Question Answering Temporal Alignment Cross Modal Reasoning Audio Visual Event Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-SA-4.0
Code license
GPL-3.0
License caution
The paper and Hugging Face card declare CC BY-NC-SA 4.0 for the benchmark, and GitHub reports GPL-3.0 for the repository. Videos are sampled from AudioSet, Video-MME, and FineVideo, so their upstream media rights and terms also require review.
Download notes
The public, ungated release contains 684 real-world videos and 1,197 English multiple-choice questions across six temporal audio-visual reasoning tasks. The helper downloads official documentation and qa.json by default; the Hugging Face API reports approximately 3.9 GB of storage, so Videos.tar requires DAILY_OMNI_DOWNLOAD_HF=1. Qwen3.5-Omni reports DailyOmni in section 5.1.4, Table 7.
Safe-first helperscripts/download/daily_omni.sh
View helper
Audiovisual & cross-modal

DAVE

DAVE: Diagnostic Benchmark for Audio Visual Evaluation

Safe-first helper
Audio Visual Alignment Multimodal Synchronization Sound Absence Detection Sound Discrimination +3 more
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT label with upstream dataset terms
Code license
MIT stated in README; no repository LICENSE file
License caution
The Hugging Face card labels DAVE as MIT and the repository README says everything is MIT, but the GitHub repository has no LICENSE file. DAVE is built on EPIC-KITCHENS and Ego4D, and its card says it inherits their risks; verify both upstream datasets' access and media terms before redistribution or commercial use.
Download notes
The public, ungated Hugging Face release has EPIC-KITCHENS- and Ego4D-derived splits with seven diagnostic task views. The helper downloads the official cards, loader, and approximately 9 MB of JSON annotations by default. Media archives are excluded because the Hugging Face API reports about 113.3 GB of repository storage; DAVE_DOWNLOAD_HF=1 explicitly opts into the full snapshot.
Safe-first helperscripts/download/dave.sh
View helper
Audio understanding, generation & events

DCASE 2024 Task 5

DCASE 2024 Task 5: Few-shot Bioacoustic Event Detection

Safe-first helper
Few Shot Bioacoustic Event Detection Sound Event Detection Animal Vocalization Detection Five Shot Learning
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
Both official Zenodo records declare Creative Commons Attribution 4.0 International. The benchmark combines multiple bioacoustic sources; retain the release attribution and review source-specific ethical or wildlife-recording constraints for downstream use.
Download notes
The official five-shot protocol provides the first five positive target events in each recording, then scores detection after the fifth event. The 2024 development release has 217 recordings: 174 training files covering 47 classes and 43 validation files covering seven classes. The official challenge reuses the 2023 evaluation release, which has 66 recordings across eight subsets. The helper downloads record metadata, class maps, and annotation-only archives by default. The current Zenodo files total approximately 20.4 GiB for development audio and 3.0 GiB for evaluation audio, so waveform archives require DCASE2024_TASK5_DOWNLOAD_AUDIO=1.
Safe-first helperscripts/download/dcase2024_task5.sh
View helper
Audio understanding, generation & events

DCASE 2024 Task 7 Sound Scene Synthesis

DCASE 2024 Task 7: Sound Scene Synthesis

Safe-first helper
Text To Audio Generation Environmental Sound Scene Synthesis Compositional Audio Generation Audio Generation Quality Evaluation
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
Zenodo declares CC BY 4.0 for the open-source dataset. The record says its audio is sourced from Freesound, so retain packaged source attribution and review per-clip provenance. The baseline repository has no LICENSE file and the task page states that it is mostly derived from the upstream AudioLDM repository; treat code terms as unspecified pending clarification.
Download notes
The public open-source release contains 310 manually composed four-second environmental sound scenes and corresponding structured text prompts. It uses only Freesound source audio and excludes the proprietary/private libraries present in the challenge reference data. The original protocol evaluates Fréchet Audio Distance with PANNs CNN14 Wavegram-Logmel embeddings plus listening tests for foreground fit, background fit, and audio quality. The challenge's 250 evaluation prompts and reference audios remain secret. The helper downloads official metadata and task documentation by default; the approximately 140 MiB public archive is an explicit opt-in.
Safe-first helperscripts/download/dcase2024_sound_scene_synthesis.sh
View helper
Enhancement, separation & quality

DCASE 2024 Task 9 LASS

DCASE 2024 Task 9: Language-Queried Audio Source Separation

Safe-first helper
Language Queried Audio Source Separation Text Conditioned Audio Separation Universal Sound Separation
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
All three official Zenodo records declare CC BY 4.0. Development audio remains in FSD50K and Clotho v2 under their own mixed or upstream terms, while validation/evaluation audio derives from Freesound; retain record attribution and inspect packaged per-clip provenance. The baseline repository has no LICENSE file or detected GitHub license and is largely derived from AudioSep, so its code terms are unspecified.
Download notes
The development release contains GPT-4-generated captions for the existing FSD50K development and evaluation clips; participants obtain the source FSD50K and Clotho v2 audio separately. The public validation release contains 3,000 synthetic mixtures built from 1,000 source clips with three captions per source. The evaluation release contains 3,000 additional synthetic mixtures plus 100 real overlapping Freesound clips annotated with two source queries each. Synthetic examples are scored with SDR; the real set uses listening tests for query relevance and overall quality. The helper downloads official task documentation, record metadata, and lightweight JSON/CSV annotations by default; approximately 1.14 GB of validation and evaluation audio is an explicit opt-in.
Safe-first helperscripts/download/dcase2024_lass.sh
View helper
Audio understanding, generation & events

DCASE 2025 Task 2 ASD

DCASE 2025 Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

Safe-first helper
Anomalous Sound Detection Machine Condition Monitoring Domain Generalization First Shot Learning
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
custom_dcase_challenge_license_v2.1
License caution
All three official Zenodo records declare CC BY-NC-SA 4.0, which prohibits commercial use and requires attribution and share-alike distribution. The evaluator repository includes a DCASE Challenge License v2.1 PDF rather than an SPDX-style open-source license; review it before reuse.
Download notes
The public first-shot protocol trains only on normal machine sounds, tests source/target-domain generalization, and uses different machine types for development versus final evaluation. The development record has seven machine types and approximately 2.36 GB of archives; the approximately 1.98 GB additional-training record and 358 MB evaluation record cover eight different types. Each evaluation section has 200 test clips, and the organizers have released labels and an evaluator. The helper downloads official task, record, and evaluator metadata by default; archives require DCASE2025_TASK2_DOWNLOAD_ARCHIVES=1 and an explicit part list.
Safe-first helperscripts/download/dcase2025_task2_asd.sh
View helper
Audio understanding, generation & events

DCASE 2025 Task 5 AudioQA

DCASE 2025 Task 5: Multi-Domain Audio Question Answering

Manual or gated
Audio Question Answering Temporal Audio Reasoning Bioacoustic Question Answering Multiple Choice Question Answering
Access pathHugging Face
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
mit_on_hugging_face_card
Code license
unspecified
License caution
The Hugging Face card metadata declares MIT, but the benchmark incorporates audio from Watkins Marine Mammal Sound Database, AudioSet, Mira, and other sources named by the organizers. Upstream audio and source-platform terms may be narrower and still apply; confirm them before redistribution or commercial use. No separate code license was found for the release scripts.
Download notes
The official English multiple-choice benchmark combines Bioacoustics QA, Temporal Soundscapes QA, and Complex QA (MMAU). The challenge page reports approximately 8.1K training and 2.4K development question-answer pairs; the released repository also includes the evaluation set. Access is public but auto-approved gated: users must sign in and provide basic identity and affiliation fields. The helper saves public challenge, paper, and repository API metadata, then prints the manual acceptance and authenticated download steps; it never downloads audio automatically.
Safe-first helperscripts/download/dcase2025_audioqa.sh
View helper
Audio understanding, generation & events

DCASE 2026 Task 1 HAC

DCASE 2026 Task 1: Heterogeneous Audio Classification

Safe-first helper
Heterogeneous Audio Classification Hierarchical Audio Classification Multimodal Audio Classification Domain Generalization
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0_with_upstream_terms
Code license
not_specified
License caution
All three official Zenodo records declare CC BY 4.0. The releases include per-sound Freesound license and uploader provenance, so users must also honor each source recording's terms. The official baseline repository has no detected license or LICENSE file; its code terms are unspecified.
Download notes
The task predicts 23 second-level Broad Sound Taxonomy categories and scores macro-averaged hierarchical F-score. Development uses the curated BSD10k-v1.2 release (about 11,000 sounds and 35 hours) and the noisier crowd-sourced BSD35k-CS release (about 35,000 sounds and 150 hours), both with text metadata and Freesound provenance. The public evaluation archive contains audio and metadata but intentionally omits labels. The helper downloads official pages, Zenodo records, READMEs, and approximately 7 MB of development metadata by default. Roughly 200 MB of CLAP features and 47 GB of audio/evaluation archives require separate explicit opt-ins.
Safe-first helperscripts/download/dcase2026_task1_hac.sh
View helper
Audiovisual & cross-modal

DCASE2025 Task 3 Stereo SELD Dataset

DCASE2025 Task 3 Stereo Sound Event Localization and Detection Dataset

Safe-first helper
Sound Event Localization And Detection Stereo Sound Source Localization Audio Visual Sound Event Localization Source Distance Estimation +2 more
Access pathZenodo
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT
Code license
mixed
License caution
Zenodo metadata and the included LICENSE identify the dataset as MIT. Sony's data-generator repository is MIT. The official baseline repository has no detected GitHub license or LICENSE file, so its code terms are not specified. Because the release is derived from STARSS23 recordings of people and rooms, review the original privacy/provenance context before sensitive visual use.
Download notes
The public Zenodo v1.1.0 release contains 30,000 labeled development clips (41.7 hours) and 10,000 unlabeled evaluation clips (13.9 hours), each five seconds long, with 24 kHz stereo audio and aligned perspective video. It is derived from STARSS23 by sampling and converting its FOA audio and 360-degree video, and adds folded azimuth, source-distance, and onscreen/offscreen labels. The helper downloads official pages, record metadata, README, license, paper page, and generator documentation by default; the approximately 15.2 MB label archive is opt-in, while the approximately 27.6 GB audio/video release remains on Zenodo.
Safe-first helperscripts/download/dcase2025_stereo_seld.sh
View helper
Audio understanding, generation & events

DESED

Domestic Environment Sound Event Detection Dataset

Safe-first helper
Sound Event Detection Audio Tagging Weakly Supervised Sound Event Detection Synthetic Soundscape Generation
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
MIT
License caution
Zenodo records for DESED real and synthetic list CC BY 4.0. The GitHub README says the Python code is MIT and that component datasets include license files at their roots; source media comes from AudioSet/YouTube, Freesound, MUSAN, SINS, and related sources, so re-check component terms before redistribution.
Download notes
The helper downloads the official repo plus Zenodo record JSON and small metadata/JAMS files by default. Real and synthetic audio archives are multi-GB and require explicit opt-in flags.
Safe-first helperscripts/download/desed.sh
View helper
Speech generation

Designed Vocalizations Dataset

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Safe-first helper
Non Human Voice Conversion Designed Vocalization Generation Timbre Transfer Sound Design Reproduction +2 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0-with-mixed-upstream-terms
Code license
not_applicable
License caution
The dataset card applies CC BY 4.0 to the compilation and authors' original metadata and annotations. Raw and designed clips retain their applicable source terms: VCTK and HiFi-TTS are CC BY 4.0, while Freesound clips are individually CC0 1.0, CC BY 3.0, or CC BY 4.0. Preserve the release NOTICE and per-row license, attribution, creator, and source fields when redistributing. The project does not publish a separate evaluation-code repository.
Download notes
The public, ungated release contains 237,574 mono 44.1 kHz WAV clips embedded in Parquet: 5,654 raw training sources, 226,160 non-parallel designed training clips, 120 test sources, and 5,640 aligned test references. The test protocol crosses source timbres seen or unseen during training with 40 seen and seven unseen effect presets. The helper downloads official documentation, API metadata, licensing notices, preset metadata, and the approximately 532 KB test-pair manifest by default. The full Hugging Face repository is approximately 37.1 GB and requires DESIGNED_VOCALIZATIONS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/designed_vocalizations.sh
View helper
Audio understanding, generation & events

DHAuDS

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

Safe-first helper
Test Time Adaptation Audio Classification Robustness Dynamic Corruption Robustness Speech Command Classification +3 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
apache-2.0_on_hugging_face_cards_with_upstream_terms
Code license
apache-2.0
License caution
All four Hugging Face cards declare Apache-2.0 and the code repository contains an Apache-2.0 LICENSE. The corrupted audio derives from Speech Commands V2, VocalSound, UrbanSound8K, ReefSet, QUT-NOISE, and DEMAND; those sources retain separate attribution, share-alike, non-commercial, or other terms. In particular, the Apache card labels should not be assumed to remove UrbanSound8K's non-commercial restriction or other upstream obligations.
Download notes
The public suite contains separately corrupted adaptation and evaluation sets derived from the held-out portions of Speech Commands V2, VocalSound, UrbanSound8K, and ReefSet. SC2-C, VS-C, and RS-C apply seven corruption categories at two severity levels; US8-C omits QUT-NOISE and DEMAND corruptions that overlap its target classes and uses four categories. The paper reports 908,196 derived samples in total and uses different random seeds for adaptation and evaluation corruptions. The helper downloads official documentation and repository metadata by default. The four Hugging Face repositories report about 50.0 GB of storage combined, so snapshots require DHAUDS_DOWNLOAD_HF=1 and an explicit DHAUDS_DATASETS selection.
Safe-first helperscripts/download/dhauds.sh
View helper
Speech recognition

Dialogs

Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus for Dialog Assistants

Safe-first helper
Expressive Text To Speech Conversational Text To Speech Speech Emotion Classification Automatic Speech Recognition +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
OpenRAIL
Code license
MIT
License caution
The dataset card links a custom OpenRAIL responsible-use license that permits use, modification, redistribution, and commercial use subject to its use-based restrictions. The paper and card state that performers gave written informed consent for public and commercial use. Review the complete LICENSE.md rather than treating OpenRAIL as an unrestricted permissive license. The linked VITS2 baseline repository is MIT.
Download notes
The public, ungated release contains 20.6 hours and 11,796 Russian utterances from face-to-face acted dialogues recorded in a studio by three professional performers. It provides transcripts, stress-marked text, speaker identifiers, and 12 style/emotion labels, with fixed 11,428/180/188 train, validation, and test splits. The paper evaluates the 188-item stratified test subset with six human-rated quality dimensions and trains a VITS2 expressive-TTS baseline. The helper downloads official documentation, API metadata, and the lightweight validation/test tables by default. The approximately 29.3 MB embedded- audio preview and 5.56 GB full Hugging Face snapshot are separate opt-ins.
Safe-first helperscripts/download/dialogs_ru.sh
View helper
Enhancement, separation & quality

Diamond Benchmark

Diamond Benchmark: 750 Real Degraded Speech Recordings for Evaluating Restoration Models

Safe-first helper
Speech Restoration Speech Enhancement Perceptual Speech Quality Evaluation Content Preservation Evaluation +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other_unspecified
Code license
not_applicable
License caution
Hugging Face metadata labels the dataset "other", but the card provides no license text, source-corpus citation, consent statement, or redistribution terms. The manifest's `emolia_id` field appears to reference an upstream collection, but the card does not identify or license it. Treat the release as evaluation-only pending clarification and verify source-recording, speaker, transcript, and redistribution rights before reuse, especially for training or commercial purposes.
Download notes
The public, ungated release contains 750 English speech clips with real-world codec, bandwidth, noise, and clipping degradation, plus reference transcripts and speaker, duration, and sample-rate metadata. Its documented protocol combines DNSMOS-P.835 for perceptual quality with ASR character error rate for content preservation. The helper downloads the official card, API metadata, and approximately 197 KB manifest by default. The Hugging Face API reports approximately 340 MB of repository storage, so the audio snapshot is an explicit opt-in.
Safe-first helperscripts/download/diamond_benchmark.sh
View helper
Speaker, identity & emotion

DIHARD III

The Third DIHARD Speech Diarization Challenge

Manual or gated
Speaker Diarization Speech Activity Detection Overlapping Speech Diarization Multisource Speech Diarization
Access pathLDC / licensed
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
ldc_user_agreement
Code license
not_applicable
License caution
LDC2022S12 and LDC2022S14 list the LDC User Agreement for Non-Members and are available through LDC membership/non-member access. Re-check active LDC terms and component source restrictions before use or redistribution.
Download notes
The LDC catalog records list web-download development and evaluation releases with user-agreement access. The helper only prints official access steps; it does not download LDC-controlled data.
Safe-first helperscripts/download/dihard_iii.sh
View helper
Enhancement, separation & quality

DNS Challenge

Deep Noise Suppression Challenge

Safe-first helper
Speech Enhancement Speech Denoising Dereverberation Personalized Speech Enhancement +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
MIT
License caution
The repository legal notice says documentation/content are CC BY 4.0 and code is MIT. DNS training data includes component sources such as AudioSet, Freesound, VCTK, VocalSet, and multilingual speech, so component/source-media terms should be re-checked before redistribution or commercial use.
Download notes
DNS5 development and blind test sets are multi-GB archives; the full training resources are hundreds of GB compressed and about 1 TB unpacked. The helper saves official README/license/downloader-script files by default and makes data archives explicit opt-ins.
Safe-first helperscripts/download/dns_challenge.sh
View helper
Representation & general suites

Doppelganger

Doppelganger: Sound Effects and Their Synthetic Twins

Safe-first helper
Synthetic Real Sound Effect Retrieval Audio Instance Matching Audio Representation Evaluation Synthetic Audio Detection
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
MIT
License caution
The Hugging Face card's MIT tag covers manifests, embeddings, and generation logs, not every audio asset. Stable Audio Open twins use the Stability AI Community License; ElevenLabs twins are redistributed under the author's ElevenLabs license; real audio retains FSD50K, UrbanSound8K, Freesound, or DCASE 2023 Task 7 source terms, and restricted real sources remain reference-by-ID only.
Download notes
The public, ungated release pairs 10,420 verified real sound-effect references across 34 Universal Category System events with Stable Audio Open synthetic twins, plus a controlled seven-class DCASE 2023 Task 7 corpus and text-only ElevenLabs controls. Real recordings are referenced by source ID and are not redistributed in bulk. The helper downloads official documentation and repository metadata by default; cloning the roughly 10 MB code/manifests repository or downloading the approximately 8.48 GB Hugging Face release requires separate opt-ins.
Safe-first helperscripts/download/doppelganger.sh
View helper
Speech understanding & dialogue

Dynamic-SUPERB

Dynamic-SUPERB: A dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech

Safe-first helper
Spoken Language Model Evaluation Instruction Following Speech Understanding Audio Understanding +4 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
not_specified
License caution
GitHub API reported no repository license on 2026-07-09. The benchmark aggregates tasks from many component datasets; use the official task metadata and each upstream corpus license before redistribution, training, or commercial use.
Download notes
The helper downloads the official README and leaderboard documentation by default and clones the benchmark repository only with DYNAMIC_SUPERB_CLONE_REPO=1. The benchmark is collaborative and spans many speech, music, and general sound tasks, so underlying task data should be checked through each component source before use.
Safe-first helperscripts/download/dynamic_superb.sh
View helper
Speech recognition

Earnings-21

Earnings-21: A Practical Benchmark for ASR in the Wild

Safe-first helper
Asr Named Entity Recognition Long Form Speech Recognition Financial Speech Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0_text
Code license
not_specified
License caution
The dataset README shows a CC BY-SA 4.0 badge, but LICENSE.md expressly covers only transcripts and associated text files used for alignment. The repository has no detected top-level license, so confirm audio rights before redistribution or commercial use.
Download notes
Earnings-21 contains 44 English-language earnings calls totaling about 39 hours, plus a representative 10-hour Eval-10 subset. The helper downloads official documentation and lightweight file/speaker metadata by default; sparse checkout of the approximately 770 MB media tree, transcripts, RTTMs, and bias lists is opt-in.
Safe-first helperscripts/download/earnings_21.sh
View helper
Speech recognition

Earnings-22

Earnings-22: A Practical Benchmark for Accents in the Wild

Safe-first helper
Asr Accented Speech Recognition Long Form Speech Recognition Financial Speech Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
not_specified
License caution
The earnings22 README shows a CC BY-SA 4.0 license badge, and LICENSE.md states that transcripts and associated text files are CC BY-SA 4.0. The top-level GitHub repository does not expose a detected repository license, and audio is stored through Git LFS, so re-check upstream terms before redistribution or commercial use.
Download notes
Earnings-22 contains 125 English-language earnings-call files totaling about 119 hours. The helper downloads README, license, and metadata by default; sparse checkout of transcripts/media is opt-in, and Git LFS audio pull is a second explicit opt-in.
Safe-first helperscripts/download/earnings_22.sh
View helper
Speaker, identity & emotion

EMO-SUPERB

EMO-SUPERB: An In-depth Look at Speech Emotion Recognition

Safe-first helper
Speech Emotion Recognition Speech Representation Evaluation Speaker Independent Cross Validation Standardized Dataset Partitioning
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_component_terms
Code license
not_specified
License caution
The official repository has no LICENSE file and GitHub reports no detected license, so the evaluation code and released partition/label files have unspecified terms. Each of the six underlying speech corpora retains its own access agreement or license; a public benchmark repository does not make their audio freely redistributable.
Download notes
The public repository provides the evaluation implementation, corpus adapters, and standardized speaker-independent partitions for IEMOCAP, CREMA-D, MSP-IMPROV, and BIIC-NNIME. EMO-SUPERB evaluates six corpora in total, also including MSP-Podcast and BIIC-Podcast, with separate primary/secondary-emotion settings for some corpora. The helper downloads official documentation by default and makes the approximately 23 MB GitHub repository clone opt-in. It does not fetch corpus audio; users must obtain each component through its official EULA, form, or repository path.
Safe-first helperscripts/download/emo_superb.sh
View helper
Audiovisual & cross-modal

EmoPrefer

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

Safe-first helper
Multimodal Emotion Preference Prediction Audio Visual Emotion Understanding Pairwise Emotion Description Evaluation Multimodal Judge Evaluation +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_non_commercial_research_terms
Code license
Apache-2.0_official_and_MIT_audit
License caution
The official EmoPrefer subdirectory includes Apache-2.0 but its README also calls the service a non-commercial research preview. The gated MER2025 card declares CC BY-NC 4.0 plus stricter academic-only, no-redistribution, and no-modification conditions that control the source media and annotations obtained there. The 2026 audit repository is MIT. Treat the narrower access terms as controlling where they conflict, and do not infer media rights from the public CSV release.
Download notes
The official repository publicly releases six small annotation tables: the original 574-pair EmoPrefer set, its V2 extension with 2,096 individual-annotator pairs, reverse-order variants for swap-consistency analysis, and variants exposing the two description-generator names for score calculation and shortcut auditing. The paired English audio/video comes from the separately gated MER2025 release and is not redistributed by the repository. The helper downloads the annotations, official documentation, licenses, and repository metadata only; users must accept the MER2025 access conditions themselves to obtain media. The 2026 audit paper and code provide reproducible content-blind, counterfactual, audio-visual judge, and ODIN-style diagnostics but intentionally distribute no private data, model weights, predictions, or checkpoints.
Safe-first helperscripts/download/emoprefer.sh
View helper
Speech generation

EmoV-DB

The Emotional Voices Database: Towards Controlling the Emotional Expressiveness in Voice Generation Systems

Safe-first helper
Emotional Speech Synthesis Expressive Tts Speech Emotion Recognition Voice Conversion
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_non_commercial
Code license
not_specified
License caution
The EmoV-DB license permits non-commercial research, teaching, scientific publication, and personal experimentation, and asks users to contact the dataset owner for commercial use. GitHub API reports license NOASSERTION/Other.
Download notes
OpenSLR SLR115 hosts per-speaker/per-emotion archives. The helper downloads OpenSLR/GitHub docs and license by default; speech archives are opt-in with EMOV_DB_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/emov_db.sh
View helper
Audiovisual & cross-modal

EPIC-SOUNDS

EPIC-SOUNDS: A Large-Scale Dataset of Actions that Sound

Safe-first helper
Egocentric Audio Event Recognition Sound Event Detection Audio Visual Understanding Action Sound Recognition
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The annotation README states all dataset files are published under Creative Commons Attribution-NonCommercial 4.0 International. Raw audio is derived from EPIC-KITCHENS-100 video recordings, so original dataset access terms and any HDF5 access approval should be checked before redistribution or commercial use.
Download notes
The helper downloads official docs and public annotation CSV files by default. Raw audio is not redistributed separately; the official README says to download EPIC-KITCHENS-100 videos and extract audio, or email the maintainers for access to an existing HDF5 file.
Safe-first helperscripts/download/epic_sounds.sh
View helper
Audio understanding, generation & events

ESC-50

ESC-50: Dataset for Environmental Sound Classification

Safe-first helper
Environmental Sound Classification Audio Tagging
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-3.0
Code license
not_specified
License caution
ESC-10 subset clips are CC BY; ESC-50 as a whole is Creative Commons Attribution-NonCommercial. Per-clip Freesound attributions are in the repository LICENSE file.
Safe-first helperscripts/download/esc_50.sh
View helper
Audio understanding, generation & events

ESCUCHA

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

Safe-first helper
Spanish Speech Understanding Audio Question Answering Audio Reasoning Multi Audio Comparison +3 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The GitHub repository has no license file or detected GitHub license. The paper is CC BY 4.0 on arXiv, but that does not license the annotations, code, or source recordings. Review the rights and platform terms for each linked recording before downloading, redistribution, or commercial use.
Download notes
The public repository releases 1,000 Spanish questions as JSON and TSV, including 900 multiple-choice and 100 audio-instruction-following items. The helper downloads those approximately 2.2 MB of annotations plus the README and scorer by default; cloning the repository is opt-in. Audio is not redistributed: the release provides source URLs and a yt-dlp script for reconstructing up to 162.9 hours from public recordings, so availability can drift and source-platform terms apply.
Safe-first helperscripts/download/escucha.sh
View helper
Speech understanding & dialogue

Europarl-ST

Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates

Safe-first helper
Speech To Text Translation Speech Translation Multilingual Asr Parliamentary Speech Translation +1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_applicable
License caution
The official README says the work carried out to construct Europarl-ST is released under CC BY-NC 4.0, while all rights of the underlying data belong to the European Union and respective copyright holders. Re-check EU/European Parliament reuse terms before redistribution or commercial use.
Download notes
The official page links release v1.1, which adds Romanian, Polish, and Dutch to German, English, Spanish, French, Italian, and Portuguese for 72 speech translation directions, plus a train-noisy set. The v1.1 archive is about 21 GB, so the helper downloads only the official page and README by default and requires EUROPARL_ST_DOWNLOAD_ARCHIVE=1 for the archive.
Safe-first helperscripts/download/europarl_st.sh
View helper
Speech recognition

Fisher English

Fisher English Training Speech and Transcripts

Manual or gated
Automatic Speech Recognition Conversational Speech Recognition Telephone Speech Recognition Speech Transcription
Access pathLDC / licensed
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_ldc_license
Code license
not_applicable
License caution
The LDC catalog pages list the LDC User Agreement for Non-Members and availability for Subscription/Standard Members and Non-Members. Re-check the current LDC agreement before use or redistribution.
Download notes
Fisher English is distributed by LDC after login/licensing. The helper only prints official access steps because the speech and transcripts are paid/licensed catalog releases, not publicly script-downloadable archives.
Safe-first helperscripts/download/fisher_english.sh
View helper
Speech recognition

FLEURS

FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Safe-first helper
Asr Speech Translation Language Identification
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
HF dataset card lists CC BY 4.0.
Safe-first helperscripts/download/fleurs.sh
View helper
Speech understanding & dialogue

Fluent Speech Commands

Fluent Speech Commands: A dataset for spoken language understanding research

Manual or gated
Spoken Language Understanding Intent Classification Slot Filling Smart Home Voice Commands
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_specified
License caution
Official Fluent.ai page says the dataset is released strictly for academic research only under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International, and not authorized for commercial use. No current public code repository was identified.
Download notes
The helper downloads the public dataset page and license PDF, then prints the manual Google Groups access path. The official page says the corpus contains 30,043 single-channel 16 kHz WAV utterances from 97 speakers with action, object, and location slot labels; do not commit granted links or downloaded audio.
Safe-first helperscripts/download/fluent_speech_commands.sh
View helper
Music

FMA

FMA: A Dataset For Music Analysis

Safe-first helper
Music Genre Classification Music Auto Tagging Music Recommendation Artist Identification +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
MIT
License caution
Official GitHub README says metadata is CC BY 4.0, code is MIT, and audio is distributed under the license chosen by each artist because the dataset maintainers do not hold audio copyright. UCI lists the dataset record as CC BY 4.0, but per-track audio licenses should be checked before redistribution or commercial use.
Download notes
The helper downloads official README/license files by default. Metadata is 342 MiB; audio archives range from 7.2 GiB for fma_small to 879 GiB for fma_full, so metadata and audio are explicit opt-ins.
Safe-first helperscripts/download/fma.sh
View helper
Audio understanding, generation & events

ForestIR

ForestIR: Physics-Informed Forest Sound Simulation for Array-Based Bioacoustic Remote Sensing

Safe-first helper
Forest Impulse Response Simulation Bioacoustic Sound Source Localization Microphone Array Design Spatial Audio Robustness +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
bundled_media_terms_not_separately_specified
Code license
MIT
License caution
The repository has an MIT LICENSE and GitHub detects MIT, but neither the paper nor README states separate licenses or complete provenance terms for bundled bird vocalizations and recorded environmental noise. Do not assume the software license clears those recordings; verify source-media attribution, wildlife-recording, and field-recording rights before redistribution or commercial use. The request-only processed validation data may carry additional terms.
Download notes
The public repository provides a physics-informed forest impulse-response simulator, command-line and Python interfaces, measured and synthetic tree/microphone/source geometry presets, example bird calls, environmental noise recordings, and manifest-producing array rendering. The helper downloads official documentation, license, repository metadata, and the paper by default; cloning the approximately 31 MB repository requires FORESTIR_CLONE_REPO=1. The paper says processed site-recorded data needed to reproduce its main validation analyses must be requested from the authors, so the public repository must not be represented as a complete release of those field measurements.
Safe-first helperscripts/download/forestir.sh
View helper
Audiovisual & cross-modal

Friend Bench

Friend Bench: Social relationship recognition from thin-slice dyadic interactions

Safe-first helper
Audio Visual Social Reasoning Social Relationship Recognition Familiar Stranger Classification Multimodal Human Behavior Understanding +1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_applicable
License caution
The dataset card declares CC BY-NC 4.0 and says the benchmark inherits that license from Seamless Interaction. The release includes recorded human interactions, transcripts, relationship labels, and anonymized rater responses with bucketed demographic fields; users should retain attribution, limit use to non-commercial purposes, and review the source corpus's privacy, consent, biometric, and responsible-use terms.
Download notes
The public, ungated validation release contains 96 approximately 20-second dyadic clips, balanced across familiar/stranger and early/late conditions. Every item provides mixed audio, side-by-side video, a turn-level transcript, and labels for binary familiarity and six-way relationship type. It also releases anonymized human judgments under audio-, video-, and text-only conditions. The helper downloads the official card and API metadata by default. Set FRIEND_BENCH_DOWNLOAD_METADATA=1 for the two lightweight JSONL tables, or FRIEND_BENCH_DOWNLOAD_HF=1 for the complete approximately 433 MB snapshot. The card says a label-held-out test split is planned but not yet released, and currently provides no paper or citation.
Safe-first helperscripts/download/friend_bench.sh
View helper
Audio understanding, generation & events

FSD50K

FSD50K: An Open Dataset of Human-Labeled Sound Events

Safe-first helper
Sound Event Classification Sound Event Tagging Audio Tagging Machine Listening
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_creative_commons
Code license
not_specified
License caution
Zenodo says individual Freesound clips keep per-clip CC0, CC-BY, CC-BY-NC, or CC Sampling+ licenses, with clip-level license mappings in the metadata JSON files. FSD50K as a curated dataset is additionally released under CC BY, but the maintainers warn that a single global license is not straightforward because items have different licenses.
Download notes
The helper downloads the small documentation, ground-truth, and metadata ZIPs by default. Audio is split across about 24.7 GiB of dev archives plus about 6.2 GiB of eval archives, so audio download is an explicit opt-in.
Safe-first helperscripts/download/fsd50k.sh
View helper
Audio understanding, generation & events

FSDKaggle2018

FSDKaggle2018: Freesound General-Purpose Audio Tagging Challenge Dataset

Safe-first helper
Audio Tagging Sound Event Classification General Purpose Audio Tagging
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_creative_commons
Code license
not_specified
License caution
The Zenodo record license id is other-at. The record states FSDKaggle2018 as a curated dataset is CC BY, while individual Freesound clips retain per-clip Creative Commons licenses listed in train_post_competition.csv and test_post_competition_scoring_clips.csv. Kaggle competition rules and Freesound source terms should also be checked for challenge use.
Download notes
The helper downloads the Zenodo record JSON plus small documentation and metadata ZIPs by default. The audio archives are about 4.6 GiB combined, so audio download is an explicit opt-in.
Safe-first helperscripts/download/fsdkaggle2018.sh
View helper
Audio understanding, generation & events

FSDKaggle2019

FSDKaggle2019: Freesound Audio Tagging 2019 / Audio Tagging with Noisy Labels and Minimal Supervision

Safe-first helper
Audio Tagging Sound Event Classification Noisy Label Learning Weakly Labeled Audio Classification +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_creative_commons
Code license
MIT
License caution
Zenodo reports license id other-at. The record says FSDKaggle2019 as a curated dataset is CC BY, while individual Freesound clips keep per-clip CC0, CC-BY, CC-BY-NC, or CC Sampling+ licenses and Flickr/YFCC noisy-train clips keep CC-BY or CC BY-SA licenses recorded in metadata CSVs. Kaggle/DCASE challenge rules and source media terms should also be checked for challenge or redistribution use.
Download notes
The helper downloads the Zenodo record JSON plus small documentation and metadata ZIPs by default. Full audio is about 25 GiB and includes a split noisy-train archive, so audio download is an explicit opt-in with part selection.
Safe-first helperscripts/download/fsdkaggle2019.sh
View helper
Enhancement, separation & quality

FUSS

FUSS: Free Universal Sound Separation Dataset

Safe-first helper
Universal Sound Separation Arbitrary Sound Separation Reverberant Source Separation Sound Event Separation
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
Apache-2.0
License caution
The official license document and Zenodo record release FUSS as a whole under CC BY 4.0. Its input audio clips are CC0 Freesound files selected using prerelease FSD50K labels; those labels are not distributed with FUSS. Google Research's sound-separation code repository is Apache-2.0.
Download notes
The public Zenodo release provides train, validation, and eval mixtures containing one to four arbitrary sound sources, with dry and reverberant references, simulated room responses, and CC0 source clips. The helper saves official documentation, repository metadata, and the small license archive by default. Data archives are about 1.9 to 8.9 GB each and require explicit selection with FUSS_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/fuss.sh
View helper
Audio understanding, generation & events

Geo-ATBench

Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context

Safe-first helper
Multi Label Audio Tagging Environmental Sound Classification Audio Geospatial Context Fusion Context Aware Audio Tagging
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0_with_upstream_terms
Code license
MIT
License caution
The Zenodo record declares CC BY 4.0 and the official repository is MIT. The paper says audio comes from Freesound and a prior geotagged-audio dataset, while contextual metadata derives from OpenStreetMap; preserve source attribution and review per-clip audio licenses and OSM attribution/database terms before redistribution.
Download notes
The public, ungated release contains 3,854 ten-second mono WAV clips totaling 10.71 hours, 28 fine-grained sound-event labels, three coarse event groups, and POI-derived geospatial semantic context over 11 OpenStreetMap feature categories. The helper downloads official documentation and repository/Zenodo metadata by default; the approximately 850 MB dataset archive is opt-in.
Safe-first helperscripts/download/geo_atbench.sh
View helper
Speech recognition

Ghana Speech Eval

Ghana Speech Eval: ASR Evaluation for 10 Ghanaian Languages

Safe-first helper
Automatic Speech Recognition Multilingual Speech Recognition Low Resource Speech Recognition African Language Speech Recognition
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_claim_with_unresolved_upstream_terms
Code license
not_applicable
License caution
The Ghana Speech Eval card declares CC BY 4.0 and says this follows the source dataset. However, the current AfriSpeech public-v1 card and API metadata do not state a repository-level license. Preserve attribution, review the original collection and speaker-consent terms, and confirm upstream reuse rights before redistribution or commercial use.
Download notes
The public, ungated release contains 9,967 read-speech clips with verbatim transcripts across Ahanta, Dagaare, Dangme, Ewe, Fante, Frafra/Gurene, Ga, Nzema, Sehwi, and Twi. It samples up to 1,000 clips per language from AfriSpeech's public African speech corpus, merges the source splits into one evaluation split, and filters clips to 3-15 seconds. The card explicitly reserves the fixed set for ASR evaluation rather than training. The helper downloads official cards and API metadata by default; the approximately 594 MB compressed benchmark snapshot requires GHANA_SPEECH_EVAL_DOWNLOAD_HF=1.
Safe-first helperscripts/download/ghana_speech_eval.sh
View helper
Speech recognition

GigaSpeech

GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Manual or gated
Asr Large Scale Speech Recognition Text To Speech
Access pathHugging Face
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
gated_non_commercial_research_educational
Code license
Apache-2.0
License caution
HF terms restrict the database to non-commercial research and educational purposes; the GitHub code repo is Apache-2.0.
Download notes
HF access is gated and the official repo also asks users to fill out the Google Form first; full HF dataset size is about 2.6 TB.
Safe-first helperscripts/download/gigaspeech.sh
View helper
Speech recognition

GigaSpeechBench

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Safe-first helper
Automatic Speech Recognition Speech To Text Translation Multilingual Speech Recognition Dialectal Speech Recognition +6 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
Neither the Hugging Face repository metadata nor the official GitHub repository declares a license, and the GitHub repository has no LICENSE file. Do not infer reuse or redistribution rights from public, ungated access; obtain clarification from SpeechColab before redistribution or commercial use.
Download notes
The paper defines a 680-hour in-the-wild benchmark with five modules covering 14 low-resource language/region sets, six Chinese dialects, six English accents, 12 Chinese and English terminology domains, and child/older-adult speech; it also reports Chinese and English translation references for 11 languages. The current public, ungated Hugging Face repository contains the Low-Resource-Languages, CH-EN-Dialects, and Vertical-Domain modules, including audio archives and JSON metadata, but no separate age-group module was visible when checked. Translation result files are public, while the exact release coverage of translation references should be verified per metadata file. The helper downloads official documentation and repository/API metadata by default; the Hugging Face API reports approximately 86.3 GB of repository storage, so the dataset snapshot requires GIGASPEECHBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/gigaspeechbench.sh
View helper
Speech recognition

Golos

Golos: Russian Dataset for Speech Research

Safe-first helper
Automatic Speech Recognition Russian Speech Recognition Far Field Speech Recognition Language Modeling
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_golos_license
Code license
not_specified
License caution
OpenSLR points to SberDevices' English and Russian PDF license files rather than an SPDX-style open license. Re-check those PDFs before redistribution, commercial use, or training release claims.
Download notes
OpenSLR SLR114 mirrors an 18 GiB Opus archive with Russian speech and transcripts, a 71 MiB QuartzNet acoustic model, and 4.7 GiB KenLM language models. The helper saves the OpenSLR page, README, checksums, and license PDFs by default and requires explicit opt-ins before downloading large archives.
Safe-first helperscripts/download/golos.sh
View helper
Music

GTZAN

GTZAN Genre Collection

Safe-first helper
Music Genre Classification Music Information Retrieval
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The reachable Hugging Face dataset card does not state a data license. Treat redistribution and commercial use as unclear until the Marsyas/original dataset terms are confirmed.
Download notes
The Hugging Face card describes 1,000 30-second mono WAV tracks across 10 genres and provides a reproducible datasets loader. The helper saves the HF dataset card by default and requires GTZAN_DOWNLOAD_HF=1 before downloading the audio snapshot.
Safe-first helperscripts/download/gtzan.sh
View helper
Speech recognition

HALAS

HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

Safe-first helper
Automatic Speech Recognition Asr Hallucination Detection Asr Error Analysis Span Level Error Detection +1 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
unknown_with_cc-by-sa-4.0_upstream
Code license
not_specified
License caution
The Hugging Face card's machine-readable license field is "unknown," and neither the GitHub nor Hugging Face release includes a standalone license file. The card says HALAS derives from CC BY-SA 4.0 Earnings-22 and describes the authors' annotations as intended for the same license unless otherwise specified, but it also tells users to consult the repository for authoritative terms. Treat annotation and code rights as unspecified pending an explicit license; Earnings-22 source terms continue to apply to separately obtained audio.
Download notes
The public, ungated release contains human-reviewed predictions from seven ASR systems for 3,611 English Earnings-22 segments, with utterance labels and character-span annotations for hallucination, looping, and hallucinated looping. Meeting-disjoint stratified splits contain 2,866 train and 745 test segments; the test split has a 22.6% hallucination rate and excludes audio at or below one second or below three words. The release contains annotations, predictions, corrected references, and source identifiers rather than redistributed audio; audio must be obtained separately from Earnings-22. The helper downloads official documentation, prompts, and repository metadata by default. The approximately 0.86 MB test CSV is opt-in, while the full approximately 7.2 MB Hugging Face snapshot requires HALAS_DOWNLOAD_HF=1.
Safe-first helperscripts/download/halas.sh
View helper
Representation & general suites

HEAR

HEAR: Holistic Evaluation of Audio Representations

Safe-first helper
Audio Representation Evaluation Speech Classification Environmental Sound Classification Music Classification +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
Apache-2.0
License caution
Zenodo record lists CC BY 4.0 but explicitly says component datasets have different open licenses and to inspect each dataset's LICENSE.txt; eval kit repo is Apache-2.0.
Download notes
The helper downloads the Zenodo record metadata and LICENSE.txt by default. Individual task archives are large and opt-in; the Zenodo record notes that TFDS-derived crema-d, GTZAN genre, and GTZAN music/speech tasks in this release were retracted because of a preprocessing bug.
Safe-first helperscripts/download/hear.sh
View helper
Speech generation

Hi-Fi TTS

Hi-Fi Multi-Speaker English TTS Dataset

Safe-first helper
Text To Speech Speech Synthesis Multi Speaker Speech Synthesis High Fidelity Speech Synthesis +1 more
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
OpenSLR SLR109 lists CC BY 4.0. The corpus is based on public-domain LibriVox audiobooks and Project Gutenberg texts, but downstream users should still preserve attribution and check packaged metadata.
Download notes
OpenSLR lists one 39-41 GiB speech/text archive. The helper saves the official OpenSLR page by default and requires HIFITTS_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/hifitts.sh
View helper
Speaker, identity & emotion

HI-MIA

HI-MIA: A Far-field Text-Dependent Speaker Verification Database and the Baselines

Safe-first helper
Text Dependent Speaker Verification Far Field Speaker Verification Wake Word Speaker Verification Speaker Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
apache-2.0
Code license
not_applicable
License caution
OpenSLR SLR85 lists Apache License v2.0. The paper says HI-MIA contains recordings from 340 people in real-room far-field settings and is extracted from AISHELL-WakeUp-1.
Download notes
OpenSLR SLR85 hosts the AISHELL Speaker Verification Challenge 2019 far-field text-dependent speaker verification data. The helper downloads the official page and filename mapping by default; train/dev/test archives are multi-GB and require HIMIA_DOWNLOAD_ARCHIVES=1.
Safe-first helperscripts/download/hi_mia.sh
View helper
Audiovisual & cross-modal

IEMOCAP

IEMOCAP: Interactive Emotional Dyadic Motion Capture Database

Manual or gated
Speech Emotion Recognition Audio Visual Emotion Recognition Multimodal Emotion Recognition Dialogue Emotion Recognition
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_research_license
Code license
not_applicable
License caution
The official release page links a USC/SAIL data release form and Google request form. Treat access as manual/form-gated and re-check the signed release terms before redistribution, commercial use, or sharing derived copies.
Download notes
IEMOCAP is released by request after reading the license and submitting the official electronic release form; there is no public one-command archive URL.
Safe-first helperscripts/download/iemocap.sh
View helper
Speech understanding & dialogue

IFEval-Audio

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Manual or gated
Audio Instruction Following Speech Instruction Following Structured Output Adherence Semantic Correctness Evaluation +2 more
Access pathHugging Face
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
apache-2.0_with_mixed_upstream_terms
Code license
cc_by_nc_unspecified_version
License caution
The Hugging Face card declares Apache-2.0, while the paper says clips are used as-is from Spoken SQuAD (CC BY-SA 4.0), TED-LIUM 3 (CC BY-NC-ND 3.0), MuChoMusic (CC BY-SA 4.0), WavCaps (academic use only), and AudioBench sources with inherited licenses. The AudioBench LICENSE says source code is Creative Commons NonCommercial without naming a version. Preserve the most restrictive source terms and verify per-clip provenance before redistribution or commercial use.
Download notes
The official release contains 280 English audio-instruction-answer triples spanning Content, Capitalization, Symbol, List Structure, Length, and Format requirements. It includes 240 speech triples and 40 music/environmental-sound triples, and reports Instruction Following Rate, Semantic Correctness Rate, and Overall Success Rate. The helper downloads public Hugging Face API metadata and official benchmark documentation by default; the approximately 45 MB compressed Hugging Face snapshot is opt-in and requires logging in and accepting the repository's access conditions.
Safe-first helperscripts/download/ifeval_audio.sh
View helper
Audiovisual & cross-modal

InCarEmo

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

Manual or gated
Speech Emotion Recognition Multimodal Emotion Recognition Fatigue Detection Driver Distraction Monitoring +1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_non_commercial_academic_research
Code license
not_specified
License caution
The paper states that InCarEmo is licensed for non-commercial academic research and use, but neither the official repository nor linked Drive landing page provides a full license text. The repository contains no LICENSE file and no released code at the time checked. ArXiv's CC BY 4.0 license covers the paper, not the feature data.
Download notes
The paper describes synchronized Chinese RGB and infrared video, in-cabin audio, and dialogue text from 25 participants, with six emotion classes plus fatigue and distraction labels. It also defines an auxiliary English setting using translated text and synthesized English speech. Because of participant privacy, the official repository releases the dataset only as feature-level data through a public Google Drive folder, not as raw audio or video. The helper saves the public repository and paper metadata, then prints the manual Drive step; it does not download participant data.
Safe-first helperscripts/download/incaremo.sh
View helper
Speech recognition

IndicContextEval

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

Safe-first helper
Contextual Asr Multilingual Asr Named Entity Recognition Context Utilisation +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
The Hugging Face card and official repository README declare the benchmark CC BY 4.0. The GitHub repository has no LICENSE file and GitHub detects no repository license, so do not assume that its evaluation outputs or forthcoming evaluation code use the same terms without clarification.
Download notes
The public, ungated release contains 16,884 natural-speech utterances totaling 55.93 hours from 555 speakers across Bengali, Gujarati, Hindi, Malayalam, Marathi, Odia, Telugu, and Urdu. It covers 23 professional domains and supplies seven controlled prompt levels: no context, language, structured metadata, natural-language description, English- script entities, native-script entities, and incorrect-domain entities. The helper downloads official documentation, repository metadata, the small published results table, and prompt-taxonomy supplements by default. The Hugging Face card reports a 6.48 GB current download and its API reports approximately 19.64 GB of repository storage including history, so embedded audio requires INDIC_CONTEXT_EVAL_DOWNLOAD_HF=1.
Safe-first helperscripts/download/indic_context_eval.sh
View helper
Speech generation

InstructTTSEval

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Safe-first helper
Controllable Text To Speech Speech Instruction Following Acoustic Parameter Control Descriptive Style Control +1 more
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mit
Code license
not_specified
License caution
The Hugging Face dataset card lists MIT, while the GitHub repository has no detected license. The paper says reference audio was curated from NCSSD plus movies, TV dramas, variety shows, and other film/television sources; it also says the dataset is solely for academic and research use. Treat the card license as insufficient to clear third-party media rights, and verify source terms before redistribution or commercial use.
Download notes
The helper downloads the official GitHub README and Hugging Face dataset card by default. The public, ungated Hugging Face snapshot contains English and Chinese Parquet splits with embedded reference audio and uses about 1.8 GB, so data download requires INSTRUCT_TTS_EVAL_DOWNLOAD_HF=1; cloning the evaluation repository is separately opt-in.
Safe-first helperscripts/download/instruct_tts_eval.sh
View helper
Audiovisual & cross-modal

InterPet4D

InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

Safe-first helper
Audio Conditioned Motion Generation Human Pet Interaction Modeling Multimodal Motion Generation Animal Behavior Understanding +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
conflicting_noncommercial_terms
Code license
not_applicable
License caution
The current public Hugging Face card declares CC BY-NC 4.0 and says participants consented to data release. The paper's ethical statement instead says users must agree to a data-use agreement prohibiting redistribution, surveillance, and biometric identification, but no such agreement or click-through gate is exposed on the current ungated dataset page. Apply the stricter paper restrictions pending clarification, do not redistribute the corpus, and do not infer that arXiv's CC BY 4.0 paper license governs the released participant recordings.
Download notes
The public, ungated v1 release contains 227 approximately 17-20-second egocentric clips from 113 interaction sessions involving 13 dogs and about 23 human participants. It provides synchronized 48 kHz stereo MP3 audio, MERT embeddings, and human-hand, human-body, dog-skeleton, and SMAL motion parameters. The helper downloads the official dataset card and API metadata by default; the Hugging Face API reports about 10.7 GB of repository storage, so the full snapshot requires INTERPET4D_DOWNLOAD_HF=1. The paper also describes 12-view and egocentric RGB video, but the current Hugging Face file inventory releases audio and motion artifacts rather than those raw videos.
Safe-first helperscripts/download/interpet4d.sh
View helper
Speech recognition

JASMIN-CGN

JASMIN-CGN: Extension of the Spoken Dutch Corpus with Speech of Elderly People, Children and Non-Natives

Manual or gated
Automatic Speech Recognition Child Speech Recognition Accented Speech Recognition Non Native Speech Recognition +3 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_signed_license_noncommercial
Code license
not_applicable
License caution
The official non-commercial product is free but requires a signed license and account login; the owner page directs commercial users to a separate commercial product. Review the current agreement before use or redistribution.
Download notes
The Dutch Language Institute distributes the approximately 115-hour Dutch/Flemish speech corpus after account login and a signed license agreement. Its official page says the corpus contains read speech and human-machine dialogues from adolescents, non-native speakers, and seniors, with WAV audio plus TXT and TextGrid annotations. The helper prints the official access steps only; it does not bypass login or download corpus audio. The 2026 evaluation paper selects 120 human-machine-interaction test utterances from children, older adults, and Flemish speakers and also reports ASR results on the full corresponding test sets.
Safe-first helperscripts/download/jasmin_cgn.sh
View helper
Speech recognition

KeSpeech

KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects

Manual or gated
Mandarin Asr Dialect Asr
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_non_commercial_no_adaptations_no_distribution
Code license
not_specified
License caution
Downloading data means accepting the custom license in dataset_license.md.
Safe-first helperscripts/download/kespeech.sh
View helper
Speaker, identity & emotion

KVoiceBench / KOpenAudioBench / KMMAU

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Safe-first helper
Korean Spoken Question Answering Speech Instruction Following Speech Reasoning Speech Safety Evaluation +4 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_Apache-2.0_and_CC-BY-NC-SA-4.0
Code license
Apache-2.0
License caution
The KVoiceBench and KOpenAudioBench cards declare Apache-2.0; KMMAU declares CC BY-NC-SA 4.0. These benchmarks adapt prior benchmark content or source speech corpora, so the declared repository licenses may not replace upstream terms. Review VoiceBench/OpenAudioBench provenance and the KSS, KMSAV, and Seoul Corpus conditions before redistribution or commercial use. Raon-Eval code is Apache-2.0.
Download notes
The three public, ungated Hugging Face releases contain 12,345 Korean test samples: 7,306 KVoiceBench spoken-QA items transferred from VoiceBench, 2,835 KOpenAudioBench items transferred from OpenAudioBench, and 2,204 KMMAU audio-understanding items derived from KSS, KMSAV, and Seoul Corpus. The helper downloads official cards, repository metadata, and the paper page by default. Full snapshots require KOREAN_SPEECHLM_BENCHMARKS_DOWNLOAD_HF=1 because Hugging Face reports approximately 4.64 GB, 608 MB, and 4.49 GB of repository storage respectively.
Safe-first helperscripts/download/korean_speechlm_benchmarks.sh
View helper
Speech recognition

L2-ARCTIC

L2-ARCTIC: A Non-native English Speech Corpus

Manual or gated
Asr Accented Speech Recognition Mispronunciation Detection Accent Conversion +2 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_applicable
License caution
Official homepage states the corpus is released under CC BY-NC 4.0; usage outside that license requires contacting the TAMU dataset owner. The current release covers 24 non-native English speakers plus suitcase-corpus material, while the Interspeech 2018 paper describes the initial v1.0 release.
Download notes
Official access requires reviewing the CC BY-NC 4.0 license terms and submitting the download form with name, email, and affiliation. The project sends a generated Google Drive link by email, so the helper prints manual access steps rather than storing or using private generated URLs.
Safe-first helperscripts/download/l2_arctic.sh
View helper
Enhancement, separation & quality

L3DAS21

L3DAS21 Challenge: Machine Learning for 3D Audio Signal Processing

Safe-first helper
Three Dimensional Speech Enhancement Speech Enhancement Sound Event Localization And Detection Sound Source Localization +2 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_with_upstream_terms
Code license
not_specified
License caution
DataCite lists CC BY 4.0 for the Zenodo dataset. The generated mixtures incorporate LibriSpeech speech and FSD50K sound events, so preserve LibriSpeech attribution and review FSD50K's per-clip Creative Commons terms, including non-commercial clips. The official GitHub repository exposes no LICENSE file or GitHub-detected license, so clarify code terms before reuse or redistribution.
Download notes
The public Zenodo V1 release is a 65-hour corpus of multi-source, multi-perspective B-format Ambisonics audio generated from impulse responses measured at 252 positions with two first-order Ambisonics microphones. Task 1 contains more than 30,000 spatial speech mixtures for 3D speech enhancement, with clean monaural speech targets. Task 2 contains 900 one-minute soundscapes for sound-event localization and detection, with up to three simultaneous events and 100 ms activity, class, and Cartesian-location targets. Both tasks have one- and two-microphone tracks. The helper downloads official pages, repository metadata, DataCite metadata, and the paper by default; cloning the code and downloading selected large archives are separate opt-ins.
Safe-first helperscripts/download/l3das21.sh
View helper
Enhancement, separation & quality

L3DAS22

L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment

Safe-first helper
Three Dimensional Speech Enhancement Speech Enhancement Sound Event Localization And Detection Sound Source Localization +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_with_upstream_terms
Code license
not_specified
License caution
Kaggle's official API lists CC BY 4.0 for L3DAS22. The generated mixtures incorporate LibriSpeech speech and FSD50K sound events; preserve LibriSpeech attribution and check FSD50K's per-clip Creative Commons licenses, including non-commercial clips. The official GitHub repository exposes no LICENSE file or GitHub-detected license, so clarify code terms before redistribution.
Download notes
The public Kaggle release contains 105,757,713,362 bytes of multi-source, multi-perspective B-format Ambisonics audio. Task 1 has more than 40,000 spatial speech mixtures totaling nearly 90 hours at 16 kHz, with up to three overlapping background noises and clean monaural speech targets. Task 2 has 900 30-second soundscapes totaling 7.5 hours at 32 kHz, with 14 office-relevant event classes, up to three overlaps, and 100 ms activity and Cartesian-location targets. Both tasks provide one- and two-microphone tracks using one or two first-order Ambisonics microphones. The helper downloads official pages, repository metadata, the paper, documentation, and Kaggle API metadata by default; cloning the code and downloading the 105.8 GB dataset are separate opt-ins. Full data download requires the Kaggle CLI and account credentials.
Safe-first helperscripts/download/l3das22.sh
View helper
Audio understanding, generation & events

LAT-Bench

LAT-Bench: Long-form Audio Temporal Awareness Benchmark

Safe-first helper
Long Form Audio Understanding Temporal Audio Reasoning Dense Audio Captioning Temporal Audio Grounding +1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0_claim_with_upstream_media_terms
Code license
Apache-2.0
License caution
The Hugging Face card declares CC BY-NC 4.0 for LAT-Bench, while the GitHub repository is Apache-2.0. The release references externally hosted source audio rather than relicensing or redistributing it; review each recording's rights and platform terms before retrieval, redistribution, or commercial use.
Download notes
The public, ungated release provides Chinese and English metadata plus task JSONL files for a human-verified 40-hour benchmark, with 25 hours of Chinese and 15 hours of English audio up to 30 minutes long. Dense Audio Captioning, Temporal Audio Grounding, and Targeted Audio Captioning annotations are included. Audio is not bundled; the metadata records source URLs, so availability can drift and source-platform terms apply. The helper downloads the approximately 2.6 MB annotations, official documentation, and repository metadata by default; cloning the evaluation repository is opt-in.
Safe-first helperscripts/download/lat_bench.sh
View helper
Speech recognition

Libri-Light

Libri-Light: A Benchmark for ASR with Limited or No Supervision

Safe-first helper
Automatic Speech Recognition Self Supervised Speech Representation Semi Supervised Asr Unsupervised Speech Representation +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
MIT
License caution
The GitHub repository says Libri-Light code is MIT, but the reviewed README and data-preparation docs do not state a standalone data license. The paper describes the data as derived from open-source LibriVox audiobooks, so re-check per-source audiobook terms and attribution requirements before redistribution.
Download notes
The helper downloads official README/license/data-preparation/evaluation docs by default. Limited-supervision finetuning data is about 0.6 GiB, ABX item data is small, and unlabeled audio ranges from 35 GiB small to 321 GiB medium and 3.05 TiB large, so all data archives require explicit opt-in.
Safe-first helperscripts/download/libri_light.sh
View helper
Speech recognition

LibriCSS

LibriCSS: Continuous Speech Separation Dataset

Safe-first helper
Continuous Speech Separation Overlapped Speech Recognition Multi Channel Asr Speaker Diarization +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_upstream_basis_data_license_not_separately_stated
Code license
MIT
License caution
The repository LICENSE applies MIT to the software and documentation and identifies the source LibriSpeech corpus as CC BY 4.0, but it does not state a separate license for the replayed LibriCSS recordings. Preserve LibriSpeech attribution and confirm recording-level rights before redistribution or commercial use.
Download notes
The public, ungated release contains ten approximately one-hour sessions made from LibriSpeech utterances replayed through eight loudspeakers and captured with a seven-channel circular microphone array in an office meeting room. Each session has six ten-minute mini-sessions spanning 0%-40% overlap, including separate short-gap and long-gap 0% conditions. The official repository provides preparation, Kaldi ASR, and WER-scoring tools for utterance-wise and continuous-input evaluation. The helper downloads official documentation and repository metadata by default; the direct Google Drive archive is 6,407,297,637 bytes (about 5.97 GiB) and requires LIBRICSS_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/libricss.sh
View helper
Enhancement, separation & quality

LibriMix

LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Safe-first helper
Speech Separation Speech Enhancement Noisy Speech Separation Source Separation
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_derived
Code license
MIT
License caution
The LibriMix repository license is MIT for code/scripts. Generated data is derived from LibriSpeech, which OpenSLR lists as CC BY 4.0, plus WHAM noise; re-check WHAM terms and cite all upstream components before redistribution or commercial use.
Download notes
The helper clones the official generator/metadata repository by default. Running generation downloads LibriSpeech clean subsets and WHAM noise, then creates many mixtures; the upstream README estimates about 430 GiB for Libri2Mix plus 332 GiB for Libri3Mix, with additional source storage, so generation requires an explicit opt-in and an external storage directory.
Safe-first helperscripts/download/librimix.sh
View helper
Speech recognition

LibriSpeech

LibriSpeech ASR corpus

Safe-first helper
Asr Speech Reconstruction
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
OpenSLR SLR12 and HF dataset card list CC BY 4.0.
Download notes
The helper downloads the OpenSLR landing page and checksums by default; corpus archives require LIBRISPEECH_DOWNLOAD_ARCHIVES=1. Qwen3-TTS section 4.1.2 evaluates tokenizer reconstruction on all 2,620 utterances in test-clean with PESQ, STOI, UTMOS, and WavLM-based speaker similarity. Qwen-Audio-VAE sections 4.1-4.3 use LibriSpeech for speech reconstruction and ablation evaluation.
Safe-first helperscripts/download/librispeech.sh
View helper
Speech generation

LibriTTS

LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Safe-first helper
Text To Speech Speech Synthesis Voice Cloning Multi Speaker Speech Synthesis
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
OpenSLR SLR60 lists CC BY 4.0. LibriTTS is derived from LibriSpeech, which in turn derives from LibriVox audio and Project Gutenberg text.
Download notes
OpenSLR lists seven archives totaling tens of GiB; the helper saves the official OpenSLR page by default and requires LIBRITTS_DOWNLOAD_ARCHIVES=1 before downloading archives.
Safe-first helperscripts/download/libritts.sh
View helper
Speech recognition

LibriWASN

LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices

Safe-first helper
Meeting Transcription Continuous Speech Separation Speaker Diarization Multi Channel Asr +2 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
MIT
License caution
Zenodo declares the LibriWASN release CC BY 4.0 and includes the license text; the official tools repository is MIT. The recordings replay LibriSpeech/LibriCSS material, so preserve that provenance and review upstream speech-data terms when redistributing derived data.
Download notes
The public release contains 20 hours of meeting-like recordings in two rooms with approximately 200 ms and 800 ms reverberation times. Five smartphones and four microphone arrays provide 29 asynchronous audio channels, with 0%-40% overlap conditions and ground-truth diarization. The same LibriSpeech sentences and speakers as LibriCSS were replayed. The helper downloads official metadata, README/license files, paper page, and repository documentation by default. Twelve per-room and per-overlap ZIP archives total about 55.8 GB and require explicit opt-in; the official repository's full downloader also fetches LibriCSS as a transcription/reference dependency.
Safe-first helperscripts/download/libriwasn.sh
View helper
Speech recognition

Live Gurbani Captioning Benchmark v1

Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan

Safe-first helper
Closed Vocabulary Singing Captioning Live Audio Tracking Lyrics Alignment Sung Scripture Identification +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_annotations
Code license
MIT
License caution
The repository LICENSE applies CC BY 4.0 to ground-truth annotations in test/ and the committed baselines, and MIT to the software and documentation. The four underlying Kirtan recordings are referenced by YouTube ID rather than bundled; the repository license does not grant rights to those recordings, and platform terms, uploader rights, performer rights, and local cultural or legal considerations remain separate.
Download notes
The public v1 repository contains ground-truth timelines for four hand-reviewed Sikh Kirtan recordings, each evaluated from 0%, 33%, and 66% start offsets, giving 12 cases and approximately 57 minutes of scored audio. It also releases a standard-library Python scorer, visualization and annotation tools, and empty, shifted, and perfect baselines. The primary metric is one-second frame accuracy with a one-second boundary collar. The helper downloads official documentation and repository metadata by default; cloning the small repository with annotations and evaluation code is opt-in. Audio is not redistributed: the official README identifies four YouTube video IDs and provides yt-dlp/ffmpeg preparation commands, so availability and source-media rights must be checked separately.
Safe-first helperscripts/download/live_gurbani_captioning_v1.sh
View helper
Speech generation

LJSpeech

The LJ Speech Dataset

Safe-first helper
Text To Speech Speech Synthesis Single Speaker Speech Synthesis Asr
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
public_domain
Code license
not_applicable
License caution
Official page says text, audio, and annotations are public domain; the HF mirror lists unlicense.
Download notes
The official archive is about 2.6 GiB. The helper saves the official dataset page by default and requires LJSPEECH_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/ljspeech.sh
View helper
Audiovisual & cross-modal

LLP

Look, Listen, and Parse: Audio-Visual Video Parsing Dataset

Safe-first helper
Audio Visual Video Parsing Temporal Event Localization Audio Event Detection Visual Event Detection +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
unclear
License caution
The official repository does not contain a license file and GitHub reports no detected license. Its README says GPLv3 but links to a GPLv3 license in an unrelated repository, so neither that statement nor the linked file establishes clear terms for LLP annotations or code. Obtain clarification before redistribution or commercial use; source videos also retain their original rights and YouTube terms.
Download notes
The official repository releases lightweight weak-label train/validation/test CSVs and dense audio and visual event annotations for validation and test. The full CSV contains 11,849 ten-second YouTube segment references; the published split files contain 10,000 train, 649 validation, and 1,200 test rows. The helper downloads documentation and all annotation CSVs by default. Extracted VGGish, ResNet-152, and R(2+1)D features remain a manual Google Drive download, while raw videos must be reconstructed from referenced YouTube segments subject to availability and platform terms.
Safe-first helperscripts/download/llp.sh
View helper
Audio understanding, generation & events

LOCATA

LOCATA: IEEE-AASP Challenge on Acoustic Source Localization and Tracking

Safe-first helper
Sound Source Localization Acoustic Source Tracking Multi Source Localization Moving Source Localization +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
ODC-By-1.0
Code license
not_specified
License caution
Zenodo identifies the final dataset release as Open Data Commons Attribution, and the official corpus page says both LOCATA and its VCTK speech material use Open Data Commons terms. Neither official MATLAB repository exposes a LICENSE file or a GitHub-detected license, so clarify software terms before redistribution.
Download notes
The open final release contains corrected development and evaluation datasets with close-talking speech, distant multichannel recordings from four microphone-array configurations, and OptiTrack ground truth for sources and microphones. Its six tasks span static and moving, single- and multi-source scenarios with static or moving arrays. The dev.zip and eval.zip archives total about 19.3 GB, so the helper downloads only official pages, documentation, Zenodo metadata, and tool metadata by default.
Safe-first helperscripts/download/locata.sh
View helper
Audiovisual & cross-modal

LRRo

LRRo: A Lip Reading Data Set for the Under-resourced Romanian Language

Safe-first helper
Visual Speech Recognition Lip Reading Isolated Word Recognition Low Resource Romanian Speech
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
not_applicable
License caution
The Zenodo record declares CC BY 4.0. Wild LRRo was collected from Internet videos and the release contains identifiable speaker imagery, so attribution, privacy, likeness, and any source-media rights should still be reviewed for redistribution, biometric use, or deployment. The documentation-only GitHub repository has no detected license.
Download notes
The public, ungated Zenodo release contains Wild LRRo, with more than 20 hours, over 35 speakers, 1,100 word instances, and a 21-word vocabulary, plus Lab LRRo, with more than 5 hours, 19 speakers, 6,400 word instances, and a 48-word vocabulary. Both isolated-word collections provide train, validation, and test subsets. VSRo-200 section 4.6 evaluates transfer on both LRRo subsets. The helper saves the official Zenodo record and repository documentation by default; the 291,504,956-byte archive is opt-in with LRRO_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/lrro.sh
View helper
Audiovisual & cross-modal

LRS2-BBC

The Oxford-BBC Lip Reading Sentences 2 Dataset

Manual or gated
Audio Visual Speech Recognition Visual Speech Recognition Lip Reading
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
non-commercial_academic_access_agreement
Code license
not_applicable
License caution
BBC R&D restricts LRS2 to non-commercial research by universities, reputable academic institutions, and relevant public organizations; companies and independent researchers are not eligible. The signed agreement controls use, and the BBC asks researchers to obtain approval before publishing sample images. Review the current agreement and BBC broadcast-content rights before use.
Download notes
The official VGG page describes 144,482 pre-train/train/validation/test utterances from BBC television, with date-disjoint validation and test broadcasts, and links a password-protected 50 GB package plus public split file links. Access requires submitting the BBC LRS2 Data Sharing Agreement from an official academic address; approved researchers receive a countersigned agreement and password. The helper saves the official pages, agreement, and paper metadata only and does not request credentials or download the corpus.
Safe-first helperscripts/download/lrs2.sh
View helper
Audiovisual & cross-modal

LRS3-TED

LRS3-TED: A Large-Scale Dataset for Visual Speech Recognition

Manual or gated
Audio Visual Speech Recognition Visual Speech Recognition Lip Reading
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
unspecified
Code license
not_applicable
License caution
No LRS3 data license is stated on the currently accessible official landing page or paper record. TED/TEDx source-video rights remain with their owners; obtain the official terms before use and do not infer reuse rights from third-party mirrors.
Download notes
The paper introduces more than 400 hours of aligned face tracks, audio, subtitles, and word boundaries from TED and TEDx for visual and audio-visual speech recognition. The official VGG landing page still identifies LRS3 and links its dataset page, but that linked page returned HTTP 404 when checked on 2026-07-22. The helper preserves the live official landing page and paper metadata and reports the unavailable official download route; it does not substitute an unofficial mirror.
Safe-first helperscripts/download/lrs3.sh
View helper
Audiovisual & cross-modal

LVOmniBench

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs

Manual or gated
Audio Visual Question Answering Long Video Understanding Cross Modal Reasoning Temporal Localization +2 more
Access pathHugging Face
Upstream termsNot specified

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The paper says every source video carries a Creative Commons license, but neither the official dataset card nor repository states a benchmark-level license or records each video's exact Creative Commons variant in public documentation. The GitHub repository has no LICENSE file or detected license. Review the access form, per-video rights, attribution requirements, and YouTube terms before reuse or redistribution; the paper's CC BY 4.0 license covers the article, not automatically the dataset or evaluation code.
Download notes
The gated Hugging Face release contains 275 English Creative Commons-licensed YouTube videos totaling 140 hours, with durations from 10 to 90 minutes, plus 1,014 manually authored four-option question-answer pairs. Questions span nine categories and require joint reasoning over speech, music, or sound with visual evidence. The helper downloads public repository documentation and API metadata by default. The Hugging Face API reports approximately 187.4 GB of repository storage, so the complete snapshot requires approval, authentication, and LVOMNIBENCH_DOWNLOAD_HF=1. OmniReasoner evaluates LVOmniBench in its main audio-visual table and reports duration-wise tool-use behavior on the benchmark.
Safe-first helperscripts/download/lvomnibench.sh
View helper
Enhancement, separation & quality

Lyra-SA

Lyra Lab Singing Assessment Dataset

Manual or gated
Singing Quality Assessment Full Song Singing Assessment Singing Score Prediction
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0
Code license
not_applicable
License caution
The official page states CC BY-NC 4.0 for non-commercial use, requires attribution to the source page and notice, reserves copyright to Tencent Music Entertainment Group, and requires separate permission for commercial use. The recordings are WeSing user performances of copyrighted songs; rely on the official authorization and application terms rather than inferring broader music or performance rights from the Creative Commons label.
Download notes
The official Tencent Music Lyra Lab page describes 1,000 complete mobile-karaoke recordings: 100 user covers for each of 10 Chinese songs, with no repeated singer. The release includes 44.1 kHz 16-bit mono WAV audio, listener-provided overall singing scores, rough singer labels, timed lyrics, and reference MIDI. Access is application-based: users must complete the official form and accept its terms, after which Lyra Lab says it emails a download link within three days. The helper saves official documentation and the recent SongSQA paper, then prints the manual application path; it never guesses or bypasses an emailed archive URL.
Safe-first helperscripts/download/lyra_sa.sh
View helper
Audio understanding, generation & events

MACS

MACS: Multi-Annotator Captioned Soundscapes

Safe-first helper
Audio Captioning Audio Tagging Multi Annotator Audio Labeling Acoustic Scene Captioning
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other-nc
Code license
not_applicable
License caution
Zenodo lists MACS as Other (Non-Commercial), and LICENSE.txt grants experimental non-commercial use with attribution to Tampere University/Machine Listening Group. Audio files come from TAU Urban Acoustic Scenes 2019, whose Zenodo record also lists Other (Non-Commercial).
Download notes
The helper downloads the MACS annotations, competence scores, license, and TAU Urban Acoustic Scenes 2019 docs/metadata by default. The upstream TAU 2019 audio shards are large, so audio download is an explicit opt-in.
Safe-first helperscripts/download/macs.sh
View helper
Music

MADB

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Safe-first helper
Music Aesthetic Assessment Music Quality Assessment Music Score Regression Multimodal Music Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0_with_upstream_audio_rights
Code license
not_specified
License caution
The Hugging Face card declares CC BY-NC 4.0 and research-only use, but also warns that audio may remain subject to original copyright restrictions. No license file is present in the official GitHub repository. Apply the non-commercial dataset terms and independently review MuChin, generator-service/output, and unidentified online-source rights before redistributing or using audio.
Download notes
The paper describes 9,999 tracks rated by 30 trained annotators, with about ten ratings per track across ten perceptual dimensions and an overall score plus comments and tags. The public, ungated Hugging Face release includes all annotations and 1,730 Suno/Levo audio tracks; its card points to the separate MuChin repository for another 4,400 tracks and says remaining tracks came from diverse online sources. The helper downloads official documentation and repository metadata by default, makes the approximately 69 MB annotation tables opt-in, and requires MADB_DOWNLOAD_HF=1 for the approximately 18.6 GB Hugging Face snapshot.
Safe-first helperscripts/download/madb.sh
View helper
Representation & general suites

MAEB

MAEB: Massive Audio Embedding Benchmark

Safe-first helper
Audio Embedding Evaluation Audio Text Retrieval Audio Classification Audio Clustering +6 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_component_terms
Code license
Apache-2.0
License caution
Apache-2.0 covers the MTEB software and benchmark registry, not the underlying MAEB task datasets. The 30 tasks draw on multiple speech, music, and environmental-audio sources with their own licenses, access controls, attribution requirements, and media rights. Review every selected task's metadata and upstream dataset terms before downloading, redistributing, or using it commercially.
Download notes
The public MTEB registry defines the 30-task MAEB beta suite across speech, music, environmental sound, and audio-text evaluation in more than 100 languages. It includes retrieval, classification, clustering, pair classification, reranking, multilabel classification, and zero-shot classification tasks. The helper downloads official documentation, repository metadata, the Apache-2.0 license, and the lightweight benchmark registry by default; cloning the approximately 55 MB MTEB source repository is opt-in. It does not fetch component datasets, which MTEB tasks acquire separately and which can be large or restricted.
Safe-first helperscripts/download/maeb.sh
View helper
Music

MAESTRO

MAESTRO: MIDI and Audio Edited for Synchronous TRacks and Organization

Safe-first helper
Automatic Music Transcription Piano Transcription Music Synthesis Symbolic Music Modeling
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
not_applicable
License caution
Official Magenta page says the dataset is made available by Google LLC under Creative Commons Attribution Non-Commercial Share-Alike 4.0.
Download notes
The helper downloads v3.0.0 CSV/JSON metadata by default. The MIDI-only archive is about 56 MiB and the full WAV+MIDI archive is about 101 GiB, so both are explicit opt-ins.
Safe-first helperscripts/download/maestro.sh
View helper
Audio understanding, generation & events

MAESTRO Real

MAESTRO Real: Multi-Annotator Estimated Strong Labels

Safe-first helper
Sound Event Detection Soft Label Sound Event Detection Multi Annotator Label Aggregation Long Form Sound Event Detection
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_noncommercial
Code license
not_applicable
License caution
The packaged Tampere University license permits copying and use only for experimental and non-commercial purposes with the full copyright notice and source acknowledgment. It explicitly prohibits commercial use, including selling or distributing results or content achieved through use of the dataset.
Download notes
The public development release contains 49 real-life recordings from five acoustic scenes with crowdsourced soft strong labels for 17 classes. Recordings are three to five minutes long and come from subsets of TUT Sound Events 2016 and 2017. DCASE 2024 Task 4 combines MAESTRO Real with DESED and evaluates 11 MAESTRO classes; its separate 26-file MAESTRO evaluation set is not included in the public development archive. The helper downloads official metadata, README, license, and the sub-megabyte annotation archive by default; the approximately 2.43 GiB audio archive requires explicit opt-in. The official Zenodo description reports 189 minutes 52 seconds total while its packaged README reports 97 minutes 4 seconds, so verify duration against the downloaded files rather than assuming either figure.
Safe-first helperscripts/download/maestro_real.sh
View helper
Speech recognition

MAGICDATA Mandarin Chinese Read Speech Corpus

Safe-first helper
Automatic Speech Recognition Speaker Recognition Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_applicable
License caution
OpenSLR lists CC BY-NC-ND 4.0 and says the corpus is freely published for non-commercial or academic use. Re-check the current official page before redistribution or commercial use.
Download notes
The helper downloads the OpenSLR page and small metadata archive by default. Speech archives are large, including about 52 GiB train, 1.0 GiB dev, and 2.2 GiB test, so archive download is an explicit opt-in.
Safe-first helperscripts/download/magicdata_mandarin.sh
View helper
Music

MagnaTagATune

MagnaTagATune: A Music Annotation Benchmark from the TagATune Game

Safe-first helper
Music Auto Tagging Music Annotation Music Similarity
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-3.0
Code license
GPL-3.0
License caution
The original TagATune details page says the data is CC BY-NC-SA 3.0 except scripts released under GPL v3, enabling non-commercial research redistribution. Audio clips are Magnatune excerpts with artist/album URLs for purchase or commercial licensing.
Download notes
City University MIRG hosts metadata, annotations, comparisons, Echo Nest features, and three 1 GiB MP3 split archives. The helper downloads CSV metadata by default and makes features/audio explicit opt-ins.
Safe-first helperscripts/download/magnatagatune.sh
View helper
Speaker, identity & emotion

MCR-Bench

MCR-Bench: Modal Conflict Resolution Benchmark for Large Audio-Language Models

Safe-first helper
Cross Modal Conflict Resolution Audio Text Robustness Audio Question Answering Speech Emotion Recognition +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_upstream_terms
Code license
apache-2.0
License caution
The repository contains an Apache-2.0 LICENSE, although its README badge says MIT. Neither statement clearly relicenses the released benchmark archive or embedded source audio. ClothoAQA/Clotho, MELD's copyrighted Friends clips, and VocalSound retain their own terms, so verify item-level provenance and upstream permissions before redistribution or commercial use.
Download notes
The public Google Drive release contains approximately 3,000 English samples across audio question answering, speech emotion recognition, and vocal-sound classification. Each audio item is paired with faithful, adversarial, irrelevant, and neutral text conditions to measure whether an audio-language model follows contradictory text instead of audio evidence. The benchmark derives its task audio from ClothoAQA, MELD, and VocalSound. The helper downloads official documentation and repository metadata by default and prints the manual Drive path; cloning the documentation-only repository is opt-in.
Safe-first helperscripts/download/mcr_bench.sh
View helper
Enhancement, separation & quality

MedleyDB

MedleyDB: A Multitrack Dataset for Annotation-Intensive MIR Research

Safe-first helper
Music Information Retrieval Melody Extraction Music Source Separation Instrument Recognition +1 more
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
MIT
License caution
The official downloads page says MedleyDB is free for non-commercial research use only and identifies the dataset as Creative Commons Attribution-NonCommercial-ShareAlike 4.0. It also asks users not to republish the dataset in full or in part without consent, even though redistribution is technically allowed under the license. The GitHub tooling repository is MIT.
Download notes
The helper saves official pages and repository license/README by default. It checks the Zenodo request records only with MEDLEYDB_CHECK_ZENODO=1, downloads the public sample archive only with MEDLEYDB_DOWNLOAD_SAMPLE=1, and clones the annotation/metadata/tooling repo only with MEDLEYDB_CLONE_REPO=1. Full MedleyDB and MedleyDB 2.0 audio require requesting access through the official Zenodo records.
Safe-first helperscripts/download/medleydb.sh
View helper
Audiovisual & cross-modal

MELD

MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Safe-first helper
Speech Emotion Recognition Multimodal Emotion Recognition Dialogue Emotion Recognition Sentiment Analysis +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
gpl-3.0
Code license
GPL-3.0
License caution
GitHub and the Hugging Face dataset card list GPL-3.0. MELD clips are derived from the Friends TV series, so downstream users should re-check media rights and any fair-use/research assumptions before redistribution or commercial use.
Download notes
The helper downloads official project, README, license, and dataset-card metadata by default. Raw audio/video tarballs and extracted feature/model tarballs are opt-in because they are larger and include TV-derived media clips.
Safe-first helperscripts/download/meld.sh
View helper
Audiovisual & cross-modal

MER2023

MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning

Manual or gated
Multimodal Emotion Recognition Discrete Emotion Classification Dimensional Emotion Recognition Modality Robustness +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0_with_additional_gated_terms
Code license
not_specified
License caution
The current Hugging Face card declares CC BY-NC 4.0 and limits access to academic research. Its gate also prohibits handing the database or derived labeling files to third parties and prohibits modification without written consent. The challenge paper describes a separate EULA with academic-only, no-editing, and no-upload conditions. The baseline repository has no LICENSE file. Clips were collected from movies and television, so underlying media, performer, privacy, and platform rights remain separate.
Download notes
The request-gated release extends CHEAVD and provides 3,373 labeled Train&Val clips, 411 MER-MULTI test clips, 412 corrupted-modality MER-NOISE test clips, and MER-SEMI with 834 labeled plus 73,148 unlabeled clips. MER-MULTI evaluates joint discrete-emotion and valence prediction, MER-NOISE tests robustness to noisy audio and blurred video, and MER-SEMI evaluates semi-supervised discrete emotion recognition. The helper saves public paper, repository, and Hugging Face API metadata only. The current Hugging Face repository is about 140 GB, password-protected, and requires approval, so benchmark files are left as a manual download.
Safe-first helperscripts/download/mer2023.sh
View helper
Audiovisual & cross-modal

MER2024

MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition

Manual or gated
Multimodal Emotion Recognition Discrete Emotion Classification Semi Supervised Emotion Recognition Modality Robustness +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0_with_additional_gated_terms
Code license
not_specified
License caution
The Hugging Face card declares CC BY-NC 4.0 and non-commercial use. Its gate prohibits transfer of the database or derived labeling files to third parties and modification without written consent. The challenge page additionally limits the dataset to academic research and prohibits uploading samples. The MER2024 README displays an Apache-2.0 badge but links to a license under the later MER2025 directory; the repository root and MER2024 directory contain no applicable LICENSE file, so the MER2024 code license is recorded as unspecified. Source-video, performer, privacy, and platform rights remain separate.
Download notes
This request-gated extension of MER2023 provides 5,030 labeled Train&Val clips and 115,595 unlabeled clips. MER-SEMI evaluates 1,169 annotated clips from the unlabeled pool, MER-NOISE evaluates 1,170 clips with additive audio noise and image blur, and MER-OV evaluates free-form emotion labels. Table 1 of the paper reports 322 MER-OV samples, while the surrounding prose says 332; the index preserves that primary-source discrepancy. The current Hugging Face tree is approximately 218.4 GB and requires approval. The helper saves only public paper, project, repository, and Hugging Face API metadata.
Safe-first helperscripts/download/mer2024.sh
View helper
Speech recognition

MInDS-14

MInDS-14: Multilingual and Cross-Lingual Intent Detection from Spoken Data

Safe-first helper
Spoken Language Understanding Intent Classification Automatic Speech Recognition Multilingual Speech Understanding
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
Hugging Face dataset card lists CC BY 4.0 and exposes 14 spoken e-banking intents across 14 language varieties. No separate code license was identified for the dataset card.
Download notes
The helper downloads the Hugging Face dataset card by default. Dataset snapshots include audio and are opt-in; choose one locale such as en-US or all with MINDS14_CONFIG.
Safe-first helperscripts/download/minds14.sh
View helper
Speech generation

Ming-Freeform-Audio-Edit

Ming-Freeform-Audio-Edit: Free-form instruction-based speech editing benchmark

Safe-first helper
Instruction Based Speech Editing Semantic Speech Editing Speech Content Deletion Speech Content Insertion +6 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
apache-2.0_with_upstream_terms
Code license
not_specified
License caution
The Hugging Face dataset card declares Apache-2.0, but the released audio is derived from multiple upstream corpora whose terms still apply. Seed-TTS Eval states no data license, LibriTTS is CC BY 4.0, and GigaSpeech uses its own agreement/access conditions. The evaluation repository has no license file or detected GitHub license; review each source corpus and clarify annotation/code rights before redistribution or commercial use.
Download notes
The public, ungated release contains Chinese and English source speech, natural-language editing instructions, and metadata for semantic deletion, insertion, and substitution plus acoustic emotion, dialect, speed, pitch, and volume changes. The paper's section 6.3 constructs the semantic set from 896 Chinese and 655 English Seed-TTS test samples and reports separate basic and full instruction versions; its acoustic sets also use Seed-TTS audio. The current dataset card additionally names LibriTTS and GigaSpeech as source corpora. The helper downloads official documentation and repository metadata by default; the Hugging Face API reports approximately 1.07 GB of repository storage, so audio and annotations require explicit opt-in.
Safe-first helperscripts/download/ming_freeform_audio_edit.sh
View helper
Speech recognition

MIR-1K vocal

MIR-1K

Safe-first helper
Singing Voice Transcription Singing Voice Separation
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified_on_official_page
Code license
not_applicable
License caution
Official MIR Lab page has direct downloads but no visible license statement; its MIR-1K.rar URL returned 404 on 2026-07-09. Figshare mirror lists CC BY 4.0.
Safe-first helperscripts/download/mir_1k_vocal.sh
View helper
Speech recognition

MLC-SLM Eval

Multilingual Conversational Speech Language Model Challenge Eval Ground Truth

Safe-first helper
Multilingual Conversational Asr Speaker Diarization Speaker Attributed Asr Long Form Speech Recognition +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0_annotation_card_label
Code license
not_specified
License caution
The Hugging Face card labels the released ground-truth repository CC BY-SA 4.0. It does not publish a separate license file or establish terms for the absent evaluation recordings. The official baseline repository has no detected license, so do not assume the annotation label licenses its code. Obtain the audio and its terms from the challenge owners before attempting full benchmark reproduction.
Download notes
The public, ungated Hugging Face release contains approximately 6.55 MB of oracle segmentation, speaker labels, and transcriptions for the challenge's 32-hour Eval-1 and 32-hour Eval-2 sets. It covers English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese; English additionally spans five accent groups. The repository currently contains 224 annotation files and no audio. Challenge participants previously received evaluation recordings, but the official paper, dataset card, and baseline repository provide no current public audio URL. The helper downloads official documentation and repository metadata by default; the lightweight annotation snapshot requires MLC_SLM_EVAL_DOWNLOAD_HF=1 and does not include audio. VibeVoice-ASR evaluates the MLC-Challenge set, while VibeVoice-ASR-BitNet section 3.1 and Table 4 report six MLC language subsets.
Safe-first helperscripts/download/mlc_slm_eval.sh
View helper
Speech recognition

MLS

MLS: A Large-Scale Multilingual Dataset for Speech Research

Safe-first helper
Multilingual Asr Asr Language Modeling Limited Supervision Asr
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
OpenSLR SLR94 lists CC BY 4.0. MLS is derived from LibriVox audiobooks and provides public-file-hosted archives plus MD5 checksums.
Download notes
OpenSLR links original FLAC and compressed OPUS archives for English, German, Dutch, French, Spanish, Italian, Portuguese, and Polish; archives range from about 1.6 GiB to multiple TiB, so the helper downloads only the page and checksums by default and requires MLS_DOWNLOAD_ARCHIVES=1 for audio.
Safe-first helperscripts/download/mls.sh
View helper
Enhancement, separation & quality

MMAE

MMAE: A Massive Multitask Audio Editing Benchmark

Safe-first helper
Instruction Based Audio Editing Multi Round Audio Editing Multi Hop Audio Editing Speech Editing +4 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
MIT
License caution
The Hugging Face card does not declare a dataset license or identify licenses for the source audio. The official GitHub repository has an MIT LICENSE for its code, but that must not be assumed to license the benchmark audio. Review source-media rights and obtain clarification before redistribution or commercial use.
Download notes
The public, ungated release contains 2,000 high-fidelity input samples across sound, speech, music, and mixtures, organized by six complexity levels, two granularities, and eight operation types. Its 17,741 rubric criteria evaluate instruction following and context consistency. The helper downloads official documentation and repository metadata by default; cloning the evaluation repository is opt-in, and the Hugging Face API reports approximately 4.43 GB of repository storage, so the audio snapshot requires MMAE_DOWNLOAD_HF=1.
Safe-first helperscripts/download/mmae.sh
View helper
Audio understanding, generation & events

MMAR

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Safe-first helper
Audio Question Answering Audio Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
GitHub repo did not expose a detected license; HF dataset card lists cc-by-nc-4.0.
Safe-first helperscripts/download/mmar.sh
View helper
Audio understanding, generation & events

MMAU

MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Safe-first helper
Audio Question Answering Audio Reasoning
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
HF cards: MMAU-test-mini is cc-by-nc-4.0; MMAU-test is mit
Code license
Apache-2.0
License caution
The two HF dataset cards list different licenses; re-check before redistribution.
Safe-first helperscripts/download/mmau.sh
View helper
Audio understanding, generation & events

MMAU-Pro

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Safe-first helper
Audio Question Answering Audio Reasoning Long Audio Understanding Multi Audio Reasoning +5 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0_with_upstream_terms
Code license
not_specified
License caution
The Hugging Face card declares CC BY-NC 4.0, but the paper says almost all audio was sourced from in-the-wild recordings and the spatial subset reuses EasyCom. Review source-media and EasyCom terms before redistribution or commercial use. The official GitHub repository has no license file or detected license, so evaluator code terms are unspecified.
Download notes
The public, ungated release contains 5,305 expert-authored multiple-choice and open-ended QA instances spanning 49 skills across speech, environmental sound, music, and their mixtures. It includes multiple-audio, spatial, instruction-following, and up-to-10-minute long-form cases. The helper downloads official documentation, repository metadata, the evaluator, and Hugging Face metadata by default; the Hugging Face API reports approximately 47.5 GB of repository storage, so the audio and test snapshot requires MMAU_PRO_DOWNLOAD_HF=1.
Safe-first helperscripts/download/mmau_pro.sh
View helper
Speech generation

MMGenre

MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

Safe-first helper
Singing Voice Synthesis Genre Conditioned Singing Voice Synthesis Singing Voice Genre Alignment Score Conditioned Singing Voice Synthesis
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0_with_generated_source_caveat
Code license
cc-by-4.0
License caution
The dataset card, dataset LICENSE, and repository LICENSE declare CC BY 4.0. The benchmark was derived from music generated with Suno V4.5 and then source-separated and automatically aligned; the release license does not independently resolve any terms or rights attached to the generation service or generated source content, so review those conditions for downstream use.
Download notes
The public, ungated Hugging Face release contains 3,152 aligned Chinese singing-voice and symbolic-score segments from 148 generated songs, totaling about 4.36 hours across 10 major genres and 26 subgenres. The paper body says 27 subgenres, but the dataset card identifies this as a counting error and treats the released 26-subgenre taxonomy as authoritative. The helper downloads official documentation and API metadata by default, can clone the lightweight score/code repository, and requires explicit opt-in for the approximately 5.54 GB Hugging Face snapshot.
Safe-first helperscripts/download/mmgenre.sh
View helper
Audiovisual & cross-modal

MMOU

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

Safe-first helper
Audio Visual Question Answering Long Video Understanding Cross Modal Reasoning Temporal Grounding +1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0 metadata label; non-commercial research restriction in paper
Code license
not_applicable
License caution
Both Hugging Face cards declare Apache-2.0, but the MMOU paper says the dataset is released solely for non-commercial research. Apply the stricter non-commercial restriction pending clarification. Videos were collected from public web platforms including YouTube, so source copyright, platform terms, availability, and any per-video rights also apply.
Download notes
The public, ungated NVIDIA release contains 20,000 English questions over 11,877 long-form web videos. The 5,000-item Test Mini split includes labels for local evaluation; answers for the main 15,000-item split are withheld for evaluator submission. The helper downloads official cards and API metadata by default, makes the approximately 48 MB question files opt-in, and requires a separate opt-in for the approximately 322.8 GB community-hosted video snapshot. Audio-Visual Flamingo evaluates MMOU as an omni-modal benchmark.
Safe-first helperscripts/download/mmou.sh
View helper
Speech understanding & dialogue

MMSU

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Safe-first helper
Spoken Language Understanding Speech Reasoning
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mit
Code license
not_specified
License caution
HF dataset card lists MIT; GitHub code repo did not expose a detected license.
Safe-first helperscripts/download/mmsu.sh
View helper
Enhancement, separation & quality

MoisesDB

MoisesDB: A Dataset for Source Separation Beyond 4-Stems

Safe-first helper
Music Source Separation Multi Stem Source Separation Instrument Source Separation Singing Voice Separation
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
cc-by-nc-sa-4.0
License caution
The official repository applies CC BY-NC-SA 4.0 to MoisesDB and its packaged loader/evaluation materials, and the Music AI page limits the dataset to non-commercial research use. Attribution and ShareAlike obligations apply; commercial use is not permitted by this release.
Download notes
The official Music AI research page provides the dataset through its browser download flow. The release contains 240 songs by 47 artists across 12 high-level genres, totaling 14 hours, 24 minutes, and 46 seconds, with mixtures and a hierarchical stem/source taxonomy extending beyond the common four-stem setup. The helper saves the official page plus repository README and license by default; it does not fetch the large audio archive, and cloning the loader/evaluation repository is opt-in.
Safe-first helperscripts/download/moisesdb.sh
View helper
Speech generation

MS-SNSD

Microsoft Scalable Noisy Speech Dataset

Safe-first helper
Speech Enhancement Speech Denoising Noisy Speech Synthesis Subjective Speech Quality Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
MIT
License caution
The README says Microsoft provides the datasets as-is under the original terms received. It lists PTDB-TUG under ODbL 1.0, Edinburgh/VoiceBank material under its DataShare license, selected Freesound noise as CC0, and DEMAND as CC BY-SA 3.0. Re-check component terms before redistribution or commercial use.
Download notes
The official GitHub repository contains clean speech, noise, noisy test, and clean test directories plus scripts for generating noisy speech at configurable SNRs. The helper saves README/license/generator files by default; cloning the large repository is an explicit opt-in.
Safe-first helperscripts/download/ms_snsd.sh
View helper
Speaker, identity & emotion

MSP-Podcast

MSP-Podcast: A Large Naturalistic Speech Emotional Dataset

Manual or gated
Speech Emotion Recognition Dimensional Emotion Recognition Categorical Emotion Recognition Speaker Independent Emotion Recognition +1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_academic_license
Code license
not_applicable
License caution
The owner page currently calls the release an Academic License and requires an institution-signed FDP data-transfer agreement. Although the page says source podcasts were chosen under permissive licenses, the signed corpus agreement controls access and reuse; review it directly before commercial use, redistribution, or sharing copies.
Download notes
Version 2.0 contains 264,705 naturalistic podcast speaking turns totaling 409 hours, with speaker-independent train, development, and three test partitions. It provides categorical emotion and activation, dominance, and valence labels; Test3 releases audio but withholds labels, speaker information, transcripts, and alignments for web-based evaluation. Access is free but institution/form-gated: an authorized institutional representative must sign the official academic agreement and send it to the corpus owner. There is no public archive URL.
Safe-first helperscripts/download/msp_podcast.sh
View helper
Speech understanding & dialogue

MSU-Bench

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios

Safe-first helper
Multi Speaker Conversation Understanding Speaker Identification Speaker Attribute Recognition Speaker Verification +4 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0_with_upstream_terms
Code license
MIT
License caution
The Hugging Face card declares CC BY-NC 4.0 and limits the release to non-commercial academic research because audio derives from third-party film/TV, telephone, meeting, and podcast sources. The repository LICENSE applies MIT only to code and explicitly gives the benchmark data separate non-commercial academic-research terms; review each source corpus and media right before redistribution. The paper reports approximately 731 hours of source corpora, not 731 released hours.
Download notes
The public, ungated release contains 2,847 English and Chinese four-choice QA items over 241 multi-speaker audio clips, including 2,223 human-verified items across 16 tasks. The helper downloads official documentation, repository metadata, and the approximately 5.8 MB test JSONL by default; the Hugging Face API reports approximately 2.5 GB of repository storage, so the audio and annotations snapshot requires MSU_BENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/msu_bench.sh
View helper
Speech understanding & dialogue

MSWC

Multilingual Spoken Words Corpus

Safe-first helper
Keyword Spotting Spoken Term Search Multilingual Speech Classification Forced Alignment
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
MLCommons and the HF card list CC BY 4.0. MSWC is derived from crowd-sourced sentence-level audio, including Common Voice, so preserve source attribution and re-check the active source terms for downstream redistribution.
Download notes
The official MLCommons page and HF card describe 50 languages, more than 340,000 keywords, 23.4 million 1-second examples, and over 6,000 hours. The helper saves official docs and the HF card by default; direct per-language audio, splits, and alignments are opt-in with MSWC_DOWNLOAD_ARCHIVES=1 because high-resource language audio archives can be many GiB.
Safe-first helperscripts/download/mswc.sh
View helper
Speech recognition

mTEDx

The Multilingual TEDx Corpus for Speech Recognition and Translation

Safe-first helper
Multilingual Asr Speech To Text Translation Speech Translation Sentence Level Alignment
Access pathOpenSLR
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_specified
License caution
OpenSLR SLR100 lists CC BY-NC-ND 4.0. The corpus is derived from TEDx Talks; respect TED/TEDx source terms as well as the corpus license.
Download notes
OpenSLR SLR100 hosts ASR-only language archives, speech-translation language-pair archives, IWSLT 2021 test sets, and a small French talk gender annotation CSV. Archives are multi-GiB, so the helper downloads only the OpenSLR page and small CSV by default and requires MTEDX_DOWNLOAD_ARCHIVES=1 for archive downloads.
Safe-first helperscripts/download/mtedx.sh
View helper
Music

MTG-Jamendo

MTG-Jamendo Dataset for Automatic Music Tagging

Safe-first helper
Music Auto Tagging Music Genre Classification Musical Instrument Recognition Music Mood Theme Recognition +1 more
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
Apache-2.0
License caution
The GitHub README says repository metadata is CC BY-NC-SA 4.0, code is Apache-2.0, audio files keep individual Creative Commons licenses listed in audio_licenses.txt, and the dataset is made available solely for non-commercial research and academic use unless Jamendo grants separate authorization.
Download notes
The helper clones or updates the official metadata/scripts repository by default and saves Zenodo record metadata. The upstream downloader can fetch very large archives: raw_30s audio is about 508 GiB, raw_30s audio-low is about 156 GiB, and autotagging_moodtheme audio-low is about 46 GiB, so media downloads require MTG_JAMENDO_DOWNLOAD_MEDIA=1.
Safe-first helperscripts/download/mtg_jamendo.sh
View helper
Speaker, identity & emotion

MUGEN

MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

Safe-first helper
Multi Audio Understanding Audio Grounding Speech Understanding Speaker And Demographic Understanding +5 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified_with_mixed_upstream_terms
Code license
MIT
License caution
The 35 Hugging Face dataset cards do not declare a data license. The repository's MIT license identifies its subject as software and associated documentation, so it should not be assumed to relicense embedded audio. The paper describes public corpora, specialized academic corpora, Mozilla Data Collective speech, and synthesized speech as sources; review each task's upstream corpus and generation terms before redistribution or commercial use.
Download notes
The public, ungated release provides 35 separate Hugging Face task repositories with 1,750 five-way audio-grounding problems and 9,250 audio clips across seven dimensions. Ten tasks add a reference clip, producing six concurrent audio inputs. The helper downloads official documentation, repository metadata, the license, and the Hugging Face collection inventory by default; set MUGEN_DOWNLOAD_TASK to one of the documented task names to explicitly download that task. The current cards report about 10.7 GB of compressed files across all task repositories and also expose reduced-candidate splits for input-scaling analysis.
Safe-first helperscripts/download/mugen.sh
View helper
Music

MulTTiPop

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

Safe-first helper
Automatic Music Transcription Multitrack Music Transcription Audio Midi Alignment Music Information Retrieval
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0_for_released_midi_and_metadata
Code license
not_applicable
License caution
The dataset card declares CC BY 4.0 and says its aligned MIDI is adapted from the CC BY 4.0 Lakh MIDI Dataset. Source audio is not licensed or redistributed; users are instructed to obtain only the referenced segments and use MulTTiPop for evaluation rather than training. YouTube availability, platform terms, and commercial-song rights remain applicable.
Download notes
The public, ungated release contains aligned multitrack MIDI and metadata for 572 commercial-pop segments (about 3.5 hours), divided into artist-disjoint development and test splits. It does not redistribute source audio; metadata provides YouTube video identifiers and segment timestamps. The helper downloads the dataset card, API metadata, and lightweight split manifests by default, while the full MIDI/metadata snapshot requires MULTTIPOP_DOWNLOAD_HF=1.
Safe-first helperscripts/download/multtipop.sh
View helper
Speaker, identity & emotion

MUSAN

MUSAN: A Music, Speech, and Noise Corpus

Safe-first helper
Voice Activity Detection Music Speech Discrimination Speech Music Noise Classification Speaker Recognition Augmentation
Access pathOpenSLR
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
OpenSLR SLR17 lists Attribution 4.0 International (CC BY 4.0). The paper describes MUSAN as music, speech, and noise recordings released under a flexible Creative Commons license.
Download notes
OpenSLR lists the corpus archive as 11 GiB. The helper downloads the OpenSLR landing page by default and requires MUSAN_DOWNLOAD_ARCHIVE=1 for the large archive.
Safe-first helperscripts/download/musan.sh
View helper
Enhancement, separation & quality

MUSDB18

MUSDB18: A Corpus for Music Separation

Safe-first helper
Music Source Separation Singing Voice Separation Audio Source Separation
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other-non-commercial
Code license
MIT
License caution
Zenodo lists "Other (Non-Commercial)" and the official pages state educational/academic use only, with commercial use requiring copyright-holder permission. Track sources include DSD100/Mixing Secrets, MedleyDB CC BY-NC-SA 4.0, Native Instruments stems, and Easton Ellises/Heise CC BY-NC-SA 3.0 material.
Download notes
The helper saves the official SigSep and Zenodo pages by default. The 4.7 GiB compressed STEMS archive and 22.7 GiB uncompressed HQ archive require explicit terms acknowledgement and opt-in.
Safe-first helperscripts/download/musdb18.sh
View helper
Audiovisual & cross-modal

MUSIC-AVQA

MUSIC-AVQA: Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Safe-first helper
Audio Visual Question Answering Multimodal Scene Understanding Spatiotemporal Reasoning Music Understanding
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
conflicting: GitHub API/LICENSE file reports MIT, while the README License section mentions GPLv3
License caution
No standalone dataset license was found on the official page or README on 2026-07-10. Raw musical-performance videos and extracted features should be treated as upstream media with terms to verify before redistribution or commercial use.
Download notes
The helper downloads the official project page, README/LICENSE, and public JSON QA annotations by default. Raw videos and large feature files are hosted through Google Drive and Baidu Drive links on the project page/README; they are not downloaded automatically.
Safe-first helperscripts/download/music_avqa.sh
View helper
Audio understanding, generation & events

MusICA-MetaBench

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music

Safe-first helper
Music Perception Question Answering Audio Music Understanding Symbolic Music Understanding Sheet Music Understanding +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_CC-BY-4.0_and_CC-BY-SA-4.0_by_source
Code license
GPL-3.0-or-later
License caution
ChoraleBricks-derived benchmark instances are CC BY 4.0, while ChoralSynth-derived instances are CC BY-SA 4.0. Original code, question templates, ontology, and configurations are GPL-3.0-or-later. Model outputs and inference logs are provided without a separate license. Source recordings and scores are downloaded separately and retain the upstream dataset terms; select the license by each benchmark file's documented source rather than treating the repository as uniformly licensed.
Download notes
The public repository releases the evaluation and benchmark-generation pipeline, question templates, ontology, configurations, inference logs, and pre-generated benchmark instances. The paper's main ChoraleBricks instance and its ChoralSynth validation instance each contain 300 five-option questions balanced across audio, symbolic, and sheet-image modalities; 20% use "none of the other options" as the correct answer. Items test pitch, rhythm, and harmony perception. The helper downloads official documentation, licenses, configs, and both sub-500 KB TSV instances by default. Cloning the approximately 49 MB GitHub repository is opt-in. Source audio and scores are not stored in that repository and must be obtained from ChoraleBricks or ChoralSynth under their respective terms.
Safe-first helperscripts/download/musica_metabench.sh
View helper
Audio understanding, generation & events

MusicCaps

MusicCaps: A Dataset of Music Captions

Safe-first helper
Music Captioning Text To Music Evaluation Music Understanding Audio Language Modeling
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
not_applicable
License caution
The Hugging Face card and Kaggle-converted metadata list CC BY-SA 4.0 for the annotation CSV. The referenced media are 10-second clips from AudioSet/YouTube, so original media copyright, platform terms, and availability still apply.
Download notes
The public release is a CSV of 5,521 music-text pairs with YouTube IDs, segment timestamps, AudioSet labels, aspect lists, and musician-written captions. Raw audio is not mirrored by the dataset and must be reconstructed from YouTube/AudioSet subject to upstream availability and terms.
Safe-first helperscripts/download/musiccaps.sh
View helper
Music

MusicNet

MusicNet: A Dataset for Music Transcription and Multi-label Classification

Safe-first helper
Music Transcription Multi Label Music Classification Note Onset Labeling Musical Instrument Recognition +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
Zenodo lists CC BY 4.0 for the MusicNet release. The record says audio recordings are Creative Commons licensed and Public Domain performances from the Isabella Stewart Gardner Museum, the European Archive Foundation, and Musopen, with per-recording provenance described in the metadata.
Download notes
The helper saves the Zenodo record JSON and 44 KiB metadata CSV by default. Reference MIDI files are a small opt-in download, while the full audio/label archive is about 10.3 GiB and requires MUSICNET_DOWNLOAD_AUDIO=1.
Safe-first helperscripts/download/musicnet.sh
View helper
Enhancement, separation & quality

NISQA

NISQA Speech Quality Corpus

Safe-first helper
Speech Quality Assessment Mean Opinion Score Prediction Non Intrusive Speech Quality Multidimensional Speech Quality +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
MIT
License caution
The GitHub README says the corpus is provided under the original terms of the source speech and noise samples, generally non-commercial research with some subsets permitting commercial use; individual README/license files inside the archive should be checked. The Zenodo record reports license id other-at, and model weights are CC BY-NC-SA 4.0.
Download notes
The helper downloads the official README, corpus wiki markdown, model-weight license, and Zenodo record JSON by default. The full NISQA_Corpus.zip archive is about 15.9 GB and requires NISQA_DOWNLOAD_CORPUS=1.
Safe-first helperscripts/download/nisqa.sh
View helper
Speech recognition

NOTSOFAR-1

NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription

Safe-first helper
Distant Automatic Speech Recognition Speaker Attributed Automatic Speech Recognition Speaker Diarization Continuous Speech Separation +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
MIT
License caution
Microsoft applies CC BY 4.0 to the released data and MIT to the baseline repository. The repository notes that challenge-only Dev-set-2 is not part of the current open release and is restricted to publications about systems developed during the challenge; use the current public subsets for new work.
Download notes
The public English recorded-meeting release currently documents 237 meetings averaging six minutes across 30 conference rooms, with 4-8 attendees and 35 speakers. It provides train, dev, an 80-meeting eval-small set matching the challenge evaluation set, and a 129-meeting eval-full set, with ground truth available for both evaluation releases. The original challenge and paper describe roughly 280 meetings; the post-challenge open release removes Dev-set-2 for legal and quality reasons and adds eval-full, so benchmark versions must be reported. Single-channel and known-geometry seven-channel tracks use speaker-attributed tcpWER for ranking. The helper downloads only official documentation and license files by default; use the repository's versioned download utilities for selected audio subsets. The Hugging Face repository has very large historical storage and must not be snapshotted wholesale.
Safe-first helperscripts/download/notsofar_1.sh
View helper
Music

NSynth

NSynth: Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders

Safe-first helper
Audio Synthesis Music Synthesis Musical Instrument Modeling Timbre Modeling
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
Apache-2.0
License caution
Official Magenta dataset page lists the dataset under CC BY 4.0; Magenta code is Apache-2.0.
Download notes
The full TFDS download is about 73 GiB. The helper saves the official dataset page by default and requires NSYNTH_DOWNLOAD_ARCHIVES=1 before downloading split archives.
Safe-first helperscripts/download/nsynth.sh
View helper
Audiovisual & cross-modal

Omni-Cloze

Omni-Cloze: Omni Detailed Captioning Benchmark

Safe-first helper
Detailed Audio Captioning Detailed Visual Captioning Detailed Audio Visual Captioning Cloze Question Answering +1 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
Neither the official Hugging Face card nor the GitHub repository states a data or code license, and GitHub reports no detected license. The release includes video files without documented source-media provenance or reuse terms; obtain clarification and review media rights before redistribution or commercial use.
Download notes
The public, ungated release contains 2,320 audio-visual files across nine main domains and 47 subcategories, with approximately 70,000 fine-grained cloze blanks. The helper downloads official documentation and evaluation scripts by default; the 25.1 MB JSONL metadata and Hugging Face snapshot are separate opt-ins. The current repository files total about 6.1 GB, while the Hugging Face API reports about 11.3 GB of repository storage including history. Qwen3.5-Omni evaluates detailed audio-visual captioning on Omni-Cloze in section 5.1.4, Table 7.
Safe-first helperscripts/download/omni_cloze.sh
View helper
Audiovisual & cross-modal

OmniBench

OmniBench: Towards the Future of Universal Omni-Language Models

Safe-first helper
Omni Modal Question Answering Audio Visual Question Answering Tri Modal Reasoning Speech Understanding +2 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The official dataset card has no license field, and the official repository has no LICENSE file or GitHub-detected license. The project page's footer links CC BY-SA 4.0 but does not clearly state that it covers the benchmark data, code, or component media. Obtain clarification and review source-image/audio rights before redistribution or commercial use.
Download notes
The public, ungated release contains 1,142 four-choice questions that jointly pair an image, audio, and text prompt across seven task types. Audio spans speech, sound events, and music. The helper downloads official documentation and repository metadata by default; the Hugging Face card reports approximately 1.26 GB of downloads, so the media-bearing Parquet snapshot requires OMNIBENCH_DOWNLOAD_HF=1. OPOD evaluates OmniBench as its omni-modal benchmark and reports accuracy in the Experiments benchmark block of arXiv:2607.20918.
Safe-first helperscripts/download/omnibench.sh
View helper
Audiovisual & cross-modal

OmniGAIA

OmniGAIA: Towards Native Omni-Modal AI Agents

Safe-first helper
Omni Modal Question Answering Audio Visual Reasoning Multi Hop Reasoning Tool Use +2 more
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
apache-2.0
Code license
MIT
License caution
The Hugging Face card declares Apache-2.0 and the GitHub repository declares MIT. The benchmark curates media from FineVideo, LongVideoBench, LongVideo-Reason, COCO 2017, and other Hugging Face sources; verify component media rights and attribution requirements before redistribution or commercial use.
Download notes
The public, ungated release contains 360 English test tasks with audio, image, and video inputs across nine domains. The helper downloads official documentation and the lightweight test metadata JSON by default; the Hugging Face API reports about 9.9 GB of repository storage, so the full media snapshot requires OMNIGAIA_DOWNLOAD_HF=1. Qwen3.5-Omni reports OmniGAIA without a thinking prompt or answer-tag formatting in section 5.1.4, Table 7.
Safe-first helperscripts/download/omnigaia.sh
View helper
Audiovisual & cross-modal

OmniRetriever-Bench

OmniRetriever-Bench: 12-Direction Audio-Video-Text Retrieval Benchmark

Safe-first helper
Audio Text Retrieval Audio Video Retrieval Audio Video Text Retrieval Cross Modal Retrieval +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_research_use
Code license
not_specified
License caution
The Hugging Face card labels the annotations Apache-2.0 but adds a biometric-identification, profiling, and surveillance prohibition, while the paper calls the release a custom research-use license. Treat the additional restriction as nonstandard and controlling pending a standalone license file. Underlying TikTok media remains owned by its uploaders and governed by platform terms. The evaluator repository has no detected license.
Download notes
The public, ungated release contains a 923,982-byte CSV with 3,782 held-out English-captioned audio-video-text triples and evaluates six single-modal plus six dual-modal retrieval directions. The helper downloads the official dataset card, CSV annotations, evaluator README, and paper page. Media is not redistributed; each row contains a TikTok source URL and clip interval, so availability depends on the source platform and users must obtain media themselves under applicable platform and uploader terms.
Safe-first helperscripts/download/omniretriever_bench.sh
View helper
Audiovisual & cross-modal

OmniVideoBench

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

Manual or gated
Audio Visual Question Answering Audio Visual Reasoning Long Video Understanding Temporal Reasoning +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
conflicting_CC-BY-NC-SA-4.0_and_CC-BY-NC-ND-4.0
Code license
not_specified
License caution
The official GitHub README says CC BY-NC-SA 4.0, while the Hugging Face card metadata declares CC BY-NC-ND 4.0 and the access form additionally requires non-commercial research use and no redistribution without permission. Apply the stricter gated terms pending clarification. The authors explicitly do not own the raw-video copyrights, so source-media rights remain separate; the GitHub repository has no detected license for its evaluation code.
Download notes
The gated Hugging Face release contains 628 English videos lasting from several seconds to 30 minutes and 1,000 manually verified question-answer pairs. Every question requires complementary audio and visual evidence and includes atomic step-by-step reasoning annotations; the paper reports 762 speech, 147 sound, and 91 music questions across 13 reasoning types. The helper downloads public repository documentation and API metadata by default. The Hugging Face API reports approximately 114 GB of repository storage, so the complete snapshot requires approval through the dataset questionnaire, authentication, and OMNIVIDEOBENCH_DOWNLOAD_HF=1. OmniReasoner evaluates OmniVideoBench in its main audio-visual benchmark table and reports duration- and audio-type-specific results.
Safe-first helperscripts/download/omnivideobench.sh
View helper
Speech recognition

Opencpop-test

Opencpop

Manual or gated
Singing Voice Transcription Mandarin Singing Voice
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_specified
License caution
License page is spelled /liscense/ on the official site.
Safe-first helperscripts/download/opencpop_test.sh
View helper
Music

OpenMIC-2018

OpenMIC-2018: An Open Dataset for Multiple Instrument Recognition

Safe-first helper
Musical Instrument Recognition Music Auto Tagging Multi Label Audio Classification Music Information Retrieval
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
Zenodo says Spotify AB releases the dataset under CC BY 4.0 and includes full license terms in the archive. The included metadata contains licenses for each audio recording, so check per-track metadata before redistribution or commercial use.
Download notes
The Zenodo archive is about 2.6 GiB and contains 10-second OGG clips, VGGish features, crowd-sourced labels, metadata with per-recording licenses, and train/test partitions. The helper saves the Zenodo record JSON and official README by default and requires OPENMIC_DOWNLOAD_ARCHIVE=1 before downloading the full archive.
Safe-first helperscripts/download/openmic_2018.sh
View helper
Speech generation

OpenSTBench

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

Safe-first helper
Speech To Text Translation Evaluation Speech To Speech Translation Evaluation Streaming Speech Translation Evaluation Translation Quality Evaluation +6 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other
Code license
MIT_with_CC-BY-SA-4.0_adapted_components
License caution
The paired-set card marks the dataset license as "other" and says users must comply with the original data and synthesis-component terms. LibriTTS is CC BY 4.0, but neither the repository's MIT license nor the paper's CC BY-NC-SA 4.0 license establishes standalone reuse terms for all translated metadata or Qwen3-TTS-generated reference audio. OpenSTBench's original code is MIT, while adapted SimulEval latency components are CC BY-SA 4.0. MSLT, RAVDESS, MCAE-SPPS, NonverbalTTS, and SynParaSpeech retain their own access and license terms.
Download notes
OpenSTBench provides a public evaluation package for offline and streaming speech-to-text and speech-to-speech translation. The paper's section 4.2 evaluates separate public source datasets for translation, speech quality, emotion, paralinguistics, temporal consistency, and latency, and constructs a 300-sample, 35-speaker LibriTTS-based paired set for speaker preservation. That paired set is public and ungated on Hugging Face and includes original and prompt LibriTTS audio, translated text, and Qwen3-TTS-synthesized reference speech. The helper downloads the paper, repository documentation, license notices, dataset card, and Hugging Face API metadata by default. Cloning the evaluation toolkit is opt-in, and downloading the approximately 511 MiB paired-set snapshot requires OPENSTBENCH_DOWNLOAD_PAIRED_SET=1. The helper does not fetch the other component datasets.
Safe-first helperscripts/download/openstbench.sh
View helper
Audiovisual & cross-modal

OV-MERD

OV-MERD: Open-Vocabulary Multimodal Emotion Recognition Dataset

Manual or gated
Open Vocabulary Multimodal Emotion Recognition Audio Visual Emotion Understanding Acoustic Emotion Cue Reasoning Free Form Multi Label Emotion Prediction +1 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0_with_additional_gated_terms
Code license
Apache-2.0_with_noncommercial_notice
License caution
The paper and Hugging Face card identify OV-MERD as CC BY-NC 4.0. MER2025's gated terms further limit use to academic research and non-commercial purposes, prohibit distribution of the dataset or derivative annotation/label files to third parties, and prohibit modification without prior written consent. The OV-MER repository includes Apache-2.0 code terms but also calls the service a non-commercial research preview. Source clips derive from MER2023 movie and television media, so underlying media rights remain separate.
Download notes
OV-MERD extends a consented subset of MER2023 movie and television clips with human-checked acoustic and visual clues, merged multimodal descriptions, and open-vocabulary emotion labels. The paper reports 236 emotion categories, one to nine labels per sample (most have two to four), and clips that are mostly one to four seconds long. The current official repository points to the gated MER2025 release, which bundles OV-MERD label and description tables with audio, video, subtitles, and face features for the broader challenge corpus. The Hugging Face API reports approximately 442 GB of repository storage. The helper downloads only public official documentation, repository metadata, the paper landing page, and the Hugging Face API response; users must request access and download any benchmark files manually.
Safe-first helperscripts/download/ov_merd.sh
View helper
Speech recognition

Pansori-TEDxKR

Pansori-TEDxKR: Korean Speech Corpus Generated from Korean Language TEDx Talks

Safe-first helper
Korean Asr Speech Transcription Subtitle Aligned Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_specified
License caution
OpenSLR lists Creative Commons BY-NC-ND 4.0. The corpus is derived from TEDx talks, so downstream use should also respect TED/TEDx source terms and original media rights.
Download notes
OpenSLR SLR58 hosts about 3 hours of Korean TEDx speech from 41 speakers with subtitle-boundary segmentation and manual alignment checks. The helper downloads the OpenSLR page, about page, info, and checksum by default; the 174 MB corpus archive is opt-in.
Safe-first helperscripts/download/pansori_tedxkr.sh
View helper
Speaker, identity & emotion

ParaPairAudioBench

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

Safe-first helper
Paralinguistic Pairwise Judgment Audio Language Model Judge Evaluation Speaking Style Comparison Speech Rate Comparison +4 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_non_commercial_and_gated
Code license
not_specified
License caution
The benchmark repository has no license file or detected GitHub license. Its README lists SVC as non-commercial academic research, EARS and Expresso as CC BY-NC 4.0, and LibriTTS as CC BY 4.0. Treat the pair annotations and builder code as rights-unspecified, obtain SVC approval for the affected age/gender rows, and preserve each source corpus's terms.
Download notes
The official repository releases pairwise JSON annotations and swapped-order variants for speech rate, emphasis, and style, plus source-pair metadata and builders for age and gender. The paper reports 5,175 pairs across five criteria, including tie and same-/cross-transcript conditions. The helper downloads official documentation and repository metadata by default; cloning the approximately 6 MB repository is opt-in and does not fetch underlying audio. Age and part of gender require manually approved SVC access; the remaining source audio comes from EARS, Expresso, and LibriTTS under their own download terms.
Safe-first helperscripts/download/parapair_audio_bench.sh
View helper
Speaker, identity & emotion

PartialEdit

PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing

Safe-first helper
Partial Speech Deepfake Detection Partial Speech Deepfake Localization Neural Speech Editing Detection Codec Artifact Analysis +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0_with_upstream_terms_and_partial_release
Code license
not_applicable
License caution
Zenodo declares CC BY 4.0 for the released record. The audio is derived from VCTK, whose official release is also CC BY 4.0, but users should preserve both provenances and review neural-editor output terms. The license does not make the withheld Audiobox-derived E3/E4 subsets public or grant rights to reconstruct them.
Download notes
The official Zenodo release contains the E1 (VoiceCraft), E1-Codec, E2 (SSR-Speech), and E2-Codec subsets derived from VCTK, plus the E1/E2 protocol CSV and modified-text metadata. The four audio archives total approximately 21.9 GB. The project and Zenodo description state that E3 (Audiobox-Speech) and E4 (Audiobox) cannot be released under Audiobox's license. The helper downloads official pages and Zenodo record metadata by default; the approximately 7.7 MB protocol/text metadata and large audio archives are separate opt-ins. SALMONN-2 section IV-E and Table VIII evaluate temporal spoof localization on PartialEdit using mean intersection over union.
Safe-first helperscripts/download/partialedit.sh
View helper
Speech recognition

PazaBench

PazaBench: A Benchmark for Automatic Speech Recognition on Low Resource Languages

Safe-first helper
Automatic Speech Recognition Low Resource Asr Multilingual Asr Asr Efficiency
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Not specified in the source record.
Code license
Not specified in the source record.
License caution
The Hugging Face Space card declares MIT, although its visible tree has no standalone LICENSE file. That declaration does not relicense the 11 source dataset groups: their displayed terms range from CC0 and CC BY to CC BY-SA, CC BY-NC-SA, gated CC BY-NC, Apache-2.0, and mixed terms. Review each upstream dataset card and access agreement before use, redistribution, or commercial deployment.
Download notes
The public Microsoft Research Africa leaderboard currently reports WER, CER, and inverse real-time factor across 61 African languages and 53 ASR/language models. It evaluates 16 kHz mono speech from 11 named public or community dataset groups, including African Next Voices, ALFFA, FLEURS, Common Voice 23.0, WAXAL, Naija Voices, and TWB Voice. The public Space exposes the leaderboard, submission interface, dataset inventory, and implementation, but its repository does not contain a standalone unified audio snapshot, frozen item manifest, or result CSV; obtain evaluation audio from each named upstream provider under that provider's access terms. The helper saves official documentation, Space metadata, implementation metadata, and the dataset inventory only; it does not download source corpora or model weights.
Safe-first helperscripts/download/pazabench.sh
View helper
Speech generation

PodEval

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Safe-first helper
Podcast Generation Evaluation Long Form Audio Generation Evaluation Dialogue Naturalness Evaluation Speech Quality Evaluation +3 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified_for_real_pod_manifest_and_linked_audio
Code license
MIT
License caution
The repository's MIT license covers the software and accompanying documentation, but no separate data license is stated for the Real-Pod manifest. The linked podcast recordings remain hosted by third parties and retain creator, publisher, platform, voice, music, and other media rights. The maintainers direct users to follow legal and ethical rules and use the reference data for research and education; public links do not grant redistribution or commercial-use rights.
Download notes
The public framework evaluates podcast generation across text, speech, and audio using objective metrics, LLM-based judging, and structured listening tests. Its Real-Pod reference manifest contains 51 topics across 17 categories, with one publicly accessible Apple Podcasts episode link per topic. The repository does not redistribute those recordings. The helper downloads official documentation, the small JSON manifest, license, repository metadata, and arXiv metadata by default; cloning the approximately 12 MB MIT toolkit is opt-in and still does not download podcast audio.
Safe-first helperscripts/download/podeval.sh
View helper
Speech recognition

Primewords Chinese Corpus Set 1

Safe-first helper
Automatic Speech Recognition Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_applicable
License caution
OpenSLR lists Attribution-NonCommercial-NoDerivatives 4.0 International and describes the corpus as free for academic use. Re-check upstream terms before redistribution or commercial use.
Download notes
OpenSLR hosts a 9.0 GiB Mandarin speech/transcript archive recorded from 296 native Chinese speakers. The helper saves the OpenSLR page by default and makes the full archive an explicit opt-in.
Safe-first helperscripts/download/primewords_chinese.sh
View helper
Enhancement, separation & quality

PVQD

Perceptual Voice Qualities Database

Safe-first helper
Clinical Voice Quality Assessment Pathological Voice Assessment Perceptual Voice Rating Cape V Prediction +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
not_applicable
License caution
Mendeley Data v4 and its DataCite DOI record declare CC BY 4.0. Preserve attribution and modification notices. The recordings contain human voices and clinical voice-quality information, and the release includes demographics; lawful and ethical handling of identifiable and health-related data remains necessary even though access is ungated.
Download notes
The public, ungated Mendeley Data v4 release contains 296 mono 44.1 kHz, 16-bit WAV recordings of sustained /a/ and /i/ vowels and six CAPE-V sentences, plus 13 XLSX files with demographics and experienced clinicians' CAPE-V and GRBAS ratings. The complete 310-file release is approximately 514.5 MiB. The helper downloads the official dataset page, DataCite DOI record, and live Mendeley file manifest by default. PVQD_DOWNLOAD_ANNOTATIONS=1 fetches the approximately 0.6 MiB PDF/XLSX documentation and labels; PVQD_DOWNLOAD_ALL=1 explicitly downloads the complete release including identifiable clinical voice recordings. The July 2026 voice-concept bottleneck paper uses an 80:20 speaker split and derives per-utterance segments with voice activity detection; those split and segment artifacts are not part of PVQD itself.
Safe-first helperscripts/download/pvqd.sh
View helper
Audiovisual & cross-modal

QIVD

Qualcomm Interactive Video Dataset: Can Vision-Language Models Answer Face to Face Questions in the Real-World?

Manual or gated
Situated Audio Visual Question Answering Real Time Audio Visual Understanding Spoken Query Understanding When To Answer Prediction +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
not_publicly_specified_account_terms_apply
Code license
not_applicable
License caution
No public dataset license text was found on the official landing page or in the paper on 2026-07-21. The paper says crowd contributors signed consent permitting research and commercial use of their video and audio, but that consent statement is not itself a downstream dataset license. Review and retain any terms shown in Qualcomm's account/download flow before use or redistribution.
Download notes
The official Qualcomm release page describes 2,900 short English video files across 13 semantic categories. Each clip contains raw audio with a spoken question, an annotated transcription, a text answer, and a timestamp indicating when enough context is available to answer. The helper saves the public landing page, then prints the manual Qualcomm account/download path because the release flow is JavaScript-driven and may present account-specific terms. Qwen3.5-Omni calls it Qualcomm IVD and evaluates audio-query interaction in section 5.1.4, Table 7.
Safe-first helperscripts/download/qivd.sh
View helper
Audiovisual & cross-modal

RAVDESS

The Ryerson Audio-Visual Database of Emotional Speech and Song

Safe-first helper
Speech Emotion Recognition Audio Visual Emotion Recognition Acted Emotional Speech Emotional Song Recognition
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-sa-4.0
Code license
not_applicable
License caution
Zenodo lists CC BY-NC-SA 4.0 for the dataset and says commercial licenses are available separately. The linked PLOS ONE paper itself is CC BY, but that article license is not the dataset license.
Download notes
The helper downloads the Zenodo metadata JSON by default. The audio-only speech and song ZIPs are about 215 MB and 198 MB respectively; video archives are much larger and are not downloaded by default.
Safe-first helperscripts/download/ravdess.sh
View helper
Speaker, identity & emotion

REAL-TSE Challenge

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Manual or gated
Target Speaker Extraction Online Target Speaker Extraction Offline Target Speaker Extraction Real Conversational Speech Separation +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
access_restricted_terms_not_publicly_specified
Code license
MIT
License caution
The official evaluation repository is MIT, but that code license does not license the challenge audio. The public challenge page restricts DEV/EVAL use to validation or final evaluation, forbids training and fine-tuning, and says access was limited to registered teams; it does not state a standalone dataset license. DEV and EVAL-1 derive from AISHELL-4, AliMeeting, AMI, DiPCo, and CHiME-6, whose upstream terms also remain applicable.
Download notes
The challenge reports 6,991 Mandarin/English mixture-enrollment trials over 2,309 real conversational mixtures and 11.3 hours of mixture audio. DEV has 1,991 REAL-T-derived pairs; EVAL-1 and EVAL-2 contain 2,000 seen and 3,000 unseen pairs without public references. The public helper saves official pages and repository metadata only. Organizers distributed password-protected data by email exclusively to registered teams, registration closed on May 31, 2026, and no public dataset URL is currently provided.
Safe-first helperscripts/download/real_tse.sh
View helper
Audio understanding, generation & events

RealDESED

RealDESED: A Real-World Domestic Sound Event Detection Benchmark

Safe-first helper
Sound Event Detection Domestic Sound Event Detection Temporal Audio Event Localization Multi Annotator Label Aggregation +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc0_or_cc-by_per_audio_file_and_cc-by-4.0_annotations
Code license
MIT
License caution
The Zenodo record is open and lists CC BY 4.0 at record level, but its description and official repository clarify that each audio recording and corresponding metadata row uses the per-file license recorded in metadata.csv, either CC0 or CC BY; remaining metadata and annotations are CC BY 4.0. Preserve creator attribution for CC BY recordings and consult metadata.csv before redistribution. The baseline repository is MIT.
Download notes
The public release contains 5,710 real-home recordings (37.85 hours) from 652 participants, with 64,430 temporal annotations across 15 domestic event classes. Multiple annotators label each recording, validation and test annotations receive additional review, and metadata covers recording devices, placement, environments, and scene descriptions. The helper saves official Zenodo/GitHub metadata and documentation by default; the train, validation, and test archives total approximately 8.74 GB and require explicit opt-in.
Safe-first helperscripts/download/realdesed.sh
View helper
Speaker, identity & emotion

RealMAN

RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization

Safe-first helper
Multichannel Speech Enhancement Speech Source Localization Dynamic Speaker Localization Variable Array Generalization +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
unspecified
License caution
The official repository README declares the dataset CC BY 4.0, but the repository has no detected standalone license file and GitHub reports no repository license. Treat CC BY 4.0 as covering the dataset release, preserve attribution and notices, and verify terms separately for the baseline code and any downstream derived artifacts.
Download notes
The public, ungated release contains 83.7 hours of 32-channel speech recorded in 32 scenes and 144.5 hours of background noise recorded in 31 scenes, with direct-path speech, transcriptions, source locations, and scene and speaker metadata. The repository lists approximately 531.4 GB of training data, 27.5 GB of validation mixtures, 39.3 GB of test mixtures, 158 GB of raw validation/test recordings, and 129 MB of dataset information; the Hugging Face API currently reports about 812.0 GB of repository storage. The helper downloads only official documentation and API metadata by default, while the complete snapshot requires explicit opt-in. A July 2026 geometry-aware enhancement paper evaluates RealMAN after resampling to 8 kHz and fixed four-second test segments, and separately tests array generalization on CHiME-4.
Safe-first helperscripts/download/realman.sh
View helper
Speech understanding & dialogue

RealSI

RealSI: Open Benchmark for Simultaneous Interpretation in Real-world Scenarios

Safe-first helper
Simultaneous Speech To Text Translation Simultaneous Speech To Speech Translation Long Form Speech Translation Streaming Latency Evaluation +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Not specified in the source record.
Code license
CC-BY-4.0_repository_license
License caution
The repository declares the dataset CC BY 4.0, but its README also says the authors do not own the copyright in the source videos and describes the release as annotations plus public video links for educational and informational use. The current repository additionally contains WAV derivatives. Treat CC BY 4.0 as covering author-created annotations, and independently review source-video copyright, platform terms, and the README disclaimer before using or redistributing audio.
Download notes
The official public repository contains timestamped Chinese-English and English-Chinese transcripts and translations for 20 natural, approximately 3-8 minute recordings across ten domains. The release totals 95 minutes 29 seconds and 778 utterance segments (431 En-to-Zh and 347 Zh-to-En), and its current tree also includes 20 WAV files. SimulS2ST-Omni section 4.1 and appendix D.3 reuse RealSI for sentence-level and long-form streaming S2TT/S2ST evaluation. The helper downloads official documentation, repository metadata, and all 20 lightweight JSON annotation files by default. The repository's WAV payload is about 351 MiB, so cloning the repository requires REALSI_CLONE_REPO=1.
Safe-first helperscripts/download/realsi.sh
View helper
Music

RUBATO

RUBATO: A Multi-Version Benchmark for Robust Music Transcription and Analysis

Safe-first helper
Automatic Music Transcription Beat Tracking Downbeat And Measure Tracking Local Key Estimation +2 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed_creative_commons
Code license
not_separately_specified
License caution
Zenodo labels the deposition CC BY 3.0, but metadata_versions.csv assigns per-recording terms that include CC0, CC BY, CC BY-SA, CC BY-ND, CC BY-NC, CC BY-NC-SA, CC BY-NC-ND, ambiguous "CC"/"CC0?", and EEF. Treat the per-recording field as controlling and review it before redistribution or commercial use; scripts inside the archive have no separate license statement.
Download notes
The open Zenodo v0.3 release contains 566 versions of 15 musical works (about 42.9 hours), including 22.05 kHz mono audio, aligned score MIDI/MuseScore/PDF/images, performance video, note/beat/measure/local-key/structure annotations, and audio-to-score warping paths. The helper downloads the 83 KB version metadata and Zenodo API record by default; the approximately 6.26 GB archive requires RUBATO_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/rubato.sh
View helper
Audio understanding, generation & events

RUL-MuchoMusic

RUL-MuchoMusic from RUListening / Are You Really Listening?

Safe-first helper
Music Question Answering Perceptual Music Understanding
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
RUL repo/HF card: mit; upstream MuChoMusic dataset: CC BY-SA 4.0
Code license
MIT
License caution
RUL-MuchoMusic derives from MuChoMusic; check upstream audio/source terms too.
Safe-first helperscripts/download/rul_muchomusic.sh
View helper
Speech recognition

S-DiverSe

S-DiverSe: Spanish Diverse Speech

Safe-first helper
Automatic Speech Recognition Pathological Speech Recognition Spanish Speech Recognition Asr Robustness +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The official repository has no license file or detected GitHub license. The paper is CC BY 4.0 on arXiv, but that does not license the annotations, reconstruction code, or linked source recordings. The manifest includes health-condition metadata and potentially identifiable speech/transcripts; review consent, privacy, research ethics, source rights, and platform terms before use or redistribution.
Download notes
The public repository releases a TSV manifest for 444 manually transcribed Spanish segments totaling 3.2 hours from 22 speakers with amyotrophic lateral sclerosis, Parkinson's disease, or post-stroke effects. Metadata includes speaker ID, sex, condition, intelligibility, source URL, timestamp, and duration. Audio is not redistributed; the repository provides a yt-dlp/ffmpeg reconstruction script for public video sources, whose availability and platform terms can change. The paper evaluates corpus-level and condition-specific WER after lowercasing, punctuation removal, digit expansion, and retention of filled pauses. The helper downloads annotations, documentation, and reconstruction code only; cloning the repository is opt-in.
Safe-first helperscripts/download/s_diverse.sh
View helper
Speaker, identity & emotion

SALMon

SALMon: A Suite for Acoustic Language Model Evaluation

Safe-first helper
Acoustic Consistency Evaluation Acoustic Semantic Alignment Speaker Consistency Speaker Gender Consistency +5 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The official repository and Hugging Face card license the SALMon dataset under CC BY-NC 4.0 because some source datasets are non-commercial. The release derives speech or acoustics from Expresso, VCTK, LJSpeech, FSD50K, and EchoThief and also includes Azure TTS output, so preserve component provenance and review upstream terms. The evaluation repository has no standalone license file and GitHub reports no detected code license.
Download notes
The public, ungated release contains 1,600 positive/negative audio pairs across eight configurations for acoustic consistency and acoustic-semantic alignment. It covers speaker identity, speaker gender, sentiment, background sound, and room impulse response, using a likelihood-ranking protocol for speech language models. The helper saves official documentation and repository metadata by default; set SALMON_DOWNLOAD_HF=1 for the approximately 562 MB Hugging Face snapshot. Google Drive provides the same benchmark as raw WAV files.
Safe-first helperscripts/download/salmon.sh
View helper
Representation & general suites

SEABAD

SEABAD: Southeast Asian Bird Activity Detection

Manual or gated
Bird Activity Detection Binary Bird Presence Detection Passive Acoustic Monitoring Tropical Bioacoustic Detection +1 more
Access pathZenodo
Upstream termsNot specified

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
Not specified in the source record.
Code license
MIT_claimed_in_readme_no_license_file
License caution
Zenodo's structured record declares CC BY 4.0 for the compilation, while the official repository says individual positive clips retain their Xeno-Canto licenses, including CC BY, CC BY-SA, CC BY-NC, and CC BY-NC-SA, and negative clips retain each source dataset's terms. Use the included provenance metadata and comply per recording; do not infer that the record-level license removes noncommercial, share-alike, or attribution requirements. The repository README calls the curation code MIT, but the repository currently has no LICENSE file and GitHub detects no license, so that code statement should be clarified before reuse.
Download notes
Zenodo v1.0.0 releases 50,000 balanced three-second, 16 kHz mono WAV clips: 25,000 bird-present clips spanning 1,677 Southeast Asian species and 25,000 bird-absent clips. Fixed stratified train, validation, and test splits contain 40,000, 5,000, and 5,000 clips. Positive recordings derive from Xeno-Canto; negatives derive from BirdVox-DCASE-20k, Freefield1010, Warblr, FSC-22, ESC-50, and DataSEC. The helper downloads official Zenodo, paper, and repository metadata by default. The single approximately 3.87 GiB mybad.zip archive requires explicit source-terms acknowledgment and an audio opt-in. DrongoNet section 7.1 evaluates the held-out 5,000-clip test split across five seeds and reports AUC, accuracy, recall, and F1.
Safe-first helperscripts/download/seabad.sh
View helper
Speech generation

Seed-TTS Eval

Seed-TTS objective zero-shot speech generation evaluation set

Safe-first helper
Zero Shot Text To Speech Voice Cloning Zero Shot Voice Conversion Speech Intelligibility Evaluation +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The official repository has no LICENSE file and GitHub reports no detected license. The objective set selects 1,000 English samples from Common Voice and 2,000 Mandarin samples from DiDiSpeech-2, so verify both component-source terms before redistribution or commercial use. Public access does not imply an open license.
Download notes
The helper downloads the official README and lightweight evaluation scripts by default; cloning the evaluation repository is opt-in. The public objective EN/ZH test set is linked through Google Drive and must be downloaded manually. The Seed-TTS paper states that the 100-sample-per-language subjective set is not released because of copyright restrictions.
Safe-first helperscripts/download/seed_tts_eval.sh
View helper
Speech generation

SILMA Open-source Arabic TTS Benchmark

Open-source Arabic TTS Benchmark

Safe-first helper
Speech Synthesis Text To Speech Arabic Speech Synthesis Dialectal Speech Synthesis +1 more
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0_declared_at_space_level
Code license
Apache-2.0_declared_at_space_level
License caution
The Space metadata declares Apache-2.0, but the repository contains no separate license file and does not document the provenance or reuse terms of its Arabic prompts. Generated clips may also remain subject to the licenses and acceptable-use terms of the evaluated TTS models. Treat the Space-level declaration as insufficient to resolve all prompt, voice, and model-output rights before redistribution or commercial use.
Download notes
SILMA's public, ungated Hugging Face Space provides fixed prompts and generated model audio for direct listening comparisons in Modern Standard Arabic, Egyptian Arabic, and Saudi Arabic. The current release has 10 MSA prompts across four systems, five Egyptian prompts across five systems, and five Saudi prompts across three systems. SILMA says this release deliberately prioritizes auditory assessment because WER, CER, speaker similarity, and UTMOS do not fully capture Arabic speech nuances; it does not publish aggregate human ratings or a formal automatic scoring protocol. The helper downloads the official README, application source, three prompt CSVs, and Space API metadata by default. Cloning the approximately 29.6 MB Space, including generated evaluation audio, requires SILMA_ARABIC_TTS_CLONE_SPACE=1.
Safe-first helperscripts/download/silma_open_source_arabic_tts.sh
View helper
Enhancement, separation & quality

SingMOS-Pro

SingMOS-Pro: A Comprehensive Benchmark for Singing Quality Assessment

Safe-first helper
Singing Quality Assessment Singing Mos Prediction Singing Voice Synthesis Evaluation Singing Voice Conversion Evaluation +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_with_upstream_source_terms
Code license
MIT
License caution
The Hugging Face card declares CC BY 4.0 and the official predictor repository is MIT. SingMOS-Pro includes ground truth and outputs from singing synthesis, conversion, resynthesis, and song-generation systems built from 12 source datasets. The paper and card do not provide a per-file license inventory, so retain source-corpus, performer, composition, model-output, and service terms rather than assuming the card clears all embedded audio rights.
Download notes
The public, ungated release contains 7,981 Chinese and Japanese singing clips totaling 11.15 hours, generated by 41 models across 12 source datasets. At least five experienced annotators rated every clip for overall MOS, and 4,155 clips additionally have lyrics and melody scores. The helper downloads official documentation, API metadata, split definitions, and system metadata by default. The approximately 11.6 MB sample/rating annotations require a separate opt-in, while the Hugging Face API reports approximately 2.83 GB of repository storage for the full audio snapshot.
Safe-first helperscripts/download/singmos_pro.sh
View helper
Enhancement, separation & quality

Slakh2100

Slakh2100: The Synthesized Lakh Dataset

Safe-first helper
Music Source Separation Multi Instrument Automatic Transcription Music Information Retrieval Synthetic Multitrack Music
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
MIT
License caution
The official Slakh page states that Slakh2100 and Flakh2100 are licensed under Creative Commons Attribution 4.0 International. The slakh-utils repository is MIT. Slakh is synthesized from Lakh MIDI Dataset v0.1, so keep source MIDI attribution/provenance in downstream use.
Download notes
The helper downloads the official Slakh page and utility README/LICENSE by default. Zenodo currently hosts the full Slakh2100 record plus a tiny prototyping subset; the helper can save the Zenodo landing pages with SLAKH_CHECK_ZENODO=1, but archive file selection should be made from the live Zenodo records because the full corpus is large.
Safe-first helperscripts/download/slakh2100.sh
View helper
Speech recognition

SLUE

SLUE: Spoken Language Understanding Evaluation

Safe-first helper
Spoken Language Understanding Automatic Speech Recognition Named Entity Recognition Named Entity Localization +1 more
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
SLUE-VoxPopuli is CC0; SLUE-VoxCeleb is CC BY 4.0; HF dataset metadata also advertises cc0-1.0 and cc-by-4.0 tags.
Code license
MIT
License caution
SLUE redistributes curated subsets of VoxPopuli and VoxCeleb plus task annotations. The VoxCeleb license notice says original and cropped video copyrights remain with the original owners.
Download notes
The helper downloads official toolkit docs and component license files by default. Hugging Face dataset snapshots are opt-in because they contain audio-derived benchmark data; use SLUE_DATASETS to choose slue, slue-phase-2, or both.
Safe-first helperscripts/download/slue.sh
View helper
Speech understanding & dialogue

SLURP

SLURP: A Spoken Language Understanding Resource Package

Safe-first helper
Spoken Language Understanding Intent Classification Slot Filling Semantic Entity Labeling
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Textual annotations are CC BY 4.0; Zenodo-hosted audio is non-commercial (Zenodo license id other-nc; GitHub README states CC BY-NC 4.0).
Code license
not_specified
License caution
GitHub README separates textual data and audio licensing and says a less-strict audio license may be available by contacting the dataset owner. No standalone code license file was found in the repo via raw GitHub on 2026-07-09.
Download notes
The helper clones or updates the official annotation/code repository by default and downloads Zenodo LICENSE.txt. Audio archives are about 3.9 GiB real plus 2.8 GiB synthetic, so audio download is an explicit opt-in.
Safe-first helperscripts/download/slurp.sh
View helper
Speech recognition

SmartGlasses Challenge 2026

SLT 2026 SmartGlasses Challenge: Egocentric Speech Interaction on AI Glasses

Manual or gated
Time Stamped Speaker Attributed Asr Meeting Transcription Speaker Diarization Spoken Language Understanding +3 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
access_restricted_terms_not_publicly_specified
Code license
not_specified
License caution
The challenge page does not publish a standalone dataset license and reserves organizer control over the participation terms. The public evaluation repository has no LICENSE file and GitHub reports no detected license, so its availability must not be treated as permission to redistribute the toolkit or corpus. Obtain permission from the organizers before reusing data or code beyond the challenge.
Download notes
The official challenge covers dyadic conversations and multi-party meetings recorded with a four-channel microphone array on smart glasses. Across train, development, and test, Track 1 reports 518 sessions and 44.95 hours, while Track 2 reports 196 sessions and 62.03 hours. Each track evaluates time-stamped speaker-attributed ASR with tcpWER and multiple-choice spoken-language understanding; public reference answers are limited to development data. The helper saves the official challenge page, public evaluation-toolkit documentation and metadata, and the July 2026 system paper. Corpus access required registration and agreement to challenge rules, download links were emailed to participating teams, registration closed in June 2026, and no current public corpus URL is provided.
Safe-first helperscripts/download/smartglasses_challenge_2026.sh
View helper
Enhancement, separation & quality

Song Describer Dataset

The Song Describer Dataset: A Corpus of Audio Captions for Music-and-Language Evaluation

Safe-first helper
Music Captioning Text To Music Generation Evaluation Music Text Retrieval Music Codec Reconstruction
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-SA-4.0
Code license
MIT
License caution
Zenodo and the repository declare CC BY-SA 4.0 for the dataset and MIT for code. The audio originates from MTG-Jamendo and retains per-track Creative Commons licenses recorded in audio_licenses.txt; preserve attribution and apply each track's terms in addition to the dataset license.
Download notes
The public, ungated release contains 706 approximately two-minute MTG-Jamendo tracks with 1,106 crowdsourced English captions. Its human-validated evaluation subset contains 546 tracks and 746 captions. The helper downloads the official annotations, per-track audio-license list, metadata, dataset documentation, and Zenodo record by default; the approximately 3.09 GiB audio archive requires explicit opt-in. Qwen-Music section 4.2.2 evaluates codec reconstruction on all 546 validated tracks. Qwen-Audio-VAE sections 4.1-4.2 also use the dataset for music reconstruction evaluation.
Safe-first helperscripts/download/song_describer.sh
View helper
Enhancement, separation & quality

SongEval

SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

Safe-first helper
Song Aesthetics Assessment Music Quality Prediction Full Song Generation Evaluation Human Preference Modeling
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-SA-4.0
Code license
Apache-2.0
License caution
The Hugging Face card declares CC BY-NC-SA 4.0 and the GitHub toolkit includes Apache-2.0. The paper says the audio includes outputs from five open and commercial song generators plus real and deliberately poor examples; generated-output, service, and any underlying music rights may still apply, and the release does not provide per-item provenance in metadata.jsonl. Review those rights before redistributing audio or relying on the card license alone.
Download notes
The public, ungated release contains 2,399 complete English and Chinese songs (about 140 hours) spanning nine mainstream genres. Sixteen musically trained annotators rated coherence, memorability, vocal breathing and phrasing naturalness, structural clarity, and overall musicality on five-point scales. The helper downloads official cards, API metadata, the approximately 1.27 MB rating JSONL, and toolkit documentation by default. The Hugging Face API reports approximately 16.1 GB of repository storage, so fetching all MP3 files requires explicit opt-in.
Safe-first helperscripts/download/songeval.sh
View helper
Audio understanding, generation & events

SongFormBench

Safe-first helper
Music Structure Analysis
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
cc-by-4.0
License caution
Dataset card lists CC BY 4.0. Audio reconstruction notes reference HarmonixSet and BigVGAN resources.
Safe-first helperscripts/download/songformbench.sh
View helper
Audiovisual & cross-modal

Sonic Seasoning

Sonic Seasoning: a Multi-Source Perceptual Dataset of Taste-Evoking Sounds

Safe-first helper
Taste From Audio Regression Taste Conditioned Music Retrieval Music Representation Evaluation Crossmodal Audio Perception
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0
Code license
Apache-2.0
License caution
The Hugging Face card declares CC BY-NC 4.0 for the compilation, ratings, splits, and redistributed clips, and the repository code is Apache-2.0. Audio provenance varies: the release combines pre-existing music, MusicGen outputs, and stimuli from 13 prior studies. The card limits redistributed clips to non-commercial research and instructs users to cite the originating studies; review underlying music, publication-stimulus, performer, and generated-output rights in addition to the compilation license.
Download notes
The public, ungated Hugging Face release contains 377 uniformly encoded WAV clips with normalized sweet, bitter, salty, sour, and spicy ratings; some subsets also provide temperature and emotion annotations. Its fixed split column contains 269 train, 68 validation, and 40 test items. The paper evaluates ten frozen audio encoders with a shared multi-task regression protocol and uses a 309-item pool for taste-conditioned retrieval. The helper downloads the official card, repository docs, API metadata, and approximately 34 KB ratings/path Parquet file by default. The Hugging Face API reports approximately 797 MB of repository storage, so the audio snapshot and the approximately 642 KB code repository are separate opt-ins. The repository README still calls the training dataset private, which conflicts with the current public, ungated Hugging Face release; the helper follows the live dataset state.
Safe-first helperscripts/download/sonic_seasoning.sh
View helper
Audio understanding, generation & events

SONYC-UST-V2

SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context

Safe-first helper
Audio Tagging Urban Sound Tagging Multilabel Sound Classification Spatiotemporal Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
Zenodo v2.3 lists CC BY 4.0 and the record README says the SONYC-UST dataset is offered under Creative Commons Attribution 4.0 International. Challenge rules also restrict private external data for reproducible task submissions.
Download notes
The helper downloads Zenodo record JSON plus README, annotations, taxonomy, and unpack script by default. Audio is split across 19 archive shards totaling about 12.8 GiB, so audio download is an explicit opt-in.
Safe-first helperscripts/download/sonyc_ust_v2.sh
View helper
Audio understanding, generation & events

Soroll-IA

Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring

Safe-first helper
Multi Label Audio Tagging Weakly Supervised Sound Event Classification Industrial Sound Monitoring Real World Environmental Audio Classification
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
Kaggle and the official benchmark README declare CC BY-NC 4.0, which prohibits commercial use and requires attribution. The benchmark repository has no LICENSE file or GitHub-detected license, so its code terms are unspecified.
Download notes
The public Kaggle release contains 7,396 FLAC clips (approximately 22 hours and 2.17 GB) recorded by two fixed sensing nodes in the Port of Valencia, with 26 weakly labeled industrial sound classes. It provides two ground-truth variants: a permissive non-cross-validated annotation set and a conservative set requiring agreement from at least two-thirds of annotators, plus five-fold assignments. The helper downloads official metadata, paper, and benchmark documentation by default; the audio and annotations require explicit opt-in and an authenticated Kaggle CLI.
Safe-first helperscripts/download/soroll_ia.sh
View helper
Audio understanding, generation & events

Spatial LibriSpeech

Spatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning

Safe-first helper
Sound Source Localization Source Distance Estimation Room Acoustics Estimation Direct To Reverberant Ratio Estimation +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
not_applicable
License caution
Apple's copyrights in the dataset are CC BY 4.0. Its license preserves upstream terms for LibriSpeech, Microsoft DNS noise, AudioSet, and CC0 Freesound material and says Apple makes no representations about those upstream rights; review component provenance before redistribution.
Download notes
The public Apple release provides more than 650 hours of 16 kHz first-order ambisonic speech and optional distractor noise, synthesized from LibriSpeech across more than 200,000 acoustic conditions and 8,000 synthetic rooms. Labels cover 3D source position and speaking direction, room geometry, C50, DRR, EDT, T20, and T30. The helper downloads official documentation by default; the approximately 365 MiB metadata Parquet file and individual FLAC samples are explicit opt-ins. The README still describes raw 19-channel audio as forthcoming and directs users to contact Apple rather than exposing a public download.
Safe-first helperscripts/download/spatial_librispeech.sh
View helper
Speech understanding & dialogue

SPEARBench

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Safe-first helper
Streaming Speech To Speech Evaluation Conversational Naturalness Evaluation Response Latency Evaluation Interruption Evaluation +4 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified_derived_from_seamless_interaction
Code license
MIT
License caution
The GitHub repository's MIT license covers the website and helper code. Neither the paper nor project page states a separate license for the downloadable benchmark audio, which is extracted from Seamless Interaction; public access does not establish redistribution or commercial-use rights, so verify the source corpus and package terms before reuse.
Download notes
The public project page links a SharePoint package containing 5,419 selected English question-answer dialogues from the Seamless Interaction development and test sets (37.33 hours including contexts, questions, and human answers). The helper downloads only official documentation, leaderboard metadata, an example submission CSV, and inference instructions; obtain the audio package manually from the project page and keep its directory structure intact.
Safe-first helperscripts/download/spearbench.sh
View helper
Speech understanding & dialogue

Speech Commands

Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Safe-first helper
Keyword Spotting Limited Vocabulary Speech Recognition Audio Classification
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
Google Research blog and Hugging Face dataset card list Creative Commons BY 4.0; HF card also asks users not to try to identify speakers.
Download notes
TensorFlow Datasets reports v0.02 as about 2.37 GiB download / 8.17 GiB extracted; avoid accidental full downloads in automated checks.
Safe-first helperscripts/download/speech_commands.sh
View helper
Speech generation

SpeechEditBench

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Safe-first helper
Instruction Guided Speech Editing Content Editing Speaker Editing Emotion Editing +5 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0_with_upstream_terms
Code license
Apache-2.0
License caution
The repository and Hugging Face card release contributor-authored code, documentation, metadata, and benchmark assets under Apache-2.0, while requiring compliance with source-corpus terms. The paper's appendix identifies mixed upstream conditions: CC BY sources, Apache-2.0 sources, CC BY-NC and CC BY-NC-SA sources, StoryTTS research-only restrictions, MagicData-RAMC custom terms, and the IEMOCAP access agreement. Apply those component restrictions to affected rows and audio rather than treating the aggregate label as overriding them.
Download notes
The public, ungated v1.1 release contains 4,700 English and Chinese source-instruction pairs and 5,400 audio files across seven atomic editing tasks and a compositional split. Evaluation separately measures target success, lexical-content preservation, and joint success. The helper downloads official documentation, release metadata, and the eight sample JSONL files by default. The Hugging Face API reports approximately 3.75 GB of repository storage, so audio requires SPEECH_EDIT_BENCH_DOWNLOAD_HF=1; cloning the roughly 3.2 MB evaluation repository is a separate opt-in.
Safe-first helperscripts/download/speech_edit_bench.sh
View helper
Speech understanding & dialogue

SpeechEQ

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

Safe-first helper
Spoken Dialogue Emotional Intelligence Paralinguistic Reasoning Multi Turn Speech Understanding Acoustic Multiple Choice +1 more
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_specified
License caution
The Hugging Face card has no license field, and neither the dataset repository nor the official code repository exposes a license file. The paper itself is CC BY-NC-SA 4.0, but that publication license must not be assumed to license the released benchmark audio, annotations, or code. Obtain clarification before redistribution or commercial use.
Download notes
The public, ungated English release contains 2,265 six-turn dialogues totaling 42 hours 23 minutes across 15 EQ-i 2.0 subscales. Evaluation selects between high- and low-EQ acoustic renditions of identical text at turns four and six, testing pitch, energy, rate, pauses, and sustained conversational context. The helper downloads official documentation and repository metadata by default. The five Parquet shards contain embedded audio and total approximately 2.45 GB, so the full Hugging Face snapshot requires SPEECHEQ_DOWNLOAD_HF=1.
Safe-first helperscripts/download/speecheq.sh
View helper
Speech understanding & dialogue

SpeechRole

SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents

Safe-first helper
Speech Role Playing Speech Dialogue Evaluation
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mit
Code license
not_specified
License caution
HF cards list MIT; GitHub repo has no separate detected license.
Safe-first helperscripts/download/speechrole.sh
View helper
Audiovisual & cross-modal

SpEmoC

SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark

Manual or gated
Speech Emotion Recognition Audio Visual Emotion Recognition Multimodal Emotion Recognition Cross Dataset Emotion Generalization
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_non_commercial_academic_eula
Code license
not_specified
License caution
The official agreement limits use to academic research, education, scientific publication, and other non-commercial research; prohibits redistribution and sharing download links; and leaves copyright in the source movie and television clips with their respective owners. The public benchmark repository has no LICENSE file or detected GitHub license, so code and public split/metadata terms remain unspecified.
Download notes
The benchmark contains 30,000 refined clips curated from 306,544 raw speaking segments across 3,100 English-language movies and television series, with aligned audio, visual, and text modalities and a near-balanced seven-emotion label distribution. The official project says the full dataset is available only after a requestor and faculty advisor or principal investigator sign the access agreement and submit it from an institutional email address. The helper saves public project, paper, repository, and agreement metadata, then prints the manual application steps; it never downloads restricted media.
Safe-first helperscripts/download/spemoc.sh
View helper
Speech recognition

SPGISpeech

SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

Manual or gated
Asr Fully Formatted Transcription Financial Speech Recognition
Access pathHugging Face
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
gated_academic_research_internal_use
Code license
not_specified
License caution
HF terms say the content is for academic research purposes and internal use only, prohibit redistribution without prior written consent, and include additional restrictions on creating competing databases/products and identifying individuals.
Download notes
Hugging Face access requires logging in and accepting Kensho terms. The HF card lists split sizes from 11 GiB for dev/test to 530 GiB for the L training subset; the helper refuses to download until SPGISPEECH_ACK_TERMS=1 is set.
Safe-first helperscripts/download/spgispeech.sh
View helper
Enhancement, separation & quality

SpInt

SpInt: A Spanish Speech Intelligibility Dataset

Safe-first helper
Speech Intelligibility Assessment Objective Intelligibility Metric Evaluation Speech Enhancement Evaluation Spanish Speech +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_for_released_SpInt_artifacts
Code license
Not specified in the source record.
License caution
Zenodo declares CC BY 4.0 for the released SpInt package. The record explicitly withholds the original clean Spanish Matrix Test recordings to respect their licensing conditions, so the release is not a standalone audio corpus and CC BY 4.0 must not be extended to those absent recordings.
Download notes
The public Zenodo v1.0 release provides behavioral intelligibility labels for 5,148 processed Spanish utterances, per-stimulus and listener-response metadata, complex speech-enhancement masks, noise signals, and a reconstruction script. The clean Spanish Matrix Test recordings are deliberately excluded because of their separate license; users must obtain that corpus independently to reconstruct the stimuli. The helper downloads the official record, README, reconstruction script, and approximately 2.7 MB JSON metadata by default. The approximately 807 MiB noise and 3.08 GiB mask archives require explicit opt-in.
Safe-first helperscripts/download/spint.sh
View helper
Speaker, identity & emotion

SpoofCeleb

SpoofCeleb: Speech Deepfake Detection and SASV in the Wild

Manual or gated
Speech Deepfake Detection Synthetic Speech Detection Spoofing Robust Speaker Verification In The Wild Anti Spoofing +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
cc-by-4.0_with_source_media_rights
Code license
not_applicable
License caution
The official project and Hugging Face tag state CC BY 4.0, while the project explicitly says copyright in the human speech files remains with the original video owners. Access is granted only after a request and agreement to Hugging Face terms. Treat the Creative Commons label as covering the released compilation and author contributions, not as clearance of every underlying video, voice, likeness, privacy, or generated-speech right.
Download notes
The gated author-owned Hugging Face release contains more than 2.5 million bona fide and synthetic utterances from 1,251 VoxCeleb1 speakers. Its 23 TTS attacks and speaker-disjoint train, validation, and evaluation partitions support both speech deepfake detection and spoofing-robust speaker verification protocols. The helper downloads the official project page, paper pages, and Hugging Face API metadata by default. The API reports approximately 268.3 GB of repository storage, so the snapshot requires author approval, local Hugging Face authentication, explicit acceptance of the terms, and both SPOOFCELEB_ACK_TERMS=1 and SPOOFCELEB_DOWNLOAD_HF=1. Section 4.1 of arXiv:2607.21127 evaluates balanced TTS attacks from all SpoofCeleb splits but does not publish its clipped row selection.
Safe-first helperscripts/download/spoofceleb.sh
View helper
Audio understanding, generation & events

SpurAudio

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

Safe-first helper
Few Shot Audio Classification Environmental Sound Classification Shortcut Learning Evaluation Background Shift Robustness +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0_with_mixed_upstream_terms
Code license
MIT
License caution
The Hugging Face card declares CC BY 4.0 for the released mixtures, but the benchmark derives audio from five upstream datasets with separate terms. In particular, ESC-50 and UrbanSound8K include non-commercial restrictions; review all component licenses before redistribution or commercial use. The evaluation repository is MIT.
Download notes
The public, ungated Hugging Face release contains 16,381 WAV files in train, validation, and test splits and reports approximately 7.69 GB of repository storage. It mixes foreground events from ESC-50, UrbanSound8K, VocalSound, WILD DESED, and USM with unrelated background textures to measure IID-versus-OOD shortcut reliance in 1-shot and 5-shot classification. The helper downloads official documentation and repository metadata by default; the audio snapshot is opt-in.
Safe-first helperscripts/download/spuraudio.sh
View helper
Speech recognition

ST-CMDS

ST-CMDS-20170001_1: Free ST Chinese Mandarin Corpus

Safe-first helper
Automatic Speech Recognition Mandarin Speech Recognition
Access pathOpenSLR
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_applicable
License caution
OpenSLR SLR38 lists Creative Commons BY-NC-ND 4.0 and asks users to cite the data as "ST-CMDS-20170001_1, Free ST Chinese Mandarin Corpus." Re-check upstream terms before redistribution or commercial use.
Download notes
OpenSLR hosts an 8.2 GiB archive with cellphone-recorded Mandarin speech, transcriptions, and metadata from 855 speakers and 102,600 utterances. The helper saves the OpenSLR page by default and requires ST_CMDS_DOWNLOAD_ARCHIVE=1 for the large archive.
Safe-first helperscripts/download/st_cmds.sh
View helper
Audio understanding, generation & events

STAR-Bench

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

Safe-first helper
Foundational Acoustic Perception Pitch Perception Loudness Perception Duration Perception +6 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0
Code license
conflicting_MIT_and_Apache-2.0_signals
License caution
The Hugging Face card declares the dataset CC BY-NC 4.0 and labels use as research-only. Preserve the terms and provenance of Clotho, FSD50K, STARSS23, and internet-sourced audio. The repository's LICENSE file is MIT, but its README badge says Apache-2.0 and describes both data and code as research-only; clarify the intended software terms before redistribution or commercial use.
Download notes
The public, ungated v1.0 release contains 2,353 English multiple-choice questions: 951 for foundational perception, 900 for temporal reasoning, and 502 for spatial reasoning. Its current metadata revises the v0.5 questions reported in the paper. Foundational audio is synthesized; temporal tasks draw on Clotho and FSD50K, while spatial tasks use STARSS23 and in-the-wild audio. The helper downloads official documentation, repository metadata, and the approximately 2 MB of JSON question metadata by default. The 2.74 GB audio archive and evaluation repository are separate opt-ins.
Safe-first helperscripts/download/star_bench.sh
View helper
Audio understanding, generation & events

STARSS22

STARSS22: Sony-TAu Realistic Spatial Soundscapes 2022

Safe-first helper
Sound Event Localization And Detection Sound Source Localization Spatial Audio Understanding Acoustic Source Tracking
Access pathZenodo
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT
Code license
not_specified
License caution
DataCite identifies the Zenodo dataset license as MIT. The baseline README says the repository and its contents use MIT, but the repository has no LICENSE file and the GitHub API reports no detected license; clarify code terms before redistribution.
Download notes
The public Zenodo v1.1.0 release contains the complete DCASE 2022 development and evaluation data (6.0 GB), with 121 labeled development recordings and 52 unlabeled evaluation recordings in 4-channel FOA and tetrahedral-microphone formats. The helper downloads official task, DOI, paper, and baseline metadata only; use the Zenodo record for the large audio archives.
Safe-first helperscripts/download/starss22.sh
View helper
Audiovisual & cross-modal

STARSS23

STARSS23: Sony-TAu Realistic Spatial Soundscapes 2023

Safe-first helper
Sound Event Localization And Detection Sound Source Localization Spatial Audio Understanding Audio Visual Sound Event Localization +2 more
Access pathZenodo
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
MIT
Code license
mixed
License caution
DataCite metadata for both Zenodo v1.0.0 and v1.1.0 identifies the dataset license as MIT. The audio-only baseline repository has no LICENSE file and the GitHub API reports no detected license; Sony's audiovisual baseline is MIT. DCASE says participants consented to the released 360-degree videos and visible faces were blurred, but the active record and privacy context should still be reviewed before sensitive visual use.
Download notes
The public Zenodo v1.1.0 release contains the complete DCASE 2023 development and evaluation data (16.3 GB), with 4-channel FOA and tetrahedral-microphone audio, synchronized 360-degree video for all but 12 audio clips, and temporal, direction-of-arrival, and distance annotations. DCASE 2024 Task 3 explicitly reuses STARSS23 rather than releasing a STARSS24 dataset, adding source-distance estimation to its audio-only and audiovisual evaluation protocol. The helper downloads both official task pages, DOI, paper, and baseline metadata only; use the Zenodo record for the large archive.
Safe-first helperscripts/download/starss23.sh
View helper
Audiovisual & cross-modal

StoryAD-QA

StoryAD-QA: Narrative-Comprehension Evaluation for Long-Form Audio Description

Safe-first helper
Long Form Audio Description Evaluation Narrative Comprehension Multiple Choice Question Answering Context Conditioned Reasoning +1 more
Access pathOfficial / other
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Not specified in the source record.
Code license
not_specified_pending_finalization
License caution
The repository's LICENSE is a placeholder that says release terms will be finalized and merely recommends MIT or Apache-2.0 for code and CC BY-NC 4.0 or another author-approved license for annotations. Those recommendations are not grants. The release contains no movie video, audio, frames, subtitles, or scripts; clip identifiers derive from ten MAD-Eval movies in LSMDC, and users must obtain lawful access to the underlying copyrighted media separately.
Download notes
The official ECCV 2026 repository releases 2,572 manually verified, five-option question-answer pairs across two tracks: 1,609 segment-only questions over 30-, 60-, 120-, and 240-second windows, and 963 context-conditioned questions using 30, 60, or 90 seconds of preceding context plus a 30-second target clip. The repository includes full annotation CSVs, question-only files, answer keys with rationales, prompts, and a local accuracy scorer. The paper reports 2,574 retained questions, while the repository README says its public release contains 2,572 after validation and cleanup; this entry uses the released-file total. The helper saves official documentation, license notice, summary, evaluator, and repository metadata by default; cloning the approximately 8 MB repository is opt-in.
Safe-first helperscripts/download/storyad_qa.sh
View helper
Speech recognition

SUPERB

SUPERB: Speech Processing Universal PERformance Benchmark

Safe-first helper
Speech Representation Evaluation Phoneme Recognition Automatic Speech Recognition Keyword Spotting +11 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mixed
Code license
Apache-2.0
License caution
S3PRL is mostly Apache-2.0, with repository notes indicating some Facebook-authored files are CC BY-NC. SUPERB tasks use multiple external datasets, so each component corpus must be downloaded and licensed through its own official source.
Download notes
SUPERB is a benchmark suite over multiple upstream corpora. The helper downloads official documentation/license files by default and only clones the S3PRL toolkit with SUPERB_CLONE_TOOLKIT=1; underlying corpora such as LibriSpeech, Speech Commands, VoxCeleb, and IEMOCAP keep their own access paths and licenses.
Safe-first helperscripts/download/superb.sh
View helper
Representation & general suites

Surge Pitch Dataset

Pitch Audio Dataset (Surge synthesizer)

Safe-first helper
Musical Pitch Classification Musical Pitch Ranking Synthesizer Preset Classification Audio Representation Evaluation
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
not_applicable
License caution
Zenodo declares CC BY 4.0 for the released dataset. The paper's publication license and the Surge synthesizer and preset licenses are separate; retain the dataset citation and review synthesizer/preset terms if regenerating or redistributing modified renders.
Download notes
The public, ungated Zenodo release contains 3.4 hours of four-second sounds generated from 2,084 human-authored Surge presets. Each preset is rendered at MIDI pitches 21 through 108 with velocity 64, a three-second note-on duration, and RMS normalization. The helper saves official record metadata and the paper by default; the approximately 7.58 GB tar archive requires explicit opt-in. NABEATs section 4.1 uses this release for downstream pitch classification under clean and constructed noisy conditions.
Safe-first helperscripts/download/surge_pitch.sh
View helper
Speech recognition

Switchboard

Switchboard-1 Release 2 conversational telephone speech corpus

Manual or gated
Automatic Speech Recognition Conversational Speech Recognition Telephone Speech Recognition Speaker Recognition
Access pathLDC / licensed
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_ldc_license
Code license
not_applicable
License caution
LDC catalog pages list membership/licensing terms and web-download access. Re-check the current LDC agreement before use or redistribution; do not treat benchmark recipes or transcripts as granting rights to the underlying audio.
Download notes
Switchboard-1 Release 2 and the 2000 HUB5 English Evaluation Speech set are distributed by LDC after login/licensing. The helper only prints official access steps because the corpus and standard evaluation audio are not publicly script-downloadable.
Safe-first helperscripts/download/switchboard.sh
View helper
Audiovisual & cross-modal

SyncBench

SyncBench: Causal-Semantic Audio-Visual Synchronization Evaluation for Generative Models

Safe-first helper
Audio Visual Synchronization Evaluation Audio Visual Generation Evaluation Video To Audio Evaluation Causal Semantic Alignment
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
MIT
License caution
The official code repository is MIT, but the Hugging Face dataset has no card, license tag, or license file. The paper is CC BY 4.0, which does not license the released generated videos or prompts. Model-output provider terms and any rights in prompt/source content may also apply; obtain clarification before redistribution or commercial use.
Download notes
The public, ungated Hugging Face repository contains 1,185 generated MP4s across six model directories plus two small evaluator-score JSON files, and reports approximately 12.9 GB of repository storage. The paper's section 4.4 defines SyncBench as 185 curated prompts across five audio-visual domains, while the current release has up to 200 numbered clips per model and does not include a dataset card or prompt manifest. The helper downloads official documentation, repository metadata, Hugging Face metadata, and the two lightweight score files by default; all videos require SYNCBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/syncbench.sh
View helper
Speaker, identity & emotion

SynSFX

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

Manual or gated
Non Speech Audio Deepfake Detection Synthetic Sound Effect Detection Unseen Generator Robustness Cross Domain Audio Forensics +1 more
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
academic_research_only
Code license
not_released
License caution
The official release page labels SynSFX "Academic research only" but does not publish a full standalone dataset license or redistribution terms. The authentic partition incorporates AudioCaps, Clotho, ESC-50, TACoS, and WavCaps material, so their source-media and per-clip terms remain applicable. The arXiv paper uses arXiv's perpetual non-exclusive publication license, which does not license the dataset. No official evaluation-code release was linked when checked.
Download notes
The official release page provides a direct private-storage download route for the academic-research-only corpus. The paper reports 43,374 clips totaling 178 hours: 16,922 authentic clips from AudioCaps, Clotho, ESC-50, TACoS, and WavCaps, plus 26,452 clips synthesized by seven text-to-audio systems. It also defines a 1,890-prompt controlled subset shared across all seven generators and train, validation, in-domain test, and unseen-generator test protocols. The official page rounds the duration to approximately 180 hours and inconsistently lists 26,460 synthetic clips, despite retaining the 43,374 total; this entry uses the internally consistent paper counts. The helper saves the official page and paper by default. Because the uncompressed-WAV archive is large and the publisher states research-only access, downloading it requires both SYNSFX_ACK_RESEARCH_ONLY=1 and SYNSFX_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/synsfx.sh
View helper
Speech recognition

Tadabur

Tadabur: A Large-Scale Quran Audio Dataset

Safe-first helper
Quranic Speech Recognition Arabic Speech Recognition Reciter Identification Word Level Alignment +2 more
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0_with_source_and_cultural_caveats
Code license
not_specified
License caution
The Hugging Face card declares CC BY-NC 4.0, limits use to research and education, and adds ethical and cultural expectations for respectful Qur'anic use. Audio was collected from public Qur'anic repositories and archives, but the release does not provide per-recording source-license provenance; confirm source rights before redistribution. The linked repository has no standalone LICENSE file, so no code license is claimed.
Download notes
The public, ungated Hugging Face release contains more than 365,000 verse-level Arabic recitation examples totaling over 1,400 hours from more than 600 reciters, with simple and Uthmani text plus automatically derived word timestamps. It exposes one training split rather than a fixed held-out evaluation split. The helper downloads official documentation and repository metadata by default; the Hugging Face API reports approximately 1.94 TB of repository storage, so the audio-bearing snapshot requires TADABUR_DOWNLOAD_HF=1.
Safe-first helperscripts/download/tadabur.sh
View helper
Audio understanding, generation & events

TAU Spatial Sound Events 2019

TAU Spatial Sound Events 2019: Ambisonic and Microphone Array

Safe-first helper
Sound Event Localization And Detection Sound Source Localization Sound Event Detection Spatial Audio Understanding
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_TAU_noncommercial
Code license
custom_TAU_noncommercial
License caution
Both Zenodo records use a custom Tampere University license permitting experimental non-commercial use with attribution and prohibiting commercial use. The source events derive from the DCASE 2016 Task 2 isolated-event dataset, so preserve that provenance. The baseline repository applies closely matching custom experimental/non-commercial terms to its code.
Download notes
The public DCASE 2019 Task 3 release has 400 development and 100 evaluation scenes of one minute each at 48 kHz, in matching 4-channel first-order Ambisonic and tetrahedral-microphone formats. Scenes use stationary sources from 11 classes, real impulse responses measured at 504 azimuth-elevation-distance combinations across five indoor locations, natural ambient noise, and zero or up to two overlapping events. Version 2 includes temporal and azimuth/elevation labels for both development and evaluation audio. The helper downloads official pages, papers, record metadata, READMEs, and license files by default; the approximately 10.1 GB audio release remains on Zenodo, while the roughly 490 KB label archives are an explicit opt-in.
Safe-first helperscripts/download/tau_spatial_sound_events_2019.sh
View helper
Audio understanding, generation & events

TAU Urban Acoustic Scenes 2019

TAU Urban Acoustic Scenes 2019 Development dataset

Safe-first helper
Acoustic Scene Classification Multi Device Audio Classification Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other_non_commercial
Code license
not_applicable
License caution
Zenodo record 2589280 lists license id "other-nc" without a more specific SPDX-style license. DCASE challenge and dataset terms should be checked before redistribution or commercial use.
Download notes
The helper downloads the Zenodo record JSON plus small doc/meta ZIPs by default. The 40-hour audio release is split across 21 ZIP files of roughly 1.3-1.8 GiB each, so audio download is an explicit opt-in.
Safe-first helperscripts/download/tau_asc_2019.sh
View helper
Audio understanding, generation & events

TAU Urban Acoustic Scenes 2020 Mobile

TAU Urban Acoustic Scenes 2020 Mobile: Development and Evaluation datasets

Safe-first helper
Acoustic Scene Classification Device Robust Audio Classification Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other_non_commercial
Code license
not_applicable
License caution
Both Zenodo development and evaluation records list "Other (Non-Commercial)" without a more specific SPDX-style license. DCASE challenge terms should be checked before redistribution or commercial use.
Download notes
The helper downloads Zenodo record JSON plus small doc/meta ZIPs by default. Development audio is split across 16 ZIP files totaling about 27.4 GiB; evaluation audio is split across 8 ZIP files totaling about 13.1 GiB, so both are explicit opt-ins.
Safe-first helperscripts/download/tau_asc_2020_mobile.sh
View helper
Audio understanding, generation & events

TAU Urban Acoustic Scenes 2022 Mobile

TAU Urban Acoustic Scenes 2022 Mobile Development and 2025 Evaluation datasets

Safe-first helper
Acoustic Scene Classification Device Robust Audio Classification Low Complexity Audio Classification
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
other_non_commercial
Code license
not_applicable
License caution
Both official Zenodo records list "Other (Non-Commercial)" without a more specific SPDX-style license. Review the record and DCASE task terms before redistribution or commercial use.
Download notes
The helper downloads both Zenodo record JSON files plus the small doc/meta ZIPs by default. The 64-hour development release contains about 25.6 GiB across 16 audio ZIPs; the DCASE 2025 evaluation release contains about 19.2 GiB across 12 audio ZIPs, so each audio collection is an explicit opt-in. DCASE 2025 Task 1 reuses a restricted subset of the 2022 development data with a new split and adds the 2025 evaluation release.
Safe-first helperscripts/download/tau_asc_2022_mobile.sh
View helper
Audio understanding, generation & events

TAU-NIGENS Spatial Sound Events 2020

Safe-first helper
Sound Event Localization And Detection Sound Source Localization Acoustic Source Tracking Sound Event Detection +1 more
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0
Code license
custom_TAU_License
License caution
Zenodo lists the dataset as CC BY-NC 4.0. Source events come from the separately released NIGENS database, so preserve its provenance and review its active terms. The baseline README places most code under a custom TAU License and only its metrics folder under MIT.
Download notes
The public v1.2 release is the complete DCASE 2020 Task 3 development and evaluation corpus: 600 development and 200 evaluation sound scenes of one minute each, in 4-channel FOA and tetrahedral-microphone formats at 24 kHz. It uses real room impulse responses from 15 enclosures, static and moving sources from 14 classes, up to two overlapping events, and direction-of-arrival trajectories plus onset/offset labels. The helper downloads official pages, paper, record metadata, and the 17 KB dataset README by default; the approximately 14.0 GB audio archives remain on Zenodo, while the roughly 1.7 MB label archives are an explicit opt-in.
Safe-first helperscripts/download/tau_nigens_sse_2020.sh
View helper
Audio understanding, generation & events

TAU-NIGENS Spatial Sound Events 2021

Safe-first helper
Sound Event Localization And Detection Sound Source Localization Acoustic Source Tracking Sound Event Detection +1 more
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-4.0
Code license
custom_TAU_License
License caution
Zenodo lists the dataset as CC BY-NC 4.0. Source events come from the separately released NIGENS database, so preserve its provenance and review its active terms. The baseline license allows experimental non-commercial use and prohibits commercial use.
Download notes
The public v1.1.0 release is the complete DCASE 2021 Task 3 development and evaluation corpus: 600 development and 200 evaluation sound scenes of one minute each, in 4-channel FOA and tetrahedral-microphone formats at 24 kHz. It adds directional non-target interferers, permits overlapping instances of the same target class, and supplies direction-of-arrival trajectories plus onset/offset labels for the development set. The helper downloads official pages, paper, record metadata, and the 23 KB dataset README by default; the approximately 14.2 GiB audio archives remain on Zenodo, while the roughly 1.8 MiB development labels are an explicit opt-in. The 200-file evaluation set intentionally has no public labels.
Safe-first helperscripts/download/tau_nigens_sse_2021.sh
View helper
Speech recognition

TEDx Spanish Corpus

TEDx Spanish Corpus: Audio and Transcripts in Spanish Taken from TEDx Talks

Safe-first helper
Spanish Asr Spontaneous Speech Recognition Speech Transcription
Access pathOpenSLR
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-nd-4.0
Code license
not_applicable
License caution
OpenSLR lists CC BY-NC-ND 4.0. The corpus is derived from TEDx Talks, so downstream use should also respect TED/TEDx source terms.
Download notes
OpenSLR SLR67 hosts a single 2.3 GiB archive with Spanish speech and transcripts from TEDx Talks. The helper saves the OpenSLR page by default and downloads the archive only with TEDX_SPANISH_DOWNLOAD_ARCHIVE=1.
Safe-first helperscripts/download/tedx_spanish.sh
View helper
Speaker, identity & emotion

TESS

Toronto Emotional Speech Set

Safe-first helper
Speech Emotion Recognition Acted Emotional Speech Auditory Emotion Perception
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_applicable
License caution
The official Borealis record lists CC BY-NC 4.0. The corpus contains identifiable human voices and permits only non-commercial reuse under that license; preserve attribution and review voice-data ethics for downstream use.
Download notes
The owner-hosted University of Toronto Dataverse release contains 2,800 WAV stimuli: 200 target words spoken by two English-speaking actresses aged 26 and 64 in seven acted emotions. The helper saves official dataset metadata by default and requires TESS_DOWNLOAD_AUDIO=1 before downloading the complete ZIP.
Safe-first helperscripts/download/tess.sh
View helper
Speech generation

Text to Audio Human Preference Benchmark

Rapidata Text to Audio Human Preference Benchmark

Safe-first helper
Text To Speech Evaluation Human Preference Evaluation Speech Naturalness Evaluation Speech Friendliness Evaluation
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified
Code license
not_applicable
License caution
The official dataset card and Hugging Face API declare no license. The rows include model labels, aggregate scores, individual votes, and annotator demographic fields; public access does not imply permission to redistribute or reuse those records, audio references, or generated outputs.
Download notes
The public, ungated release contains 4,269 pairwise comparison rows and about 32,000 human responses judging generated voices for friendliness and naturalness. It stores audio references as strings rather than embedding audio. The helper downloads the dataset card and API metadata by default; the approximately 0.8 MB repository snapshot is opt-in.
Safe-first helperscripts/download/rapidata_tts_preference.sh
View helper
Speech recognition

THCHS-30

THCHS-30: A Free Chinese Speech Corpus

Safe-first helper
Automatic Speech Recognition Mandarin Speech Recognition Noisy Speech Recognition
Access pathOpenSLR
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Apache-2.0
Code license
not_applicable
License caution
OpenSLR lists Apache License v2.0 and the resource description says the database is free to academic users. The original CSLT URL linked from OpenSLR returned 404 when checked, so use the current OpenSLR page and paper for access/provenance.
Download notes
OpenSLR SLR18 hosts a 6.4 GiB speech/transcript archive, a 1.9 GiB 0 dB noisy test archive, and a 24 MiB supplementary resource archive with lexicon/noise samples. The helper saves the OpenSLR page by default and only downloads selected archives through THCHS30_DOWNLOAD_PARTS.
Safe-first helperscripts/download/thchs_30.sh
View helper
Speaker, identity & emotion

TidyVoice

TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice

Manual or gated
Multilingual Speaker Verification Cross Lingual Speaker Verification Speaker Recognition Language Mismatch Robustness +1 more
Access pathOfficial / other
Upstream termsOpen / attribution signals

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
CC0-1.0_with_use_restrictions
Code license
Apache-2.0
License caution
Mozilla Data Collective labels TidyVoiceX_ASV CC0-1.0 but also states that it must only be used for speaker verification and forbids speaker identification or attempts to recover speaker identity. Treat those owner-stated usage rules and the current Common Voice terms as binding access conditions despite the permissive license label. Apache-2.0 covers the WeSpeaker baseline repository, not any separate model or derived artifact rights.
Download notes
The public Mozilla Data Collective release contains 321,711 utterances (457 hours) from 4,474 multilingual speakers across 40 languages, with training and development splits, pseudonymized speaker IDs, language metadata, and same-/cross-language target and non-target trial lists. The current archive is approximately 36.72 GB. Download requires a Data Collective account and API key, so the helper saves official public documentation and prints the owner-supported access path without accepting credentials or fetching audio. The January paper also describes the broader Tidy-M monolingual condition across 81 languages; this entry's reproducible download pointer is the released TidyVoiceX_ASV challenge package. AMECxSV section 4.1 evaluates a deterministic speaker-disjoint split derived from the 12-million-trial TidyVoiceX development protocol, not the challenge's hidden official evaluation set.
Safe-first helperscripts/download/tidyvoice.sh
View helper
Audio understanding, generation & events

TimeGround-1M

Safe-first helper
Temporal Audio Grounding Temporal Audio Localization Timestamped Audio Description Timestamped Audio Summarization +1 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-3.0_on_hugging_face_card
Code license
not_specified
License caution
The official dataset card declares CC BY 3.0. Its recordings are selected from English YODAS2 YouTube-derived shards, so source-video rights, availability, attribution, and platform terms still require review. No separate license for generation or evaluation code is provided on the dataset card.
Download notes
The public, ungated English release contains separate train and test splits for temporal localization, temporal description, timed summarization, and recording-level nested annotations. The official card reports about 59,000 training and 4,200 test recordings totaling roughly 14,200 hours across duration buckets from under 10 minutes to 120 minutes. The GigaChat 3.1 Audio paper evaluates these generated tasks by duration bucket in section 4.1. The helper downloads the dataset card, repository API metadata, paper page, and model card by default; the Hugging Face API reports about 1.50 TB of repository storage, so the full snapshot requires explicit opt-in.
Safe-first helperscripts/download/timeground_1m.sh
View helper
Speech recognition

TIMIT

TIMIT Acoustic-Phonetic Continuous Speech Corpus

Manual or gated
Automatic Speech Recognition Phone Recognition Acoustic Phonetic Analysis Speaker Dialect Coverage
Access pathLDC / licensed
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
custom_ldc_license
Code license
not_applicable
License caution
LDC catalog pages list licensing instructions for Subscription/Standard Members and Non-Members, web download media, and fee visibility after login. Portions are copyright 1993 Trustees of the University of Pennsylvania; consult the current LDC agreement before use or redistribution.
Download notes
LDC distributes TIMIT by web download after login/licensing. The helper only prints official access steps because the corpus is paid/licensed and not publicly script-downloadable.
Safe-first helperscripts/download/timit.sh
View helper
Speech recognition

TORGO

TORGO Database of Acoustic and Articulatory Speech from Speakers with Dysarthria

Safe-first helper
Dysarthria Detection Pathological Speech Recognition Speech Intelligibility Assessment Acoustic Articulatory Modeling
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_academic_nonprofit_only
Code license
not_applicable
License caution
The owner page says use is free for academic, non-profit purposes and requires citation of at least one listed TORGO paper. It supplies the data as-is and does not identify a standard open-data license or grant commercial use. The recordings contain identifiable voices and disability and health information, so ethical and privacy review remains necessary.
Download notes
The public University of Toronto release contains aligned 16 kHz acoustic recordings and measured 3D articulatory features from eight English speakers with cerebral palsy or amyotrophic lateral sclerosis and seven matched controls. Stimuli include non-words, isolated words, restricted sentences, and spontaneous descriptions. Four BZip2 archives are organized as female dysarthric (F), female control (FC), male dysarthric (M), and male control (MC); they total approximately 8.9 GiB compressed and 18 GB uncompressed. The helper downloads the official page, correction spreadsheet, and coil-location documentation by default. Archive downloads require explicit terms acknowledgment and a selected group list. The July 2026 voice-concept bottleneck paper evaluates only headMic recordings with leave-one-speaker-out cross-validation; its exact derived split is not separately released.
Safe-first helperscripts/download/torgo.sh
View helper
Audio understanding, generation & events

TREA

Temporal Reasoning Evaluation of Audio

Safe-first helper
Audio Question Answering Temporal Audio Reasoning Audio Event Ordering Audio Event Counting +2 more
Access pathOfficial / other
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC0-1.0_repository_with_ESC-50_upstream_terms
Code license
CC0-1.0
License caution
The repository applies a CC0-1.0 LICENSE and GitHub detects CC0-1.0, but the paper states that every TREA audio file combines recordings from ESC-50, whose dataset is CC BY-NC 3.0 and whose ESC-10 subset clips are CC BY. Treat the restrictive upstream terms and clip attribution as surviving the derived release rather than assuming the repository-level CC0 waiver clears all source-audio rights.
Download notes
TREA is a public, ungated 600-item temporal-reasoning benchmark derived by combining ESC-50 clips. Its TREA-O, TREA-C, and TREA-D subsets each contain 200 ordering, counting, or duration questions. The July 2026 Audio-Zero paper evaluates both Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA alongside MMAU Test-mini and MMAR. The repository releases both four-option multiple-choice and open-text answer formats, audio, metadata, evaluation code, and uncertainty perturbation scripts. The helper downloads official documentation, repository metadata, license, paper page, and the lightweight CSV annotations by default; set TREA_CLONE_REPO=1 to clone the approximately 688 MiB GitHub repository and its audio.
Safe-first helperscripts/download/trea.sh
View helper
Speech generation

TTS Multilingual Test Set

MiniMaxAI TTS Multilingual Test Set

Safe-first helper
Multilingual Text To Speech Zero Shot Voice Cloning Cross Lingual Voice Cloning Speech Intelligibility Evaluation +2 more
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
not_applicable
License caution
The official Hugging Face dataset card lists CC BY-SA 4.0. Its 48 speaker prompts are selected from Mozilla Common Voice, whose data is CC0-1.0; retain benchmark attribution and share adaptations under the card's stated terms.
Download notes
The public, ungated Hugging Face repository contains 100 test sentences and two Common Voice-derived speaker prompts (one female and one male) for each of 24 languages. The helper downloads the official dataset card by default; the approximately 7.3 MB snapshot requires TTS_MULTILINGUAL_TEST_SET_DOWNLOAD_HF=1. Qwen3-TTS evaluates a 10-language subset for zero-shot multilingual and target-speaker generation, but the report does not identify the exact text rows used.
Safe-first helperscripts/download/tts_multilingual_test_set.sh
View helper
Audio understanding, generation & events

TUT Sound Events 2017

TUT Sound Events 2017: Sound Event Detection in Real-Life Audio

Safe-first helper
Sound Event Detection Polyphonic Sound Event Detection Temporal Audio Event Localization Street Sound Event Detection
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_noncommercial
Code license
not_applicable
License caution
Zenodo labels both releases Other (Non-Commercial), and each documentation archive contains a controlling EULA. Review that packaged agreement before use or redistribution; the generic Zenodo label is not a permissive Creative Commons grant.
Download notes
The public version-2 development release contains 24 street recordings totaling 1:32:08 with verified strong annotations for six overlapping event classes and an official four-fold cross-validation setup. The public evaluation release contains eight recordings totaling 29:09 and now includes reference metadata. DCASE 2017 Task 3 ranks systems by one-second segment-based error rate. The helper downloads Zenodo record JSON, documentation, and small annotation archives by default; the approximately 1.55 GiB of 24-bit, 44.1 kHz audio requires explicit opt-in.
Safe-first helperscripts/download/tut_sound_events_2017.sh
View helper
Audio understanding, generation & events

UrBAN

UrBAN: Urban Beehive Acoustics and PheNotyping Dataset

Safe-first helper
Environmental Sound Classification Beehive Acoustic Monitoring Colony Strength Regression Hive Health Monitoring +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
not_specified
License caution
The FRDR dataset record explicitly lists CC BY 4.0. The GitHub repository has no detected license, so the analysis notebooks and scripts should not be assumed to use the dataset license. The Scientific Data article itself is CC BY-NC-ND 4.0, distinct from the dataset terms.
Download notes
The public FRDR release contains longitudinal 2021-2022 raw 16 kHz beehive audio plus inspection, temperature, humidity, and weather metadata from a ten-hive Montréal rooftop apiary. The Scientific Data descriptor reports more than 3,000 hours, while the older FRDR record and repository README say more than 2,000 hours; the current FRDR landing page reports approximately 1.265 TB of files. Its benchmark protocols include random-split and hive-independent colony-strength regression. A 2026 follow-up evaluates modulation-tensorgram models on nine hives and emphasizes cross-hive generalization. The helper saves official landing pages, repository documentation, and API metadata only. Full data transfer remains a manual FRDR Globus workflow because of the corpus size and may require a Globus account and client.
Safe-first helperscripts/download/urban_beehive.sh
View helper
Audio understanding, generation & events

UrbanSound8K

UrbanSound8K: A Dataset and Taxonomy for Urban Sound Research

Safe-first helper
Urban Sound Classification Environmental Sound Classification Audio Tagging
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc
Code license
not_applicable
License caution
Zenodo lists CC BY-NC 4.0. The official Urban Sound site says UrbanSound/UrbanSound8K are free for non-commercial use under Creative Commons BY-NC 3.0; Freesound attributions are included in the dataset.
Download notes
The archive is about 6 GiB and contains 8732 WAV clips pre-sorted into 10 official folds. The helper downloads citation/license metadata by default and requires URBANSOUND8K_DOWNLOAD_AUDIO=1 for the full archive.
Safe-first helperscripts/download/urbansound8k.sh
View helper
Speech understanding & dialogue

URO-Bench-pro

URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models

Safe-first helper
Spoken Dialogue Model Evaluation Speech To Speech Evaluation
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
mit
Code license
MIT
License caution
Qwen uses the pro track. HF card and GitHub repo list MIT.
Safe-first helperscripts/download/uro_bench_pro.sh
View helper
Representation & general suites

User-Intent Queries (UIQ)

User-Intent Queries benchmark from Omni-Embed-Audio

Safe-first helper
User Intent Audio Retrieval Language Based Audio Retrieval Query Reformulation Robustness Exclusionary Query Understanding +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0
Code license
MIT
License caution
CC BY 4.0 applies only to the released UIQ text queries. AudioCaps, Clotho, and MECAT audio is not redistributed and retains its original source terms; the top-level Omni-Embed-Audio code repository is MIT.
Download notes
The public, ungated release contains 13,053 text-query records over the AudioCaps test, Clotho evaluation, and MECAT pools: question, imperative, tagging, paraphrase, and exclusionary negative variants. The helper downloads the approximately 12 MiB of query JSONL files plus the benchmark README and license; it does not download source audio. Fusion Embedding section 6.3 independently reuses UIQ on the 1,045-clip Clotho pool and reports only the four positive query formulations.
Safe-first helperscripts/download/uiq.sh
View helper
Speech generation

VCTK

CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit

Safe-first helper
Text To Speech Speech Synthesis Voice Cloning Multi Speaker Speech Synthesis +1 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
The official VCTK README and DataShare license_text identify Creative Commons Attribution 4.0 International. The newspaper text source was used with permission from Herald & Times Group.
Download notes
The official DataShare ZIP is about 10.94 GiB. The helper saves the official README and license text by default and requires VCTK_DOWNLOAD_ARCHIVE=1 before downloading the archive.
Safe-first helperscripts/download/vctk.sh
View helper
Audiovisual & cross-modal

VGGSound

VGGSound: A Large-scale Audio-Visual Dataset

Safe-first helper
Audio Visual Event Classification Audio Event Classification Audio Tagging
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
Official VGG page and repository license file list the dataset as CC BY 4.0 for commercial/research use, while copyright remains with original video owners. Re-check YouTube availability and upstream media terms before reconstructing clips.
Download notes
The official VGG page currently says the original dataset download links are no longer available from that website. The helper downloads the official CSV metadata, license, and optional pretrained model files only; it does not fetch or redistribute YouTube media.
Safe-first helperscripts/download/vggsound.sh
View helper
Audiovisual & cross-modal

Video-MME

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Safe-first helper
Audio Visual Question Answering Long Video Understanding Multimodal Reasoning Audio Enabled Video Understanding
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
custom_academic_research_only
Code license
not_specified
License caution
The official README prohibits commercial use and, without prior approval, distribution, publication, copying, dissemination, or modification of Video-MME in whole or in part. Video copyrights remain with their owners. The GitHub repository has no detected license; obtain approval and re-check source-video rights before reuse beyond the stated academic evaluation context.
Download notes
The public, ungated release contains 900 videos totaling 254 hours and 2,700 human-annotated question-answer pairs, with audio and subtitles available as evaluation modalities. The helper downloads only official documentation by default. The Hugging Face API reports about 389 GB of repository storage, so the media snapshot requires both VIDEO_MME_ACK_TERMS=1 and VIDEO_MME_DOWNLOAD_HF=1. Qwen3.5-Omni evaluates Video-MME with use_audio_in_video=True in section 5.1.4, Table 7.
Safe-first helperscripts/download/video_mme.sh
View helper
Audiovisual & cross-modal

video-SALMONN 2 Caption Benchmark

video-SALMONN 2 Human-Annotated Audio-Visual Caption Benchmark

Safe-first helper
Audio Visual Video Captioning Detailed Video Captioning Audio Visual Event Understanding Caption Completeness Evaluation +1 more
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
apache-2.0_card_label
Code license
Apache-2.0
License caution
The Hugging Face card labels the dataset Apache-2.0, and the official GitHub repository contains an Apache-2.0 LICENSE. The release does not document per-video provenance or underlying media licenses, so the card label must not be assumed to clear third-party video, audio, speech, music, likeness, or platform rights. Review source-media rights before redistribution or commercial use.
Download notes
The public, ungated test set contains 483 audio-bearing videos, each 30-60 seconds long, with a human-annotated detailed caption and manually refined visual, speech, and non-speech atomic events. The released evaluator uses an LLM to report missing-event, incorrect-event, hallucination, and total error rates. The helper downloads official documentation, API metadata, the approximately 3.5 MB annotation JSON, and evaluator by default. The current Hugging Face files total approximately 1.70 GB, so the 483 MP4 files require VIDEO_SALMONN2_DOWNLOAD_HF=1. ReMo evaluates this test set as video-SALMONN2 / video-SAL2 in section 5.1 of arXiv:2607.21179.
Safe-first helperscripts/download/video_salmonn2_caption.sh
View helper
Music

VocalSet

VocalSet: A Singing Voice Dataset

Safe-first helper
Singing Voice Analysis Vocal Technique Classification Vowel Classification Singing Voice Synthesis +1 more
Access pathZenodo
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
The Zenodo record lists CC BY 4.0 and open access. Re-check subject-consent and attribution expectations before redistributing derivative voice data.
Download notes
Zenodo hosts a single VocalSet.zip archive of about 2.1 GB with 10.1 hours of monophonic professional singing from 20 singers, covering all five vowels across standard and extended vocal techniques. The helper saves the Zenodo record metadata by default and requires VOCALSET_DOWNLOAD_ARCHIVE=1 before downloading the full archive.
Safe-first helperscripts/download/vocalset.sh
View helper
Audio understanding, generation & events

VocalSound

VocalSound: A Dataset for Improving Human Vocal Sounds Recognition

Safe-first helper
Human Vocal Sound Classification Vocalization Recognition Audio Classification Demographic Bias Evaluation
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-sa-4.0
Code license
not_specified
License caution
The official README includes a Creative Commons BY-SA 4.0 notice for the VocalSound dataset. GitHub API reports no repository-level license, so the code/baseline license is not specified; re-check before redistributing code or derived data.
Download notes
The official README describes 21,024 crowdsourced recordings from 3,365 subjects covering laughter, sighs, coughs, throat clearing, sneezes, and sniffs, with speaker metadata such as age, gender, native language, country, and health condition. The helper saves official documentation by default and requires VOCALSOUND_DOWNLOAD_ARCHIVE=1 before downloading the 1.7 GiB 16 kHz or 4.5 GiB 44.1 kHz ZIP.
Safe-first helperscripts/download/vocalsound.sh
View helper
Enhancement, separation & quality

VoiceBank-DEMAND

Noisy speech database for training speech enhancement algorithms and TTS models

Safe-first helper
Speech Enhancement Speech Denoising Noise Robust Tts Clean Noisy Parallel Speech
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_applicable
License caution
Edinburgh DataShare metadata lists Creative Commons Attribution 4.0 International Public License. The corpus derives clean speech from VCTK and noises from DEMAND plus speech-shaped/babble sources; re-check component/source terms before redistribution.
Download notes
The DataShare record exposes paired clean/noisy train and test ZIPs plus text/log files. The helper saves public metadata and license by default; text files and multi-GB audio archives are explicit opt-ins.
Safe-first helperscripts/download/voicebank_demand.sh
View helper
Speech understanding & dialogue

VoiceBench

VoiceBench: Benchmarking LLM-Based Voice Assistants

Safe-first helper
Voice Assistant Evaluation Spoken Instruction Following
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
apache-2.0
Code license
Apache-2.0
License caution
HF dataset card and GitHub repo both list Apache-2.0.
Safe-first helperscripts/download/voicebench.sh
View helper
Speech recognition

VoiceCodeBench

VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition

Safe-first helper
Asr Structured Token Recovery Entity Recovery Workplace Speech Recognition
Access pathHugging Face
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
Not specified in the source record.
Code license
Not specified in the source record.
License caution
The repository and dataset card declare MIT. The card says paid contributors consented to dataset use and release, but the audio contains identifiable voice characteristics and its stated intended-use guidance excludes speaker identification, biometric modeling, voice cloning, demographic profiling, and model training or post-training.
Download notes
The public, ungated test-only release contains 300 human-recorded English workplace-speech segments totaling 5.587 hours, with 85 anonymized speakers and 1,482 audited targets across 26 structured entity types. Its primary Canonical Token/Entity Match and Task Success Rate metrics test exact recovery of values such as email addresses, phone numbers, URLs, command-line flags, file paths, identifiers, dates, and measurements. The helper downloads official documentation, license, paper, API metadata, and the approximately 1.1 MB annotation JSONL by default. The complete Hugging Face repository is approximately 1.83 GiB and requires VOICECODEBENCH_DOWNLOAD_HF=1.
Safe-first helperscripts/download/voicecodebench.sh
View helper
Speaker, identity & emotion

VoiceMOS Challenge 2026

VoiceMOS Challenge 2026: Automatic Prediction of Human Ratings of Speech

Manual or gated
Mean Opinion Score Prediction Speech Quality Assessment Comparative Category Rating Prediction Emotional Speech Naturalness Assessment +3 more
Access pathOfficial / other
Upstream termsMixed / custom — review

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
not_publicly_specified
Code license
Apache-2.0
License caution
The public challenge page and baseline README do not state dataset reuse or redistribution terms. Apache-2.0 covers the baseline repository only, not challenge audio, listener ratings, URGENT material, CodecMOS-Accent, or emotional-speech source data.
Download notes
The official site says training data were released to registered participants through a CodaBench page sent by email, with evaluation data scheduled for July 31, 2026. Track 1 covers 840 multilingual utterances in nine languages from six URGENT speech-enhancement systems; Track 2 covers emotional TTS and human speech; Track 3 uses 4,000 CodecMOS-Accent samples from 24 codec-resynthesis and TTS systems, 32 speakers, and ten accents. The helper saves public challenge and baseline documentation, then prints the registration path; it does not guess or expose the emailed CodaBench URL.
Safe-first helperscripts/download/voicemos_challenge_2026.sh
View helper
Audiovisual & cross-modal

VoxBlink2

VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Safe-first helper
Speaker Verification Open Set Speaker Identification Speaker Recognition Audio Visual Speaker Recognition
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-NC-SA-4.0
Code license
not_specified
License caution
The repository states that released annotation data is CC BY-NC-SA 4.0, but does not separately license its software. YouTube source-media rights, platform terms, privacy considerations, and local law remain separate and are not granted by the annotation license.
Download notes
The official release provides annotations, YouTube links, timestamps, speaker labels, ASR outputs, speaker metadata, and evaluation protocols rather than redistributing audio or video. The corpus describes approximately 10 million segments, more than 110,000 speakers, and 16,000 hours across more than 15 language families. The helper downloads official documentation and license text by default; the Google Drive resource bundle remains a manual download, and cloning evaluation/data-construction code is opt-in. Source media must be obtained separately and may be unavailable or removed.
Safe-first helperscripts/download/voxblink2.sh
View helper
Audiovisual & cross-modal

VoxCeleb

VoxCeleb speaker recognition datasets

Safe-first helper
Speaker Identification Speaker Verification Speaker Recognition Audio Visual Speaker Recognition
Access pathOpenSLR
Upstream termsNot specified

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
not_specified_for_original_media
Code license
not_applicable
License caution
Official VGG pages say provided VoxCeleb/VoxCeleb2 metadata is CC BY-SA 4.0 and the corpora consist of YouTube URLs with timestamps; original media rights and privacy terms remain with upstream owners. OpenSLR SLR49 lists its small metadata resource as not copyrighted.
Download notes
Official VGG pages currently say VoxCeleb1 and VoxCeleb2 audio, URL/timestamp, and identifying metadata files are no longer available from that website. The helper downloads small OpenSLR speaker-recognition recipe metadata and trial lists only; it does not fetch the original audio/video.
Safe-first helperscripts/download/voxceleb.sh
View helper
Audiovisual & cross-modal

VoxConverse

VoxConverse: A Large Scale Audio-Visual Diarisation Dataset

Safe-first helper
Speaker Diarization Audio Visual Diarization Overlapping Speech Diarization
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0
Code license
not_specified
License caution
The official page and repository README say VoxConverse is available for research purposes under CC BY 4.0, while copyright remains with the original video owners. The GitHub repository does not expose a standalone license file through the API.
Download notes
The helper clones or updates the official annotation repository and saves the official page by default. The official page lists dev/test WAV ZIPs with MD5 checksums; audio downloads are explicit opt-ins because the dev ZIP is about 1.9 GiB and the test ZIP is also large.
Safe-first helperscripts/download/voxconverse.sh
View helper
Speaker, identity & emotion

VoxENES 2026

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

Safe-first helper
Speech Spoofing Detection Audio Deepfake Detection Synthetic Speech Detection Voice Conversion Detection +2 more
Access pathOfficial / other
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-4.0-with-upstream-terms
Code license
not_applicable
License caution
Kaggle declares CC BY 4.0 for the release. Bona fide speech derives from LibriSpeech and VoxPopuli, and synthetic samples incorporate source speech, speaker references, and outputs from multiple TTS/VC systems; review those upstream terms, voice-data rights, and model-output policies before redistribution or commercial use. The paper's CC BY 4.0 license applies to the paper, not by itself to every incorporated recording.
Download notes
The public Kaggle release contains 53,628 standardized 16 kHz mono WAV samples across English and Spanish, including 3,028 bona fide samples, 4,600 original synthetic samples from seven TTS and three voice-conversion systems, and 46,000 post-processed variants. The helper downloads Kaggle metadata by default; the approximately 23.3 GB dataset requires explicit opt-in and an authenticated Kaggle CLI.
Safe-first helperscripts/download/voxenes_2026.sh
View helper
Speech understanding & dialogue

VoxLingua107

VoxLingua107: a Dataset for Spoken Language Recognition

Safe-first helper
Spoken Language Identification Language Recognition Speech Classification
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The TalTechNLP Hugging Face dataset card lists cc-by-nc-4.0. The dataset is built from YouTube-derived speech segments, so source-media availability and platform terms still apply; the SpeechBrain recipe repository did not expose a detected license.
Download notes
The paper reports 6628 hours across 107 languages plus a 1609-utterance verified evaluation set. The helper downloads small Hugging Face metadata files by default and requires VOXLINGUA107_DOWNLOAD_HF=1 before attempting the larger mirrored dataset snapshot. The original TalTech host was not reliably reachable during the 2026-07-09 check, so verify upstream availability before large downloads.
Safe-first helperscripts/download/voxlingua107.sh
View helper
Speech recognition

VoxPopuli

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Safe-first helper
Multilingual Asr Speech To Text Translation Self Supervised Speech Representation Learning Accented Speech Recognition
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc0-1.0
Code license
cc-by-nc-4.0
License caution
Official repo lists VoxPopuli data as CC0 and points users to the European Parliament legal notice for raw data; code and pretrained models are CC BY-NC 4.0.
Download notes
HF hosts converted Parquet shards and is about 673 GiB total; select a language/config and split before downloading.
Safe-first helperscripts/download/voxpopuli.sh
View helper
Audio understanding, generation & events

WABAD

WABAD: A World Annotated Bird Acoustic Dataset for Passive Acoustic Monitoring

Safe-first helper
Bird Species Detection Passive Acoustic Monitoring Temporal Audio Event Localization Time Frequency Event Localization +1 more
Access pathZenodo
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
conflicting_zenodo_metadata_treat_as_cc-by-nc-4.0
Code license
not_applicable
License caution
The Zenodo structured license field says CC BY 4.0, but the record's human-readable description explicitly says Creative Commons Attribution-NonCommercial 4.0. Treat the release as CC BY-NC 4.0 pending clarification from the maintainers; retain attribution and do not assume commercial-use permission from the structured field alone.
Download notes
The public, ungated release contains 5,047 minutes of passive-acoustic audio with 91,931 time-frequency-bounded vocalizations from 1,192 bird species, collected at 72 sites in 29 recording locations across 13 biomes. MetaPerch evaluates WABAD as an 84-hour multi-species detection benchmark in its results section. The helper downloads the Zenodo record, README, site metadata, pooled annotations, and species list by default; the 72 site archives total approximately 19.8 GiB and require explicit site-level opt-in.
Safe-first helperscripts/download/wabad.sh
View helper
Audio understanding, generation & events

WavCaps

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Safe-first helper
Audio Captioning Audio Language Retrieval Audio Language Modeling Zero Shot Audio Classification
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
academic_only
Code license
not_specified
License caution
The GitHub README and Hugging Face card say only academic uses are allowed for WavCaps audio. The HF metadata advertises CC BY 4.0, but the dataset card also points users to component source terms for FreeSound, BBC Sound Effects, SoundBible, and AudioSet; re-check those source licenses before redistribution or commercial use. Provided models are described as non-commercial research under a UK data copyright exemption.
Download notes
The Hugging Face repository exposes JSON metadata and split FLAC waveform ZIPs for FreeSound, BBC Sound Effects, SoundBible, and AudioSet SL. The full repository is hundreds of GiB, so the helper downloads README/JSON metadata by default and requires WAVCAPS_DOWNLOAD_ZIPS=1 plus WAVCAPS_ZIP_SOURCES for waveform archives.
Safe-first helperscripts/download/wavcaps.sh
View helper
Speech recognition

WenetSpeech

Manual or gated
Mandarin Asr
Access pathOfficial / other
Upstream termsNon-commercial / research terms

Start with the helper: it prints the required form, password, license, or access-acknowledgement steps.

Access, terms & download helper
Data license / terms
non-commercial use under CC BY 4.0
Code license
Apache-2.0
License caution
Official site says WenetSpeech does not own audio copyright; original audio copyrights remain with owners.
Safe-first helperscripts/download/wenetspeech.sh
View helper
Enhancement, separation & quality

WHAM! / WHAMR!

WSJ0 Hipster Ambient Mixtures and WHAMR!: Noisy and Reverberant Single-Channel Speech Separation

Safe-first helper
Noisy Speech Separation Speech Enhancement Reverberant Speech Separation Source Separation
Access pathOfficial / other
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
cc-by-nc-4.0
Code license
not_specified
License caution
The official WHAM page states the WHAM! and WHAM!48kHz noise datasets are CC BY-NC 4.0. Generated mixtures also depend on WSJ0/wsj0-2mix licensing, so redistribution or commercial use requires checking those upstream terms too.
Download notes
The helper downloads the official landing page and small WHAM!/WHAMR! generation script archives by default. WHAM! noise is 17 GiB compressed and WHAM!48kHz is 68.1 GiB compressed, so those archives are explicit opt-ins. Building full WHAM!/WHAMR! mixtures also requires separately licensed WSJ0/wsj0-2mix access.
Safe-first helperscripts/download/wham_whamr.sh
View helper
Speech recognition

Whisper-RIR-Mega

Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics

Safe-first helper
Automatic Speech Recognition Reverberant Speech Recognition Asr Robustness Room Acoustics Robustness
Access pathHugging Face
Upstream termsNon-commercial / research terms

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC-BY-4.0_with_CC-BY-NC-4.0_upstream_terms
Code license
not_specified_currently_unavailable
License caution
The benchmark dataset card declares CC BY 4.0 and identifies LibriSpeech as CC BY 4.0, but the current RIR-Mega v2 card declares CC BY-NC 4.0 for its RIR audio. Apply the stricter non-commercial upstream terms to the derived reverberant audio unless the owner clarifies otherwise. The benchmark card says its curation repository is MIT, but the linked repository was unavailable, so that code license could not be independently verified.
Download notes
The public, ungated release contains 2,000 English LibriSpeech test-clean utterances, each paired with a 16 kHz reverberant version made using one RIR-Mega room impulse response. Its deterministic, acoustically stratified split has 400 validation and 1,600 test pairs; evaluation reports clean/reverberant WER and CER plus the reverb penalty, with RT60 and DRR metadata when available. The helper saves the dataset card, API metadata, paper, and small leaderboard files by default. The Hugging Face API reports about 1.13 GB of repository storage, so the complete audio and Arrow snapshot requires WHISPER_RIRMEGA_DOWNLOAD_HF=1. The paper's cited GitHub code repository returned HTTP 404 when checked on 2026-07-22.
Safe-first helperscripts/download/whisper_rirmega.sh
View helper
Speech understanding & dialogue

WildSpeech-Bench

WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild

Safe-first helper
Speech To Speech Evaluation Natural Speech Conversation
Access pathHugging Face
Upstream termsOpen / attribution signals

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
CC BY 4.0, except third-party datasets with their own terms
Code license
CC BY 4.0, except third-party datasets with their own terms
License caution
License.txt says users must comply with original licenses for third-party datasets.
Safe-first helperscripts/download/wildspeech_bench.sh
View helper
Audiovisual & cross-modal

WorldSense

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Safe-first helper
Audio Visual Question Answering Omni Modal Video Understanding Cross Modal Reasoning Audio Visual Perception
Access pathHugging Face
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
conflicting_cc_by_nc_sa_4_0_and_cc_by_4_0
Code license
not_specified
License caution
The WorldSense paper v3 Appendix G states CC BY-NC-SA 4.0, while the official repository README and Hugging Face card state CC BY 4.0. Apply the more restrictive CC BY-NC-SA 4.0 interpretation until the maintainers resolve the conflict. Videos are sourced primarily from FineVideo with selected MUSIC-AVQA material, so component-media terms and rights also require review. GitHub reports no detected repository license.
Download notes
The public, ungated release contains 1,662 synchronized audio-visual videos and 3,172 multiple-choice question-answer pairs across 26 tasks. The helper downloads official documentation and the approximately 4.3 MB QA JSON by default; the Hugging Face API reports approximately 18.1 GB of repository storage, so video and subtitle archives require WORLDSENSE_DOWNLOAD_HF=1. Qwen3.5-Omni reports WorldSense in section 5.1.4, Table 7.
Safe-first helperscripts/download/worldsense.sh
View helper
Speaker, identity & emotion

WSJ0-2mix / wsj0-mix

wsj0-mix: Single-channel multi-speaker speech separation mixtures from WSJ0

Safe-first helper
Speech Separation Multi Speaker Speech Separation Source Separation Cocktail Party Speech Separation
Access pathLDC / licensed
Upstream termsMixed / custom — review

The helper starts with public documentation or metadata and keeps large or restricted downloads opt-in.

Access, terms & download helper
Data license / terms
ldc_restricted_derived
Code license
MERL script license not specified on the reachable page; pywsj0-mix is MIT.
License caution
The generated mixtures derive from the LDC CSR-I WSJ0 corpus, so access, use, and redistribution must follow the active LDC agreement. The MERL page provides scripts but does not publish the audio mixtures or a standalone data license.
Download notes
The helper downloads the official MERL page and generation scripts by default and can clone the MIT-licensed Python generator. It does not download WSJ0 audio; generation requires an already licensed local WSJ0 corpus from LDC and explicit WSJ0_2MIX_RUN_GENERATION=1. TF-MossFormer sections 3.1-3.3 use the standard 8 kHz two-speaker setup with 20,000 training, 5,000 validation, and 3,000 speaker-disjoint test mixtures and report SI-SDRi and SDRi; that paper adds no new mixture release.
Safe-first helperscripts/download/wsj0_2mix.sh
View helper
From catalog to evidence

Use a benchmark to make a model decision

Choose a compatible benchmark here, turn examples into evaluation cases, and make small task-specific changes to the adapter and scoring rubric. Open Audio Judge’s ASR and TTS demos show the same case-to-report workflow end to end.

01

Choose a benchmark

Match the task, data access, and upstream terms to your use case.

02

Prepare cases

Map benchmark examples into references, audio inputs, and metadata.

03

Adapt the rubric

Keep the workflow; adjust the adapter and scoring criteria for the task.

04

Compare & inspect

Run candidate models, summarize slices, and open case-level evidence.