VibeVoice ASR
96.5 semantic score; 33/35 accurate and 2 needing attention.
Inspect this modelCompare three MLX ASR models by semantic accuracy, understand where each one fails, and open the exact transcript evidence behind every score.
Turn public-safe text seeds into local audio under ignored runs/.
Run each MLX ASR model through oaj autojudge-mlx-asr.
Gemini scores candidate transcripts against references and semantic rubrics.
Combine model reports into a leaderboard page with category slices.
| Model | Wrapper | Status | Why This Model |
|---|---|---|---|
mlx-community/whisper-large-v3-turbo-asr-fp16 | mlx-audio | 35-case run passed | Strong Whisper-family baseline with MLX conversion. |
mlx-community/Qwen3-ASR-1.7B-8bit | mlx-audio | 35-case run passed | Recent ASR-specific model for comparison against Whisper. |
mlx-community/VibeVoice-ASR-4bit | mlx-audio | 35-case run passed | Compact quantized ASR model for semantic-quality comparison. |
Which model best preserves meaning on this small, controlled ASR set? 3 MLX Community ASR models transcribe the same research-guided eval set. Scores below are Gemini semantic-judge scores (higher is better), not word error rate.
96.5 semantic score; 33/35 accurate and 2 needing attention.
Inspect this model96.6 on acoustic-noise robustness.
Inspect noise cases3 of 3 models need attention on this case.
Inspect shared failure| Rank | Model | Semantic score ↑ | Accurate cases | Needs attention | Weakest category | Evidence |
|---|---|---|---|---|---|---|
| #1 | VibeVoice ASRmlx-community/VibeVoice-ASR-4bit | 96.5 | 33/3535/35 ok | 21 review, 1 inaccurate | Entity 90.2 | View cases |
| #2 | Qwen3 ASR 1.7Bmlx-community/Qwen3-ASR-1.7B-8bit | 95.2 | 31/3535/35 ok | 43 review, 1 inaccurate | Entity 89.0 | View cases |
| #3 | Whisper Large v3 Turbomlx-community/whisper-large-v3-turbo-asr-fp16 | 94.3 | 31/3535/35 ok | 42 review, 2 inaccurate | Entity 88.6 | View cases |
Use this matrix to choose for a workload. Each populated cell averages semantic scores over 5 cases; it is not WER. Best-in-column cells are marked.
| Model | Surface Transcription | Numeric/Unit | Negation/Modality | Temporal | Entity | Paraphrase | Acoustic Noise |
|---|---|---|---|---|---|---|---|
VibeVoice ASRmlx-community/VibeVoice-ASR-4bit | Best96.25/5 accurate5 accurate | Best99.25/5 accurate5 accurate | 99.65/5 accurate5 accurate | 98.85/5 accurate5 accurate | Best90.24/5 accurate · 1 flagged4 accurate, 1 inaccurate | 98.45/5 accurate5 accurate | 92.84/5 accurate · 1 flagged4 accurate, 1 needs_review |
Qwen3 ASR 1.7Bmlx-community/Qwen3-ASR-1.7B-8bit | 92.84/5 accurate · 1 flagged4 accurate, 1 needs_review | 92.04/5 accurate · 1 flagged4 accurate, 1 needs_review | Best100.05/5 accurate5 accurate | Best100.05/5 accurate5 accurate | 89.04/5 accurate · 1 flagged4 accurate, 1 inaccurate | Best100.05/5 accurate5 accurate | 92.44/5 accurate · 1 flagged4 accurate, 1 needs_review |
Whisper Large v3 Turbomlx-community/whisper-large-v3-turbo-asr-fp16 | 94.84/5 accurate · 1 flagged4 accurate, 1 needs_review | 92.04/5 accurate · 1 flagged4 accurate, 1 needs_review | Best100.05/5 accurate5 accurate | 90.64/5 accurate · 1 flagged4 accurate, 1 inaccurate | 88.64/5 accurate · 1 flagged4 accurate, 1 inaccurate | 97.65/5 accurate5 accurate | Best96.65/5 accurate5 accurate |
3/3 models flagged. The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
2/3 models flagged. The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.
2/3 models flagged. The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
35 English clips generated with 1 synthetic voice were evaluated across 3 models (105 model-case evaluations; 5 cases per category). Each result uses up to 3 Gemini judge attempts with prompt 0.2.0. Semantic scores measure meaning preservation and downstream usability; they do not measure word error rate, latency, memory, or production robustness. Validate on representative real speech, accents, devices, and domain data before deployment.
Judge score policy: Semantic scores average successful judge attempts. Failed attempts are excluded only when at least one attempt succeeds; an evaluation with no successful attempt retains failure status instead of receiving a quality average.
Partial judge failures affected 1 evaluation; inspect the linked evidence:
The 7 research categories are acoustic_noise_robustness, entity_factual_integrity, negation_modality_scope, numeric_unit_integrity, semantic_paraphrase_preservation, temporal_scheduling_accuracy, transcription_accuracy_wer. Results are directional evidence, not deployment proof.
Total Gemini judge samples: 315. Refresh this block with .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py after rerunning the verified ASR model jobs. The combined local report is runs/asr-leaderboard/full-35-combined/report.html and the committed summary artifact is docs/asr-leaderboard-summary.json. The generated refresh report is docs/asr-leaderboard-refresh-report.md, and the generated shell playbook is docs/asr-leaderboard-refresh-commands.sh. The machine-readable workflow is docs/asr-leaderboard-refresh-workflow.json. The opt-in live model refresh script is docs/asr-leaderboard-live-refresh.sh. The committed run manifest is docs/asr-leaderboard-run-manifest.json, with coverage validation in docs/asr-leaderboard-manifest-validation.json and seed-manifest validation in docs/asr-seed-manifest-validation.json. The next-refresh plan is docs/asr-leaderboard-next-runs.json, and the hosted artifact manifest is docs/asr-leaderboard-hosted-manifest.json. The artifact bundle index is docs/asr-leaderboard-artifacts.json. Runtime readiness is tracked in docs/asr-leaderboard-runtime-status.json, the cron refresh decision is recorded in docs/asr-leaderboard-refresh-decision.json, the Telegram-ready next-action note is docs/asr-leaderboard-next-action.md, the compact cron status is docs/asr-leaderboard-cron-status.json, the human cron handoff is docs/asr-leaderboard-cron-handoff.md, and source selection is recorded in docs/asr-leaderboard-source-selection.json. The compact artifact-bundle digest is docs/asr-leaderboard-bundle-status.json; together they include the source result files, complete model/category matrix, missing-cell guidance, runtime-gated next action, hosted copy map, and reproducible refresh workflow. Pass ASR_LEADERBOARD_HOSTED_DIR with --hosted-dir-from-env to copy the same verified artifacts into the hosted Pages checkout. Use docs/asr-leaderboard-report-index.md as the generated map from the demo page to the combined benchmark report and per-source run reports; use docs/asr-leaderboard-report-links.json for the same map in machine-readable form, including the source artifact list behind each model/category cell. Use docs/asr-leaderboard-report-bundle.json as the single automation entry point for the combined report, source reports, hosted URLs, selected source files, and refresh provenance.
These commands are generated from the same workflow metadata written to docs/asr-leaderboard-summary.json and docs/asr-leaderboard-refresh-report.md.
| Step | Command |
|---|---|
| Preflight refresh inputs | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only |
| Write preflight summary | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --require-generated-fresh --check-summary-out runs/asr-leaderboard/preflight-summary.json |
| Require audio manifest readiness | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --require-audio-ready |
| Validate seed manifest | .venv/bin/python scripts/validate_asr_seed_manifest.py --summary-out docs/asr-seed-manifest-validation.json |
| Materialize audio | .venv/bin/python scripts/synthesize_tts_cases.py --cases examples/asr_research_cases.jsonl --out runs/asr-research-audio --discard-text-sidecars --summary-out runs/asr-research-audio/summary.json |
| Refresh runtime status | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --check-mlx-runtime |
| Require runtime readiness | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --check-mlx-runtime --require-runtime-ready |
| Full refresh readiness check | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --require-generated-fresh --require-audio-ready --check-summary-out runs/asr-leaderboard/preflight-summary.json |
| Cron refresh rehearsal | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --require-generated-fresh --require-audio-ready --check-mlx-runtime --check-summary-out runs/asr-leaderboard/preflight-summary.json |
| Check MLX ASR runtime | PYTHONPATH=src .venv/bin/python -m open_audio_judge.cli check-mlx-asr-runtime --python-bin .venv/bin/python --model mlx-community/whisper-large-v3-turbo-asr-fp16 |
| Run one MLX ASR model | .venv/bin/oaj autojudge-mlx-asr --python-bin .venv/bin/python --cases runs/asr-research-audio/tts_audio_cases.jsonl --model <mlx-community/model-id> --judge-provider gemini --judge-samples 3 --out runs/asr-leaderboard/<run-name> |
| Discover latest complete runs | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --discover-complete-model-runs --update-run-manifest --source-selection-summary-out docs/asr-leaderboard-source-selection.json |
| Refresh committed artifacts | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --source-selection-summary-out docs/asr-leaderboard-source-selection.json |
| Run refresh shell playbook | bash docs/asr-leaderboard-refresh-commands.sh |
| Run live model refresh script | bash docs/asr-leaderboard-live-refresh.sh |
| Review blocked model log | tail -n 20 runs/asr-leaderboard/blocked-models.jsonl |
| Check generated page | .venv/bin/python scripts/check_asr_leaderboard_page.py |
| Verify generated artifacts are fresh | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --require-generated-fresh |
| Run commit verification | .venv/bin/python scripts/verify_asr_leaderboard_commit.py |
| Run hosted commit verification | .venv/bin/python scripts/verify_asr_leaderboard_commit.py --hosted-dir-from-env |
| Sync hosted artifacts | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --hosted-dir-from-env |
| Check hosted mirror | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --hosted-dir-from-env --require-hosted-current |
| Write hosted drift report | .venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py --check-only --hosted-dir-from-env --hosted-status-out runs/asr-leaderboard/hosted-status.json |
| Stage | Commands | Behavior |
|---|---|---|
preflight | cron_rehearsal_command, runtime_ready_check_command | no live model calls; writes committed artifacts |
live_refresh | local_secret_env_command, model_run_commands | runs live models; validation only |
artifact_refresh | discover_refresh_command, combine_refresh_command, manifest_refresh_command | no live model calls; writes committed artifacts |
verification | page_validation_command, report_bundle_check_command, freshness_check_command, commit_verification_command, cron_commit_verification_command | no live model calls; validation only |
hosted_sync | hosted_artifact_command, hosted_validation_command, hosted_status_command, hosted_commit_verification_command | no live model calls; validation only; requires ASR_LEADERBOARD_HOSTED_DIR |
Load the Gemini secret only in the local shell before running live judge calls: source "${OPEN_AUDIO_JUDGE_GEMINI_ENV_FILE:?Set_OPEN_AUDIO_JUDGE_GEMINI_ENV_FILE_to_your_local_Gemini_environment_file}".
| Model | Run Command |
|---|---|
mlx-community/whisper-large-v3-turbo-asr-fp16 | .venv/bin/oaj autojudge-mlx-asr --python-bin .venv/bin/python --cases runs/asr-research-audio/tts_audio_cases.jsonl --model mlx-community/whisper-large-v3-turbo-asr-fp16 --judge-provider gemini --judge-samples 3 --out runs/asr-leaderboard/whisper-large-v3-turbo-refresh |
mlx-community/Qwen3-ASR-1.7B-8bit | .venv/bin/oaj autojudge-mlx-asr --python-bin .venv/bin/python --cases runs/asr-research-audio/tts_audio_cases.jsonl --model mlx-community/Qwen3-ASR-1.7B-8bit --judge-provider gemini --judge-samples 3 --out runs/asr-leaderboard/qwen3-asr-1.7b-refresh |
mlx-community/VibeVoice-ASR-4bit | .venv/bin/oaj autojudge-mlx-asr --python-bin .venv/bin/python --cases runs/asr-research-audio/tts_audio_cases.jsonl --model mlx-community/VibeVoice-ASR-4bit --judge-provider gemini --judge-samples 3 --out runs/asr-leaderboard/vibevoice-asr-refresh |
If a primary MLX ASR model is unsupported locally, record that blocked state in the run notes before trying the documented fallbacks: mlx-community/whisper-small.en-asr-4bit, mlx-community/parakeet-rnnt-0.6b, mlx-community/GLM-ASR-Nano-2512-4bit.
| Path | Purpose |
|---|---|
runs/asr-leaderboard/full-35-combined/results.jsonl | Combined ASR judge results used by the generated page and report. |
runs/asr-leaderboard/full-35-combined/report.html | Local combined HTML report with per-case judge details. |
docs/asr-leaderboard-summary.json | Machine-readable leaderboard summary and reproducible refresh workflow. |
docs/asr-leaderboard-refresh-report.md | Human-readable coverage, score, source-file, and command report. |
docs/asr-leaderboard-report-index.md | Human-readable index linking the demo page, combined report, and source run reports. |
docs/asr-leaderboard-report-links.json | Machine-readable map linking the demo page to combined and source ASR reports. |
docs/asr-leaderboard-report-bundle.json | Single machine-readable entry point for ASR report URLs, source reports, and refresh provenance. |
docs/asr-leaderboard-refresh-commands.sh | Generated shell playbook for repeatable ASR leaderboard refreshes. |
docs/asr-leaderboard-refresh-workflow.json | Machine-readable generated workflow for ASR refresh automation. |
docs/asr-leaderboard-live-refresh.sh | Opt-in generated shell script for live MLX ASR/Gemini refreshes. |
docs/asr-leaderboard-run-manifest.json | Committed source result manifest for manifest-based refreshes. |
docs/asr-leaderboard-manifest-validation.json | Coverage validation for the model/category result matrix. |
docs/asr-seed-manifest-validation.json | Seed-manifest validation proving public-safe ASR cases keep exact category coverage. |
docs/asr-leaderboard-next-runs.json | Machine-readable next-refresh plan for missing ASR model/category cells. |
docs/asr-leaderboard-hosted-manifest.json | Machine-readable manifest of ASR demo artifacts mirrored to the hosted Pages checkout. |
docs/asr-leaderboard-artifacts.json | Single machine-readable index for the ASR leaderboard artifact bundle. |
docs/asr-leaderboard-runtime-status.json | Machine-readable MLX ASR and Gemini readiness status for refresh automation. |
docs/asr-leaderboard-refresh-decision.json | Machine-readable runtime-gated decision for the next ASR refresh action. |
docs/asr-leaderboard-next-action.md | Telegram-ready Markdown note summarizing the runtime-gated next ASR action. |
docs/asr-leaderboard-cron-status.json | Compact machine-readable cron handoff with action, coverage, and runtime gate status. |
docs/asr-leaderboard-cron-handoff.md | Human-readable cron handoff summary for scheduled ASR refresh turns. |
docs/asr-leaderboard-source-selection.json | Machine-readable record of selected ASR source result files for the last refresh. |
docs/asr-leaderboard-bundle-status.json | Compact digest of ASR leaderboard artifact, hosted, runtime, and decision status. |
Each source run keeps its own local report alongside the JSONL file that feeds the combined leaderboard; hosted refreshes mirror available source reports under the same ASR leaderboard path.
| Result File | Local Report | Hosted Report | Cases | Categories |
|---|---|---|---|---|
runs/asr-leaderboard/whisper-large-v3-turbo-smoke/judge-report/results.jsonl | runs/asr-leaderboard/whisper-large-v3-turbo-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/whisper-large-v3-turbo-smoke/judge-report/report.html | 3/3 ok | transcription_accuracy_wer: 3 |
runs/asr-leaderboard/whisper-large-v3-turbo-full-gap/judge-report/results.jsonl | runs/asr-leaderboard/whisper-large-v3-turbo-full-gap/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/whisper-large-v3-turbo-full-gap/judge-report/report.html | 12/12 ok | negation_modality_scope: 3, numeric_unit_integrity: 3, temporal_scheduling_accuracy: 4, transcription_accuracy_wer: 2 |
runs/asr-leaderboard/whisper-large-v3-turbo-semantic-smoke/judge-report/results.jsonl | runs/asr-leaderboard/whisper-large-v3-turbo-semantic-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/whisper-large-v3-turbo-semantic-smoke/judge-report/report.html | 5/5 ok | negation_modality_scope: 2, numeric_unit_integrity: 2, temporal_scheduling_accuracy: 1 |
runs/asr-leaderboard/whisper-large-v3-turbo-entity-smoke/judge-report/results.jsonl | runs/asr-leaderboard/whisper-large-v3-turbo-entity-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/whisper-large-v3-turbo-entity-smoke/judge-report/report.html | 5/5 ok | entity_factual_integrity: 5 |
runs/asr-leaderboard/whisper-large-v3-turbo-paraphrase-smoke/judge-report/results.jsonl | runs/asr-leaderboard/whisper-large-v3-turbo-paraphrase-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/whisper-large-v3-turbo-paraphrase-smoke/judge-report/report.html | 5/5 ok | semantic_paraphrase_preservation: 5 |
runs/asr-leaderboard/whisper-large-v3-turbo-noise-smoke/judge-report/results.jsonl | runs/asr-leaderboard/whisper-large-v3-turbo-noise-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/whisper-large-v3-turbo-noise-smoke/judge-report/report.html | 5/5 ok | acoustic_noise_robustness: 5 |
runs/asr-leaderboard/qwen3-asr-1.7b-smoke/judge-report/results.jsonl | runs/asr-leaderboard/qwen3-asr-1.7b-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/qwen3-asr-1.7b-smoke/judge-report/report.html | 3/3 ok | transcription_accuracy_wer: 3 |
runs/asr-leaderboard/qwen3-asr-1.7b-full-gap/judge-report/results.jsonl | runs/asr-leaderboard/qwen3-asr-1.7b-full-gap/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/qwen3-asr-1.7b-full-gap/judge-report/report.html | 12/12 ok | negation_modality_scope: 3, numeric_unit_integrity: 3, temporal_scheduling_accuracy: 4, transcription_accuracy_wer: 2 |
runs/asr-leaderboard/qwen3-asr-1.7b-semantic-smoke/judge-report/results.jsonl | runs/asr-leaderboard/qwen3-asr-1.7b-semantic-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/qwen3-asr-1.7b-semantic-smoke/judge-report/report.html | 5/5 ok | negation_modality_scope: 2, numeric_unit_integrity: 2, temporal_scheduling_accuracy: 1 |
runs/asr-leaderboard/qwen3-asr-1.7b-entity-smoke/judge-report/results.jsonl | runs/asr-leaderboard/qwen3-asr-1.7b-entity-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/qwen3-asr-1.7b-entity-smoke/judge-report/report.html | 5/5 ok | entity_factual_integrity: 5 |
runs/asr-leaderboard/qwen3-asr-1.7b-paraphrase-smoke/judge-report/results.jsonl | runs/asr-leaderboard/qwen3-asr-1.7b-paraphrase-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/qwen3-asr-1.7b-paraphrase-smoke/judge-report/report.html | 5/5 ok | semantic_paraphrase_preservation: 5 |
runs/asr-leaderboard/qwen3-asr-1.7b-noise-smoke/judge-report/results.jsonl | runs/asr-leaderboard/qwen3-asr-1.7b-noise-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/qwen3-asr-1.7b-noise-smoke/judge-report/report.html | 5/5 ok | acoustic_noise_robustness: 5 |
runs/asr-leaderboard/vibevoice-asr-smoke/judge-report/results.jsonl | runs/asr-leaderboard/vibevoice-asr-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/vibevoice-asr-smoke/judge-report/report.html | 3/3 ok | transcription_accuracy_wer: 3 |
runs/asr-leaderboard/vibevoice-asr-full-gap/judge-report/results.jsonl | runs/asr-leaderboard/vibevoice-asr-full-gap/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/vibevoice-asr-full-gap/judge-report/report.html | 12/12 ok | negation_modality_scope: 3, numeric_unit_integrity: 3, temporal_scheduling_accuracy: 4, transcription_accuracy_wer: 2 |
runs/asr-leaderboard/vibevoice-asr-semantic-smoke/judge-report/results.jsonl | runs/asr-leaderboard/vibevoice-asr-semantic-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/vibevoice-asr-semantic-smoke/judge-report/report.html | 5/5 ok | negation_modality_scope: 2, numeric_unit_integrity: 2, temporal_scheduling_accuracy: 1 |
runs/asr-leaderboard/vibevoice-asr-entity-smoke/judge-report/results.jsonl | runs/asr-leaderboard/vibevoice-asr-entity-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/vibevoice-asr-entity-smoke/judge-report/report.html | 5/5 ok | entity_factual_integrity: 5 |
runs/asr-leaderboard/vibevoice-asr-paraphrase-smoke/judge-report/results.jsonl | runs/asr-leaderboard/vibevoice-asr-paraphrase-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/vibevoice-asr-paraphrase-smoke/judge-report/report.html | 5/5 ok | semantic_paraphrase_preservation: 5 |
runs/asr-leaderboard/vibevoice-asr-noise-smoke/judge-report/results.jsonl | runs/asr-leaderboard/vibevoice-asr-noise-smoke/judge-report/report.html | https://kennethli319.github.io/open-audio-judge/asr-leaderboard/vibevoice-asr-noise-smoke/judge-report/report.html | 5/5 ok | acoustic_noise_robustness: 5 |
Classic edit-distance calibration: substitutions, deletions, insertions, homophones, punctuation, and disfluency policy.
Names, addresses, organizations, product labels, and alphanumeric identifiers.
Amounts, dosages, decimals, percentages, account numbers, and value-changing numeric pairs.
Negation, permission, prohibition, modality, only-scope, and exception boundaries.
Dates, times, durations, AM/PM, relative ordering, and deadlines.
Events, causal relations, metric tradeoffs, coreference, and conditional instructions.
Cafe noise, vehicle noise, industrial background noise, reverberant rooms, and overlapping speech.
: "${OPEN_AUDIO_JUDGE_GEMINI_ENV_FILE:?Set this to your local Gemini environment file}"
source "$OPEN_AUDIO_JUDGE_GEMINI_ENV_FILE"
oaj autojudge-mlx-asr \
--python-bin .venv/bin/python \
--cases runs/asr-research-audio/tts_audio_cases.jsonl \
--model mlx-community/whisper-large-v3-turbo-asr-fp16 \
--judge-provider gemini \
--judge-samples 3 \
--out runs/asr-leaderboard/whisper-large-v3-turbo-full-gap
oaj autojudge-mlx-asr \
--python-bin .venv/bin/python \
--cases runs/asr-research-audio/tts_audio_cases.jsonl \
--model mlx-community/Qwen3-ASR-1.7B-8bit \
--judge-provider gemini \
--judge-samples 3 \
--out runs/asr-leaderboard/qwen3-asr-1.7b-full-gap
oaj autojudge-mlx-asr \
--python-bin .venv/bin/python \
--cases runs/asr-research-audio/tts_audio_cases.jsonl \
--model mlx-community/VibeVoice-ASR-4bit \
--judge-provider gemini \
--judge-samples 3 \
--out runs/asr-leaderboard/vibevoice-asr-full-gap
.venv/bin/python scripts/refresh_asr_leaderboard_artifacts.py
| Path | Purpose |
|---|---|
candidate_cases.jsonl | Source cases plus MLX-generated candidate transcripts. |
model_summary.json | Model id, transcriber, category/slice coverage, and candidate coverage. |
judge-report/results.jsonl | Gemini structured scores, reasons, semantic categories, and sample aggregates. |
judge-report/report.html | Per-model HTML report. |
combined/report.html | Combined ASR leaderboard report. |
{
"case_id": "asr-numeric-transfer-001",
"provider": "gemini",
"overall_score": 42,
"meaning_preservation": "major_loss",
"error_categories": ["number_error", "contrast_error"],
"metadata": {
"candidate_model": "mlx-community/whisper-large-v3-turbo-asr-fp16",
"eval_category": "numeric_unit_integrity",
"asr_slice": "amount_contrast",
"judge_sample_count": 3
}
}
See docs/asr-eval-taxonomy.md for the category rationale and examples/asr_research_cases.jsonl for the seed manifest. These leaderboard results are verified local artifacts from all seven categories, with private/generated audio kept out of the repository.
Choose a compatible benchmark from the Audio Benchmark Index, then reuse Open Audio Judge's case-to-report workflow to rate speech models with small task-specific changes to the case adapter and scoring rubric.