Open Audio Judge ASR report

ASR model comparison

105 evaluated cases · 7 evaluation categories · en · semantic judge score; higher is better
Cases105
Semantic score95.3
Median100.0
Accurate95
Needs Review6
Inaccurate4
Decision brief Review 10 of 105 cases before relying on this model

95/105 cases are labeled accurate. Use the lowest-scoring evidence below to decide whether the remaining errors matter for your workload. Also verify 1 case with incomplete judge coverage.

Inspect asr-entity-ticket-001-local-tts firstScore 53: The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
Recommended next stepThe model struggled with the alphanumeric ticket ID, merging the word 'ticket' with the identifier and misrecognizing the digits '4187' as '457'.
314/315 judge attempts succeeded 1 failed attempt excluded from quality scores 4 cases with 20+ point judge spread

Priority Cases

Start here: these cases have the largest semantic impact, a high-impact error, a non-accurate label, or incomplete judge coverage.

Cases that need attention, ordered by judge status, semantic severity, score, and judge coverage
Case Score Meaning Category Summary
asr-entity-ticket-001-local-tts 53 partial loss entity error, number error The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
asr-entity-ticket-001-local-tts 57 partial loss entity error, number error The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.
asr-temporal-duration-001-local-tts 53 partial loss number error The pump's rest duration was incorrectly transcribed as 15 minutes instead of fifty minutes.
asr-entity-ticket-001-local-tts 55 partial loss entity error The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.
asr-numeric-account-001-local-tts 60 partial loss number error The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
asr-negation-refund-001-local-tts 100 preserved 1 judge attempt failed No semantic errors are present as the candidate transcript perfectly preserves the spoken meaning.

Score Distribution

1-20
0
21-40
0
41-60
6
61-80
4
81-100
95

Semantic Diagnostics

Meaning Preservation
  • preserved 93
  • partial loss 8
  • minor loss 4
Error Categories
  • no error 46
  • formatting only 46
  • substitution 14
  • number error 7
  • entity error 5
High-Impact Errors
  • number error 7
  • entity error 5
Actionable Notes

Weakest Segments

Compare slices with enough variation to reveal where quality drops and what to improve next.

Model mlx-community/whisper-large-v3-turbo-asr-fp16
avg 94.3 / n 35 / range 53-100
Source basis
WER remains the standard ASR edit-distance baseline; include low-level substitutions before semantic judging.WER catches token-level confusions while semantic review determines downstream severity.
Likely fix areas
formatting onlysubstitutiontext faithfulnessentity error
Issue categories
formatting only x15substitution x5number error x3entity error x2
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 53 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorformatting only
    The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
  • asr-temporal-duration-001-local-tts 53 / inaccurate
    status: ok
    source: asr-temporal-duration-001
    issues: number errorsubstitution
    The pump's rest duration was incorrectly transcribed as 15 minutes instead of fifty minutes.
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
Model mlx-community/Qwen3-ASR-1.7B-8bit
avg 95.2 / n 35 / range 55-100
Source basis
WER remains the standard ASR edit-distance baseline; include low-level substitutions before semantic judging.WER catches token-level confusions while semantic review determines downstream severity.
Likely fix areas
formatting onlysubstitutiontext faithfulnessentity error
Issue categories
formatting only x14substitution x6number error x2entity error x2
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 55 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errorsubstitution
    The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The candidate incorrectly transcribes the final digit of the account number, substituting '8' with the letter 'E'.
  • asr-noise-reverberant-room-001-local-tts 62 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.
Model mlx-community/VibeVoice-ASR-4bit
avg 96.5 / n 35 / range 57-100
Source basis
WER remains the standard ASR edit-distance baseline; include low-level substitutions before semantic judging.WER catches token-level confusions while semantic review determines downstream severity.
Likely fix areas
formatting onlysubstitutiontext faithfulnessentity error
Issue categories
formatting only x17substitution x3number error x2entity error x1
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 57 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorsubstitution
    The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.
  • asr-noise-reverberant-room-001-local-tts 67 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number '412' was incorrectly transcribed as '42012', creating a highly misleading location error.
  • asr-wer-function-word-001-local-tts 85 / accurate
    status: ok
    source: asr-wer-function-word-001
    issues: substitution
    The substitution of 'patch is for' with 'patches for' creates a minor grammatical slip but does not change the core message.
Evaluation Category entity factual integrity
avg 89.3 / n 15 / range 53-100
Source basis
Entity recall and named-entity error analysis are common ASR quality slices beyond aggregate WER.Custom vocabulary and entity recall are critical for product, contact, and organization names.
Likely fix areas
formatting onlyentity errorsubstitutiontext faithfulness
Issue categories
formatting only x7entity error x5substitution x4number error x2
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 53 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorformatting only
    The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
  • asr-entity-ticket-001-local-tts 55 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errorsubstitution
    The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.
  • asr-entity-ticket-001-local-tts 57 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorsubstitution
    The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.
Evaluation Category acoustic noise robustness
avg 93.9 / n 15 / range 62-100
Source basis
CHiME-style and far-field ASR benchmarks stress recognition under real background noise, not only clean-speech WER.Robust ASR evaluation includes adverse acoustic environments such as vehicle noise and reverberation.
Likely fix areas
formatting onlytext faithfulnesssubstitution
Issue categories
formatting only x4number error x2substitution x2
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-noise-reverberant-room-001-local-tts 62 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.
  • asr-noise-reverberant-room-001-local-tts 67 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number '412' was incorrectly transcribed as '42012', creating a highly misleading location error.
  • asr-noise-reverberant-room-001-local-tts 83 / accurate
    status: ok
    source: asr-noise-reverberant-room-001
    issues: No judge issue categories
    The candidate correctly transcribes the spoken room number 'four twenty twelve' which deviates from the reference text due to a TTS pronunciation error.
Evaluation Category numeric unit integrity
avg 94.4 / n 15 / range 60-100
Source basis
Numeric normalization should distinguish equivalent formatting from value-changing recognition errors.Digit-sequence accuracy is a practical ASR slice for forms, support, and voice agents.
Likely fix areas
formatting onlysubstitutiontext faithfulness
Issue categories
formatting only x9substitution x2number error x2
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The candidate incorrectly transcribes the final digit of the account number, substituting '8' with the letter 'E'.
  • asr-numeric-medication-001-local-tts 98 / accurate
    status: ok
    source: asr-numeric-medication-001
    issues: No judge issue categories
    There are no semantic errors; the transcription is perfectly accurate.
ASR Slice alphanumeric identifier
avg 55.0 / n 3 / range 53-57
Source basis
Identifier errors can have high task impact despite low edit distance.
Likely fix areas
entity errortext faithfulnesssubstitutionformatting only
Issue categories
entity error x3number error x2substitution x2formatting only x1
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 53 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorformatting only
    The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
  • asr-entity-ticket-001-local-tts 55 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errorsubstitution
    The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.
  • asr-entity-ticket-001-local-tts 57 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorsubstitution
    The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.
ASR Slice reverberant meeting room
avg 70.7 / n 3 / range 62-83
Source basis
Far-field and meeting ASR research tracks reverberation effects separately from clean microphone transcripts.
Likely fix areas
text faithfulnesssubstitution
Issue categories
number error x2substitution x2
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-noise-reverberant-room-001-local-tts 62 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.
  • asr-noise-reverberant-room-001-local-tts 67 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number '412' was incorrectly transcribed as '42012', creating a highly misleading location error.
  • asr-noise-reverberant-room-001-local-tts 83 / accurate
    status: ok
    source: asr-noise-reverberant-room-001
    issues: No judge issue categories
    The candidate correctly transcribes the spoken room number 'four twenty twelve' which deviates from the reference text due to a TTS pronunciation error.
ASR Slice digit sequence
avg 73.3 / n 3 / range 60-100
Source basis
Digit-sequence accuracy is a practical ASR slice for forms, support, and voice agents.
Likely fix areas
substitutiontext faithfulnessformatting only
Issue categories
substitution x2number error x2formatting only x1
Evaluation failures
No failed evaluations
Representative low-score samples
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The candidate incorrectly transcribes the final digit of the account number, substituting '8' with the letter 'E'.
  • asr-numeric-account-001-local-tts 100 / accurate
    status: ok
    source: asr-numeric-account-001
    issues: formatting only
    No semantic errors; the candidate transcript is fully correct with a minor formatting variation in the numeric representation.

Case Results

Lowest scores appear first. Search the judge rationale, model, category, slice, or issue.

105 / 105 shown
transcription accuracy wer · function word substitution

asr-wer-function-word-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors; the transcription is perfectly accurate.

no error
Recommended action
  • The ASR model performed perfectly on this utterance with 100% accuracy and perfect and punctuation and casing.

Judge transcript

The release notes say the patch is for the beta channel, not the stable channel.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly matches the spoken audio and the reference transcript, preserving the exact wording and meaning with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcription is perfectly accurate.
Judge transcriptThe release notes say the patch is for the beta channel, not the stable channel.
Categories
no_error
Researcher notes
  • The ASR model performed perfectly on this utterance with 100% accuracy and perfect and punctuation and casing.

Provenance

  • Category transcription accuracy wer
  • ASR slice function word substitution
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-function-word-001
  • Text SHA-256 db881f12982b7e1bcdb006f5bcb9ce39ddcb8ce29bfa65b29176ffcd5301d3a7
  • Duration 4.12s
  • Bytes 197684

Evaluation status: ok

transcription accuracy wer · homophone confusion

asr-wer-homophone-001-local-tts

Whisper Large v3 Turbo (FP16)
76 needs review

Semantic impact

The homophone error 'Write Team' instead of 'right team' slightly alters the meaning regarding which team is reviewing the policy.

substitutionformatting only
Recommended action
  • The model struggled with homophone disambiguation, choosing the verb-spelled 'Write' over the adjective 'right' in the context of 'right team'.
  • Language model context training could help resolve 'right team' vs 'write team' based on surrounding syntactic cues.

Judge transcript

Please write the new access policy before the right team reviews it.
Key differences
  • Used 'Write Team' instead of 'right team'
  • Capitalized 'Access Policy' instead of lowercase 'access policy'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 76.00 (scores: 75, 75, 78). Representative reason: The candidate transcript contains a homophone substitution, transcribing 'right team' (the correct or appropriate team) as 'Write Team'. This introduces a minor semantic error by implying a specific team name or function (writing) rather than the intended sense of 'correctness', along with unnecessary capitalization of 'Access Policy'.

judge samples: 75, 75, 78; avg 76.00
Meaning
minor_loss
Semantic impact
The homophone error 'Write Team' instead of 'right team' slightly alters the meaning regarding which team is reviewing the policy.
Judge transcriptPlease write the new access policy before the right team reviews it.
Key differences
  • Used 'Write Team' instead of 'right team'
  • Capitalized 'Access Policy' instead of lowercase 'access policy'
Categories
substitution, formatting_only
Researcher notes
  • The model struggled with homophone disambiguation, choosing the verb-spelled 'Write' over the adjective 'right' in the context of 'right team'.
  • Language model context training could help resolve 'right team' vs 'write team' based on surrounding syntactic cues.

Provenance

  • Category transcription accuracy wer
  • ASR slice homophone confusion
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-homophone-001
  • Text SHA-256 e98bb2cafe5fd5b2ad5f26020053ff27ca22a0e7c17c904647d2449af7b4657e
  • Duration 3.24s
  • Bytes 155444

Evaluation status: ok

transcription accuracy wer · short command edit distance

asr-wer-short-utterance-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The transcript preserves the meaning perfectly, with only minor formatting differences.

formatting only
Recommended action
  • The model performs perfectly on this utterance. The change from 'gate four' to 'Gate 4' is an acceptable formatting variation.

Judge transcript

Turn left at Pine Street, then park by gate four.
Key differences
  • gate four formatted as Gate 4
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate, preserving the spoken meaning perfectly. The only differences are minor formatting choices, specifically capitalizing 'Gate' and writing 'four' as the digit '4'.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript preserves the meaning perfectly, with only minor formatting differences.
Judge transcriptTurn left at Pine Street, then park by gate four.
Key differences
  • gate four formatted as Gate 4
Categories
formatting_only
Researcher notes
  • The model performs perfectly on this utterance. The change from 'gate four' to 'Gate 4' is an acceptable formatting variation.

Provenance

  • Category transcription accuracy wer
  • ASR slice short command edit distance
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-short-utterance-001
  • Text SHA-256 f0270a081ebe4d9733c17ae6c19c100028c61723a25cf10de9aa04db0e29e736
  • Duration 2.96s
  • Bytes 142004

Evaluation status: ok

transcription accuracy wer · disfluency token accounting

asr-wer-disfluency-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors exist as the candidate transcript is identical to the reference.

no error
Recommended action
  • The model successfully handles self-correction and disfluencies ('I need the, sorry, the') without deleting them or hallucinating incorrect text.

Judge transcript

I need the, sorry, the final invoice attached to the March ticket.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact match of the reference and successfully captures the entire phrase, including the disfluency 'sorry', with perfect accuracy and punctuation.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors exist as the candidate transcript is identical to the reference.
Judge transcriptI need the, sorry, the final invoice attached to the March ticket.
Categories
no_error
Researcher notes
  • The model successfully handles self-correction and disfluencies ('I need the, sorry, the') without deleting them or hallucinating incorrect text.

Provenance

  • Category transcription accuracy wer
  • ASR slice disfluency token accounting
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-disfluency-001
  • Text SHA-256 d7554e1f20e8046ff869c7e16f61ebdc8335e982164130cb2d3c6c4bf9916eee
  • Duration 3.92s
  • Bytes 188084

Evaluation status: ok

transcription accuracy wer · sentence boundary punctuation

asr-wer-punctuation-boundary-001-local-tts

Whisper Large v3 Turbo (FP16)
98 accurate

Semantic impact

The transcript preserves the exact meaning of the audio, with only minor punctuation variations.

formatting only
Recommended action
  • The ASR model successfully recognized all words correctly.
  • Punctuation formatting could be improved to avoid trailing commas and comma splices, though meaning is fully preserved.

Judge transcript

No refunds after Friday. Exceptions require manager approval.
Key differences
  • Used a comma instead of a period after 'Friday'
  • Added a trailing comma at the end of the sentence
Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.00 (scores: 98, 98, 98). Representative reason: The candidate transcript accurately captures all spoken words. The only differences are punctuation-related, specifically using a comma instead of a period to separate the sentences and adding a trailing comma, neither of which impacts the overall meaning.

judge samples: 98, 98, 98; avg 98.00
Meaning
preserved
Semantic impact
The transcript preserves the exact meaning of the audio, with only minor punctuation variations.
Judge transcriptNo refunds after Friday. Exceptions require manager approval.
Key differences
  • Used a comma instead of a period after 'Friday'
  • Added a trailing comma at the end of the sentence
Categories
formatting_only
Researcher notes
  • The ASR model successfully recognized all words correctly.
  • Punctuation formatting could be improved to avoid trailing commas and comma splices, though meaning is fully preserved.

Provenance

  • Category transcription accuracy wer
  • ASR slice sentence boundary punctuation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-punctuation-boundary-001
  • Text SHA-256 30ce2ce13eaebf18c2d60223c27e6e644e491d4d24806ee559d1a36d9b6f76aa
  • Duration 3.64s
  • Bytes 174644

Evaluation status: ok

numeric unit integrity · decimal measurement

asr-numeric-measurement-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors are present in the candidate transcript.

formatting only
Recommended action
  • The candidate's choice to output numeric digits for measurements is highly readable and appropriate.
  • No errors detected.

Judge transcript

Set the incubator to 37.5 degrees Celsius.
Key differences
  • The number 'thirty-seven point five' is represented numerically as '37.5'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate and preserves the meaning of the audio. The only difference from the reference transcript is the formatting of the temperature as digits (37.5) instead of words, which is a standard and often preferred representation.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors are present in the candidate transcript.
Judge transcriptSet the incubator to 37.5 degrees Celsius.
Key differences
  • The number 'thirty-seven point five' is represented numerically as '37.5'.
Categories
formatting_only
Researcher notes
  • The candidate's choice to output numeric digits for measurements is highly readable and appropriate.
  • No errors detected.

Provenance

  • Category numeric unit integrity
  • ASR slice decimal measurement
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-measurement-001
  • Text SHA-256 1ba0e91325222fab47ab079798743fd1e6c43e845377315a824e82f8afe516d5
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

numeric unit integrity · digit sequence

asr-numeric-account-001-local-tts

Whisper Large v3 Turbo (FP16)
60 needs review

Semantic impact

The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.

substitutionnumber error
Recommended action
  • The model confused the spoken word 'eight' with the letter 'E' at the end of the account number. Improving acoustic modeling for digits vs. letters in alphanumeric strings could help prevent this.

Judge transcript

Read back account number 4902 7718 before submitting the form.
Key differences
  • Substituted '4902771E' for '4902 7718'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 60.00 (scores: 60, 55, 65). Representative reason: The candidate transcript correctly transcribed the entire sentence except for the final digit of the account number, where it substituted the number '8' with the capital letter 'E' ('4902771E' instead of '4902 7718'). While the rest of the sentence is perfectly accurate, an error in a critical identifier like an account number represents a partial semantic loss for downstream applications.

judge samples: 60, 55, 65; avg 60.00
Meaning
partial_loss
Semantic impact
The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
Judge transcriptRead back account number 4902 7718 before submitting the form.
Key differences
  • Substituted '4902771E' for '4902 7718'
Categories
substitution, number_error
Researcher notes
  • The model confused the spoken word 'eight' with the letter 'E' at the end of the account number. Improving acoustic modeling for digits vs. letters in alphanumeric strings could help prevent this.

Provenance

  • Category numeric unit integrity
  • ASR slice digit sequence
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-account-001
  • Text SHA-256 1b8373fa92691331a8512bd9fdb9eaa080397a75a368a630b825fb74f6d0f8ed
  • Duration 4.84s
  • Bytes 232244

Evaluation status: ok

numeric unit integrity · percentage minimal pair

asr-numeric-percentage-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors; formatting differences only.

formatting only
Recommended action
  • The model successfully captured the numerical values '90' and '19' and correctly associated them with percent symbols.
  • This is a high-quality transcript with acceptable formatting-only differences.

Judge transcript

The battery dropped from ninety percent to nineteen percent during the test.
Key differences
  • Spelled out 'ninety percent' represented as '90%'
  • Spelled out 'nineteen percent' represented as '19%'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate and preserves the exact meaning of the spoken audio. The only differences are standard formatting variations, representing written numbers and percentage symbols instead of spelled-out words.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; formatting differences only.
Judge transcriptThe battery dropped from ninety percent to nineteen percent during the test.
Key differences
  • Spelled out 'ninety percent' represented as '90%'
  • Spelled out 'nineteen percent' represented as '19%'
Categories
formatting_only
Researcher notes
  • The model successfully captured the numerical values '90' and '19' and correctly associated them with percent symbols.
  • This is a high-quality transcript with acceptable formatting-only differences.

Provenance

  • Category numeric unit integrity
  • ASR slice percentage minimal pair
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-percentage-001
  • Text SHA-256 aa7d137ea1db996c1e57892ce44ca8be6dc6df047decaa3e34ddc150a5335c24
  • Duration 3.88s
  • Bytes 186164

Evaluation status: ok

negation modality scope · clinical negation

asr-negation-medical-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors occurred; the transcript is perfectly accurate.

no error
Recommended action
  • Excellent transcription of medical terms and negation.

Judge transcript

The patient denies chest pain but reports shortness of breath.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript matches the audio and reference transcript perfectly with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors occurred; the transcript is perfectly accurate.
Judge transcriptThe patient denies chest pain but reports shortness of breath.
Categories
no_error
Researcher notes
  • Excellent transcription of medical terms and negation.

Provenance

  • Category negation modality scope
  • ASR slice clinical negation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-medical-001
  • Text SHA-256 dbdcbaeba42a61f71f6c3283ff950b99d754363b182b2ef052eff60f552edd82
  • Duration 3.44s
  • Bytes 165044

Evaluation status: ok

negation modality scope · permission scope

asr-negation-policy-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors were found as the transcription is completely accurate with only minor formatting/punctuation differences.

formatting only
Recommended action
  • The model performed perfectly. The punctuation difference is completely acceptable and does not alter the meaning or grammatical structure of the sentence.

Judge transcript

Only administrators may export records. Contractors may view records only.
Key differences
  • Semicolon in reference is replaced by a period in the candidate transcript.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate and captures the spoken words exactly. The only difference is the use of a period instead of a semicolon to split the two clauses, which is a benign formatting difference that has no impact on meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors were found as the transcription is completely accurate with only minor formatting/punctuation differences.
Judge transcriptOnly administrators may export records. Contractors may view records only.
Key differences
  • Semicolon in reference is replaced by a period in the candidate transcript.
Categories
formatting_only
Researcher notes
  • The model performed perfectly. The punctuation difference is completely acceptable and does not alter the meaning or grammatical structure of the sentence.

Provenance

  • Category negation modality scope
  • ASR slice permission scope
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-policy-001
  • Text SHA-256 599ef83653daccc61f7cd8f7daed7b7e910ab4c4919442065b793630923fb099
  • Duration 4.52s
  • Bytes 216884

Evaluation status: ok

negation modality scope · unless condition

asr-negation-condition-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The candidate transcript contains no semantic errors and perfectly captures the spoken message.

no error
Recommended action
  • The ASR system performed flawlessly on this conditional statement, capturing the critical words 'unless' and 'fails' accurately.

Judge transcript

Ship the replacement unless the inspection fails tomorrow morning.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match of the spoken audio and reference transcript, maintaining the exact conditional meaning of the instruction with zero errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The candidate transcript contains no semantic errors and perfectly captures the spoken message.
Judge transcriptShip the replacement unless the inspection fails tomorrow morning.
Categories
no_error
Researcher notes
  • The ASR system performed flawlessly on this conditional statement, capturing the critical words 'unless' and 'fails' accurately.

Provenance

  • Category negation modality scope
  • ASR slice unless condition
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-condition-001
  • Text SHA-256 82c70b514883f60d665a4cf44698b4dcb754c2636052cca7a7165104e40b4eb6
  • Duration 3.64s
  • Bytes 174644

Evaluation status: ok

temporal scheduling accuracy · am pm window

asr-temporal-window-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

There are no semantic errors or omissions in the candidate transcript.

formatting only
Recommended action
  • This is a flawless transcription from a semantic standpoint. The minor formatting variation in '10pm' is acceptable.

Judge transcript

Start the migration window tonight at 10 pm after traffic drops.
Key differences
  • The candidate writes '10pm' without a space, whereas the reference/audio implies '10 pm'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate and captures the audio exactly. The only difference between the candidate and the reference is a minor formatting choice, writing '10pm' as a single word instead of '10 pm', which has zero impact on meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors or omissions in the candidate transcript.
Judge transcriptStart the migration window tonight at 10 pm after traffic drops.
Key differences
  • The candidate writes '10pm' without a space, whereas the reference/audio implies '10 pm'.
Categories
formatting_only
Researcher notes
  • This is a flawless transcription from a semantic standpoint. The minor formatting variation in '10pm' is acceptable.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice am pm window
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-window-001
  • Text SHA-256 66ed82094ed39a1d36cd5387c7648754c1c6dc7807e73c30aee7e0773552281c
  • Duration 3.84s
  • Bytes 184244

Evaluation status: ok

temporal scheduling accuracy · relative time relation

asr-temporal-relative-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is completely accurate.

no error
Recommended action
  • Excellent performance with perfect alignment, proper punctuation, and correct spelling.

Judge transcript

Book the follow up two weeks after the first infusion, not before it.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact match to the audio and the reference transcript, with no errors of any kind.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is completely accurate.
Judge transcriptBook the follow up two weeks after the first infusion, not before it.
Categories
no_error
Researcher notes
  • Excellent performance with perfect alignment, proper punctuation, and correct spelling.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice relative time relation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-relative-001
  • Text SHA-256 683eb39c6dccaa56d07fc3c6fef174a36de3b354f5a5c17ad0cc0cc6dc211101
  • Duration 3.96s
  • Bytes 190004

Evaluation status: ok

temporal scheduling accuracy · ordinal date minimal pair

asr-temporal-date-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors exist; only formatting differences are present.

formatting only
Recommended action
  • The ASR output is correct. Converting spoken numbers/dates into standard digit format is expected behavior for many downstream tasks.

Judge transcript

Move the inspection from July fourth to July fourteenth.
Key differences
  • Used 'July 4th' instead of 'July fourth'
  • Used 'July 14th' instead of 'July fourteenth'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an accurate transcription of the audio, with the only differences being formatting-based (representing the dates with digits and ordinals rather than writing them out in text). This is a completely acceptable variation and preserves the meaning perfectly.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors exist; only formatting differences are present.
Judge transcriptMove the inspection from July fourth to July fourteenth.
Key differences
  • Used 'July 4th' instead of 'July fourth'
  • Used 'July 14th' instead of 'July fourteenth'
Categories
formatting_only
Researcher notes
  • The ASR output is correct. Converting spoken numbers/dates into standard digit format is expected behavior for many downstream tasks.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice ordinal date minimal pair
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-date-001
  • Text SHA-256 e79ea5b3319ccee0d2c33d7f53824ce7a080f06aa60c47f39f725a4209609d4a
  • Duration 3.20s
  • Bytes 153524

Evaluation status: ok

temporal scheduling accuracy · duration contrast

asr-temporal-duration-001-local-tts

Whisper Large v3 Turbo (FP16)
53 inaccurate

Semantic impact

The pump's rest duration was incorrectly transcribed as 15 minutes instead of fifty minutes.

number errorsubstitution
Recommended action
  • The model confused 'fifty' with '15' ('fifteen'), which is a classic phonetic ambiguity. Acoustic modeling or language model priors for numeric sequences should be improved to better distinguish teen/ty suffixes.

Judge transcript

Run the pump for fifteen minutes, then let it rest for fifty minutes.
Key differences
  • Candidate transcribed '15 minutes' instead of the spoken 'fifty minutes' (50) for the resting duration.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 53.33 (scores: 55, 55, 50). Representative reason: The candidate transcript contains a critical numeric error, transcribing the resting duration as '15 minutes' instead of '50 minutes' ('fifty minutes'). This changes the instruction's meaning, potentially causing downstream operational issues or equipment damage if implemented.

judge samples: 55, 55, 50; avg 53.33
Meaning
partial_loss
Semantic impact
The pump's rest duration was incorrectly transcribed as 15 minutes instead of fifty minutes.
Judge transcriptRun the pump for fifteen minutes, then let it rest for fifty minutes.
Key differences
  • Candidate transcribed '15 minutes' instead of the spoken 'fifty minutes' (50) for the resting duration.
Categories
number_error, substitution
Researcher notes
  • The model confused 'fifty' with '15' ('fifteen'), which is a classic phonetic ambiguity. Acoustic modeling or language model priors for numeric sequences should be improved to better distinguish teen/ty suffixes.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice duration contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-duration-001
  • Text SHA-256 46846ed936375faff65199b095184a0ce00cea6cb427f8e34c78a31cc4374a2c
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

numeric unit integrity · amount contrast

asr-numeric-transfer-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors exist; the changes are strictly formatting differences.

formatting only
Recommended action
  • The model correctly recognized the distinction between '$15' and '$50', which can be easily confused in audio.
  • The text-to-number/currency conversion is highly accurate and appropriate for downstream applications.

Judge transcript

Transfer fifteen dollars to Maya today, not fifty dollars.
Key differences
  • The spoken phrase 'fifteen dollars' is written as '$15'.
  • The spoken phrase 'fifty dollars' is written as '$50'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate, with the only difference being the clean formatting of 'fifteen dollars' to '$15' and 'fifty dollars' to '$50'. The semantic meaning is fully preserved without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors exist; the changes are strictly formatting differences.
Judge transcriptTransfer fifteen dollars to Maya today, not fifty dollars.
Key differences
  • The spoken phrase 'fifteen dollars' is written as '$15'.
  • The spoken phrase 'fifty dollars' is written as '$50'.
Categories
formatting_only
Researcher notes
  • The model correctly recognized the distinction between '$15' and '$50', which can be easily confused in audio.
  • The text-to-number/currency conversion is highly accurate and appropriate for downstream applications.

Provenance

  • Category numeric unit integrity
  • ASR slice amount contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-transfer-001
  • Text SHA-256 387a5443dbcbe43becc7f50aee52cd98ecd1503a46eb68183fbf58045f00b49c
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

numeric unit integrity · dosage unit

asr-numeric-medication-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The transcript is fully accurate with only a minor formatting difference in the representation of the number five.

formatting only
Recommended action
  • The model correctly transcribed the critical medication name 'melatonin' and dosage '5 milligrams'.
  • Digit normalization (5 vs five) is acceptable and should not be penalized as an error.

Judge transcript

Give the patient five milligrams of melatonin before bedtime.
Key differences
  • The word 'five' is formatted as the digit '5'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly preserves the meaning of the spoken audio. The only difference is the representation of the number 'five' as the digit '5', which is a standard formatting variation and does not affect the semantic accuracy of the medication dosage.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript is fully accurate with only a minor formatting difference in the representation of the number five.
Judge transcriptGive the patient five milligrams of melatonin before bedtime.
Key differences
  • The word 'five' is formatted as the digit '5'.
Categories
formatting_only
Researcher notes
  • The model correctly transcribed the critical medication name 'melatonin' and dosage '5 milligrams'.
  • Digit normalization (5 vs five) is acceptable and should not be penalized as an error.

Provenance

  • Category numeric unit integrity
  • ASR slice dosage unit
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-medication-001
  • Text SHA-256 566164fc8bcb2a1adb3b5929af80f8b1755e68f61d2967650ec925d8043b6a8f
  • Duration 3.76s
  • Bytes 180404

Evaluation status: ok

negation modality scope · negated action

asr-negation-server-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

There are no semantic errors; the transcription is perfectly accurate.

no error
Recommended action
  • The model performed perfectly on this clean, clear TTS audio sample containing a critical negation instruction.

Judge transcript

Do not restart the server until the database backup finishes.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match with the spoken audio and the reference transcript, preserving the exact meaning with zero errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the transcription is perfectly accurate.
Judge transcriptDo not restart the server until the database backup finishes.
Categories
no_error
Researcher notes
  • The model performed perfectly on this clean, clear TTS audio sample containing a critical negation instruction.

Provenance

  • Category negation modality scope
  • ASR slice negated action
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-server-001
  • Text SHA-256 9a47d14232d15170b7f3a734e6179db244d25afef06121759447c4f8a318d276
  • Duration 3.28s
  • Bytes 157364

Evaluation status: ok

negation modality scope · permission prohibition

asr-negation-refund-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors are present as the candidate transcript perfectly preserves the spoken meaning.

formatting only
Recommended action
  • Excellent transcript with zero semantic errors.

Judge transcript

The customer can cancel the trial, but cannot receive a refund after Friday.
Key differences
  • omission of a comma before 'but'
Judge rationale, vote details, and provenance

Average of 2 successful judge samples (1 failed attempt excluded): 100.00 (scores: 100, 100). Representative reason: The candidate transcript is completely accurate and matches the spoken audio and reference transcript perfectly, with only a minor, harmless punctuation difference (the omission of a comma).

judge samples: 100, 100; avg 100.00; 2/3 attempts succeeded; 1 failed attempt excluded
Meaning
preserved
Semantic impact
No semantic errors are present as the candidate transcript perfectly preserves the spoken meaning.
Judge transcriptThe customer can cancel the trial, but cannot receive a refund after Friday.
Key differences
  • omission of a comma before 'but'
Categories
formatting_only
Researcher notes
  • Excellent transcript with zero semantic errors.

Provenance

  • Category negation modality scope
  • ASR slice permission prohibition
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-refund-001
  • Text SHA-256 55a42c2151a5880e0aa2f00f13860c2acf87f536e24d9cec099342e0da50f211
  • Duration 3.80s
  • Bytes 182324

Evaluation status: ok

temporal scheduling accuracy · deadline day time

asr-temporal-deadline-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

There are no semantic errors; the candidate is completely accurate with only a minor formatting variation.

formatting only
Recommended action
  • Perfect transcription. The minor punctuation difference in the time period indicator ('a.m.' vs 'am') has zero impact on downstream semantic tasks.

Judge transcript

The benefits enrollment deadline is Monday at 9 a.m.
Key differences
  • The candidate uses 'a.m.' whereas the reference uses 'am'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate, with only a minor, harmless formatting difference in the abbreviation of 'a.m.' versus 'am'. The meaning of the audio is fully preserved.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate is completely accurate with only a minor formatting variation.
Judge transcriptThe benefits enrollment deadline is Monday at 9 a.m.
Key differences
  • The candidate uses 'a.m.' whereas the reference uses 'am'.
Categories
formatting_only
Researcher notes
  • Perfect transcription. The minor punctuation difference in the time period indicator ('a.m.' vs 'am') has zero impact on downstream semantic tasks.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice deadline day time
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-deadline-001
  • Text SHA-256 b118c3c195ab7d6fd58def5f41f4f899aa49aba40735806279b2f743868df37c
  • Duration 3.04s
  • Bytes 145844

Evaluation status: ok

entity factual integrity · person place entity

asr-entity-doctor-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The transcript is completely accurate with only a minor, harmless capitalization difference.

formatting only
Recommended action
  • Excellent performance. The transcript accurately captures the doctor's name, 'Rao', and the location 'Cedar Avenue clinic' without any errors.

Judge transcript

Schedule the follow-up with Dr. Rao at the Cedar Avenue clinic.
Key differences
  • Capitalized 'Clinic' instead of lowercase 'clinic'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is identical to the reference and audio, with only a minor capitalization difference on the word 'Clinic' which has no impact on meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript is completely accurate with only a minor, harmless capitalization difference.
Judge transcriptSchedule the follow-up with Dr. Rao at the Cedar Avenue clinic.
Key differences
  • Capitalized 'Clinic' instead of lowercase 'clinic'
Categories
formatting_only
Researcher notes
  • Excellent performance. The transcript accurately captures the doctor's name, 'Rao', and the location 'Cedar Avenue clinic' without any errors.

Provenance

  • Category entity factual integrity
  • ASR slice person place entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-doctor-001
  • Text SHA-256 9ff5a0befafca6c4d8752ec9817f898caed7717e40beb76cd3eaa239a8cf9473
  • Duration 3.40s
  • Bytes 163124

Evaluation status: ok

entity factual integrity · product name contrast

asr-entity-product-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors; the transcription is perfectly accurate.

no error
Recommended action
  • Perfect transcription with no issues to report.

Judge transcript

Compare the Zephyr Mini plan with the Zephyr Max plan before renewal.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact match to both the reference transcript and the spoken audio, preserving all semantic meaning with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcription is perfectly accurate.
Judge transcriptCompare the Zephyr Mini plan with the Zephyr Max plan before renewal.
Categories
no_error
Researcher notes
  • Perfect transcription with no issues to report.

Provenance

  • Category entity factual integrity
  • ASR slice product name contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-product-001
  • Text SHA-256 919dd10bf1c803fb423936e4f48c18d6469b7753fa5383a94ff26fc32bf6d745
  • Duration 3.84s
  • Bytes 184244

Evaluation status: ok

entity factual integrity · alphanumeric identifier

asr-entity-ticket-001-local-tts

Whisper Large v3 Turbo (FP16)
53 inaccurate

Semantic impact

The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.

entity errornumber errorformatting only
Recommended action
  • The model struggled with the alphanumeric ticket ID, merging the word 'ticket' with the identifier and misrecognizing the digits '4187' as '457'.
  • Ensure better training or language model biasing for alphanumeric codes and digit sequences.

Judge transcript

Attach the call summary to ticket PX4187 before closing the case.
Key differences
  • transcribed 'TicketPX457' instead of 'ticket PX-4187'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 53.33 (scores: 50, 55, 55). Representative reason: The candidate transcript contains a critical entity error in the ticket ID, transcribing 'ticket PX-4187' as 'TicketPX457'. This error alters a key identifier (missing the digits '1' and '8', and incorrectly inserting '5'), which would cause downstream systems to attach the call summary to the wrong ticket or fail entirely due to a non-existent ID.

judge samples: 50, 55, 55; avg 53.33
Meaning
partial_loss
Semantic impact
The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
Judge transcriptAttach the call summary to ticket PX4187 before closing the case.
Key differences
  • transcribed 'TicketPX457' instead of 'ticket PX-4187'
Categories
entity_error, number_error, formatting_only
Researcher notes
  • The model struggled with the alphanumeric ticket ID, merging the word 'ticket' with the identifier and misrecognizing the digits '4187' as '457'.
  • Ensure better training or language model biasing for alphanumeric codes and digit sequences.

Provenance

  • Category entity factual integrity
  • ASR slice alphanumeric identifier
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-ticket-001
  • Text SHA-256 845a5229d7c30013391f3297d1a9cc591ee108b93c4cf22218cdb749913fbf6e
  • Duration 4.60s
  • Bytes 220724

Evaluation status: ok

entity factual integrity · organization contact entity

asr-entity-organization-001-local-tts

Whisper Large v3 Turbo (FP16)
90 accurate

Semantic impact

The candidate transcript contains a minor name spelling variation ('Alina' vs 'Elena') but preserves all other information perfectly.

substitutionentity error
Recommended action
  • The substitution of 'Alina' for 'Elena' is phonetically plausible based on the audio. Consider utilizing context-aware entity correction or user-specific contact lists to disambiguate similar-sounding names.

Judge transcript

Send the signed form to Northstar Benefits and copy Elena Morales.
Key differences
  • Candidate spelled the name as 'Alina' while the reference/audio indicates 'Elena'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 90.00 (scores: 80, 95, 95). Representative reason: The candidate transcript is extremely accurate, with the only discrepancy being a minor spelling variation of the name 'Alina' instead of 'Elena'. This phonetically similar substitution does not affect the overall meaning or utility of the transcript.

judge samples: 80, 95, 95; avg 90.00
Meaning
preserved
Semantic impact
The candidate transcript contains a minor name spelling variation ('Alina' vs 'Elena') but preserves all other information perfectly.
Judge transcriptSend the signed form to Northstar Benefits and copy Elena Morales.
Key differences
  • Candidate spelled the name as 'Alina' while the reference/audio indicates 'Elena'.
Categories
substitution, entity_error
Researcher notes
  • The substitution of 'Alina' for 'Elena' is phonetically plausible based on the audio. Consider utilizing context-aware entity correction or user-specific contact lists to disambiguate similar-sounding names.

Provenance

  • Category entity factual integrity
  • ASR slice organization contact entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-organization-001
  • Text SHA-256 2c9213d5e33f7fc5a5924f6f72b4f91505c0b4c71b26848cb000e2d00ef961bc
  • Duration 3.80s
  • Bytes 182324

Evaluation status: ok

entity factual integrity · address entity

asr-entity-address-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The candidate transcript is perfectly accurate and preserves all meaning.

no error
Recommended action
  • Perfect transcription of the address and numbers. No improvements needed.

Judge transcript

The delivery address is 812 Ashwood Lane, apartment 5B.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact and perfect match of the spoken audio, correctly capturing the entire address including numbers and apartment details without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The candidate transcript is perfectly accurate and preserves all meaning.
Judge transcriptThe delivery address is 812 Ashwood Lane, apartment 5B.
Categories
no_error
Researcher notes
  • Perfect transcription of the address and numbers. No improvements needed.

Provenance

  • Category entity factual integrity
  • ASR slice address entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-address-001
  • Text SHA-256 5cf2d55b0dddffcc8b60dc856eb7816a36e745257daaf10abd184d3fcd40a9dc
  • Duration 4.24s
  • Bytes 203444

Evaluation status: ok

semantic paraphrase preservation · event state preservation

asr-semantic-paraphrase-001-local-tts

Whisper Large v3 Turbo (FP16)
97 accurate

Semantic impact

The candidate transcript uses 'stop' instead of 'stopped', which is a negligible grammatical difference that preserves the full meaning of the sentence.

substitution
Recommended action
  • The model failed to capture the final word's past-tense inflection, which could be addressed by tuning the language model or improving acoustic modeling of word endings.

Judge transcript

The technician replaced the cracked valve and confirmed the leak stopped.
Key differences
  • transcribed 'stop' instead of 'stopped' at the end of the sentence
Judge rationale, vote details, and provenance

Average of 3 judge samples: 97.00 (scores: 95, 98, 98). Representative reason: The candidate transcript is almost entirely accurate, only missing the past-tense suffix '-ed' on the final word, transcribing 'stopped' as 'stop'. This minor grammatical variation does not impact the overall meaning or usability of the transcript.

judge samples: 95, 98, 98; avg 97.00
Meaning
preserved
Semantic impact
The candidate transcript uses 'stop' instead of 'stopped', which is a negligible grammatical difference that preserves the full meaning of the sentence.
Judge transcriptThe technician replaced the cracked valve and confirmed the leak stopped.
Key differences
  • transcribed 'stop' instead of 'stopped' at the end of the sentence
Categories
substitution
Researcher notes
  • The model failed to capture the final word's past-tense inflection, which could be addressed by tuning the language model or improving acoustic modeling of word endings.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice event state preservation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-paraphrase-001
  • Text SHA-256 459210fe6f6f4d41259ceba39ab9c8fc608abfb927c5a49c8750ac89f86e21d0
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

semantic paraphrase preservation · causal relation

asr-semantic-causal-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

No semantic errors were detected; the transcription is completely accurate.

no error
Recommended action
  • The ASR system performed flawlessly on this clean text-to-speech audio sample, successfully capturing specialized technical terminology like 'token expired' and 'client retried after refresh'.

Judge transcript

The upload failed because the token expired, so the client retried after refresh.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly matches the spoken audio and the reference transcript, with no omissions, insertions, or substitutions.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors were detected; the transcription is completely accurate.
Judge transcriptThe upload failed because the token expired, so the client retried after refresh.
Categories
no_error
Researcher notes
  • The ASR system performed flawlessly on this clean text-to-speech audio sample, successfully capturing specialized technical terminology like 'token expired' and 'client retried after refresh'.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice causal relation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-causal-001
  • Text SHA-256 caeb5d404f654e7d980c3f61cc1576806deb874cf24af2120f26700bb7aa471e
  • Duration 4.72s
  • Bytes 226484

Evaluation status: ok

semantic paraphrase preservation · metric tradeoff

asr-semantic-comparison-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

There are no semantic errors; the transcription is perfect.

no error
Recommended action
  • The model achieved 100% accuracy on this audio clip.

Judge transcript

the candidate model improved recall but reduced precision on noisy calls

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match of the spoken audio, preserving all semantic meaning with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the transcription is perfect.
Judge transcriptthe candidate model improved recall but reduced precision on noisy calls
Categories
no_error
Researcher notes
  • The model achieved 100% accuracy on this audio clip.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice metric tradeoff
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-comparison-001
  • Text SHA-256 fbaadae669f95ffc77afc2da5b9661b0a48486f7e71ada1bd05170df2649a739
  • Duration 4.32s
  • Bytes 207284

Evaluation status: ok

semantic paraphrase preservation · coreference ambiguity

asr-semantic-coreference-001-local-tts

Whisper Large v3 Turbo (FP16)
91 accurate

Semantic impact

The candidate matches the spoken audio perfectly, with only a minor name discrepancy relative to the reference transcript.

no error
Recommended action
  • The audio actually says 'Rhea' / 'Ria' instead of 'Maria'. The candidate's transcription is correct relative to the audio. The reference transcript may contain a typo or discrepancy compared to the actual TTS audio generation.

Judge transcript

Rhea called Priya after she reviewed the contract changes.
Key differences
  • Candidate transcribes 'Rhea' matching the audio, whereas the reference transcript uses 'Maria'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 91.00 (scores: 75, 98, 100). Representative reason: The candidate transcript is an extremely accurate transcription of the audio. Although the reference transcript notes the name as 'Maria', the audio clearly pronounces 'Rhea' (or 'Ria'), which the candidate captures perfectly. There is no loss of meaning compared to the spoken audio.

judge samples: 75, 98, 100; avg 91.00
Meaning
preserved
Semantic impact
The candidate matches the spoken audio perfectly, with only a minor name discrepancy relative to the reference transcript.
Judge transcriptRhea called Priya after she reviewed the contract changes.
Key differences
  • Candidate transcribes 'Rhea' matching the audio, whereas the reference transcript uses 'Maria'.
Categories
no_error
Researcher notes
  • The audio actually says 'Rhea' / 'Ria' instead of 'Maria'. The candidate's transcription is correct relative to the audio. The reference transcript may contain a typo or discrepancy compared to the actual TTS audio generation.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice coreference ambiguity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-coreference-001
  • Text SHA-256 7ec151a16318fda290762dbf9f9ad83e2ce7c5df725aeba3043965cfd596e437
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

semantic paraphrase preservation · conditional instruction

asr-semantic-instruction-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The candidate transcript is a perfect transcription of the audio with no errors.

no error
Recommended action
  • Perfect match. The system recognized all words, phrasing, and punctuation correctly.

Judge transcript

If the first backup succeeds, skip the manual export and notify operations.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact, flawless match to the spoken audio and the reference transcript, preserving the meaning perfectly with zero errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The candidate transcript is a perfect transcription of the audio with no errors.
Judge transcriptIf the first backup succeeds, skip the manual export and notify operations.
Categories
no_error
Researcher notes
  • Perfect match. The system recognized all words, phrasing, and punctuation correctly.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice conditional instruction
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-instruction-001
  • Text SHA-256 c095d3a9fadb208cbc7ae9c12d4168e37dff40b73108cd2b28c219e12511a104
  • Duration 4.48s
  • Bytes 214964

Evaluation status: ok

acoustic noise robustness · cafe background command

asr-noise-cafe-order-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

None. The candidate transcript is perfectly accurate.

no error
Recommended action
  • The ASR model transcribed the audio flawlessly, preserving all intent and entity information.

Judge transcript

Add oat milk to the latte and leave out the caramel drizzle.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an absolute match to the reference transcript and the spoken audio. It perfectly captures the user's request without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
None. The candidate transcript is perfectly accurate.
Judge transcriptAdd oat milk to the latte and leave out the caramel drizzle.
Categories
no_error
Researcher notes
  • The ASR model transcribed the audio flawlessly, preserving all intent and entity information.

Provenance

  • Category acoustic noise robustness
  • ASR slice cafe background command
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-cafe-order-001
  • Text SHA-256 613cfd4c60424169e80cba7ab2ecc6f0fcb35f677cdfbfd9746b4efaebf165ef
  • Duration 3.28s
  • Bytes 157364

Evaluation status: ok

acoustic noise robustness · vehicle noise navigation

asr-noise-driving-route-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

There are no semantic errors; the candidate transcript is 100% accurate.

no error
Recommended action
  • The model achieved perfect accuracy on this short navigational instruction.

Judge transcript

Take the second exit after the bridge, then stay in the right lane.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect match with the audio and the reference transcript, capturing every word flawlessly with no semantic or formatting errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate transcript is 100% accurate.
Judge transcriptTake the second exit after the bridge, then stay in the right lane.
Categories
no_error
Researcher notes
  • The model achieved perfect accuracy on this short navigational instruction.

Provenance

  • Category acoustic noise robustness
  • ASR slice vehicle noise navigation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-driving-route-001
  • Text SHA-256 c3ce8e7e94fd6ea30d703cbca09fc8f6c03373f0ea379bfc5d94dd49aa48ce26
  • Duration 3.72s
  • Bytes 178484

Evaluation status: ok

acoustic noise robustness · industrial noise identifier

asr-noise-warehouse-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The candidate transcript preserves the meaning perfectly with only a minor formatting difference in the representation of the number three.

formatting only
Recommended action
  • The ASR output is highly accurate. Representing 'three' as '3' is a standard and acceptable variation in ASR systems.

Judge transcript

Scan bin A17 before moving the pallet to loading dock three.
Key differences
  • The word 'three' is formatted as the digit '3'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate, preserving the meaning of the spoken audio with the only difference being the formatting of the number 'three' as the digit '3'.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The candidate transcript preserves the meaning perfectly with only a minor formatting difference in the representation of the number three.
Judge transcriptScan bin A17 before moving the pallet to loading dock three.
Key differences
  • The word 'three' is formatted as the digit '3'
Categories
formatting_only
Researcher notes
  • The ASR output is highly accurate. Representing 'three' as '3' is a standard and acceptable variation in ASR systems.

Provenance

  • Category acoustic noise robustness
  • ASR slice industrial noise identifier
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-warehouse-001
  • Text SHA-256 f7c26924015438634f3eb4be84b886d690e9009b39ecd8f4fa88408ff3ec2e11
  • Duration 4.08s
  • Bytes 195764

Evaluation status: ok

acoustic noise robustness · reverberant meeting room

asr-noise-reverberant-room-001-local-tts

Whisper Large v3 Turbo (FP16)
83 accurate

Semantic impact

The candidate correctly transcribes the spoken room number 'four twenty twelve' which deviates from the reference text due to a TTS pronunciation error.

no error
Recommended action
  • The TTS model generated 'four twenty twelve' instead of 'four twelve' or 'four one two' for '412'. The ASR model performed perfectly by transcribing what was actually uttered.

Judge transcript

The quarterly review starts in room four twenty twelve at half past three.
Key differences
  • Candidate transcribes '4-20-12' matching the spoken 'four twenty twelve', while the reference text lists '412'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 82.67 (scores: 55, 98, 95). Representative reason: The candidate transcript is an extremely accurate transcription of the actual audio. The TTS synthesis apparently mispronounced '412' as 'four twenty twelve', which the candidate correctly transcribed as '4-20-12'. The candidate should not be penalized for faithfully transcribing what was actually spoken instead of matching the erroneous reference.

judge samples: 55, 98, 95; avg 82.67
Meaning
preserved
Semantic impact
The candidate correctly transcribes the spoken room number 'four twenty twelve' which deviates from the reference text due to a TTS pronunciation error.
Judge transcriptThe quarterly review starts in room four twenty twelve at half past three.
Key differences
  • Candidate transcribes '4-20-12' matching the spoken 'four twenty twelve', while the reference text lists '412'.
Categories
no_error
Researcher notes
  • The TTS model generated 'four twenty twelve' instead of 'four twelve' or 'four one two' for '412'. The ASR model performed perfectly by transcribing what was actually uttered.

Provenance

  • Category acoustic noise robustness
  • ASR slice reverberant meeting room
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-reverberant-room-001
  • Text SHA-256 f0f9e458a7bf2804ef37a0a709c4a7882c50036351a8f11cd69c590f1805b4d8
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

acoustic noise robustness · overlapping speech target

asr-noise-overlapping-speech-001-local-tts

Whisper Large v3 Turbo (FP16)
100 accurate

Semantic impact

The candidate transcript preserves the meaning of the audio perfectly.

no error
Recommended action
  • The model performed perfectly on this clean utterance.

Judge transcript

Ignore the side conversation and record the final answer as option C.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect transcription of the spoken audio, matching both the reference and the audio exactly without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The candidate transcript preserves the meaning of the audio perfectly.
Judge transcriptIgnore the side conversation and record the final answer as option C.
Categories
no_error
Researcher notes
  • The model performed perfectly on this clean utterance.

Provenance

  • Category acoustic noise robustness
  • ASR slice overlapping speech target
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/whisper-large-v3-turbo-asr-fp16
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-overlapping-speech-001
  • Text SHA-256 e72cbb8fa9f427ff6a5ac7dcb1cf1e153ec37baeeba0a80c6be4033169b3f66b
  • Duration 4.08s
  • Bytes 195764

Evaluation status: ok

transcription accuracy wer · function word substitution

asr-wer-function-word-001-local-tts

Qwen3 ASR 1.7B (8-bit)
87 accurate

Semantic impact

The candidate substituted 'patch is' with 'patches', slightly altering the grammar but preserving the primary meaning.

substitution
Recommended action
  • The model collapsed the noun + auxiliary verb 'patch is' into the phonetically similar plural noun 'patches'. Improving acoustic modeling for weak auxiliary verbs following sibilant-ending words would help prevent this error.

Judge transcript

The release notes say the patch is for the beta channel, not the stable channel.
Key differences
  • Substituted 'patch is' with 'patches'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 86.67 (scores: 80, 85, 95). Representative reason: The candidate transcript substitutes the singular noun and verb 'patch is' with the plural noun 'patches'. While this introduces a minor grammatical shift, the overall meaning—that the update is meant for the beta channel instead of the stable channel—remains completely clear and unambiguous.

judge samples: 80, 85, 95; avg 86.67
Meaning
minor_loss
Semantic impact
The candidate substituted 'patch is' with 'patches', slightly altering the grammar but preserving the primary meaning.
Judge transcriptThe release notes say the patch is for the beta channel, not the stable channel.
Key differences
  • Substituted 'patch is' with 'patches'
Categories
substitution
Researcher notes
  • The model collapsed the noun + auxiliary verb 'patch is' into the phonetically similar plural noun 'patches'. Improving acoustic modeling for weak auxiliary verbs following sibilant-ending words would help prevent this error.

Provenance

  • Category transcription accuracy wer
  • ASR slice function word substitution
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-function-word-001
  • Text SHA-256 db881f12982b7e1bcdb006f5bcb9ce39ddcb8ce29bfa65b29176ffcd5301d3a7
  • Duration 4.12s
  • Bytes 197684

Evaluation status: ok

transcription accuracy wer · homophone confusion

asr-wer-homophone-001-local-tts

Qwen3 ASR 1.7B (8-bit)
77 needs review

Semantic impact

The homophone error 'write' instead of 'right' slightly degrades the semantic clarity of the sentence by creating an illogical phrase 'write team'.

substitution
Recommended action
  • Improve context-aware language modeling to correctly distinguish homophones like 'write' and 'right' when modifying nouns like 'team'.

Judge transcript

Please write the new access policy before the right team reviews it.
Key differences
  • Used 'write team' instead of 'right team'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 76.67 (scores: 75, 75, 80). Representative reason: The candidate transcript contains a homophone substitution, writing 'write' instead of 'right' in the phrase 'right team reviews it'. While the overall sentence structure and other words are perfectly captured, this spelling mistake changes the meaning of the noun modifier from 'correct/appropriate' to 'writing'.

judge samples: 75, 75, 80; avg 76.67
Meaning
minor_loss
Semantic impact
The homophone error 'write' instead of 'right' slightly degrades the semantic clarity of the sentence by creating an illogical phrase 'write team'.
Judge transcriptPlease write the new access policy before the right team reviews it.
Key differences
  • Used 'write team' instead of 'right team'
Categories
substitution
Researcher notes
  • Improve context-aware language modeling to correctly distinguish homophones like 'write' and 'right' when modifying nouns like 'team'.

Provenance

  • Category transcription accuracy wer
  • ASR slice homophone confusion
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-homophone-001
  • Text SHA-256 e98bb2cafe5fd5b2ad5f26020053ff27ca22a0e7c17c904647d2449af7b4657e
  • Duration 3.24s
  • Bytes 155444

Evaluation status: ok

transcription accuracy wer · short command edit distance

asr-wer-short-utterance-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic loss; only minor capitalization differences.

formatting only
Recommended action
  • The model performed perfectly, with only insignificant capitalization differences.

Judge transcript

Turn left at Pine Street, then park by gate four.
Key differences
  • Capitalization of 'Gate Four' instead of 'gate four'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is identical to the audio and reference, with only a minor capitalization difference ('Gate Four' instead of 'gate four'), which does not affect the meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic loss; only minor capitalization differences.
Judge transcriptTurn left at Pine Street, then park by gate four.
Key differences
  • Capitalization of 'Gate Four' instead of 'gate four'
Categories
formatting_only
Researcher notes
  • The model performed perfectly, with only insignificant capitalization differences.

Provenance

  • Category transcription accuracy wer
  • ASR slice short command edit distance
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-short-utterance-001
  • Text SHA-256 f0270a081ebe4d9733c17ae6c19c100028c61723a25cf10de9aa04db0e29e736
  • Duration 2.96s
  • Bytes 142004

Evaluation status: ok

transcription accuracy wer · disfluency token accounting

asr-wer-disfluency-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors in the candidate transcript; it is a word-for-word correct transcription.

formatting only
Recommended action
  • The candidate transcript successfully captures disfluencies ('sorry') without deleting words or hallucinating. The error is strictly a minor punctuation difference.

Judge transcript

I need the, sorry, the final invoice attached to the March ticket.
Key differences
  • missing commas around the word 'sorry'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly captures the spoken words, including the disfluency, with the only difference being the omission of commas, which has no impact on meaning preservation.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors in the candidate transcript; it is a word-for-word correct transcription.
Judge transcriptI need the, sorry, the final invoice attached to the March ticket.
Key differences
  • missing commas around the word 'sorry'
Categories
formatting_only
Researcher notes
  • The candidate transcript successfully captures disfluencies ('sorry') without deleting words or hallucinating. The error is strictly a minor punctuation difference.

Provenance

  • Category transcription accuracy wer
  • ASR slice disfluency token accounting
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-disfluency-001
  • Text SHA-256 d7554e1f20e8046ff869c7e16f61ebdc8335e982164130cb2d3c6c4bf9916eee
  • Duration 3.92s
  • Bytes 188084

Evaluation status: ok

transcription accuracy wer · sentence boundary punctuation

asr-wer-punctuation-boundary-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

The transcript is completely accurate with no semantic loss.

no error
Recommended action
  • Perfect transcription, including capitalization and punctuation.

Judge transcript

No refunds after Friday. Exceptions require manager approval.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript matches the reference and the audio perfectly, with absolutely no errors in transcription, punctuation, or spelling.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript is completely accurate with no semantic loss.
Judge transcriptNo refunds after Friday. Exceptions require manager approval.
Categories
no_error
Researcher notes
  • Perfect transcription, including capitalization and punctuation.

Provenance

  • Category transcription accuracy wer
  • ASR slice sentence boundary punctuation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-punctuation-boundary-001
  • Text SHA-256 30ce2ce13eaebf18c2d60223c27e6e644e491d4d24806ee559d1a36d9b6f76aa
  • Duration 3.64s
  • Bytes 174644

Evaluation status: ok

numeric unit integrity · decimal measurement

asr-numeric-measurement-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the spoken words and numeric representations match in meaning perfectly.

formatting only
Recommended action
  • The model correctly transcribed the measurement. The numeric formatting (37.5) is often preferred for downstream tasks over the fully spelled-out words.

Judge transcript

Set the incubator to 37.5 degrees Celsius.
Key differences
  • The candidate transcript uses digits '37.5' instead of spelled-out words 'thirty-seven point five'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and preserves the meaning perfectly, with only a minor formatting difference where the spoken number 'thirty-seven point five' is represented numerically as '37.5'.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the spoken words and numeric representations match in meaning perfectly.
Judge transcriptSet the incubator to 37.5 degrees Celsius.
Key differences
  • The candidate transcript uses digits '37.5' instead of spelled-out words 'thirty-seven point five'.
Categories
formatting_only
Researcher notes
  • The model correctly transcribed the measurement. The numeric formatting (37.5) is often preferred for downstream tasks over the fully spelled-out words.

Provenance

  • Category numeric unit integrity
  • ASR slice decimal measurement
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-measurement-001
  • Text SHA-256 1ba0e91325222fab47ab079798743fd1e6c43e845377315a824e82f8afe516d5
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

numeric unit integrity · digit sequence

asr-numeric-account-001-local-tts

Qwen3 ASR 1.7B (8-bit)
60 needs review

Semantic impact

The candidate incorrectly transcribes the final digit of the account number, substituting '8' with the letter 'E'.

substitutionnumber error
Recommended action
  • The model confused the phonetically similar sounds of 'eight' (/eɪt/) and 'E' (/iː/) at the end of the numeric sequence.
  • Ensure acoustic models are better trained on distinguishing cardinal numbers from letters in alphanumeric contexts.

Judge transcript

Read back account number 4902 7718 before submitting the form.
Key differences
  • Candidate substituted '8' with 'E' at the end of the account number.
  • Candidate omitted the space separating the two blocks of the account number.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 60.00 (scores: 70, 60, 50). Representative reason: The candidate transcript accurately captures almost all of the spoken sentence, but it misrecognizes the final digit of the account number, transcribing '8' as 'E'. In the context of an account number, this single character substitution is a critical error that would lead to downstream failures (e.g., failed lookup or validation), representing a partial semantic loss.

judge samples: 70, 60, 50; avg 60.00
Meaning
partial_loss
Semantic impact
The candidate incorrectly transcribes the final digit of the account number, substituting '8' with the letter 'E'.
Judge transcriptRead back account number 4902 7718 before submitting the form.
Key differences
  • Candidate substituted '8' with 'E' at the end of the account number.
  • Candidate omitted the space separating the two blocks of the account number.
Categories
substitution, number_error
Researcher notes
  • The model confused the phonetically similar sounds of 'eight' (/eɪt/) and 'E' (/iː/) at the end of the numeric sequence.
  • Ensure acoustic models are better trained on distinguishing cardinal numbers from letters in alphanumeric contexts.

Provenance

  • Category numeric unit integrity
  • ASR slice digit sequence
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-account-001
  • Text SHA-256 1b8373fa92691331a8512bd9fdb9eaa080397a75a368a630b825fb74f6d0f8ed
  • Duration 4.84s
  • Bytes 232244

Evaluation status: ok

numeric unit integrity · percentage minimal pair

asr-numeric-percentage-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

The transcript preserves the meaning perfectly with only standard formatting differences.

formatting only
Recommended action
  • The model correctly recognized the numbers '90' and '19' and applied clean, readable formatting using digits and percent signs.

Judge transcript

The battery dropped from ninety percent to nineteen percent during the test.
Key differences
  • ninety percent is written as 90%
  • nineteen percent is written as 19%
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate and preserves the spoken meaning entirely. The only differences are formatting choices, where the spoken numbers and percent terms are written as digits and symbols (90% and 19% instead of ninety percent and nineteen percent), which is highly readable and often preferred.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript preserves the meaning perfectly with only standard formatting differences.
Judge transcriptThe battery dropped from ninety percent to nineteen percent during the test.
Key differences
  • ninety percent is written as 90%
  • nineteen percent is written as 19%
Categories
formatting_only
Researcher notes
  • The model correctly recognized the numbers '90' and '19' and applied clean, readable formatting using digits and percent signs.

Provenance

  • Category numeric unit integrity
  • ASR slice percentage minimal pair
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-percentage-001
  • Text SHA-256 aa7d137ea1db996c1e57892ce44ca8be6dc6df047decaa3e34ddc150a5335c24
  • Duration 3.88s
  • Bytes 186164

Evaluation status: ok

negation modality scope · clinical negation

asr-negation-medical-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the spoken meaning is fully and accurately preserved.

formatting only
Recommended action
  • Excellent transcription. The minor punctuation change does not affect readability or downstream processing.

Judge transcript

The patient denies chest pain but reports shortness of breath.
Key differences
  • Added a comma after 'chest pain' in the candidate transcript.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate, with only a minor, grammatically correct comma added before the conjunction 'but' which has no impact on the spoken meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the spoken meaning is fully and accurately preserved.
Judge transcriptThe patient denies chest pain but reports shortness of breath.
Key differences
  • Added a comma after 'chest pain' in the candidate transcript.
Categories
formatting_only
Researcher notes
  • Excellent transcription. The minor punctuation change does not affect readability or downstream processing.

Provenance

  • Category negation modality scope
  • ASR slice clinical negation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-medical-001
  • Text SHA-256 dbdcbaeba42a61f71f6c3283ff950b99d754363b182b2ef052eff60f552edd82
  • Duration 3.44s
  • Bytes 165044

Evaluation status: ok

negation modality scope · permission scope

asr-negation-policy-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the spoken meaning is fully preserved.

formatting only
Recommended action
  • The transcript is perfect. The punctuation difference is completely acceptable and preserves all semantic information.

Judge transcript

Only administrators may export records. Contractors may view records only.
Key differences
  • Candidate split the sentence with a period instead of using a semicolon as in the reference.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and preserves the exact meaning of the spoken audio. The only difference is minor formatting, specifically using a period and a new sentence instead of a semicolon, which does not impact comprehension or downstream usability.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the spoken meaning is fully preserved.
Judge transcriptOnly administrators may export records. Contractors may view records only.
Key differences
  • Candidate split the sentence with a period instead of using a semicolon as in the reference.
Categories
formatting_only
Researcher notes
  • The transcript is perfect. The punctuation difference is completely acceptable and preserves all semantic information.

Provenance

  • Category negation modality scope
  • ASR slice permission scope
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-policy-001
  • Text SHA-256 599ef83653daccc61f7cd8f7daed7b7e910ab4c4919442065b793630923fb099
  • Duration 4.52s
  • Bytes 216884

Evaluation status: ok

negation modality scope · unless condition

asr-negation-condition-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the transcription is perfectly accurate.

no error
Recommended action
  • Perfect transcription. Correctly captured the conditional logic and negation structure 'unless the inspection fails'.

Judge transcript

Ship the replacement unless the inspection fails tomorrow morning.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect match to the spoken audio and the reference transcript, preserving the exact meaning and details with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcription is perfectly accurate.
Judge transcriptShip the replacement unless the inspection fails tomorrow morning.
Categories
no_error
Researcher notes
  • Perfect transcription. Correctly captured the conditional logic and negation structure 'unless the inspection fails'.

Provenance

  • Category negation modality scope
  • ASR slice unless condition
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-condition-001
  • Text SHA-256 82c70b514883f60d665a4cf44698b4dcb754c2636052cca7a7165104e40b4eb6
  • Duration 3.64s
  • Bytes 174644

Evaluation status: ok

temporal scheduling accuracy · am pm window

asr-temporal-window-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the transcript is fully accurate with only minor formatting variation.

formatting only
Recommended action
  • Perfect transcription. The ASR model handled the content flawlessly, including specialized terms like 'migration window'.

Judge transcript

Start the migration window tonight at 10 pm after traffic drops.
Key differences
  • Candidate writes 'p.m.' with periods, while the reference/audio interpretation is 'pm'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect representation of the audio, with the only variation being the standard abbreviation formatting of 'p.m.' instead of 'pm', which has zero impact on the meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcript is fully accurate with only minor formatting variation.
Judge transcriptStart the migration window tonight at 10 pm after traffic drops.
Key differences
  • Candidate writes 'p.m.' with periods, while the reference/audio interpretation is 'pm'.
Categories
formatting_only
Researcher notes
  • Perfect transcription. The ASR model handled the content flawlessly, including specialized terms like 'migration window'.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice am pm window
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-window-001
  • Text SHA-256 66ed82094ed39a1d36cd5387c7648754c1c6dc7807e73c30aee7e0773552281c
  • Duration 3.84s
  • Bytes 184244

Evaluation status: ok

temporal scheduling accuracy · relative time relation

asr-temporal-relative-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is perfectly accurate.

no error
Recommended action
  • The ASR model performed flawlessly on this temporal relative constraint instruction, capturing all key elements including the negation at the end.

Judge transcript

Book the follow-up two weeks after the first infusion, not before it.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly matches the spoken audio and the reference transcript with zero errors, fully preserving the meaning and all critical details.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is perfectly accurate.
Judge transcriptBook the follow-up two weeks after the first infusion, not before it.
Categories
no_error
Researcher notes
  • The ASR model performed flawlessly on this temporal relative constraint instruction, capturing all key elements including the negation at the end.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice relative time relation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-relative-001
  • Text SHA-256 683eb39c6dccaa56d07fc3c6fef174a36de3b354f5a5c17ad0cc0cc6dc211101
  • Duration 3.96s
  • Bytes 190004

Evaluation status: ok

temporal scheduling accuracy · ordinal date minimal pair

asr-temporal-date-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the candidate transcript accurately captures the audio and only differs from the reference in numeric formatting.

formatting only
Recommended action
  • The model correctly recognized the speech. The difference is purely in the text normalization style (digits vs. words).

Judge transcript

Move the inspection from July fourth to July fourteenth.
Key differences
  • July fourth is written as July 4th
  • July fourteenth is written as July 14th
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and preserves the full meaning of the audio, with only a benign formatting difference where the dates are written as digits and ordinals rather than being spelled out as words.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate transcript accurately captures the audio and only differs from the reference in numeric formatting.
Judge transcriptMove the inspection from July fourth to July fourteenth.
Key differences
  • July fourth is written as July 4th
  • July fourteenth is written as July 14th
Categories
formatting_only
Researcher notes
  • The model correctly recognized the speech. The difference is purely in the text normalization style (digits vs. words).

Provenance

  • Category temporal scheduling accuracy
  • ASR slice ordinal date minimal pair
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-date-001
  • Text SHA-256 e79ea5b3319ccee0d2c33d7f53824ce7a080f06aa60c47f39f725a4209609d4a
  • Duration 3.20s
  • Bytes 153524

Evaluation status: ok

temporal scheduling accuracy · duration contrast

asr-temporal-duration-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the candidate represents numbers as digits but preserves the exact meaning.

formatting only
Recommended action
  • The model correctly transcribed and distinguished the highly similar sounding numbers 'fifteen' and 'fifty', representing them accurately in digit format.

Judge transcript

Run the pump for fifteen minutes then let it rest for fifty minutes.
Key differences
  • The word 'fifteen' is written as the digit '15'
  • The word 'fifty' is written as the digit '50'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is semantically perfect, with the only difference being the formatting of the numbers 'fifteen' and 'fifty' as digits '15' and '50', which does not alter the meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate represents numbers as digits but preserves the exact meaning.
Judge transcriptRun the pump for fifteen minutes then let it rest for fifty minutes.
Key differences
  • The word 'fifteen' is written as the digit '15'
  • The word 'fifty' is written as the digit '50'
Categories
formatting_only
Researcher notes
  • The model correctly transcribed and distinguished the highly similar sounding numbers 'fifteen' and 'fifty', representing them accurately in digit format.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice duration contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-duration-001
  • Text SHA-256 46846ed936375faff65199b095184a0ce00cea6cb427f8e34c78a31cc4374a2c
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

numeric unit integrity · amount contrast

asr-numeric-transfer-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; meaning is perfectly preserved.

formatting only
Recommended action
  • Excellent performance. The model correctly captured the distinction between 'fifteen' and 'fifty', as well as the negation and name 'Maya'.

Judge transcript

Transfer fifteen dollars to Maya today, not fifty dollars.
Key differences
  • The candidate splits the sentence with a period instead of a comma after 'today'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and perfectly preserves the meaning of the spoken audio. The only difference is a minor punctuation variation where the comma is replaced by a period, which has zero impact on semantic understanding.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; meaning is perfectly preserved.
Judge transcriptTransfer fifteen dollars to Maya today, not fifty dollars.
Key differences
  • The candidate splits the sentence with a period instead of a comma after 'today'.
Categories
formatting_only
Researcher notes
  • Excellent performance. The model correctly captured the distinction between 'fifteen' and 'fifty', as well as the negation and name 'Maya'.

Provenance

  • Category numeric unit integrity
  • ASR slice amount contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-transfer-001
  • Text SHA-256 387a5443dbcbe43becc7f50aee52cd98ecd1503a46eb68183fbf58045f00b49c
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

numeric unit integrity · dosage unit

asr-numeric-medication-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the candidate transcript is completely accurate.

no error
Recommended action
  • The model recognized the medication 'melatonin' and the dosage 'five milligrams' perfectly.

Judge transcript

Give the patient 5 milligrams of melatonin before bedtime.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript matches the spoken audio and the reference transcript perfectly, preserving all semantic details including the medication, dosage, and instructions.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate transcript is completely accurate.
Judge transcriptGive the patient 5 milligrams of melatonin before bedtime.
Categories
no_error
Researcher notes
  • The model recognized the medication 'melatonin' and the dosage 'five milligrams' perfectly.

Provenance

  • Category numeric unit integrity
  • ASR slice dosage unit
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-medication-001
  • Text SHA-256 566164fc8bcb2a1adb3b5929af80f8b1755e68f61d2967650ec925d8043b6a8f
  • Duration 3.76s
  • Bytes 180404

Evaluation status: ok

negation modality scope · negated action

asr-negation-server-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

None

no error
Recommended action
  • The model performed perfectly with no errors.

Judge transcript

Do not restart the server until the database backup finishes.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect match to the audio and reference transcript, preserving the meaning completely without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
None
Judge transcriptDo not restart the server until the database backup finishes.
Categories
no_error
Researcher notes
  • The model performed perfectly with no errors.

Provenance

  • Category negation modality scope
  • ASR slice negated action
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-server-001
  • Text SHA-256 9a47d14232d15170b7f3a734e6179db244d25afef06121759447c4f8a318d276
  • Duration 3.28s
  • Bytes 157364

Evaluation status: ok

negation modality scope · permission prohibition

asr-negation-refund-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the candidate transcript is 100% accurate.

no error
Recommended action
  • Perfect match. The system correctly handled the negation 'cannot' and key terms 'trial', 'refund', and 'Friday' with correct punctuation and spelling.

Judge transcript

The customer can cancel the trial, but cannot receive a refund after Friday.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match with the spoken audio and the reference transcript, preserving the exact meaning and all details including negation and entities.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate transcript is 100% accurate.
Judge transcriptThe customer can cancel the trial, but cannot receive a refund after Friday.
Categories
no_error
Researcher notes
  • Perfect match. The system correctly handled the negation 'cannot' and key terms 'trial', 'refund', and 'Friday' with correct punctuation and spelling.

Provenance

  • Category negation modality scope
  • ASR slice permission prohibition
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-refund-001
  • Text SHA-256 55a42c2151a5880e0aa2f00f13860c2acf87f536e24d9cec099342e0da50f211
  • Duration 3.80s
  • Bytes 182324

Evaluation status: ok

temporal scheduling accuracy · deadline day time

asr-temporal-deadline-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is completely accurate.

formatting only
Recommended action
  • The transcription is highly accurate. The minor formatting variation in 'a.m.' is acceptable and does not impact downstream processing.

Judge transcript

The benefits enrollment deadline is Monday at 9 a.m.
Key differences
  • Formatting of the time suffix ('a.m.' vs 'am')
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate with only a minor formatting difference ('a.m.' instead of 'am'), preserving the exact meaning of the spoken audio.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is completely accurate.
Judge transcriptThe benefits enrollment deadline is Monday at 9 a.m.
Key differences
  • Formatting of the time suffix ('a.m.' vs 'am')
Categories
formatting_only
Researcher notes
  • The transcription is highly accurate. The minor formatting variation in 'a.m.' is acceptable and does not impact downstream processing.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice deadline day time
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-deadline-001
  • Text SHA-256 b118c3c195ab7d6fd58def5f41f4f899aa49aba40735806279b2f743868df37c
  • Duration 3.04s
  • Bytes 145844

Evaluation status: ok

entity factual integrity · person place entity

asr-entity-doctor-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

The transcript preserves the meaning perfectly with no semantic errors.

formatting only
Recommended action
  • The model performs exceptionally well here. Minor formatting variations like hyphenation and capitalization of 'clinic' do not impact meaning preservation.

Judge transcript

Schedule the follow up with Dr. Rao at the Cedar Avenue Clinic.
Key differences
  • The candidate omits the hyphen in 'follow-up' (written as 'follow up')
  • The candidate capitalizes 'Clinic' while the reference does not
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is identical to the audio and reference transcript in meaning, with only minor, harmless formatting differences (the omission of a hyphen in 'follow-up' and capitalization of 'Clinic'). All critical entities and information are perfectly preserved.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript preserves the meaning perfectly with no semantic errors.
Judge transcriptSchedule the follow up with Dr. Rao at the Cedar Avenue Clinic.
Key differences
  • The candidate omits the hyphen in 'follow-up' (written as 'follow up')
  • The candidate capitalizes 'Clinic' while the reference does not
Categories
formatting_only
Researcher notes
  • The model performs exceptionally well here. Minor formatting variations like hyphenation and capitalization of 'clinic' do not impact meaning preservation.

Provenance

  • Category entity factual integrity
  • ASR slice person place entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-doctor-001
  • Text SHA-256 9ff5a0befafca6c4d8752ec9817f898caed7717e40beb76cd3eaa239a8cf9473
  • Duration 3.40s
  • Bytes 163124

Evaluation status: ok

entity factual integrity · product name contrast

asr-entity-product-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors found; the transcript is perfectly accurate.

no error
Recommended action
  • The model successfully recognized and transcribed the entity names 'Zephyr Mini' and 'Zephyr Max' correctly.

Judge transcript

Compare the Zephyr Mini plan with the Zephyr Max plan before renewal.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match with the spoken audio and reference transcript, accurately capturing the product names and intent with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors found; the transcript is perfectly accurate.
Judge transcriptCompare the Zephyr Mini plan with the Zephyr Max plan before renewal.
Categories
no_error
Researcher notes
  • The model successfully recognized and transcribed the entity names 'Zephyr Mini' and 'Zephyr Max' correctly.

Provenance

  • Category entity factual integrity
  • ASR slice product name contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-product-001
  • Text SHA-256 919dd10bf1c803fb423936e4f48c18d6469b7753fa5383a94ff26fc32bf6d745
  • Duration 3.84s
  • Bytes 184244

Evaluation status: ok

entity factual integrity · alphanumeric identifier

asr-entity-ticket-001-local-tts

Qwen3 ASR 1.7B (8-bit)
55 inaccurate

Semantic impact

The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.

entity errorsubstitution
Recommended action
  • The model failed to correctly transcribe the alphanumeric sequence 'PX-4187', outputting 'PX4Y7'. Improving spelling/character-level acoustic modeling for numbers and alphanumeric code patterns is recommended.

Judge transcript

Attach the call summary to ticket PX-4187 before closing the case.
Key differences
  • Candidate has 'PX4Y7' instead of 'PX-4187'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 55.00 (scores: 55, 55, 55). Representative reason: The candidate transcript misrecognizes the ticket identifier 'PX-4187' as 'PX4Y7'. While the rest of the sentence is perfectly captured, this entity error is highly detrimental for downstream ticketing systems, resulting in partial semantic loss.

judge samples: 55, 55, 55; avg 55.00
Meaning
partial_loss
Semantic impact
The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.
Judge transcriptAttach the call summary to ticket PX-4187 before closing the case.
Key differences
  • Candidate has 'PX4Y7' instead of 'PX-4187'
Categories
entity_error, substitution
Researcher notes
  • The model failed to correctly transcribe the alphanumeric sequence 'PX-4187', outputting 'PX4Y7'. Improving spelling/character-level acoustic modeling for numbers and alphanumeric code patterns is recommended.

Provenance

  • Category entity factual integrity
  • ASR slice alphanumeric identifier
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-ticket-001
  • Text SHA-256 845a5229d7c30013391f3297d1a9cc591ee108b93c4cf22218cdb749913fbf6e
  • Duration 4.60s
  • Bytes 220724

Evaluation status: ok

entity factual integrity · organization contact entity

asr-entity-organization-001-local-tts

Qwen3 ASR 1.7B (8-bit)
90 accurate

Semantic impact

The candidate transcript contains a minor phonetic substitution of the proper name 'Elena' with 'Alina'.

substitutionentity error
Recommended action
  • The model substituted 'Elena' with 'Alina', which are homophones or near-homophones in many English dialects. This is a common phonetic substitution and can be mitigated with contextual entity resolution or personalized language models.

Judge transcript

Send the signed form to Northstar Benefits and copy Elena Morales.
Key differences
  • Candidate transcribed the name 'Elena' as 'Alina'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 90.00 (scores: 85, 95, 90). Representative reason: The candidate transcript is highly accurate, with the only discrepancy being a minor phonetic spelling variation of the recipient's name, transcribing 'Elena' as 'Alina'. This minor entity substitution has negligible impact on the overall comprehension of the spoken instruction.

judge samples: 85, 95, 90; avg 90.00
Meaning
minor_loss
Semantic impact
The candidate transcript contains a minor phonetic substitution of the proper name 'Elena' with 'Alina'.
Judge transcriptSend the signed form to Northstar Benefits and copy Elena Morales.
Key differences
  • Candidate transcribed the name 'Elena' as 'Alina'.
Categories
substitution, entity_error
Researcher notes
  • The model substituted 'Elena' with 'Alina', which are homophones or near-homophones in many English dialects. This is a common phonetic substitution and can be mitigated with contextual entity resolution or personalized language models.

Provenance

  • Category entity factual integrity
  • ASR slice organization contact entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-organization-001
  • Text SHA-256 2c9213d5e33f7fc5a5924f6f72b4f91505c0b4c71b26848cb000e2d00ef961bc
  • Duration 3.80s
  • Bytes 182324

Evaluation status: ok

entity factual integrity · address entity

asr-entity-address-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript has perfect accuracy with a minor capitalization change.

formatting only
Recommended action
  • The transcription is highly accurate, matching the reference exactly except for capitalization of 'Apartment'.

Judge transcript

The delivery address is 812 Ashwood Lane, apartment 5B.
Key differences
  • Capitalization of 'Apartment' in the candidate vs 'apartment' in the reference.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate, with only a minor capitalization difference ('Apartment' instead of 'apartment') which does not affect the meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript has perfect accuracy with a minor capitalization change.
Judge transcriptThe delivery address is 812 Ashwood Lane, apartment 5B.
Key differences
  • Capitalization of 'Apartment' in the candidate vs 'apartment' in the reference.
Categories
formatting_only
Researcher notes
  • The transcription is highly accurate, matching the reference exactly except for capitalization of 'Apartment'.

Provenance

  • Category entity factual integrity
  • ASR slice address entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-address-001
  • Text SHA-256 5cf2d55b0dddffcc8b60dc856eb7816a36e745257daaf10abd184d3fcd40a9dc
  • Duration 4.24s
  • Bytes 203444

Evaluation status: ok

semantic paraphrase preservation · event state preservation

asr-semantic-paraphrase-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the transcription is 100% accurate.

no error
Recommended action
  • The model achieved perfect accuracy on this sample, capturing all words correctly with correct spelling and punctuation.

Judge transcript

The technician replaced the cracked valve and confirmed the leak stopped.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect match to both the spoken audio and the reference transcript, containing no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the transcription is 100% accurate.
Judge transcriptThe technician replaced the cracked valve and confirmed the leak stopped.
Categories
no_error
Researcher notes
  • The model achieved perfect accuracy on this sample, capturing all words correctly with correct spelling and punctuation.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice event state preservation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-paraphrase-001
  • Text SHA-256 459210fe6f6f4d41259ceba39ab9c8fc608abfb927c5a49c8750ac89f86e21d0
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

semantic paraphrase preservation · causal relation

asr-semantic-causal-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors detected as the transcript is perfectly accurate.

no error
Recommended action
  • The model transcribed this technical sentence flawlessly, showing robust handling of technical terminology like 'token expired' and 'retried after refresh'.

Judge transcript

The upload failed because the token expired, so the client retried after refresh.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect match to both the audio and the reference transcript, preserving the spoken meaning completely without any errors or omissions.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors detected as the transcript is perfectly accurate.
Judge transcriptThe upload failed because the token expired, so the client retried after refresh.
Categories
no_error
Researcher notes
  • The model transcribed this technical sentence flawlessly, showing robust handling of technical terminology like 'token expired' and 'retried after refresh'.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice causal relation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-causal-001
  • Text SHA-256 caeb5d404f654e7d980c3f61cc1576806deb874cf24af2120f26700bb7aa471e
  • Duration 4.72s
  • Bytes 226484

Evaluation status: ok

semantic paraphrase preservation · metric tradeoff

asr-semantic-comparison-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is completely accurate.

no error
Recommended action
  • The model achieved 100% accuracy on this segment, properly capturing technical machine learning terms like 'recall' and 'precision'.

Judge transcript

The candidate model improved recall but reduced precision on noisy calls.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact word-for-word match of the spoken audio, preserving the complete meaning, entities, and technical terminology without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is completely accurate.
Judge transcriptThe candidate model improved recall but reduced precision on noisy calls.
Categories
no_error
Researcher notes
  • The model achieved 100% accuracy on this segment, properly capturing technical machine learning terms like 'recall' and 'precision'.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice metric tradeoff
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-comparison-001
  • Text SHA-256 fbaadae669f95ffc77afc2da5b9661b0a48486f7e71ada1bd05170df2649a739
  • Duration 4.32s
  • Bytes 207284

Evaluation status: ok

semantic paraphrase preservation · coreference ambiguity

asr-semantic-coreference-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

The candidate transcript is completely accurate to the audio and has no errors.

no error
Recommended action
  • The candidate performed perfectly. The reference transcript contains a slight transcription error by writing 'Maria' instead of 'Ria'.

Judge transcript

Ria called Priya after she reviewed the contract changes.
Key differences
  • The reference transcript has 'Maria' while both the audio and candidate transcript have 'Ria'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly matches the spoken audio. The reference transcript contains a minor discrepancy ('Maria' instead of 'Ria'), but the candidate correctly captured the name spoken in the audio.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The candidate transcript is completely accurate to the audio and has no errors.
Judge transcriptRia called Priya after she reviewed the contract changes.
Key differences
  • The reference transcript has 'Maria' while both the audio and candidate transcript have 'Ria'.
Categories
no_error
Researcher notes
  • The candidate performed perfectly. The reference transcript contains a slight transcription error by writing 'Maria' instead of 'Ria'.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice coreference ambiguity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-coreference-001
  • Text SHA-256 7ec151a16318fda290762dbf9f9ad83e2ce7c5df725aeba3043965cfd596e437
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

semantic paraphrase preservation · conditional instruction

asr-semantic-instruction-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the transcript is perfectly accurate.

no error
Recommended action
  • The ASR model performed perfectly on this utterance.

Judge transcript

If the first backup succeeds, skip the manual export and notify operations.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect match to the audio and reference transcript, with no spelling, grammatical, or semantic errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcript is perfectly accurate.
Judge transcriptIf the first backup succeeds, skip the manual export and notify operations.
Categories
no_error
Researcher notes
  • The ASR model performed perfectly on this utterance.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice conditional instruction
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-instruction-001
  • Text SHA-256 c095d3a9fadb208cbc7ae9c12d4168e37dff40b73108cd2b28c219e12511a104
  • Duration 4.48s
  • Bytes 214964

Evaluation status: ok

acoustic noise robustness · cafe background command

asr-noise-cafe-order-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors; the candidate transcript is completely accurate.

no error
Recommended action
  • The ASR model successfully transcribed the audio with 100% accuracy, showing no issues with names, ingredients, or instruction structure.

Judge transcript

Add oat milk to the latte and leave out the caramel drizzle.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an absolute perfect match to both the audio and the reference transcript, with no errors whatsoever.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate transcript is completely accurate.
Judge transcriptAdd oat milk to the latte and leave out the caramel drizzle.
Categories
no_error
Researcher notes
  • The ASR model successfully transcribed the audio with 100% accuracy, showing no issues with names, ingredients, or instruction structure.

Provenance

  • Category acoustic noise robustness
  • ASR slice cafe background command
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-cafe-order-001
  • Text SHA-256 613cfd4c60424169e80cba7ab2ecc6f0fcb35f677cdfbfd9746b4efaebf165ef
  • Duration 3.28s
  • Bytes 157364

Evaluation status: ok

acoustic noise robustness · vehicle noise navigation

asr-noise-driving-route-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

There are no semantic errors in the candidate transcript; the meaning is fully preserved.

formatting only
Recommended action
  • Excellent transcription accuracy with zero word errors. The sentence boundary detection is the only minor variation, which is a benign formatting difference.

Judge transcript

Take the second exit after the bridge, then stay in the right lane.
Key differences
  • The candidate transcript splits the instruction into two sentences using a period and capitalization ('bridge. Then') instead of a comma ('bridge, then').
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and preserves the spoken meaning perfectly. The only difference from the reference is a minor punctuation change, where a comma is replaced with a period, splitting the sentence. This has no impact on semantic understanding.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors in the candidate transcript; the meaning is fully preserved.
Judge transcriptTake the second exit after the bridge, then stay in the right lane.
Key differences
  • The candidate transcript splits the instruction into two sentences using a period and capitalization ('bridge. Then') instead of a comma ('bridge, then').
Categories
formatting_only
Researcher notes
  • Excellent transcription accuracy with zero word errors. The sentence boundary detection is the only minor variation, which is a benign formatting difference.

Provenance

  • Category acoustic noise robustness
  • ASR slice vehicle noise navigation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-driving-route-001
  • Text SHA-256 c3ce8e7e94fd6ea30d703cbca09fc8f6c03373f0ea379bfc5d94dd49aa48ce26
  • Duration 3.72s
  • Bytes 178484

Evaluation status: ok

acoustic noise robustness · industrial noise identifier

asr-noise-warehouse-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is completely accurate.

no error
Recommended action
  • The ASR model performed perfectly on this short, clear warehouse command, correctly transcribing the alphanumeric bin ID 'A17' and the destination 'loading dock three'.

Judge transcript

Scan bin A17 before moving the pallet to loading dock three.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match to both the audio and the reference transcript, preserving the spoken meaning completely without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is completely accurate.
Judge transcriptScan bin A17 before moving the pallet to loading dock three.
Categories
no_error
Researcher notes
  • The ASR model performed perfectly on this short, clear warehouse command, correctly transcribing the alphanumeric bin ID 'A17' and the destination 'loading dock three'.

Provenance

  • Category acoustic noise robustness
  • ASR slice industrial noise identifier
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-warehouse-001
  • Text SHA-256 f7c26924015438634f3eb4be84b886d690e9009b39ecd8f4fa88408ff3ec2e11
  • Duration 4.08s
  • Bytes 195764

Evaluation status: ok

acoustic noise robustness · reverberant meeting room

asr-noise-reverberant-room-001-local-tts

Qwen3 ASR 1.7B (8-bit)
62 needs review

Semantic impact

The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.

number errorsubstitution
Recommended action
  • The model inserted an extra digit '2' into the room number, misinterpreting 'four twelve' as 'four two one two' or similar. Focus on fine-tuning number recognition, especially in slightly reverberant conditions.

Judge transcript

The quarterly review starts in room 412 at half past three.
Key differences
  • Room '412' was transcribed as '4212'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 61.67 (scores: 70, 55, 60). Representative reason: The candidate transcript incorrectly transcribes the room number as '4212' instead of '412', which is a significant entity/number error that could cause confusion for downstream tasks like scheduling or navigation.

judge samples: 70, 55, 60; avg 61.67
Meaning
partial_loss
Semantic impact
The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.
Judge transcriptThe quarterly review starts in room 412 at half past three.
Key differences
  • Room '412' was transcribed as '4212'
Categories
number_error, substitution
Researcher notes
  • The model inserted an extra digit '2' into the room number, misinterpreting 'four twelve' as 'four two one two' or similar. Focus on fine-tuning number recognition, especially in slightly reverberant conditions.

Provenance

  • Category acoustic noise robustness
  • ASR slice reverberant meeting room
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-reverberant-room-001
  • Text SHA-256 f0f9e458a7bf2804ef37a0a709c4a7882c50036351a8f11cd69c590f1805b4d8
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

acoustic noise robustness · overlapping speech target

asr-noise-overlapping-speech-001-local-tts

Qwen3 ASR 1.7B (8-bit)
100 accurate

Semantic impact

No semantic errors; the transcript is perfectly accurate.

no error
Recommended action
  • The ASR model performed flawlessly on this sample with no issues in transcription, formatting, or punctuation.

Judge transcript

Ignore the side conversation and record the final answer as option C.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect, word-for-word match to both the reference transcript and the spoken audio, preserving the exact meaning with zero errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcript is perfectly accurate.
Judge transcriptIgnore the side conversation and record the final answer as option C.
Categories
no_error
Researcher notes
  • The ASR model performed flawlessly on this sample with no issues in transcription, formatting, or punctuation.

Provenance

  • Category acoustic noise robustness
  • ASR slice overlapping speech target
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/Qwen3-ASR-1.7B-8bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-overlapping-speech-001
  • Text SHA-256 e72cbb8fa9f427ff6a5ac7dcb1cf1e153ec37baeeba0a80c6be4033169b3f66b
  • Duration 4.08s
  • Bytes 195764

Evaluation status: ok

transcription accuracy wer · function word substitution

asr-wer-function-word-001-local-tts

VibeVoice ASR (4-bit)
85 accurate

Semantic impact

The substitution of 'patch is for' with 'patches for' creates a minor grammatical slip but does not change the core message.

substitution
Recommended action
  • The phonetic similarity between 'patch is' (/pætʃ ɪz/) and 'patches' (/pætʃɪz/) caused a minor transcription error. Language model constraints could help prevent grammatically incomplete sequences like 'the patches for... not the stable channel'.

Judge transcript

The release notes say the patch is for the beta channel, not the stable channel.
Key differences
  • Substituted 'patches for' for 'patch is for'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 85.00 (scores: 85, 85, 85). Representative reason: The candidate transcript contains a minor acoustic substitution, transcribing 'patch is' as 'patches', which slightly impacts the grammar of the sentence but preserves the overall semantic meaning and intent.

judge samples: 85, 85, 85; avg 85.00
Meaning
preserved
Semantic impact
The substitution of 'patch is for' with 'patches for' creates a minor grammatical slip but does not change the core message.
Judge transcriptThe release notes say the patch is for the beta channel, not the stable channel.
Key differences
  • Substituted 'patches for' for 'patch is for'
Categories
substitution
Researcher notes
  • The phonetic similarity between 'patch is' (/pætʃ ɪz/) and 'patches' (/pætʃɪz/) caused a minor transcription error. Language model constraints could help prevent grammatically incomplete sequences like 'the patches for... not the stable channel'.

Provenance

  • Category transcription accuracy wer
  • ASR slice function word substitution
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-function-word-001
  • Text SHA-256 db881f12982b7e1bcdb006f5bcb9ce39ddcb8ce29bfa65b29176ffcd5301d3a7
  • Duration 4.12s
  • Bytes 197684

Evaluation status: ok

transcription accuracy wer · homophone confusion

asr-wer-homophone-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors occurred; the transcript is perfectly accurate.

no error
Recommended action
  • Excellent transcription performance with perfect semantic and lexical matching.

Judge transcript

Please write the new access policy before the right team reviews it.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact match to both the spoken audio and the reference transcript, preserving the original meaning perfectly with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors occurred; the transcript is perfectly accurate.
Judge transcriptPlease write the new access policy before the right team reviews it.
Categories
no_error
Researcher notes
  • Excellent transcription performance with perfect semantic and lexical matching.

Provenance

  • Category transcription accuracy wer
  • ASR slice homophone confusion
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-homophone-001
  • Text SHA-256 e98bb2cafe5fd5b2ad5f26020053ff27ca22a0e7c17c904647d2449af7b4657e
  • Duration 3.24s
  • Bytes 155444

Evaluation status: ok

transcription accuracy wer · short command edit distance

asr-wer-short-utterance-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

None

no error
Recommended action
  • Excellent transcription with no errors.

Judge transcript

Turn left at Pine Street, then park by gate four.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.33 (scores: 100, 95, 100). Representative reason: The candidate transcript is completely accurate and perfectly matches the spoken audio and the reference transcript with no errors.

judge samples: 100, 95, 100; avg 98.33
Meaning
preserved
Semantic impact
None
Judge transcriptTurn left at Pine Street, then park by gate four.
Categories
no_error
Researcher notes
  • Excellent transcription with no errors.

Provenance

  • Category transcription accuracy wer
  • ASR slice short command edit distance
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-short-utterance-001
  • Text SHA-256 f0270a081ebe4d9733c17ae6c19c100028c61723a25cf10de9aa04db0e29e736
  • Duration 2.96s
  • Bytes 142004

Evaluation status: ok

transcription accuracy wer · disfluency token accounting

asr-wer-disfluency-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; the transcript is a perfect match to the audio.

no error
Recommended action
  • The model successfully captured the self-correction 'sorry' and the repeated definite article 'the', showing robust handling of conversational disfluency.

Judge transcript

I need the, sorry, the final invoice attached to the March ticket.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly matches the spoken audio, capturing the words, the disfluency, and the self-correction exactly as spoken.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcript is a perfect match to the audio.
Judge transcriptI need the, sorry, the final invoice attached to the March ticket.
Categories
no_error
Researcher notes
  • The model successfully captured the self-correction 'sorry' and the repeated definite article 'the', showing robust handling of conversational disfluency.

Provenance

  • Category transcription accuracy wer
  • ASR slice disfluency token accounting
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-disfluency-001
  • Text SHA-256 d7554e1f20e8046ff869c7e16f61ebdc8335e982164130cb2d3c6c4bf9916eee
  • Duration 3.92s
  • Bytes 188084

Evaluation status: ok

transcription accuracy wer · sentence boundary punctuation

asr-wer-punctuation-boundary-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

The candidate transcript is completely accurate and preserves all meaning.

no error
Recommended action
  • The ASR output is flawless. The transcription is wrapped in a JSON structure containing timestamps and speaker labels, which is accurate and helpful for downstream tasks.

Judge transcript

No refunds after Friday. Exceptions require manager approval.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.33 (scores: 95, 100, 100). Representative reason: The candidate transcript matches the spoken audio and reference transcript perfectly with zero errors.

judge samples: 95, 100, 100; avg 98.33
Meaning
preserved
Semantic impact
The candidate transcript is completely accurate and preserves all meaning.
Judge transcriptNo refunds after Friday. Exceptions require manager approval.
Categories
no_error
Researcher notes
  • The ASR output is flawless. The transcription is wrapped in a JSON structure containing timestamps and speaker labels, which is accurate and helpful for downstream tasks.

Provenance

  • Category transcription accuracy wer
  • ASR slice sentence boundary punctuation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-wer-punctuation-boundary-001
  • Text SHA-256 30ce2ce13eaebf18c2d60223c27e6e644e491d4d24806ee559d1a36d9b6f76aa
  • Duration 3.64s
  • Bytes 174644

Evaluation status: ok

numeric unit integrity · decimal measurement

asr-numeric-measurement-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

The transcript is entirely accurate with no semantic errors.

no error
Recommended action
  • Excellent transcription with perfect word accuracy and formatting.

Judge transcript

Set the incubator to thirty-seven point five degrees Celsius.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact word-for-word match with the spoken audio and the reference transcript, preserving the meaning perfectly.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript is entirely accurate with no semantic errors.
Judge transcriptSet the incubator to thirty-seven point five degrees Celsius.
Categories
no_error
Researcher notes
  • Excellent transcription with perfect word accuracy and formatting.

Provenance

  • Category numeric unit integrity
  • ASR slice decimal measurement
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-measurement-001
  • Text SHA-256 1ba0e91325222fab47ab079798743fd1e6c43e845377315a824e82f8afe516d5
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

numeric unit integrity · digit sequence

asr-numeric-account-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is fully correct with a minor formatting variation in the numeric representation.

formatting only
Recommended action
  • The model correctly transcribed the acoustic signals. An inverse text normalization (ITN) post-processing step would be needed to convert the spelled-out numbers into digits like the reference transcript.

Judge transcript

Read back account number forty nine zero two seven seven one eight before submitting the form.
Key differences
  • The account number '4902 7718' is represented as spoken words 'forty nine zero two seven seven one eight'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate and preserves the spoken meaning without any errors. The only difference is that the candidate spells out the digits in the account number as words ('forty nine zero two seven seven one eight') rather than using numeric digits ('4902 7718'), which is a formatting preference rather than a semantic error.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is fully correct with a minor formatting variation in the numeric representation.
Judge transcriptRead back account number forty nine zero two seven seven one eight before submitting the form.
Key differences
  • The account number '4902 7718' is represented as spoken words 'forty nine zero two seven seven one eight'.
Categories
formatting_only
Researcher notes
  • The model correctly transcribed the acoustic signals. An inverse text normalization (ITN) post-processing step would be needed to convert the spelled-out numbers into digits like the reference transcript.

Provenance

  • Category numeric unit integrity
  • ASR slice digit sequence
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-account-001
  • Text SHA-256 1b8373fa92691331a8512bd9fdb9eaa080397a75a368a630b825fb74f6d0f8ed
  • Duration 4.84s
  • Bytes 232244

Evaluation status: ok

numeric unit integrity · percentage minimal pair

asr-numeric-percentage-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

The transcript is perfectly accurate with no semantic errors.

no error
Recommended action
  • Perfect transcription of numbers and percentages. No action needed.

Judge transcript

The battery dropped from ninety percent to nineteen percent during the test.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and matches the spoken audio and reference transcript perfectly with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript is perfectly accurate with no semantic errors.
Judge transcriptThe battery dropped from ninety percent to nineteen percent during the test.
Categories
no_error
Researcher notes
  • Perfect transcription of numbers and percentages. No action needed.

Provenance

  • Category numeric unit integrity
  • ASR slice percentage minimal pair
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-percentage-001
  • Text SHA-256 aa7d137ea1db996c1e57892ce44ca8be6dc6df047decaa3e34ddc150a5335c24
  • Duration 3.88s
  • Bytes 186164

Evaluation status: ok

negation modality scope · clinical negation

asr-negation-medical-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; the transcript is completely accurate.

no error
Recommended action
  • The model achieved 100% accuracy on this short medical statement, correctly capturing both the negation 'denies chest pain' and the reported symptom 'shortness of breath'.

Judge transcript

The patient denies chest pain but reports shortness of breath.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript matches the audio and reference transcript perfectly with no errors or omissions.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the transcript is completely accurate.
Judge transcriptThe patient denies chest pain but reports shortness of breath.
Categories
no_error
Researcher notes
  • The model achieved 100% accuracy on this short medical statement, correctly capturing both the negation 'denies chest pain' and the reported symptom 'shortness of breath'.

Provenance

  • Category negation modality scope
  • ASR slice clinical negation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-medical-001
  • Text SHA-256 dbdcbaeba42a61f71f6c3283ff950b99d754363b182b2ef052eff60f552edd82
  • Duration 3.44s
  • Bytes 165044

Evaluation status: ok

negation modality scope · permission scope

asr-negation-policy-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

There are no semantic errors; the candidate transcript perfectly preserves the meaning.

formatting only
Recommended action
  • The ASR model performed perfectly on this audio segment.
  • Punctuation and capitalization choices do not affect the downstream usability or interpretation of this policy statement.

Judge transcript

Only administrators may export records. Contractors may view records only.
Key differences
  • The candidate uses a period and starts a new sentence instead of using a semicolon as in the reference transcript.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate and preserves the meaning of the spoken audio perfectly. The only difference is the use of a period instead of a semicolon, which is a benign formatting variation that has no semantic impact.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the candidate transcript perfectly preserves the meaning.
Judge transcriptOnly administrators may export records. Contractors may view records only.
Key differences
  • The candidate uses a period and starts a new sentence instead of using a semicolon as in the reference transcript.
Categories
formatting_only
Researcher notes
  • The ASR model performed perfectly on this audio segment.
  • Punctuation and capitalization choices do not affect the downstream usability or interpretation of this policy statement.

Provenance

  • Category negation modality scope
  • ASR slice permission scope
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-policy-001
  • Text SHA-256 599ef83653daccc61f7cd8f7daed7b7e910ab4c4919442065b793630923fb099
  • Duration 4.52s
  • Bytes 216884

Evaluation status: ok

negation modality scope · unless condition

asr-negation-condition-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript matches the spoken audio perfectly.

no error
Recommended action
  • Excellent transcription with perfect accuracy and zero deviation from the spoken text.

Judge transcript

Ship the replacement unless the inspection fails tomorrow morning.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact match to both the audio and reference transcript, preserving the spoken meaning perfectly without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript matches the spoken audio perfectly.
Judge transcriptShip the replacement unless the inspection fails tomorrow morning.
Categories
no_error
Researcher notes
  • Excellent transcription with perfect accuracy and zero deviation from the spoken text.

Provenance

  • Category negation modality scope
  • ASR slice unless condition
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-condition-001
  • Text SHA-256 82c70b514883f60d665a4cf44698b4dcb754c2636052cca7a7165104e40b4eb6
  • Duration 3.64s
  • Bytes 174644

Evaluation status: ok

temporal scheduling accuracy · am pm window

asr-temporal-window-001-local-tts

VibeVoice ASR (4-bit)
96 accurate

Semantic impact

There are no semantic errors; the spoken words are transcribed perfectly.

formatting only
Recommended action
  • The speech recognition is flawless. The formatting contains metadata (Start, End, Speaker, Content) which may need to be parsed depending on the downstream application, but the text content itself is 100% correct.

Judge transcript

Start the migration window tonight at 10 p.m. after traffic drops.
Key differences
  • Candidate transcript is wrapped in a JSON structure containing timing and speaker metadata.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 96.00 (scores: 95, 98, 95). Representative reason: The candidate transcript captures the spoken audio perfectly with 100% semantic accuracy. The only discrepancy is that the candidate output is wrapped in a JSON segment format (including timestamps and speaker tags), which is a formatting difference and does not impact the meaning preservation.

judge samples: 95, 98, 95; avg 96.00
Meaning
preserved
Semantic impact
There are no semantic errors; the spoken words are transcribed perfectly.
Judge transcriptStart the migration window tonight at 10 p.m. after traffic drops.
Key differences
  • Candidate transcript is wrapped in a JSON structure containing timing and speaker metadata.
Categories
formatting_only
Researcher notes
  • The speech recognition is flawless. The formatting contains metadata (Start, End, Speaker, Content) which may need to be parsed depending on the downstream application, but the text content itself is 100% correct.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice am pm window
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-window-001
  • Text SHA-256 66ed82094ed39a1d36cd5387c7648754c1c6dc7807e73c30aee7e0773552281c
  • Duration 3.84s
  • Bytes 184244

Evaluation status: ok

temporal scheduling accuracy · relative time relation

asr-temporal-relative-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

None, the spoken content is fully and accurately captured.

formatting only
Recommended action
  • The speech recognition itself is flawless. The formatting contains additional speaker and timing metadata which may need to be parsed depending on the downstream application.

Judge transcript

Book the follow-up two weeks after the first infusion, not before it.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing speaker and timestamp metadata, but the textual content is identical to the reference.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate, matching the spoken audio and the reference transcript word-for-word. The only difference is the JSON formatting wrapping the output, which does not affect semantic understanding.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
None, the spoken content is fully and accurately captured.
Judge transcriptBook the follow-up two weeks after the first infusion, not before it.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing speaker and timestamp metadata, but the textual content is identical to the reference.
Categories
formatting_only
Researcher notes
  • The speech recognition itself is flawless. The formatting contains additional speaker and timing metadata which may need to be parsed depending on the downstream application.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice relative time relation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-relative-001
  • Text SHA-256 683eb39c6dccaa56d07fc3c6fef174a36de3b354f5a5c17ad0cc0cc6dc211101
  • Duration 3.96s
  • Bytes 190004

Evaluation status: ok

temporal scheduling accuracy · ordinal date minimal pair

asr-temporal-date-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; date representations differ only by text normalization formatting.

formatting only
Recommended action
  • The model correctly transcribed the spoken text, applying standard text normalization for dates. This is a correct transcription.

Judge transcript

Move the inspection from July fourth to July fourteenth.
Key differences
  • Used 'July 4th' instead of 'July fourth'
  • Used 'July 14th' instead of 'July fourteenth'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate, with the only differences being formatting preferences: using numerical ordinals '4th' and '14th' instead of the spelled-out words 'fourth' and 'fourteenth'. Meaning is perfectly preserved.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; date representations differ only by text normalization formatting.
Judge transcriptMove the inspection from July fourth to July fourteenth.
Key differences
  • Used 'July 4th' instead of 'July fourth'
  • Used 'July 14th' instead of 'July fourteenth'
Categories
formatting_only
Researcher notes
  • The model correctly transcribed the spoken text, applying standard text normalization for dates. This is a correct transcription.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice ordinal date minimal pair
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-date-001
  • Text SHA-256 e79ea5b3319ccee0d2c33d7f53824ce7a080f06aa60c47f39f725a4209609d4a
  • Duration 3.20s
  • Bytes 153524

Evaluation status: ok

temporal scheduling accuracy · duration contrast

asr-temporal-duration-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

The transcript is completely accurate with no semantic errors, only containing formatting metadata wrapper.

formatting only
Recommended action
  • The ASR output is highly accurate. If a plain-text output is desired, a post-processing step to extract the 'Content' field from the JSON structure should be applied.

Judge transcript

Run the pump for fifteen minutes, then let it rest for fifty minutes.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing start, end, and speaker metadata.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 97.67 (scores: 95, 100, 98). Representative reason: The candidate transcript perfectly captures the spoken audio with 100% accuracy. The only difference is the inclusion of metadata (timestamps and speaker labels) wrapped in a JSON array format, which does not affect semantic meaning preservation.

judge samples: 95, 100, 98; avg 97.67
Meaning
preserved
Semantic impact
The transcript is completely accurate with no semantic errors, only containing formatting metadata wrapper.
Judge transcriptRun the pump for fifteen minutes, then let it rest for fifty minutes.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing start, end, and speaker metadata.
Categories
formatting_only
Researcher notes
  • The ASR output is highly accurate. If a plain-text output is desired, a post-processing step to extract the 'Content' field from the JSON structure should be applied.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice duration contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-duration-001
  • Text SHA-256 46846ed936375faff65199b095184a0ce00cea6cb427f8e34c78a31cc4374a2c
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

numeric unit integrity · amount contrast

asr-numeric-transfer-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

The transcript is fully accurate, with only a formatting-level difference due to JSON container wrapping.

formatting only
Recommended action
  • The model successfully transcribed the numeric distinction ('fifteen' vs 'fifty') and the entity name 'Maya' perfectly.
  • Post-processing can easily strip the JSON wrapper if only the raw text is required for downstream tasks.

Judge transcript

Transfer fifteen dollars to Maya today, not fifty dollars.
Key differences
  • Candidate transcript is wrapped in a JSON array containing Start, End, Speaker, and Content keys.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 97.67 (scores: 98, 95, 100). Representative reason: The candidate transcript is perfectly accurate word-for-word, correctly identifying the critical distinction between 'fifteen' and 'fifty' dollars. The text is wrapped in a JSON segment structure with timestamps and speaker tags, which is a formatting difference but does not degrade the semantic quality of the transcription itself.

judge samples: 98, 95, 100; avg 97.67
Meaning
preserved
Semantic impact
The transcript is fully accurate, with only a formatting-level difference due to JSON container wrapping.
Judge transcriptTransfer fifteen dollars to Maya today, not fifty dollars.
Key differences
  • Candidate transcript is wrapped in a JSON array containing Start, End, Speaker, and Content keys.
Categories
formatting_only
Researcher notes
  • The model successfully transcribed the numeric distinction ('fifteen' vs 'fifty') and the entity name 'Maya' perfectly.
  • Post-processing can easily strip the JSON wrapper if only the raw text is required for downstream tasks.

Provenance

  • Category numeric unit integrity
  • ASR slice amount contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-transfer-001
  • Text SHA-256 387a5443dbcbe43becc7f50aee52cd98ecd1503a46eb68183fbf58045f00b49c
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

numeric unit integrity · dosage unit

asr-numeric-medication-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

There are no semantic errors; the transcription is perfectly accurate.

no error
Recommended action
  • The model correctly transcribed both the numerical dosage ('five milligrams') and the specific medication ('melatonin') without any errors.

Judge transcript

Give the patient five milligrams of melatonin before bedtime.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.33 (scores: 95, 100, 100). Representative reason: The candidate transcript is an exact match to the audio and the reference transcript. All medical terminology, dosages, and instructions are perfectly transcribed.

judge samples: 95, 100, 100; avg 98.33
Meaning
preserved
Semantic impact
There are no semantic errors; the transcription is perfectly accurate.
Judge transcriptGive the patient five milligrams of melatonin before bedtime.
Categories
no_error
Researcher notes
  • The model correctly transcribed both the numerical dosage ('five milligrams') and the specific medication ('melatonin') without any errors.

Provenance

  • Category numeric unit integrity
  • ASR slice dosage unit
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-numeric-medication-001
  • Text SHA-256 566164fc8bcb2a1adb3b5929af80f8b1755e68f61d2967650ec925d8043b6a8f
  • Duration 3.76s
  • Bytes 180404

Evaluation status: ok

negation modality scope · negated action

asr-negation-server-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

There are no semantic errors; the meaning is fully preserved.

no error
Recommended action
  • The model achieved perfect accuracy, correctly capturing negations, entities, and actions.

Judge transcript

Do not restart the server until the database backup finishes.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a flawless transcription of the audio, capturing the instruction perfectly and matching the reference transcript exactly without any errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors; the meaning is fully preserved.
Judge transcriptDo not restart the server until the database backup finishes.
Categories
no_error
Researcher notes
  • The model achieved perfect accuracy, correctly capturing negations, entities, and actions.

Provenance

  • Category negation modality scope
  • ASR slice negated action
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-server-001
  • Text SHA-256 9a47d14232d15170b7f3a734e6179db244d25afef06121759447c4f8a318d276
  • Duration 3.28s
  • Bytes 157364

Evaluation status: ok

negation modality scope · permission prohibition

asr-negation-refund-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

The candidate transcript preserves the meaning perfectly with only a minor formatting difference.

formatting only
Recommended action
  • The transcription is highly accurate, capturing the negation ('cannot') and crucial details flawlessly. No semantic errors present.

Judge transcript

The customer can cancel the trial, but cannot receive a refund after Friday.
Key differences
  • Omitted comma after 'trial'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.33 (scores: 95, 100, 100). Representative reason: The candidate transcript is perfectly accurate, matching the audio and reference transcript word-for-word. The only difference is the omission of a comma, which is a minor punctuation variant and does not affect the meaning or downstream usability.

judge samples: 95, 100, 100; avg 98.33
Meaning
preserved
Semantic impact
The candidate transcript preserves the meaning perfectly with only a minor formatting difference.
Judge transcriptThe customer can cancel the trial, but cannot receive a refund after Friday.
Key differences
  • Omitted comma after 'trial'
Categories
formatting_only
Researcher notes
  • The transcription is highly accurate, capturing the negation ('cannot') and crucial details flawlessly. No semantic errors present.

Provenance

  • Category negation modality scope
  • ASR slice permission prohibition
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-negation-refund-001
  • Text SHA-256 55a42c2151a5880e0aa2f00f13860c2acf87f536e24d9cec099342e0da50f211
  • Duration 3.80s
  • Bytes 182324

Evaluation status: ok

temporal scheduling accuracy · deadline day time

asr-temporal-deadline-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

There are no semantic errors, only a minor formatting variation for the time.

formatting only
Recommended action
  • The transcription is highly accurate. Orthographic variation in time representations is benign and does not affect semantic understanding.

Judge transcript

The benefits enrollment deadline is Monday at 9 a.m.
Key differences
  • The candidate writes '9:00 a.m.' while the reference writes '9 am'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is perfectly accurate, with only a minor formatting difference in how the time is represented ('9:00 a.m.' instead of '9 am'), preserving the spoken meaning completely.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors, only a minor formatting variation for the time.
Judge transcriptThe benefits enrollment deadline is Monday at 9 a.m.
Key differences
  • The candidate writes '9:00 a.m.' while the reference writes '9 am'.
Categories
formatting_only
Researcher notes
  • The transcription is highly accurate. Orthographic variation in time representations is benign and does not affect semantic understanding.

Provenance

  • Category temporal scheduling accuracy
  • ASR slice deadline day time
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-temporal-deadline-001
  • Text SHA-256 b118c3c195ab7d6fd58def5f41f4f899aa49aba40735806279b2f743868df37c
  • Duration 3.04s
  • Bytes 145844

Evaluation status: ok

entity factual integrity · person place entity

asr-entity-doctor-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

No semantic errors; the spoken words are perfectly captured.

formatting only
Recommended action
  • The model performs exceptionally well, accurately spelling the proper noun 'Rao' and the location 'Cedar Avenue'.

Judge transcript

Schedule the follow-up with Dr. Rao at the Cedar Avenue clinic.
Key differences
  • The candidate capitalizes 'Clinic' as 'Clinic' instead of 'clinic'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.33 (scores: 100, 95, 100). Representative reason: The candidate transcript is perfectly accurate and preserves the meaning completely, with the only difference being a minor capitalization of 'Clinic' and the segment block format.

judge samples: 100, 95, 100; avg 98.33
Meaning
preserved
Semantic impact
No semantic errors; the spoken words are perfectly captured.
Judge transcriptSchedule the follow-up with Dr. Rao at the Cedar Avenue clinic.
Key differences
  • The candidate capitalizes 'Clinic' as 'Clinic' instead of 'clinic'.
Categories
formatting_only
Researcher notes
  • The model performs exceptionally well, accurately spelling the proper noun 'Rao' and the location 'Cedar Avenue'.

Provenance

  • Category entity factual integrity
  • ASR slice person place entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-doctor-001
  • Text SHA-256 9ff5a0befafca6c4d8752ec9817f898caed7717e40beb76cd3eaa239a8cf9473
  • Duration 3.40s
  • Bytes 163124

Evaluation status: ok

entity factual integrity · product name contrast

asr-entity-product-001-local-tts

VibeVoice ASR (4-bit)
99 accurate

Semantic impact

No semantic errors; the transcription is perfectly accurate.

no error
Recommended action
  • Perfect transcription of the audio, accurately capturing product names like Zephyr Mini and Zephyr Max.

Judge transcript

Compare the Zephyr Mini plan with the Zephyr Max plan before renewal.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 99.33 (scores: 100, 100, 98). Representative reason: The candidate transcript is a perfect match with the spoken audio and reference transcript, preserving the exact meaning and all entities with no errors.

judge samples: 100, 100, 98; avg 99.33
Meaning
preserved
Semantic impact
No semantic errors; the transcription is perfectly accurate.
Judge transcriptCompare the Zephyr Mini plan with the Zephyr Max plan before renewal.
Categories
no_error
Researcher notes
  • Perfect transcription of the audio, accurately capturing product names like Zephyr Mini and Zephyr Max.

Provenance

  • Category entity factual integrity
  • ASR slice product name contrast
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-product-001
  • Text SHA-256 919dd10bf1c803fb423936e4f48c18d6469b7753fa5383a94ff26fc32bf6d745
  • Duration 3.84s
  • Bytes 184244

Evaluation status: ok

entity factual integrity · alphanumeric identifier

asr-entity-ticket-001-local-tts

VibeVoice ASR (4-bit)
57 inaccurate

Semantic impact

The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.

entity errornumber errorsubstitution
Recommended action
  • The model misheard 'one eight' (/wʌn eɪt/) as 'y' (/waɪ/) in the alphanumeric sequence 'PX-4187'. Alphanumeric sequence recognition is a known failure point that needs targeted training data or specialized language model biasing.

Judge transcript

Attach the call summary to ticket PX-4187 before closing the case.
Key differences
  • Candidate has 'px4y7' instead of 'PX-4187'
Judge rationale, vote details, and provenance

Average of 3 judge samples: 56.67 (scores: 55, 55, 60). Representative reason: The candidate transcript mistranscribes the critical alphanumeric ticket ID 'PX-4187' as 'px4y7'. While the rest of the command is perfectly captured, this entity error would prevent downstream ticketing systems from routing or associating the call summary correctly.

judge samples: 55, 55, 60; avg 56.67
Meaning
partial_loss
Semantic impact
The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.
Judge transcriptAttach the call summary to ticket PX-4187 before closing the case.
Key differences
  • Candidate has 'px4y7' instead of 'PX-4187'
Categories
entity_error, number_error, substitution
Researcher notes
  • The model misheard 'one eight' (/wʌn eɪt/) as 'y' (/waɪ/) in the alphanumeric sequence 'PX-4187'. Alphanumeric sequence recognition is a known failure point that needs targeted training data or specialized language model biasing.

Provenance

  • Category entity factual integrity
  • ASR slice alphanumeric identifier
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-ticket-001
  • Text SHA-256 845a5229d7c30013391f3297d1a9cc591ee108b93c4cf22218cdb749913fbf6e
  • Duration 4.60s
  • Bytes 220724

Evaluation status: ok

entity factual integrity · organization contact entity

asr-entity-organization-001-local-tts

VibeVoice ASR (4-bit)
97 accurate

Semantic impact

The candidate transcript contains only a minor spacing difference in the organization's name, preserving the full meaning of the audio.

formatting only
Recommended action
  • Consider incorporating common organizational entities or compound names into the language model's lexicon to prevent unnecessary splitting of words like 'Northstar'.

Judge transcript

Send the signed form to Northstar Benefits and copy Elena Morales.
Key differences
  • The compound name 'Northstar' is transcribed as 'North Star' (two words).
Judge rationale, vote details, and provenance

Average of 3 judge samples: 97.00 (scores: 95, 98, 98). Representative reason: The candidate transcript is exceptionally accurate, capturing all words, entities, and punctuation correctly. The only discrepancy is a minor formatting difference where 'Northstar' is split into two words 'North Star', which has zero impact on the semantic comprehension of the sentence.

judge samples: 95, 98, 98; avg 97.00
Meaning
preserved
Semantic impact
The candidate transcript contains only a minor spacing difference in the organization's name, preserving the full meaning of the audio.
Judge transcriptSend the signed form to Northstar Benefits and copy Elena Morales.
Key differences
  • The compound name 'Northstar' is transcribed as 'North Star' (two words).
Categories
formatting_only
Researcher notes
  • Consider incorporating common organizational entities or compound names into the language model's lexicon to prevent unnecessary splitting of words like 'Northstar'.

Provenance

  • Category entity factual integrity
  • ASR slice organization contact entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-organization-001
  • Text SHA-256 2c9213d5e33f7fc5a5924f6f72b4f91505c0b4c71b26848cb000e2d00ef961bc
  • Duration 3.80s
  • Bytes 182324

Evaluation status: ok

entity factual integrity · address entity

asr-entity-address-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is completely accurate.

formatting only
Recommended action
  • The model performed perfectly, accurately transcribing both numerical digits and letters ('812', '5B') as well as the street name.

Judge transcript

The delivery address is 812 Ashwood Lane, apartment 5B.
Key differences
  • The word 'apartment' is capitalized as 'Apartment' in the candidate transcript.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is exceptionally accurate, matching the audio and reference transcripts perfectly. The only difference is the capitalization of the word 'Apartment', which is a harmless formatting difference with absolutely no semantic impact.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is completely accurate.
Judge transcriptThe delivery address is 812 Ashwood Lane, apartment 5B.
Key differences
  • The word 'apartment' is capitalized as 'Apartment' in the candidate transcript.
Categories
formatting_only
Researcher notes
  • The model performed perfectly, accurately transcribing both numerical digits and letters ('812', '5B') as well as the street name.

Provenance

  • Category entity factual integrity
  • ASR slice address entity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-entity-address-001
  • Text SHA-256 5cf2d55b0dddffcc8b60dc856eb7816a36e745257daaf10abd184d3fcd40a9dc
  • Duration 4.24s
  • Bytes 203444

Evaluation status: ok

semantic paraphrase preservation · event state preservation

asr-semantic-paraphrase-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

No semantic errors found; the transcript is perfectly accurate.

no error
Recommended action
  • The transcript is completely accurate and matches the reference and audio perfectly.

Judge transcript

The technician replaced the cracked valve and confirmed the leak stopped.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.33 (scores: 95, 100, 100). Representative reason: The candidate transcript perfectly captures the spoken audio with no errors, omissions, or substitutions. The text matches the reference exactly, with only minor segment formatting wrappers.

judge samples: 95, 100, 100; avg 98.33
Meaning
preserved
Semantic impact
No semantic errors found; the transcript is perfectly accurate.
Judge transcriptThe technician replaced the cracked valve and confirmed the leak stopped.
Categories
no_error
Researcher notes
  • The transcript is completely accurate and matches the reference and audio perfectly.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice event state preservation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-paraphrase-001
  • Text SHA-256 459210fe6f6f4d41259ceba39ab9c8fc608abfb927c5a49c8750ac89f86e21d0
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

semantic paraphrase preservation · causal relation

asr-semantic-causal-001-local-tts

VibeVoice ASR (4-bit)
98 accurate

Semantic impact

No semantic errors are present; the text is fully accurate despite the metadata wrapper.

formatting only
Recommended action
  • The transcription accuracy is 100%. The ASR pipeline outputted metadata formatting which might need to be parsed or stripped depending on the downstream application's expectations.

Judge transcript

The upload failed because the token expired, so the client retried after refresh.
Key differences
  • The transcript is contained within a JSON block with 'Start', 'End', 'Speaker', and 'Content' keys instead of being a plain string.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 97.67 (scores: 95, 100, 98). Representative reason: The candidate transcript successfully captures the spoken audio with perfect accuracy. The only difference is that the text is wrapped in a JSON format containing timestamp and speaker metadata, which does not degrade the semantic meaning of the spoken words.

judge samples: 95, 100, 98; avg 97.67
Meaning
preserved
Semantic impact
No semantic errors are present; the text is fully accurate despite the metadata wrapper.
Judge transcriptThe upload failed because the token expired, so the client retried after refresh.
Key differences
  • The transcript is contained within a JSON block with 'Start', 'End', 'Speaker', and 'Content' keys instead of being a plain string.
Categories
formatting_only
Researcher notes
  • The transcription accuracy is 100%. The ASR pipeline outputted metadata formatting which might need to be parsed or stripped depending on the downstream application's expectations.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice causal relation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-causal-001
  • Text SHA-256 caeb5d404f654e7d980c3f61cc1576806deb874cf24af2120f26700bb7aa471e
  • Duration 4.72s
  • Bytes 226484

Evaluation status: ok

semantic paraphrase preservation · metric tradeoff

asr-semantic-comparison-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors occurred; the words are perfectly transcribed.

formatting only
Recommended action
  • ASR transcription is flawless.
  • If raw text is required, a simple post-processing parser can extract the content field from the JSON structure.

Judge transcript

The candidate model improved recall but reduced precision on noisy calls.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing segment metadata ('Start', 'End', 'Speaker', 'Content').
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript perfectly captures the spoken audio with 100% accuracy, matching the reference text exactly. The only difference is the JSON formatting wrapping the output, which contains metadata like timestamps and speaker tags, but does not impact meaning preservation.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors occurred; the words are perfectly transcribed.
Judge transcriptThe candidate model improved recall but reduced precision on noisy calls.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing segment metadata ('Start', 'End', 'Speaker', 'Content').
Categories
formatting_only
Researcher notes
  • ASR transcription is flawless.
  • If raw text is required, a simple post-processing parser can extract the content field from the JSON structure.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice metric tradeoff
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-comparison-001
  • Text SHA-256 fbaadae669f95ffc77afc2da5b9661b0a48486f7e71ada1bd05170df2649a739
  • Duration 4.32s
  • Bytes 207284

Evaluation status: ok

semantic paraphrase preservation · coreference ambiguity

asr-semantic-coreference-001-local-tts

VibeVoice ASR (4-bit)
99 accurate

Semantic impact

The candidate transcript matches the audio perfectly, though it differs from the reference transcript on the first name (Rhea vs. Maria).

no error
Recommended action
  • The candidate transcript is more accurate to the actual audio than the reference transcript in this case. The reference transcript has 'Maria' but the audio clearly says 'Rhea'.

Judge transcript

Rhea called Priya after she reviewed the contract changes.
Key differences
  • Candidate uses 'Rhea' matching the audio, while the reference transcript uses 'Maria'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 98.67 (scores: 98, 98, 100). Representative reason: The candidate transcript is highly accurate and perfectly matches the spoken audio, capturing 'Rhea' instead of the reference's 'Maria'. The rest of the sentence is identical and fully preserves the meaning.

judge samples: 98, 98, 100; avg 98.67
Meaning
preserved
Semantic impact
The candidate transcript matches the audio perfectly, though it differs from the reference transcript on the first name (Rhea vs. Maria).
Judge transcriptRhea called Priya after she reviewed the contract changes.
Key differences
  • Candidate uses 'Rhea' matching the audio, while the reference transcript uses 'Maria'.
Categories
no_error
Researcher notes
  • The candidate transcript is more accurate to the actual audio than the reference transcript in this case. The reference transcript has 'Maria' but the audio clearly says 'Rhea'.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice coreference ambiguity
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-coreference-001
  • Text SHA-256 7ec151a16318fda290762dbf9f9ad83e2ce7c5df725aeba3043965cfd596e437
  • Duration 3.52s
  • Bytes 168884

Evaluation status: ok

semantic paraphrase preservation · conditional instruction

asr-semantic-instruction-001-local-tts

VibeVoice ASR (4-bit)
97 accurate

Semantic impact

The meaning is fully preserved with zero transcription errors, though the output format contains structured JSON metadata.

formatting only
Recommended action
  • The core ASR engine performed flawlessly. The formatting wrapper (JSON with timestamps and speaker diarization) should be parsed or stripped depending on downstream application requirements.

Judge transcript

If the first backup succeeds, skip the manual export and notify operations.
Key differences
  • The candidate transcript is formatted as a JSON array with 'Start', 'End', 'Speaker', and 'Content' keys instead of plain text.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 96.67 (scores: 95, 100, 95). Representative reason: The candidate transcript captures the spoken audio perfectly with 100% accuracy. However, it is wrapped in a JSON formatting structure containing timestamp and speaker metadata instead of being returned as plain text.

judge samples: 95, 100, 95; avg 96.67
Meaning
preserved
Semantic impact
The meaning is fully preserved with zero transcription errors, though the output format contains structured JSON metadata.
Judge transcriptIf the first backup succeeds, skip the manual export and notify operations.
Key differences
  • The candidate transcript is formatted as a JSON array with 'Start', 'End', 'Speaker', and 'Content' keys instead of plain text.
Categories
formatting_only
Researcher notes
  • The core ASR engine performed flawlessly. The formatting wrapper (JSON with timestamps and speaker diarization) should be parsed or stripped depending on downstream application requirements.

Provenance

  • Category semantic paraphrase preservation
  • ASR slice conditional instruction
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-semantic-instruction-001
  • Text SHA-256 c095d3a9fadb208cbc7ae9c12d4168e37dff40b73108cd2b28c219e12511a104
  • Duration 4.48s
  • Bytes 214964

Evaluation status: ok

acoustic noise robustness · cafe background command

asr-noise-cafe-order-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

The transcript is perfectly accurate with no semantic errors.

no error
Recommended action
  • The ASR engine captured all entities ('oat milk', 'latte', 'caramel drizzle') and commands perfectly.
  • The output contains structured timestamp and speaker information which is accurate.

Judge transcript

Add oat milk to the latte and leave out the caramel drizzle.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is an exact word-for-word representation of the audio, preserving the entire meaning perfectly. The text is flawless, with the only difference being the structured JSON wrapper in the candidate.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
The transcript is perfectly accurate with no semantic errors.
Judge transcriptAdd oat milk to the latte and leave out the caramel drizzle.
Categories
no_error
Researcher notes
  • The ASR engine captured all entities ('oat milk', 'latte', 'caramel drizzle') and commands perfectly.
  • The output contains structured timestamp and speaker information which is accurate.

Provenance

  • Category acoustic noise robustness
  • ASR slice cafe background command
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-cafe-order-001
  • Text SHA-256 613cfd4c60424169e80cba7ab2ecc6f0fcb35f677cdfbfd9746b4efaebf165ef
  • Duration 3.28s
  • Bytes 157364

Evaluation status: ok

acoustic noise robustness · vehicle noise navigation

asr-noise-driving-route-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

There are no semantic errors in the candidate transcript.

no error
Recommended action
  • The ASR output successfully captured the entire phrase with perfect accuracy. The metadata formatting (timestamps and speaker tag) contains the correct text verbatim.

Judge transcript

Take the second exit after the bridge, then stay in the right lane.

No key differences recorded.

Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is a perfect word-for-word match to the audio, preserving the spoken instruction completely with no errors.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
There are no semantic errors in the candidate transcript.
Judge transcriptTake the second exit after the bridge, then stay in the right lane.
Categories
no_error
Researcher notes
  • The ASR output successfully captured the entire phrase with perfect accuracy. The metadata formatting (timestamps and speaker tag) contains the correct text verbatim.

Provenance

  • Category acoustic noise robustness
  • ASR slice vehicle noise navigation
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-driving-route-001
  • Text SHA-256 c3ce8e7e94fd6ea30d703cbca09fc8f6c03373f0ea379bfc5d94dd49aa48ce26
  • Duration 3.72s
  • Bytes 178484

Evaluation status: ok

acoustic noise robustness · industrial noise identifier

asr-noise-warehouse-001-local-tts

VibeVoice ASR (4-bit)
100 accurate

Semantic impact

No semantic errors; the candidate transcript is perfectly accurate.

formatting only
Recommended action
  • Excellent transcription quality. The only variation is a standard text normalization difference (digit vs word for number).

Judge transcript

Scan bin A17 before moving the pallet to loading dock three.
Key differences
  • The word 'three' is written as the digit '3'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 100.00 (scores: 100, 100, 100). Representative reason: The candidate transcript is completely accurate, with the only difference from the reference being the formatting of the number 'three' as '3', which does not affect semantic meaning.

judge samples: 100, 100, 100; avg 100.00
Meaning
preserved
Semantic impact
No semantic errors; the candidate transcript is perfectly accurate.
Judge transcriptScan bin A17 before moving the pallet to loading dock three.
Key differences
  • The word 'three' is written as the digit '3'.
Categories
formatting_only
Researcher notes
  • Excellent transcription quality. The only variation is a standard text normalization difference (digit vs word for number).

Provenance

  • Category acoustic noise robustness
  • ASR slice industrial noise identifier
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-warehouse-001
  • Text SHA-256 f7c26924015438634f3eb4be84b886d690e9009b39ecd8f4fa88408ff3ec2e11
  • Duration 4.08s
  • Bytes 195764

Evaluation status: ok

acoustic noise robustness · reverberant meeting room

asr-noise-reverberant-room-001-local-tts

VibeVoice ASR (4-bit)
67 needs review

Semantic impact

The room number '412' was incorrectly transcribed as '42012', creating a highly misleading location error.

number errorsubstitution
Recommended action
  • The model inserted extra digits into the room number ('42012' instead of '412'). Better digit-level acoustic modeling is needed to prevent number insertion errors.

Judge transcript

The quarterly review starts in room 412 at half past three.
Key differences
  • Candidate transcribed '412' as '42012'.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 66.67 (scores: 75, 70, 55). Representative reason: The candidate transcript accurately captures the entire sentence except for the room number, transcribing '412' as '42012'. This error in the room number is critical as it would mislead someone trying to find the meeting.

judge samples: 75, 70, 55; avg 66.67
Meaning
partial_loss
Semantic impact
The room number '412' was incorrectly transcribed as '42012', creating a highly misleading location error.
Judge transcriptThe quarterly review starts in room 412 at half past three.
Key differences
  • Candidate transcribed '412' as '42012'.
Categories
number_error, substitution
Researcher notes
  • The model inserted extra digits into the room number ('42012' instead of '412'). Better digit-level acoustic modeling is needed to prevent number insertion errors.

Provenance

  • Category acoustic noise robustness
  • ASR slice reverberant meeting room
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-reverberant-room-001
  • Text SHA-256 f0f9e458a7bf2804ef37a0a709c4a7882c50036351a8f11cd69c590f1805b4d8
  • Duration 3.56s
  • Bytes 170804

Evaluation status: ok

acoustic noise robustness · overlapping speech target

asr-noise-overlapping-speech-001-local-tts

VibeVoice ASR (4-bit)
97 accurate

Semantic impact

The spoken meaning is perfectly preserved, with the only discrepancy being the JSON metadata format in the candidate output.

formatting only
Recommended action
  • The transcription content is flawless. If plain text is expected, the post-processing pipeline should strip the JSON wrapper, or the model should be instructed to avoid metadata output.

Judge transcript

Ignore the side conversation and record the final answer as option C.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing 'Start', 'End', 'Speaker', and 'Content' metadata fields instead of plain text.
Judge rationale, vote details, and provenance

Average of 3 judge samples: 96.67 (scores: 95, 100, 95). Representative reason: The candidate transcript is 100% accurate to the spoken audio, preserving the exact wording and meaning. The only discrepancy is that the candidate output is wrapped in a JSON-style array containing metadata (Start, End, Speaker, and Content) instead of presenting as a plain text string.

judge samples: 95, 100, 95; avg 96.67
Meaning
preserved
Semantic impact
The spoken meaning is perfectly preserved, with the only discrepancy being the JSON metadata format in the candidate output.
Judge transcriptIgnore the side conversation and record the final answer as option C.
Key differences
  • The candidate transcript is wrapped in a JSON structure containing 'Start', 'End', 'Speaker', and 'Content' metadata fields instead of plain text.
Categories
formatting_only
Researcher notes
  • The transcription content is flawless. If plain text is expected, the post-processing pipeline should strip the JSON wrapper, or the model should be instructed to avoid metadata output.

Provenance

  • Category acoustic noise robustness
  • ASR slice overlapping speech target
  • Model mlx-community/chatterbox-turbo-6bit
  • Candidate model mlx-community/VibeVoice-ASR-4bit
  • Voice af heart
  • Language en
  • Lang code en
  • Sample local synthetic tts
  • Source case asr-noise-overlapping-speech-001
  • Text SHA-256 e72cbb8fa9f427ff6a5ac7dcb1cf1e153ec37baeeba0a80c6be4033169b3f66b
  • Duration 4.08s
  • Bytes 195764

Evaluation status: ok

Dataset coverage, metadata, and calibration

Candidate Metadata

ASR Slice
  • function word substitution 3
  • homophone confusion 3
  • short command edit distance 3
  • disfluency token accounting 3
  • sentence boundary punctuation 3
  • decimal measurement 3
  • digit sequence 3
  • percentage minimal pair 3
  • clinical negation 3
  • permission scope 3
  • unless condition 3
  • am pm window 3
  • relative time relation 3
  • ordinal date minimal pair 3
  • duration contrast 3
  • amount contrast 3
  • dosage unit 3
  • negated action 3
  • permission prohibition 3
  • deadline day time 3
  • person place entity 3
  • product name contrast 3
  • alphanumeric identifier 3
  • organization contact entity 3
  • address entity 3
  • event state preservation 3
  • causal relation 3
  • metric tradeoff 3
  • coreference ambiguity 3
  • conditional instruction 3
  • cafe background command 3
  • vehicle noise navigation 3
  • industrial noise identifier 3
  • reverberant meeting room 3
  • overlapping speech target 3
ASR Model
  • mlx-community/whisper-large-v3-turbo-asr-fp16 35
  • mlx-community/Qwen3-ASR-1.7B-8bit 35
  • mlx-community/VibeVoice-ASR-4bit 35
Synthesis Voice
  • af heart 105
Language
  • en 105
Evaluation Category
  • transcription accuracy wer 15
  • numeric unit integrity 15
  • negation modality scope 15
  • temporal scheduling accuracy 15
  • entity factual integrity 15
  • semantic paraphrase preservation 15
  • acoustic noise robustness 15
Sample Kind
  • local synthetic tts 105
Issues By Category
  • temporal scheduling accuracy / formatting only 12
  • numeric unit integrity / formatting only 9
  • entity factual integrity / formatting only 7
  • negation modality scope / formatting only 6
  • transcription accuracy wer / formatting only 5
  • entity factual integrity / entity error 5
  • transcription accuracy wer / substitution 4
  • entity factual integrity / substitution 4
Issues By ASR Slice
  • permission scope / formatting only 3
  • am pm window / formatting only 3
  • ordinal date minimal pair / formatting only 3
  • amount contrast / formatting only 3
  • deadline day time / formatting only 3
  • person place entity / formatting only 3
  • alphanumeric identifier / entity error 3
  • homophone confusion / substitution 2
Issues By Model
  • mlx-community/VibeVoice-ASR-4bit / formatting only 17
  • mlx-community/whisper-large-v3-turbo-asr-fp16 / formatting only 15
  • mlx-community/Qwen3-ASR-1.7B-8bit / formatting only 14
  • mlx-community/Qwen3-ASR-1.7B-8bit / substitution 6
  • mlx-community/whisper-large-v3-turbo-asr-fp16 / substitution 5
  • mlx-community/whisper-large-v3-turbo-asr-fp16 / number error 3
  • mlx-community/VibeVoice-ASR-4bit / substitution 3
  • mlx-community/whisper-large-v3-turbo-asr-fp16 / entity error 2
Issues By Voice
  • af heart / formatting only 46
  • af heart / substitution 14
  • af heart / number error 7
  • af heart / entity error 5
Issues By Language
  • en / formatting only 46
  • en / substitution 14
  • en / number error 7
  • en / entity error 5
Scores By Category
  • entity factual integrity avg 89.3 / n 15 / 53-100
  • acoustic noise robustness avg 93.9 / n 15 / 62-100
  • numeric unit integrity avg 94.4 / n 15 / 60-100
  • transcription accuracy wer avg 94.6 / n 15 / 76-100
  • temporal scheduling accuracy avg 96.5 / n 15 / 53-100
  • semantic paraphrase preservation avg 98.7 / n 15 / 91-100
  • negation modality scope avg 99.9 / n 15 / 98-100
Scores By ASR Slice
  • alphanumeric identifier avg 55.0 / n 3 / 53-57
  • reverberant meeting room avg 70.7 / n 3 / 62-83
  • digit sequence avg 73.3 / n 3 / 60-100
  • duration contrast avg 83.7 / n 3 / 53-100
  • homophone confusion avg 84.3 / n 3 / 76-100
  • function word substitution avg 90.7 / n 3 / 85-100
  • organization contact entity avg 92.3 / n 3 / 90-97
  • coreference ambiguity avg 96.7 / n 3 / 91-100
Scores By Model
  • mlx-community/whisper-large-v3-turbo-asr-fp16 avg 94.3 / n 35 / 53-100
  • mlx-community/Qwen3-ASR-1.7B-8bit avg 95.2 / n 35 / 55-100
  • mlx-community/VibeVoice-ASR-4bit avg 96.5 / n 35 / 57-100
Scores By Voice
  • af heart avg 95.3 / n 105 / 53-100
Scores By Language
  • en avg 95.3 / n 105 / 53-100
Failures By CategoryNo category failures
Failures By ASR SliceNo slice failures
Failures By ModelNo model failures
Failures By VoiceNo voice failures
Failures By LanguageNo language failures
Failures By Sample KindNo sample-kind failures

Calibration Checks

All calibration expectations matched
Advanced comparison analysis

Model-Category Action Matrix

entity factual integrity

Whisper Large v3 Turbo (FP16)

88.6n=5 · range 53-100
Likely Fix Areas
formatting onlyentity errortext faithfulnesssubstitution
Category Guidance
Source basis
Entity recall and named-entity error analysis are common ASR quality slices beyond aggregate WER.Custom vocabulary and entity recall are critical for product, contact, and organization names.
Issues
formatting only x2entity error x2number error x1substitution x1
Failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 53 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorformatting only
    The unique ticket ID was incorrectly transcribed with incorrect numbers and missing spacing, causing a failure in downstream entity mapping.
  • asr-entity-organization-001-local-tts 90 / accurate
    status: ok
    source: asr-entity-organization-001
    issues: substitutionentity error
    The candidate transcript contains a minor name spelling variation ('Alina' vs 'Elena') but preserves all other information perfectly.
entity factual integrity

Qwen3 ASR 1.7B (8-bit)

89.0n=5 · range 55-100
Likely Fix Areas
formatting onlyentity errorsubstitution
Category Guidance
Source basis
Entity recall and named-entity error analysis are common ASR quality slices beyond aggregate WER.Custom vocabulary and entity recall are critical for product, contact, and organization names.
Issues
formatting only x2entity error x2substitution x2
Failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 55 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errorsubstitution
    The critical ticket ID entity was misrecognized, preventing successful mapping in downstream applications.
  • asr-entity-organization-001-local-tts 90 / accurate
    status: ok
    source: asr-entity-organization-001
    issues: substitutionentity error
    The candidate transcript contains a minor phonetic substitution of the proper name 'Elena' with 'Alina'.
entity factual integrity

VibeVoice ASR (4-bit)

90.2n=5 · range 57-100
Likely Fix Areas
formatting onlyentity errortext faithfulnesssubstitution
Category Guidance
Source basis
Entity recall and named-entity error analysis are common ASR quality slices beyond aggregate WER.Custom vocabulary and entity recall are critical for product, contact, and organization names.
Issues
formatting only x3entity error x1number error x1substitution x1
Failures
No failed evaluations
Representative low-score samples
  • asr-entity-ticket-001-local-tts 57 / inaccurate
    status: ok
    source: asr-entity-ticket-001
    issues: entity errornumber errorsubstitution
    The ticket ID entity is corrupted, making it impossible to correctly identify the target ticket.
  • asr-entity-organization-001-local-tts 97 / accurate
    status: ok
    source: asr-entity-organization-001
    issues: formatting only
    The candidate transcript contains only a minor spacing difference in the organization's name, preserving the full meaning of the audio.
temporal scheduling accuracy

Whisper Large v3 Turbo (FP16)

90.6n=5 · range 53-100
Likely Fix Areas
formatting onlytext faithfulnesssubstitution
Category Guidance
Source basis
AM/PM and temporal relation errors can be severe even when token edit distance is small.Temporal relation preservation belongs in semantic ASR evaluation beyond WER.
Issues
formatting only x3number error x1substitution x1
Failures
No failed evaluations
Representative low-score samples
  • asr-temporal-duration-001-local-tts 53 / inaccurate
    status: ok
    source: asr-temporal-duration-001
    issues: number errorsubstitution
    The pump's rest duration was incorrectly transcribed as 15 minutes instead of fifty minutes.
  • asr-temporal-date-001-local-tts 100 / accurate
    status: ok
    source: asr-temporal-date-001
    issues: formatting only
    No semantic errors exist; only formatting differences are present.
numeric unit integrity

Qwen3 ASR 1.7B (8-bit)

92.0n=5 · range 60-100
Likely Fix Areas
formatting onlysubstitutiontext faithfulness
Category Guidance
Source basis
Numeric normalization should distinguish equivalent formatting from value-changing recognition errors.Digit-sequence accuracy is a practical ASR slice for forms, support, and voice agents.
Issues
formatting only x3substitution x1number error x1
Failures
No failed evaluations
Representative low-score samples
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The candidate incorrectly transcribes the final digit of the account number, substituting '8' with the letter 'E'.
  • asr-numeric-measurement-001-local-tts 100 / accurate
    status: ok
    source: asr-numeric-measurement-001
    issues: formatting only
    There are no semantic errors; the spoken words and numeric representations match in meaning perfectly.
numeric unit integrity

Whisper Large v3 Turbo (FP16)

92.0n=5 · range 60-100
Likely Fix Areas
formatting onlysubstitutiontext faithfulness
Category Guidance
Source basis
Numeric normalization should distinguish equivalent formatting from value-changing recognition errors.Digit-sequence accuracy is a practical ASR slice for forms, support, and voice agents.
Issues
formatting only x4substitution x1number error x1
Failures
No failed evaluations
Representative low-score samples
  • asr-numeric-account-001-local-tts 60 / needs review
    status: ok
    source: asr-numeric-account-001
    issues: substitutionnumber error
    The last digit of the account number was incorrectly transcribed as the letter 'E' instead of the number '8'.
  • asr-numeric-measurement-001-local-tts 100 / accurate
    status: ok
    source: asr-numeric-measurement-001
    issues: formatting only
    No semantic errors are present in the candidate transcript.
acoustic noise robustness

Qwen3 ASR 1.7B (8-bit)

92.4n=5 · range 62-100
Likely Fix Areas
formatting onlytext faithfulnesssubstitution
Category Guidance
Source basis
CHiME-style and far-field ASR benchmarks stress recognition under real background noise, not only clean-speech WER.Robust ASR evaluation includes adverse acoustic environments such as vehicle noise and reverberation.
Issues
formatting only x1number error x1substitution x1
Failures
No failed evaluations
Representative low-score samples
  • asr-noise-reverberant-room-001-local-tts 62 / needs review
    status: ok
    source: asr-noise-reverberant-room-001
    issues: number errorsubstitution
    The room number is incorrectly transcribed as 4212 instead of 412, altering the meeting location.
  • asr-noise-cafe-order-001-local-tts 100 / accurate
    status: ok
    source: asr-noise-cafe-order-001
    issues: No judge issue categories
    There are no semantic errors; the candidate transcript is completely accurate.
transcription accuracy wer

Qwen3 ASR 1.7B (8-bit)

92.8n=5 · range 77-100
Likely Fix Areas
substitutionformatting only
Category Guidance
Source basis
WER remains the standard ASR edit-distance baseline; include low-level substitutions before semantic judging.WER catches token-level confusions while semantic review determines downstream severity.
Issues
substitution x2formatting only x2
Failures
No failed evaluations
Representative low-score samples
  • asr-wer-homophone-001-local-tts 77 / needs review
    status: ok
    source: asr-wer-homophone-001
    issues: substitution
    The homophone error 'write' instead of 'right' slightly degrades the semantic clarity of the sentence by creating an illogical phrase 'write team'.
  • asr-wer-function-word-001-local-tts 87 / accurate
    status: ok
    source: asr-wer-function-word-001
    issues: substitution
    The candidate substituted 'patch is' with 'patches', slightly altering the grammar but preserving the primary meaning.