Repository navigation
fix(whisperx): keep the transcript when diarization fails, report real errors - #12427
Conversation
|
@mudler code review looks good at The diarization fallback preserves the available transcript, while transcription failures now return a gRPC error. The missing-token precondition remains separate. I added documentation for the fallback and error behavior. Validation: Static review and git diff --check passed. Python tests could not run because the installed Python executable requires a runtime loader unavailable in this environment. The documentation commit is pushed to this PR branch. I left its human DCO sign-off to the contributor. No merge performed. |
…l errors AudioTranscription caught every exception and returned an empty TranscriptResult. A failed diarization step therefore discarded a transcript that was already finished: with an HF token that has not accepted the terms of the gated pyannote pipeline, the download fails with 403 and every transcription came back as an empty text with HTTP 200. Diarization now degrades: if it fails, the transcript is returned without speaker labels and the reason is logged. Any other failure aborts the call with INTERNAL instead of pretending success with an empty text. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
d07bf5f to
13307ea
Compare
|
Thanks for the review and for adding the documentation. That doc commit had no DCO sign-off, so the DCO check stayed red. I folded equivalent documentation (reworded and checked against the code) into the signed-off fix commit and force-pushed the branch. The code and tests are unchanged from the reviewed state. Sorry for leaving the docs out of the original submission. |
Description
AudioTranscriptionin the whisperx backend caught every exception and returned an emptyTranscriptResult. A failure in the diarization step therefore discarded a transcript that was already finished. Since/v1/audio/transcriptionstreats a missingdiarizefield as true, every request reaches that step; when the HF token has not accepted the terms of the gatedpyannote/speaker-diarization-community-1pipeline, the download fails with 403 and every transcription came back as an empty text with HTTP 200 (backend log: language detected, alignment done, thenGatedRepoError, result empty).Diarization now degrades: if it fails, the transcript is returned without speaker labels and the reason is logged. Any other failure aborts the call with
INTERNALinstead of returning an empty, successful-looking result.test_transcript_utils.pycovers the newdiarize_or_keephelper: a refused diarization keeps the transcript and logs the error; a successful one is returned.Notes for Reviewers
A related inconsistency I did not change: the endpoint enables diarization unless
diarize=falseis sent (relied on by parakeet-cpp since #12335), while whisperx rejects diarization withoutHF_TOKEN(#8744de44d). So a whisperx install without a token fails every plain transcription withFailedPrecondition. Making whisperx skip diarization when it was not explicitly requested would need the request to carry whether the field was set; I left that decision to you.Signed commits
🤖 Generated with Claude Code