Skip to content

feat(audio): remember speakers from diarization - #12414

Merged
mudler merged 11 commits into
masterfrom
feat/diarization-profiles
Oct 2, 2026
Merged

mudler merged 11 commits into
masterfrom
feat/diarization-profiles

Conversation

@mudler-agent

@mudler-agent mudler-agent commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Remember speakers from arbitrary multi-speaker recordings without preparing
separate clean enrollment clips. Diarization exports one profile per usable
speaker, and users preview clean speech before explicitly naming and saving it.

Follow-up to #12382 and mudler/parakeet.cpp#79; no issue-closing claim.
Dependency: mudler/parakeet.cpp#80 is merged.
The backend pins the merged native commit
bee7c14dfcc23613df58176c59a40459e7b47095.

  • Add opt-in include_speaker_profiles to diarization and transport the
    versioned profiles through the parakeet backend. Normal responses stay
    unchanged. Unsupported profile export fails explicitly rather than
    silently falling back.
  • Extend the existing /v1/voice/register endpoint for explicit profile
    enrollment while preserving audio enrollment. Validate profiles against
    identity and dimension supplied by the trusted loaded encoder, not the
    caller. Exact GGUF bytes determine encoder identity.
  • Use one profile-capable diarization for slots, clean spans, and names;
    assign timestamped ASR words without a second diarization. Associate
    profiles by raw speaker slot, including zero and sparse slots, never by
    display order or name.
  • Add a Diarization page in Studio with clean-interval previews and a
    “Name and remember” action. Relabel only after successful enrollment,
    reject stale results after recording/model changes, and keep independent
    registration IDs even when display names match.
  • Require voice-recognition permission for export/enrollment alongside
    existing diarization and model-access checks. Exclude diarization and
    registration exchanges before API trace body capture, including the
    diarization alias and JSON audio. This flow does not persist vectors or
    recordings in browser storage.
  • Update user documentation and Swagger artifacts.

Notes for Reviewers

Verification

The recorded focused Go suites pass. With protobuf tooling installed, run
from the repository root:

make protogen-go
GOMAXPROCS=2 go test -count=1 -p 2 \
  ./pkg/grpc/proto \
  ./core/schema \
  ./core/services/voicerecognition \
  ./core/backend \
  ./backend/go/parakeet-cpp \
  ./core/http/endpoints/localai \
  ./core/http/auth \
  ./core/http/middleware

Coverage includes profile validation, trusted encoder compatibility,
permissions, raw-slot mapping, duplicate-name registrations, and trace
privacy. These tests use synthetic data or mocked inference.

The UI production build passes under Node 22. The focused Playwright
suite passes 10 tests, with mocked API responses, not real inference.
To reproduce with Chromium installed, use two terminals:

# Terminal 1, using Node 22
cd core/http/react-ui
npm ci
npm run build
npm run dev -- --port 8089
# Terminal 2
cd core/http/react-ui
PLAYWRIGHT_EXTERNAL_SERVER=1 PW_WORKERS=1 \
  npx playwright test e2e/diarization-profiles.spec.js --project=chromium

Swagger JSON, YAML, and docs.go paths/definitions were checked for semantic
consistency. Generator execution was blocked by a missing go.sum
dependency; this is not a successful regeneration claim.

Limitations and compatibility

  • The existing voice registry is global and ephemeral per process, not
    per-user durable storage. Registrations disappear on restart and are not
    synchronized across independent frontends.
  • Profiles are sensitive, unsigned biometric data, not proof of identity
    or consent. Enrollment remains explicit; exporting never registers anyone.
  • Real-model end-to-end enrollment/recognition, listening checks, and
    visual inspection remain unverified. No full make build or full test
    suite success is claimed.
  • Native verification reports four passing tests and four model-dependent
    skips. See the dependency PR for its CMake/CTest reproduction commands.
  • Existing UI lint/style debt remains unchanged; this is not a claim that
    the entire UI passes lint. The production build retains chunk-size warnings.
  • Existing opt-out API behavior remains compatible. The new export workflow
    requires the pinned native symbols and a configured speaker encoder.

Signed commits

  • Yes, I signed my commits.
  • Documentation updated (docs/content/) for user-facing changes, or not applicable

The commits are not cryptographically signed; the checkbox remains unchecked.

mudler added 10 commits October 2, 2026 00:22
Add the versioned profile schema for explicit speaker enrollment.
Validate compatibility against separately supplied loaded-encoder metadata.
Reject unusable speakers, invalid vectors, and inconsistent clean spans.

This slice does not change HTTP routes, backend integration, or the UI.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Export opt-in speaker profiles and trusted encoder metadata.
Replay registrations by ID so duplicate display names keep independent
vectors.

Use one profile-capable diarization for slots, names, and clean spans.
Assign timestamped ASR words to those slots without a second diarization.
Preserve legacy opt-out and no-ASR behavior, and propagate failures.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Gate profile exports with voice-recognition permission and validate
registration against metadata from the loaded encoder. Preserve audio
enrollment and independent registrations with duplicate display names.

Exclude diarization and registration exchanges before API trace capture
so persisted traces cannot retain profile vectors or JSON audio.

Defer candidate dimensions to trusted loaded metadata. Sort candidates
by registration ID so incompatible profiles cannot suppress legacy voices
through registry iteration order. Keep portable identity checks closed
when trusted metadata is unavailable.

Test persisted traces, explicit slot zero, and selection through offline
and live transport. Document privacy and the ephemeral registry lifecycle.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add a Studio page for diarization and opt-in speaker profiles. Preview
clean intervals from the original recording before explicit registration.

Join profiles by raw speaker labels, preserve duplicate names, and relabel
turns only after a successful save. Discard stale results when the model
or recording changes. Share registration metadata with voice management
without storing vectors or recordings from this flow.

Document permissions and the global, ephemeral registry. Cover enrollment,
permissions, previews, and asynchronous races with mocked Playwright tests.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Replace the stale enrollment limitation with the current HTTP workflow.
Distinguish native transport from explicit registration and link its docs.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Use the merged commit from mudler/parakeet.cpp#80.
Its tree matches the previously accepted native pin.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Connect the existing gallery modes to the speaker enrollment workflow.
Show installation, private profile export, explicit raw-slot registration,
and later recognition without another export.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Put the diarization walkthrough on the LocalAI website in the feature PR.
Cover the three gallery modes, explicit enrollment, and privacy limits.
Link setup instructions and keep availability conditional on feature support.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Explain what users can do with recordings before the setup steps.
Replace the technical walkthrough with a short Studio guide and link
readers to the existing reference for model names and developer use.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Present speaker recognition through everyday uses and a short UI flow.
Keep technical reference details in the existing documentation.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
@mudler
mudler force-pushed the feat/diarization-profiles branch from 2d333cc to 9a3aa17 Compare October 2, 2026 00:22
Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions.

Assisted-by: nib:gpt-5.6-sol

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
@mudler
mudler merged commit 9eb5a9e into master Oct 2, 2026
79 checks passed
@mudler
mudler deleted the feat/diarization-profiles branch October 2, 2026 06:14
@mudler-agent mudler-agent added the enhancement New feature or request label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants