Skip to content

feat(voice): dictation in the tui, desktop, and vs code - #47

Merged
Zfinix merged 4 commits into
mainfrom
feat/voice-dictation
Sep 29, 2026
Merged

Zfinix merged 4 commits into
mainfrom
feat/voice-dictation

Conversation

@Zfinix

@Zfinix Zfinix commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Steps 1 and 2 of #46: dictation in the TUI, the desktop app and VS Code.

What changes

  • aster-voice crate. Records from the default microphone, encodes mono 16-bit WAV, and transcribes with ElevenLabs Scribe (ELEVENLABS_API_KEY) or OpenAI gpt-4o-transcribe (OPENAI_API_KEY), in that order.
  • TUI. ctrl+r starts listening, a second ctrl+r transcribes into the composer, and Esc discards. The footer shows ● listening · ctrl+r to stop, then transcribing…. Failures appear as one plain sentence with the provider detail below it.
  • aster dictate. A new command that records until stdin gets a line or closes, then prints NDJSON: listening, transcribing, then transcript (text) or error (message, detail). The desktop app and VS Code use it. Neither has a channel to the CLI before a message is sent, so this leaves the --stream and ACP contracts alone.
  • Desktop. A mic button in the chat composer, with errors shown as a toast. New Tauri commands start_dictation, stop_dictation and cancel_dictation. The bundle gains the com.apple.security.device.audio-input entitlement and NSMicrophoneUsageDescription, without which the hardened runtime kills the process when it opens the mic.
  • VS Code. A mic button in the webview composer, editor only. Errors show as a VS Code notification, with the detail in the Aster output channel. A new dictation message pair is added to protocol.ts, and the dev host relays it too.
  • Keys and docs. aster key list gains a Voice group; --json rows gain "group": "Voice", which is additive. docs/CONFIG.md documents voice input.
  • Separate commit. The pre-commit hook clears GIT_DIR, GIT_INDEX_FILE, and related vars. Committing from a worktree let environment_note_snapshots_git_state commit into the real repository and set core.bare=true.

New dependency: cpal

Microphone capture needs a native audio API. cpal is only a target dependency on macOS and Windows, where CoreAudio and WASAPI need no system libraries. On Linux it would pull in ALSA headers and break the static musl and Android release builds, so Linux reports that voice input is not available yet.

Context and limits

Nothing is added to model-visible context: a transcript lands in the composer and is sent like typed text. Recordings stop after 5 minutes (MAX_RECORDING), and the sample buffer is capped to match.

Testing

  • make check and make desktop-check pass. VS Code typechecks, and vitest passes (325 tests), including new src/dictation.test.ts.
  • aster dictate ran on macOS with the real microphone:
    • With no key, it prints the key message.
    • With a key, it refuses a 0.1 second clip before uploading.
    • With a spoken sentence, it recorded and uploaded to OpenAI. The account has no credits, so the reply was a 429 rather than a transcript.
  • The VS Code composer ran in the dev host against the real CLI: click, listening, click, transcribing, then the error, with button states and labels checked at each step.
  • Not verified: a successful transcript from either provider, and the desktop button inside the built app.

Found, not fixed

desktop/src-tauri/src/lib.rs: the chat command takes the child's stdin into a local and never stores it, so answer_approval never reaches the CLI.

🤖 Generated with Claude Code

Git exports GIT_DIR and GIT_INDEX_FILE to hooks. environment_note_snapshots_git_state runs git init and git commit in a temp dir, and with those set from a worktree it committed into the real repository and set core.bare=true in the shared config.
Adds the aster-voice crate: default microphone capture through cpal on macOS and Windows, WAV encoding, and ElevenLabs Scribe and OpenAI transcription. The chat TUI toggles listening with ctrl+r and inserts the transcript into the composer.
Adds aster dictate, which records until stdin gets a line and prints NDJSON events. The desktop app and the VS Code extension run it from a new composer mic button. The desktop bundle declares microphone use.
The plain sentence becomes the notification and the provider detail goes to the Aster output channel. The dev host relays dictation so the mic button works there too.
@Zfinix Zfinix changed the title feat(voice): dictate into the chat composer with ctrl+r feat(voice): dictation in the tui, desktop, and vs code Sep 29, 2026
@Zfinix
Zfinix merged commit 131e081 into main Sep 29, 2026
1 check passed
@Zfinix
Zfinix deleted the feat/voice-dictation branch September 29, 2026 14:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant