mirror of
https://github.com/SerTimBerrners-Lee/talkis.git
synced 2026-10-03 04:36:55 +00:00
release: prepare v0.3.13
This commit is contained in:
parent
957f7c4f7f
commit
49a59ca573
79 changed files with 11648 additions and 2051 deletions
|
|
@ -46,6 +46,82 @@ Widget call mode
|
|||
-> file transcription pipeline
|
||||
```
|
||||
|
||||
Live translation:
|
||||
|
||||
```text
|
||||
Widget live-translation button
|
||||
-> live_translation.rs
|
||||
-> call_capture.rs system-audio stream plus optional cpal microphone stream
|
||||
-> mono / 16 kHz / PCM16 / 100 ms chunks
|
||||
-> bounded channels
|
||||
-> Talkis Cloud (short-lived OpenAI secret) or configured OpenAI/Gemini API
|
||||
-> separate realtime sessions per enabled audio source
|
||||
-> normalized partial/final text events
|
||||
-> existing widget-text overlay and history
|
||||
```
|
||||
|
||||
## Live Streaming And Translation
|
||||
|
||||
Realtime audio stays in Rust. Never route PCM through React/Tauri events and do
|
||||
not use ffmpeg in this hot path. System audio is always translated; microphone
|
||||
audio is opt-in to avoid sending speaker echo as the user's voice. When enabled,
|
||||
microphone and system audio are independent channels with separate translation
|
||||
connections. Audio callbacks use bounded
|
||||
non-blocking queues and may drop old audio instead of blocking capture.
|
||||
|
||||
Remote streaming STT is supported only for adapters whose exact key, model, and
|
||||
endpoint configuration passed a realtime handshake. A transport failure must not
|
||||
stop ordinary recording: the existing batch transcription remains the stop-time
|
||||
fallback. Local Whisper and WhisperX are not realtime translation adapters.
|
||||
In Cloud mode the desktop authenticates to `talkis-proxy` with its device token;
|
||||
the proxy checks the subscription and returns a short-lived OpenAI Realtime
|
||||
client secret. The primary cloud API key must never be returned to or stored by
|
||||
the desktop. API mode continues to use the user's verified adapter directly.
|
||||
|
||||
Live translation reconnects after 0.5/1/2 seconds and replays no more than two
|
||||
seconds of recent PCM. Gemini `goAway` events rotate the session. Provider audio
|
||||
responses are ignored unless synchronous voice playback is explicitly enabled.
|
||||
OpenAI translation uses the official `gpt-realtime`/`gpt-realtime-mini` session
|
||||
models with text-only output by default. Voice playback switches the same session
|
||||
to audio output and uses `response.output_audio_transcript.*` for the visible
|
||||
translated text, so speech and overlay do not come from separate translation
|
||||
requests. PCM16/24 kHz audio deltas stay in Rust and feed a bounded native
|
||||
playback queue; stale queued speech is discarded instead of increasing live
|
||||
latency without limit. Voice playback is currently enabled only on macOS, where
|
||||
the global Core Audio tap excludes the Talkis process from capture. Other
|
||||
platforms must not enable playback until their system-audio path can provide
|
||||
equivalent self-process exclusion. When original-audio ducking is enabled, the
|
||||
macOS tap uses `MutedWhenTapped` so the source remains available to Talkis at
|
||||
full capture level without also playing at full volume. The captured 16 kHz
|
||||
monitor feed is replayed through the self-excluded Talkis playback stream at low
|
||||
gain and mixed with the translated voice, so the original remains quietly
|
||||
audible without entering the translation loop;
|
||||
legacy `gpt-realtime-translate` settings are resolved to `gpt-realtime` at the
|
||||
runtime boundary. Continuous video speech must not be treated as one unbounded
|
||||
turn. Talkis disables server-side turn detection for this path and commits voiced
|
||||
PCM in short, roughly two-second segments. Only one text response may be active
|
||||
at a time; audio received while a response is active stays in the next input
|
||||
buffer and is committed as soon as the
|
||||
previous response finishes. Session shutdown commits the final voiced fragment
|
||||
and waits briefly for its response instead of dropping it.
|
||||
Provider commit boundaries are transport details, not speaker boundaries. The
|
||||
frontend coalesces consecutive final and partial segments from the same channel
|
||||
into one visible turn; a new turn starts when the channel changes between the
|
||||
system audio and the optional microphone.
|
||||
OpenAI text deltas are accumulated by `response_id` without trimming or adding
|
||||
synthetic spaces. A final event promotes only that response's draft; late deltas
|
||||
for a finalized response are ignored. The frontend keeps finalized text stable,
|
||||
renders only the active draft with reduced emphasis, coalesces rapid partial
|
||||
updates to at most one render per 100 ms, and serializes overlay commands so an
|
||||
older render cannot overwrite a newer one.
|
||||
API keys stay in the Rust request path and must never appear in events or logs.
|
||||
|
||||
When `saveRecordingAudio` is enabled, `system.wav` is written and `mic.wav` is
|
||||
added only when live microphone translation is enabled. When saving is disabled,
|
||||
no audio file is created and the
|
||||
whole session must not be retained in memory. Live translation, call capture, and
|
||||
ordinary dictation mutually exclude each other.
|
||||
|
||||
## Voice Dictation
|
||||
|
||||
The primary voice path is native Rust capture:
|
||||
|
|
@ -181,7 +257,9 @@ Call capture has two different tracks:
|
|||
|
||||
Current system-audio support:
|
||||
|
||||
- macOS: implemented via Core Audio process tap / aggregate device in `call_capture.rs`;
|
||||
- macOS: implemented via a stereo global Core Audio tap / aggregate device in
|
||||
`call_capture.rs`, so output is captured even when an app routes audio through
|
||||
a non-default device stream;
|
||||
- Windows: implemented via `cpal` WASAPI loopback on the default output device in `call_capture.rs`;
|
||||
- Linux: implemented via PipeWire default output monitor capture in `call_capture.rs`.
|
||||
|
||||
|
|
@ -197,6 +275,20 @@ System audio capture level: max=... dBFS, frames_above_noise_floor=...
|
|||
|
||||
If `max=-120.0 dBFS` and `frames_above_noise_floor=0`, the system track is silent. The transcript should not treat that as usable remote-speaker audio.
|
||||
|
||||
## macOS Development Permissions
|
||||
|
||||
Run the macOS dev binary from `Talkis Dev.app` with a stable Apple Development
|
||||
code-signing identity. An ad-hoc signature uses a designated requirement tied
|
||||
to the binary's changing CDHash, so rebuilding makes TCC treat the next process
|
||||
as a different application even when the old checkbox remains visible.
|
||||
|
||||
Microphone status should be read from AVFoundation without starting an audio
|
||||
session. System-audio access must first be confirmed with the existing short
|
||||
Core Audio capture probe. Only a successful probe may write the versioned
|
||||
verification marker used after relaunch with the stable signing identity;
|
||||
legacy or unversioned flags must be ignored, and a failed probe clears the
|
||||
current marker.
|
||||
|
||||
## Speaker Diarization
|
||||
|
||||
Local speaker diarization uses:
|
||||
|
|
|
|||
79
docs/release/review-v0.3.13.md
Normal file
79
docs/release/review-v0.3.13.md
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
# Release Review v0.3.13
|
||||
|
||||
## Release
|
||||
|
||||
- Version: 0.3.13
|
||||
- Release branch: release/v0.3.13
|
||||
- Target tag: v0.3.13
|
||||
- Reviewer: Codex
|
||||
- Date: 2026-07-15
|
||||
|
||||
## Scope
|
||||
|
||||
- Key changes included in this release:
|
||||
- Adds live system-audio translation through Talkis Cloud or verified OpenAI/Gemini Realtime API adapters, with optional translated voice playback on macOS.
|
||||
- Adds realtime dictation for supported local and API STT models with live partial text and a final batch fallback.
|
||||
- Reworks dictation and selected-text hotkeys into one transactional capture and runtime registration path with conflict checks, rollback, and stale-request protection.
|
||||
- Adds macOS, Windows, and Linux system-audio capture for call transcription and stores live-translation sessions and audio tracks in local history.
|
||||
- Improves the floating widget, text overlay, permission recovery, history action menus, settings persistence, and macOS development app runner.
|
||||
- Adds the authenticated Talkis Cloud Realtime client-secret endpoint in `talkis-proxy` v0.1.5 without exposing the provider API key to the desktop app.
|
||||
- User-facing changes:
|
||||
- Live translation starts from the widget and displays compact translated speaker turns while audio is still playing.
|
||||
- Selected text can be translated with a separate configurable global shortcut.
|
||||
- Streaming text replaces stale overlay content immediately; completed overlays can be dragged and close automatically after ten seconds.
|
||||
- Live-translation history keeps its saved audio track available after the full entry loads.
|
||||
- Risky areas:
|
||||
- Native system-audio capture differs across Core Audio, WASAPI loopback, and PipeWire.
|
||||
- Realtime provider protocols, reconnect/commit timing, optional voice playback, and feedback prevention.
|
||||
- Runtime replacement of two global hotkeys and settings rollback.
|
||||
- macOS TCC permission persistence and signed development app relaunch.
|
||||
- Cross-platform release bundles must pass GitHub Release Preflight on the exact release commit.
|
||||
|
||||
## Checks run
|
||||
|
||||
- `git diff --check`: passed.
|
||||
- Secret-pattern scan of changed source and scripts: passed; no credentials found.
|
||||
- `node --check scripts/run-tauri.mjs` and `zsh -n scripts/run-macos-dev-app.sh`: passed.
|
||||
- `bun test`: passed, 193 tests.
|
||||
- `cargo test --manifest-path src-tauri/Cargo.toml --lib`: passed, 89 tests.
|
||||
- `bun run check:release`: passed.
|
||||
- Version sync passed for `package.json`, `src-tauri/Cargo.toml`, and `src-tauri/tauri.conf.json` at 0.3.13.
|
||||
- Sidecar preparation, `bunx tsc --noEmit`, `cargo check`, hotkey smoke tests, and Vite production build passed.
|
||||
- `TAURI_SIGNING_PRIVATE_KEY_PATH=/Users/trixter/.tauri/talkis-updater.key bun run build:release:macos`: passed.
|
||||
- Built release sidecars, frontend, Rust binary, `Talkis.app`, updater `.app.tar.gz`, updater `.sig`, and `Talkis_0.3.13_aarch64.dmg`.
|
||||
- Talkis Cloud proxy: `go test ./...` passed; commit `1f366a7` was pushed to `main` and tagged `v0.1.5`.
|
||||
- Deploy run `29450364496` failed at `Configure SSH` because `ssh-keyscan` could not reach the configured `VDS_HOST` from the GitHub runner; no code was transferred and the existing production proxy stayed running.
|
||||
- GitHub Release Preflight: pending for `Preflight macos`, `Preflight windows`, and `Preflight linux` on the final release commit.
|
||||
- Native/GitHub Windows build: pending Release Preflight.
|
||||
- Native/GitHub Linux build: pending Release Preflight.
|
||||
- Additional manual checks:
|
||||
- The user confirmed live translation streaming behavior during development.
|
||||
- A final short smoke test is required for the latest hotkey transaction, Cloud deployment, history audio retention, and restart/permission fixes before `main` merge and tag publish.
|
||||
|
||||
## Manual review
|
||||
|
||||
- Hotkey flow: capture state machine, physical-key normalization, conflict, registration, persistence, rollback, stale request, and restart behavior are covered by frontend and Rust tests; final OS-level smoke test pending.
|
||||
- Onboarding permissions: native microphone/accessibility checks and signed dev-app relaunch path reviewed; final restart smoke test pending.
|
||||
- Widget position and notice behavior: edge padding, wider streaming layout, immediate replacement, dragging, terminal auto-dismiss, dark theme, and bottom-edge menu positioning reviewed and covered by focused tests where practical.
|
||||
- Transcription quality and short-utterance handling: existing filters remain in place; realtime commit, reconnect, replay, partial/final merge, and batch fallback tests pass.
|
||||
- README refreshed: yes, `README.md` and `README.ru.md` document v0.3.13 behavior and supported platforms.
|
||||
|
||||
## Findings
|
||||
|
||||
- Blockers:
|
||||
- Restore proxy deployment access and deploy `talkis-proxy` v0.1.5 before publishing the desktop tag; otherwise Cloud live translation cannot obtain a Realtime client secret.
|
||||
- Do not merge or tag until the final manual smoke test is confirmed.
|
||||
- Do not merge or tag until Release Preflight is green for macOS, Windows, and Linux on the exact release commit.
|
||||
- Non-blocking issues:
|
||||
- Local Node.js is 22.6.0 while Vite recommends 20.19+ or 22.12+; both production frontend builds still completed successfully.
|
||||
- Vite reports the existing large main chunk and mixed static/dynamic `SummaryModal.tsx` import warnings.
|
||||
- The local Beads Dolt database is unavailable because its journal is corrupted, so release tracking could not be updated through `bd`.
|
||||
- Follow-ups after release:
|
||||
- Consider migrating OpenAI live translation to the dedicated Realtime translation endpoint after the current generic Realtime flow has shipped and can be regression-tested separately.
|
||||
- Optimize frontend chunk splitting without coupling it to this functional release.
|
||||
|
||||
## Decision
|
||||
|
||||
- Ready for `main` merge: yes after the manual smoke test and final Release Preflight pass.
|
||||
- Release preflight green on exact tag commit: pending.
|
||||
- Ready for tag publish: yes after the same two gates pass.
|
||||
Loading…
Add table
Add a link
Reference in a new issue