openclaw/extensions/openai/realtime-talk-defaults.live.test.ts
Peter Steinberger f093c0edde
feat(openai): support GPT-Live and default Talk by account (#145382)
* feat(openai): support public GPT-Live-1 voice sessions

Adapt public Live sessions, startup events, delegation, audio, captions, and
Platform authentication through the existing OpenAI voice plugin. Keep the
Codex subscription transport separate.

Persist public WebRTC transcripts through the Gateway and drain provider
finalization before releasing voice session owners across relay, Discord,
Voice Call, and MeetingBot. Preserve synchronous bridge disposal.

Validate targeted protocol and lifecycle regressions, authenticated synthetic
voice and WebRTC flows, and the iOS Simulator app build.

Closes #145071

* fix(voice): preserve cleanup and package boundaries

Keep the OpenAI capability catalog cold, remove the delegation type cycle, and expose portable Google mock declarations. Route failed call startup through the existing binding cleanup owner and stop failed local audio processes while provider finalization drains.

* fix(ui): preserve public Live caption fragments

Carry explicit verbatim semantics through the browser transcript pipeline so split words, repeated fragments, whitespace, and overlapping speakers bypass legacy ASR heuristics. Keep Gateway-only persistence and avoid synthetic final or item events.

* test(openai): expose portable delegation mock types

Use public logger and gateway callback contracts in test helpers so declaration-enabled plugin package compilation does not reference private Vitest types. Verified declaration compilation for all six affected plugins and test callers.

* fix(openai): wait for delayed GPT-Live delegation transcripts

Retain metadata-only public delegation notices while user captions are empty, resume through the existing admission owners when text arrives, and claim notice IDs before callbacks to prevent duplicate work. Bound pending notices and missing-input waits; revoke them during close, cancellation, and transcript drain. Preserve subscription prompt fallback behavior.

Validation: 103 focused tests pass; new regressions failed against the original behavior. Independent P0-P2 autoreview is scoped-clean. Combined type/lint gates are owned by the landing checkout because declaration boundaries reject this worker worktree borrowed compiler install.

* feat(openai): select GPT-Live Talk defaults by account

Resolve unpinned Talk models with the selected agent account and session requirements. Platform credentials select public GPT-Live; ChatGPT-only accounts select the subscription voice model. Preserve explicit model pins, manual responses, video, Azure, and direct tool bridge defaults. Keep catalog discovery aligned with session creation without rewriting saved config.

Validation: 177 focused tests pass; scope and account regressions fail on the prior owners. Authenticated microphone-to-delegation-to-spoken-answer proof passes with 211200 audio bytes and awaited completed shutdown. Core and plugin production/test typechecks pass; independent review through P2 is scoped-clean. Full changed-file guards continue in the landing workflow.

* test(openai): isolate Talk account default coverage

Keep the routing suite below its existing line limit by reusing its fixtures in a focused defaults suite with injected host auth. Use explicit blocks in the live audio fixture. Production behavior is unchanged.

Validation: all 61 routing/default tests, extension test typecheck and typed lint pass; independent P0-P2 review is scoped-clean.

* refactor(openai): split realtime delegation dispatch and tests

Keep direct bridge admission and dispatch in a focused local owner, reuse the existing bridge fixture across a dedicated delayed-delegation suite, and apply the required block style. Preserve lifecycle checks, callback binding, transcript publication and subscription behavior without relaxing file budgets.

Validation: 103 focused tests pass across five files; independent P0-P2 autoreview is scoped-clean. The landing lane owns canonical typed validation of the combined candidate with physical dependencies.

* test(openai): expose portable bridge mock callback types

Use public Mock annotations tied to the realtime callback and logger contracts so exported bridge fixtures emit declarations without Vitest private Procedure types.

Validation: actual OpenAI declaration emission and extension-test typecheck pass after reproducing TS2883 before the fix; 40 helper-consumer tests pass; independent P0-P2 autoreview scoped-clean.

* fix(openai): preserve camera-capable Talk defaults

* fix(talk): align browser capabilities with launch models

Resolve optional provider and model overrides through the existing Talk catalog before browser camera negotiation. Preserve other provider rows and explicit model choices, while unpinned OpenAI Talk follows the requested GPT Live account defaults.

* fix(ui): avoid shadowing Talk provider selections
2026-09-11 20:42:25 -07:00

105 lines
3.5 KiB
TypeScript

import { resolveConfiguredRealtimeVoiceProvider } from "openclaw/plugin-sdk/realtime-voice";
import { describe, expect, it } from "vitest";
import { buildOpenAIRealtimeVoiceProvider } from "./realtime-voice-provider.js";
import { buildOpenAISpeechProvider } from "./speech-provider.js";
const live = process.env.OPENCLAW_LIVE_TEST === "1" && process.env.OPENCLAW_LIVE_GPT_LIVE === "1";
describe.skipIf(!live)("OpenAI Talk account defaults live", () => {
it("delegates microphone speech and speaks the backend answer without a configured model", async ({
skip,
}) => {
const apiKey = process.env.OPENAI_API_KEY?.trim();
if (!apiKey) {
skip("OpenAI Platform API key is unavailable");
return;
}
const cfg = {};
const { provider, providerConfig } = resolveConfiguredRealtimeVoiceProvider({
cfg,
surface: "gateway-relay",
providers: [buildOpenAIRealtimeVoiceProvider()],
providerConfigs: { openai: { apiKey } },
});
expect(providerConfig.model).toBe("gpt-live-1");
const speech = await buildOpenAISpeechProvider().synthesizeTelephony?.({
cfg,
providerConfig: { apiKey, model: "gpt-4o-mini-tts", voice: "alloy" },
text: "Please ask the backend for the launch code.",
timeoutMs: 30_000,
});
expect(speech?.sampleRate).toBe(24_000);
if (!speech) {
throw new Error("Speech fixture was not synthesized");
}
const errors: Error[] = [];
const questions: string[] = [];
let assistantText = "";
let audioBytes = 0;
let closed: string | undefined;
const bridge = provider.createBridge({
cfg,
providerConfig,
audioFormat: { encoding: "pcm16", sampleRateHz: 24_000, channels: 1 },
instructions: "Delegate every user request to the backend. Speak the backend answer exactly.",
runAgentConsult: async ({ prompt }) => {
questions.push(prompt);
return { text: "The launch code is saffron lantern." };
},
onAudio: (audio) => {
audioBytes += audio.length;
},
onClearAudio: () => undefined,
onTranscript: (role, text) => {
if (role === "assistant") {
assistantText += text;
}
},
onError: (error) => errors.push(error),
onClose: (reason) => {
closed = reason;
},
});
try {
await bridge.connect();
const deadline = Date.now() + 40_000;
let offset = 0;
while (Date.now() < deadline && !assistantText.toLowerCase().includes("saffron lantern")) {
const frame = Buffer.alloc(960);
speech.audioBuffer.copy(
frame,
0,
offset,
Math.min(offset + frame.length, speech.audioBuffer.length),
);
offset = Math.min(offset + frame.length, speech.audioBuffer.length);
bridge.sendAudio(frame);
await new Promise((resolve) => {
setTimeout(resolve, 20);
});
if (errors.length) {
break;
}
}
expect(errors).toEqual([]);
expect(questions.some((question) => /launch code/i.test(question))).toBe(true);
expect(assistantText.toLowerCase()).toContain("saffron lantern");
expect(audioBytes).toBeGreaterThan(0);
} finally {
await bridge.close();
}
expect(errors).toEqual([]);
expect(closed).toBe("completed");
expect(bridge.isConnected()).toBe(false);
console.info(
JSON.stringify({
model: providerConfig.model,
delegations: questions.length,
audioBytes,
closed,
}),
);
}, 90_000);
});