Sinua
Connect your appVoice

OpenAI Realtime

Follow an OpenAI Realtime voice session over WebRTC.

The OpenAI source opens a Realtime session over WebRTC with a short-lived client secret from your server, plays the agent's voice, and reports state and audio to the visual. When the user starts speaking over the agent, the visual flashes: the server cancels the response and the source reports the interruption.

openai-realtime-web.ts
import { OpenAIRealtimeVoiceSource } from "@sinua/voice/openai";

// Your backend mints a short-lived `ek_…` key (POST /v1/realtime/client_secrets with
// your API key, which never reaches the browser). An `ek_` is single-use, so give the
// source a function: it asks again on every reconnect.
export const voice = new OpenAIRealtimeVoiceSource({
  getCredential: async () => {
    const res = await fetch("/api/openai-realtime-key", { method: "POST" });
    return (await res.json()).key as string;
  },
  instructions: "You are a concise voice assistant.",
});

Credentials

Your server mints an ephemeral client secret (ek_…) with your API key and the session's configuration (model, voice, instructions), and returns only that value. It expires quickly, and the session's model, voice and instructions are set when your server mints it. Give the source a function that fetches a fresh one, so reconnects work.

Packages

PlatformPackage
WebOpenAIRealtimeVoiceSource from @sinua/voice/openai
iOSSinuaOpenAI
Android:sinua-openai

Good to know

  • One secret per connection. An ek_… secret works once. Pass a function that fetches a fresh one (getCredential on the web, credentialProvider on iOS and Android), not a fixed string, or a dropped call can't reconnect.
  • No second WebRTC on native. The iOS and Android sources use the WebRTC build that ships with LiveKit's SDK, so an app that has both doesn't carry two copies.
  • Start from a user gesture on the web, since the browser only allows the microphone and audio playback after one. On iOS, add NSMicrophoneUsageDescription; on Android, declare RECORD_AUDIO. Say where the audio goes: the user's voice is sent to OpenAI, so your iOS purpose string (and, on Android, the explanation you show before the system prompt, which carries no app text) should say so. For example: "Your voice is sent to OpenAI so the assistant can hear you." App Store review expects the purpose string to explain the use.

On this page