Sinua
Connect your appVoice

Gemini Live

Follow a Gemini Live voice session.

The Gemini source streams audio to Gemini Live and plays the model's reply, reporting state and audio to the visual. Long conversations cross Gemini's connection limit transparently through session resumption.

gemini-live-web.ts
import { GeminiLiveVoiceSource } from "@sinua/voice/gemini";

// Your backend mints an ephemeral token (`auth_tokens/…`) with the Gemini API key,
// which never reaches the browser.
export async function geminiVoice() {
  const { token } = await (await fetch("/api/gemini-live-token", { method: "POST" })).json();
  return new GeminiLiveVoiceSource({ credential: token, instructions: "Keep answers short." });
}

Credentials

Your server mints an ephemeral auth token (auth_tokens/…) with your API key. Lock the model and setup on the server side so the client can't change them, and hand the token to the source.

Packages

PlatformPackage
WebGeminiLiveVoiceSource from @sinua/voice/gemini
iOSSinuaGeminiLive
Android:sinua-gemini

Good to know

  • Ephemeral token in production. With an auth_tokens/… token, the source uses Gemini's constrained endpoint automatically. A raw API key also works, but only for local development: never ship one.
  • Reconnects are handled. When Gemini announces the end of a connection (goAway) or the socket closes, the source resumes the session on a new connection.
  • Start from a user gesture on the web, since the browser only allows the microphone and audio playback after one. On iOS, add NSMicrophoneUsageDescription; on Android, declare RECORD_AUDIO. Say where the audio goes: the user's voice is sent to Google, so your iOS purpose string (and, on Android, the explanation you show before the system prompt, which carries no app text) should say so. For example: "Your voice is sent to Google so the assistant can hear you." App Store review expects the purpose string to explain the use.

On this page