Calls
Joining a call from your system
Play one side of a call from your own media server, contact-center platform or bot, over two WebSockets.
A browser isn't the only way to take part in a call. A media server, a contact-center platform or a bot can play either side over two WebSockets: one to send that side's voice, one to hear the other side interpreted. Spiiksi's own phone bridges work this way.
Get the sockets#
Use the key from the side's link (the part after #k=):
curl https://api.spiiksi.fi/calls/public/$CALL_ID -H "X-Call-Key: $KEY"{"id": "6f0c…", "role": "customer", "status": "open", "ready": true,
"my_language": "es", "other_language": "fi",
"stream": {"path": "/ws/stream/…", "channel_id": "…", "source_id": "…", "token": "…"},
"listen": {"path": "/ws/listen/…", "channel_id": "…", "unlock": "…"}, …}stream and listen are given once both sides' languages are known ("ready": true). If the customer chooses their language on joining, poll until then, or set it yourself with PATCH /calls/public/{id}/me and {"language": "es"}.
Send this side's voice#
Open wss://api.spiiksi.fi{stream.path}?token={stream.token}&source_id={stream.source_id} and send binary frames of mono, 16-bit little-endian PCM at 24 kHz. About 4,096 samples (170 ms) a frame works well.
- Keep sending through pauses, with silence: the interpreter keeps its sense of time.
- Text frames come back as JSON notices:
starting(the interpreter is warming up),live,paused(quiet for a minute; it resumes on the next sound).
Hear the other side#
Open wss://api.spiiksi.fi{listen.path}?unlock={listen.unlock}. Binary frames are WAV chunks (24 kHz mono PCM16, each with its own header): play them in order, one after another. Text frames are JSON: captions of the interpretation (caption, caption_delta, caption_done) and a provenance notice marking the audio as AI-generated.
Keep the interpretation out of your microphone#
If what you play can reach what you send (a speaker near a microphone, a phone line's echo), send silence while it plays, plus a short tail of about 300 ms. Otherwise the interpretation is interpreted again and sent back.
Example: a minimal client in Python#
import asyncio, json, wave
import httpx, websockets
API, WS = "https://api.spiiksi.fi", "wss://api.spiiksi.fi"
async def play_side(call_id: str, key: str, wav_in: str):
info = httpx.get(f"{API}/calls/public/{call_id}", headers={"X-Call-Key": key}).json()
assert info["ready"], "wait until both languages are known"
s, l = info["stream"], info["listen"]
stream_url = f"{WS}{s['path']}?token={s['token']}&source_id={s['source_id']}"
listen_url = f"{WS}{l['path']}?unlock={l['unlock']}"
async with websockets.connect(stream_url) as stream, websockets.connect(listen_url) as listen:
async def hear():
async for message in listen:
if isinstance(message, bytes):
pass # a WAV chunk: play it, after the previous one
else:
print(json.loads(message)) # captions
hearing = asyncio.create_task(hear())
with wave.open(wav_in) as audio: # mono, 16-bit, 24 kHz
while frames := audio.readframes(4096):
await stream.send(frames)
await asyncio.sleep(4096 / 24000) # real time
hearing.cancel()Phone audio#
Telephone lines carry 8 kHz audio (μ-law on most trunks). Resample to 24 kHz 16-bit PCM before sending, and back to 8 kHz before playing on the line. For Twilio and Asterisk, Spiiksi does this for you: see Phone lines and SIP with AudioSocket.