Documentation menu

Calls

Joining a call from your system

Play one side of a call from your own media server, contact-center platform or bot, over two WebSockets.

A browser isn't the only way to take part in a call. A media server, a contact-center platform or a bot can play either side over two WebSockets: one to send that side's voice, one to hear the other side interpreted. Spiiksi's own phone bridges work this way.

Get the sockets#

Use the key from the side's link (the part after #k=):

Shell
curl https://api.spiiksi.fi/calls/public/$CALL_ID -H "X-Call-Key: $KEY"
JSON
{"id": "6f0c…", "role": "customer", "status": "open", "ready": true,
 "my_language": "es", "other_language": "fi",
 "stream": {"path": "/ws/stream/…", "channel_id": "…", "source_id": "…", "token": "…"},
 "listen": {"path": "/ws/listen/…", "channel_id": "…", "unlock": "…"}, …}

stream and listen are given once both sides' languages are known ("ready": true). If the customer chooses their language on joining, poll until then, or set it yourself with PATCH /calls/public/{id}/me and {"language": "es"}.

Send this side's voice#

Open wss://api.spiiksi.fi{stream.path}?token={stream.token}&source_id={stream.source_id} and send binary frames of mono, 16-bit little-endian PCM at 24 kHz. About 4,096 samples (170 ms) a frame works well.

  • Keep sending through pauses, with silence: the interpreter keeps its sense of time.
  • Text frames come back as JSON notices: starting (the interpreter is warming up), live, paused (quiet for a minute; it resumes on the next sound).

Hear the other side#

Open wss://api.spiiksi.fi{listen.path}?unlock={listen.unlock}. Binary frames are WAV chunks (24 kHz mono PCM16, each with its own header): play them in order, one after another. Text frames are JSON: captions of the interpretation (caption, caption_delta, caption_done) and a provenance notice marking the audio as AI-generated.

Keep the interpretation out of your microphone#

If what you play can reach what you send (a speaker near a microphone, a phone line's echo), send silence while it plays, plus a short tail of about 300 ms. Otherwise the interpretation is interpreted again and sent back.

Example: a minimal client in Python#

Python
import asyncio, json, wave
import httpx, websockets

API, WS = "https://api.spiiksi.fi", "wss://api.spiiksi.fi"

async def play_side(call_id: str, key: str, wav_in: str):
    info = httpx.get(f"{API}/calls/public/{call_id}", headers={"X-Call-Key": key}).json()
    assert info["ready"], "wait until both languages are known"
    s, l = info["stream"], info["listen"]
    stream_url = f"{WS}{s['path']}?token={s['token']}&source_id={s['source_id']}"
    listen_url = f"{WS}{l['path']}?unlock={l['unlock']}"

    async with websockets.connect(stream_url) as stream, websockets.connect(listen_url) as listen:
        async def hear():
            async for message in listen:
                if isinstance(message, bytes):
                    pass  # a WAV chunk: play it, after the previous one
                else:
                    print(json.loads(message))  # captions

        hearing = asyncio.create_task(hear())
        with wave.open(wav_in) as audio:  # mono, 16-bit, 24 kHz
            while frames := audio.readframes(4096):
                await stream.send(frames)
                await asyncio.sleep(4096 / 24000)  # real time
        hearing.cancel()

Phone audio#

Telephone lines carry 8 kHz audio (μ-law on most trunks). Resample to 24 kHz 16-bit PCM before sending, and back to 8 kHz before playing on the line. For Twilio and Asterisk, Spiiksi does this for you: see Phone lines and SIP with AudioSocket.