v1.0.0
AsyncAPI 3.0.0

MyLab TTS Streaming API

Synthesize speech in a registered voice and receive the audio as it is generated, over a single WebSocket connection.

Overview

A WebSocket connection is a session, intended to span an entire call: open it when the call starts, send one synthesize request per agent turn, and close it when the call ends. Inactivity does not close a session.

Prerequisites: an API key issued to your account (a secret; store it only on your server) and a voice ID (voiceProfileId) of a voice registered on the MyLab TTS Voices page.

Authentication

Authentication is a bearer API key, sent as shown under Connect. API keys do not expire; contact MyLab to revoke one.

Connect

wss://api-lab.myorder.ai/api/tts/stream?voiceProfileId=<voice ID>
Parameter Description
voiceProfileId (query, required) The voice ID (UUID) to synthesize with.
Authorization (header, required) Bearer <api key>. Browsers cannot set request headers on a WebSocket; browser clients send the key as the subprotocol bearer.<api key>, which the server echoes in the handshake response.

The request is validated before the WebSocket upgrade. A rejection is an HTTP response with the body {"error": {"code": "...", "message": "..."}}; no WebSocket is opened. WebSocket libraries surface it as a connection failure that includes the HTTP status (ws: the unexpected-response event; websockets: InvalidStatus). Browser clients receive only close code 1006.

HTTP status error.code Description
401 INVALID_API_KEY The API key is missing or invalid.
404 VOICE_NOT_REGISTERED The voice is not registered for TTS. Register it on the TTS Voices page.
422 VALIDATION voiceProfileId is missing or not a UUID.
426 UPGRADE_REQUIRED The request is not a WebSocket upgrade.
429 TOO_MANY_CONNECTIONS The account's open-session limit, or the service's capacity, has been reached.
500 INTERNAL_ERROR Server error.
503 TTS_UNAVAILABLE The TTS service is unavailable or did not answer in time.

Session lifecycle

Session lifecycle: connect, session-ready, then per turn synthesize → audio × N → done or error; reconnect-soon 60 s before the 4-hour close 1001

The server sends no close reason; the close code identifies the cause:

Close code Description
1000 Normal closure, or no pong received within 60 seconds.
1001 The session reached its 4-hour maximum. Open a new session.
1011 The TTS service became unavailable.

Reconnect using exponential backoff with jitter, from 500 ms up to 30 seconds. Apply the same backoff to 429, 500 and 503 responses on connection.

Audio format

audio frames are binary WebSocket frames of raw PCM, 16-bit signed little-endian, mono, 48 kHz: no header, no container; 1 ms of audio is 96 bytes. Frames are delivered in order and vary in size.

Limits

Limit Value
text per request 1–384 Unicode characters
styleHint per request up to 128 Unicode characters
Requests in flight per session 1
Open sessions per account 100
Session length 4 hours

Limits are measured in Unicode characters; each Thai vowel and tone mark counts as one. Send one or two sentences per request and split longer text at sentence boundaries. Open several sessions for parallel synthesis.

Billing

Usage is metered per request on the audio generated, in whole seconds.

Quickstart

JavaScript (Node.js, ws):

import WebSocket from 'ws';

const ws = new WebSocket(
  `wss://api-lab.myorder.ai/api/tts/stream?voiceProfileId=${VOICE_ID}`,
  { headers: { Authorization: `Bearer ${API_KEY}` } },
);
const pcm = [];

ws.on('message', (data, isBinary) => {
  if (isBinary) return pcm.push(data);
  const frame = JSON.parse(data);
  if (frame.type === 'session-ready') {
    ws.send(JSON.stringify({ type: 'synthesize', text: 'Hello, thank you for calling.' }));
  } else if (frame.type === 'done') {
    ws.close(1000);
  }
});

Python (websockets 14 or later):

import asyncio, json, websockets

async def main():
    url = f"wss://api-lab.myorder.ai/api/tts/stream?voiceProfileId={VOICE_ID}"
    async with websockets.connect(url, additional_headers={"Authorization": f"Bearer {API_KEY}"}) as ws:
        pcm = bytearray()
        async for message in ws:
            if isinstance(message, bytes):
                pcm += message
                continue
            frame = json.loads(message)
            if frame["type"] == "session-ready":
                await ws.send(json.dumps({"type": "synthesize", "text": "Hello, thank you for calling."}))
            elif frame["type"] == "done":
                break

asyncio.run(main())
Server:wss://api-lab.myorder.ai

TTS session

​
Servers:
production
Protocols:
wss

One WebSocket connection is one session. See Connect and Session lifecycle.

receive

Session ready

​

First message after the upgrade; carries the session ID

Session ready

Protocols:
wss
​
send

Synthesize

​

Synthesis request; answered by audio frames and done, or by error

Synthesize request

Protocols:
wss
​
receive

Audio

​

Audio chunk, streamed in order until done

Audio chunk

Protocols:
wss
​
receive

Done

​

End of a successful request

Done

Protocols:
wss
​
receive

Error

​

End of a failed request; the session stays open

Error

Protocols:
wss
​
receive

Reconnect soon

​

Notice sent 60 seconds before the 4-hour maximum

Reconnect soon

Protocols:
wss
​

Models