Zeli AvatarDeveloper docs

Guides / Your LiveKit room

Add a Zeli Avatar to your LiveKit room

Already running a voice agent on LiveKit? Give it a face. The Zeli Avatar joins your room as its own participant, lip syncs every word your agent says, and publishes the video and audio on your agent's behalf. Your agent, your room, your LiveKit project. We only bring the face.

How it works

 Your LiveKit room
 ┌──────────────────────────────────────────────────────────────────┐
 │                                                                  │
 │ Your agent       ── TTS audio (lk.audio_stream) ──►  zeli-avatar │
 │ (livekit-agents) ◄── lk.playback_finished ────────   (Zeli box)  │
 │                                                         │        │
 │ Users  ◄────────── lip synced video and audio ──────────┘        │
 │                                                                  │
 └──────────────────────────────────────────────────────────────────┘
      ▲ 1. token for zeli-avatar,          ▲ 3. joins the room
      │    signed with YOUR LiveKit secret │    with that token
      │                                    │
 Your worker ── 2. POST /livekit/join ──► Zeli Avatar API
  1. Your worker mints a LiveKit token for the avatar's participant identity, signed with your own LiveKit API secret. The secret never leaves your worker.
  2. It sends that token to POST /livekit/join, authenticated with your Zeli API key.
  3. The avatar box joins your room with the token and publishes its video.
  4. Your agent's audio output is pointed at the avatar (DataStreamAudioOutput), so its speech streams to the box instead of straight into the room. The box speaks it with a moving face, and reports back when each segment finishes playing, so interruptions stay in sync.

This is the same remote avatar protocol LiveKit's own avatar plugins use. If you have wired up an avatar plugin before, you already know the shape.

Prerequisites

  • A LiveKit project (LiveKit Cloud or self hosted), with its URL, API key and API secret.
  • A voice agent built on livekit-agents 1.5 (Python) or any agent that can send audio with the livekit-agents avatar data stream.
  • A Zeli API key from the portal: API keys. It starts zsk_live_. Short lived session tokens (zsk_temp_) are refused on these routes, because a join spends a GPU and takes a room token, so it belongs on a server.
  • Optionally an avatar id. Without one the box uses its active avatar.

SDK or API?

Python SDK (zeli.livekit)Raw API
Best forA livekit-agents worker in PythonNode agents, other frameworks, custom brokers
Token mintingDone for youYou mint it (any LiveKit server SDK)
Audio routingDone for youYou point your agent's audio at the avatar
Leave on shutdownAutomatic, on job shutdownYou call POST /livekit/leave
ErrorsTyped exceptions (AvatarBusyError, ...)JSON error codes
Tone signalavatar.send_tone("friendly")Send a text stream yourself

Pick the SDK if your agent is Python. Reach for the API when it is not, or when something other than the agent (a dispatcher, a broker) decides which box joins.

Quickstart with the Python SDK

Install the SDK with its LiveKit extra:

pip install "zeli-avatar[livekit]"

Set your credentials:

export ZELI_SERVER_URL=https://avatar.zeligate.ai
export ZELI_API_KEY=zsk_live_...
export LIVEKIT_URL=wss://your-project.livekit.cloud
export LIVEKIT_API_KEY=...
export LIVEKIT_API_SECRET=...

Then add the avatar to your worker. This is the complete, runnable example that ships with the SDK as examples/livekit_agent.py:

import logging
import os
 
from livekit.agents import Agent, AgentServer, AgentSession, JobContext, cli
 
from zeli import ZeliError
from zeli.livekit import AvatarBusyError, AvatarSession, AvatarUnavailableError
 
logger = logging.getLogger("zeli-avatar-agent")
 
server = AgentServer()
 
 
@server.rtc_session()
async def entrypoint(ctx: JobContext) -> None:
    session = AgentSession(
        stt="deepgram/nova-3",
        llm="openai/gpt-4.1-mini",
        tts="cartesia/sonic-2",
    )
 
    # Before session.start(): the avatar takes over the session's audio output.
    avatar = AvatarSession(
        os.environ.get("ZELI_AVATAR_ID") or None,
        api_key=os.environ.get("ZELI_API_KEY"),
    )
    try:
        await avatar.start(session, room=ctx.room)
        logger.info("Zeli Avatar joined, session %s", avatar.session_id)
    except (AvatarBusyError, AvatarUnavailableError) as err:
        # One avatar box serves one room at a time. Carry on audio only.
        logger.warning("Zeli Avatar unavailable (%s), continuing with voice only", err.code)
    except ZeliError as err:
        logger.error("Zeli Avatar could not join: %s", err)
 
    await session.start(
        agent=Agent(instructions="You are a friendly assistant. Keep answers short."),
        room=ctx.room,
    )
    await session.generate_reply(instructions="Greet the user and offer your help.")
 
 
if __name__ == "__main__":
    cli.run_app(server)

Run it with python livekit_agent.py dev, open your LiveKit playground, and say hello. The session uses LiveKit Inference model strings, which need LiveKit Cloud; any STT, LLM and TTS plugin works the same way.

Call start() before session.start()

avatar.start() replaces the session's audio output. Started after the session, the first words may already have gone to the room without a face. If a join is refused, the audio output is left exactly as it was, so your agent simply carries on as voice only.

AvatarSession options

avatar_idstr | NoneOptional

The avatar to show. None uses the box's active avatar.

api_keystrOptional

Your Zeli API key. Defaults to ZELI_API_KEY.

base_urlstrOptional

The Zeli Avatar API. Defaults to ZELI_SERVER_URL.

avatar_participant_identitystrOptional

The avatar's identity in your room. Default zeli-avatar. Change it if another participant already uses that identity: LiveKit keeps one participant per identity.

max_minutesfloatOptional

A ceiling on this session's length. The box lowers it to its own ceiling.

idle_minutesfloatOptional

End the session after this long idle. Lowered to the box's ceiling.

expected_minutesfloatOptional

How long you expect the call to run. A box that will shut down sooner refuses up front instead of dropping your call halfway.

join_timeoutfloatOptional

Seconds to wait for the join. Default 55, just under the 60 seconds the gateway in front of the box allows. Loading a face takes about 20.

start(agent_session, room, *, livekit_url=None, livekit_api_key=None, livekit_api_secret=None) reads LIVEKIT_URL, LIVEKIT_API_KEY and LIVEKIT_API_SECRET for anything you leave out. aclose() takes the avatar out of the room; inside a job it runs for you at shutdown.

Step by step with the API

Every route lives under your environment's host and takes your API key in X-Api-Key.

1. Mint a token for the avatar

A LiveKit access token for identity zeli-avatar, kind agent, allowed to join your room, with lk.publish_on_behalf naming your agent so clients show the avatar's tracks as the agent's.

from datetime import timedelta
from livekit import api
 
token = (
    api.AccessToken(api_key=LIVEKIT_API_KEY, api_secret=LIVEKIT_API_SECRET)
    .with_kind("agent")
    .with_identity("zeli-avatar")
    .with_name("Zeli Avatar")
    .with_ttl(timedelta(hours=2))
    .with_grants(api.VideoGrants(room_join=True, room=room_name))
    .with_attributes({"lk.publish_on_behalf": agent_identity})
    .to_jwt()
)

2. Ask the avatar to join

curl -X POST https://avatar.zeligate.ai/livekit/join \
  -H "X-Api-Key: $ZELI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "livekit_url": "wss://your-project.livekit.cloud",
    "livekit_token": "<the token from step 1>",
    "room_name": "support-room-1",
    "identity": "zeli-avatar",
    "engine_identity": "<your agent identity>",
    "avatar": "your-avatar-id",
    "usage": { "meeting_id": "your-call-id" }
  }'

The call returns once the avatar has joined and published its video, usually in about 20 seconds:

{
  "ok": true,
  "identity": "zeli-avatar",
  "room": "support-room-1",
  "avatar": "your-avatar-id",
  "silence": "your-avatar-id__silence",
  "idle_stand_in": false,
  "limits": { "max_minutes": 60.0, "idle_minutes": 5.0, "expected_minutes": 45.0, "clamped": [] },
  "session_id": "0192f3a4b5c6d9f1e2a3b4c5d"
}

Every field is in the API reference.

3. Route your agent's audio to the avatar

Point your agent's audio output at identity zeli-avatar with the livekit-agents avatar data stream. In Python that is one assignment:

from livekit import rtc
from livekit.agents.voice.avatar import DataStreamAudioOutput
 
session.output.audio = DataStreamAudioOutput(
    room=ctx.room,
    destination_identity="zeli-avatar",
    sample_rate=24000,
    wait_remote_track=rtc.TrackKind.KIND_VIDEO,
)

Not on Python? The same protocol ships in LiveKit's Node agents as DataStreamAudioOutput (agents-js source). The wire protocol itself is small: PCM audio on the byte stream topic lk.audio_stream addressed to the avatar, the lk.clear_buffer RPC to interrupt, and the lk.playback_started and lk.playback_finished RPCs back from the avatar (reference implementation).

4. Leave when the call ends

curl -X POST https://avatar.zeligate.ai/livekit/leave \
  -H "X-Api-Key: $ZELI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"room_name": "support-room-1"}'

Always name the room. A leave for a room the avatar is still joining cancels that join cleanly. Only a key from the account that joined the room (or the box operator) can end it; anyone else gets 403 livekit_not_owner. The same holds for POST /disconnect, POST /connect and POST /api/avatar/delete while your room is live, so nobody else's session can end it by a side door, and nobody else can drive it either: POST /api/chat answers 403 and the microphone socket closes with 4003.

Tone, session id and limits

Tone (optional). The avatar picks a facial tone for each reply from the audio by itself. If your agent already knows the mood, tell it: a text stream on topic zeli.avatar_tone, sent to the avatar just before the reply's audio.

{ "v": 1, "tone": "friendly", "pace": "normal", "speech_id": "optional-id" }

tone is one of neutral, friendly, calm, cheerful, serious, encouraging, thoughtful or empathetic (another word falls back to neutral). pace is slow, normal or fast. The most recent tone applies to the next reply only. With the SDK it is await avatar.send_tone("friendly"). The stream must come from your agent: the identity you sent as engine_identity, or the first agent participant in the room.

Session id and usage. Every join is metered: the box keeps one usage record per meeting, billed to the account that owns the API key, and answers the join with its session_id. Keep it with your call records. GET /livekit/status shows it to you while the call is live, and POST /livekit/leave echoes it. To tie the record to your own call, send "usage": {"meeting_id": "your-call-id"} with the join. The SDK exposes the id as avatar.session_id (None on an older box that does not report one).

Recording and audit log. Every meeting records the avatar's own video and the voice it speaks, and nothing else: nobody else in the room is recorded. It exists so you can check how the avatar looked and sounded, and only the account that owns the meeting can see it. The box also keeps an audit log of the meeting (when it joined, each utterance with its tone and first frame latency, barge ins, warnings, and the leave). Open the meeting on the Sessions page to play the recording and read the log, or download the log as JSON. Once the meeting has ended you can delete its recording and audit log there too (or with DELETE /auth/sessions/{session_id}/recording). The meeting record itself stays in your usage history.

Limits. A session ends on its own when it runs past its maximum (60 minutes by default) or sits idle (5 minutes by default). Ask for different values with max_minutes and idle_minutes; the box lowers them to its ceilings (120 and 20 minutes by default) and lists any it lowered in limits.clamped.

One room per avatar box

An avatar box renders one face for one room at a time. While it serves a call, every other join gets 409 livekit_session_in_use. Check GET /livekit/status first if you want to route around a busy box, and plan an audio only fallback.

Errors and fallbacks

A refusal from these routes is {"ok": false, "error": {"code", "message"}}. The key check in front of them answers {"code", "message"} instead. Either way, the SDK raises a typed exception carrying the same code.

StatusCodeWhat happenedWhat to doSDK error
400livekit_bad_requestA required field is missing, a limit is not a number greater than zero, or expected_minutes is over the ceiling or your max_minutesFix the requestConfigurationError
401missing_api_key, invalid_api_key, insufficient_scopeNo key, a bad key, or a session tokenUse a full API keyAuthenticationError
403livekit_not_ownerA leave, disconnect, connect or avatar delete while a room another account joined is liveUse your own account's keyAuthenticationError
409livekit_session_in_useThe box is serving another callRetry later, or go voice onlyAvatarBusyError
409livekit_join_abortedA leave for this room arrived mid joinNothing, it was cancelledAvatarBusyError
409livekit_not_in_roomA leave named a room the box is not inNothing was touchedAvatarBusyError
502livekit_join_failedThe box could not join your roomCheck the URL and the token grantsAvatarJoinError
502livekit_identity_mismatchThe token's identity is not the identity you sentMint the token for the same identityAvatarJoinError
503box_drainingThe box shuts down before your call would endTry another box, or a shorter expected_minutesAvatarUnavailableError
503livekit_unavailableThis box cannot join LiveKit roomsCheck readiness, contact supportAvatarUnavailableError
503livekit_owner_unverifiedThe box could not look your key up just then, so it ended nothingRetry the leaveAvatarUnavailableError
502, 503, 504none (a gateway page)The gateway in front of the box answered, not the boxSend a leave naming the room, then retryAvatarUnavailableError with code gateway_error

Timeouts. If your client gives up on a join, the box may still finish it. Send a leave naming the room: it cancels a join in flight and ends one that just landed. The SDK does this for you on anything other than the box's own 409 or 503 (a timeout, a dropped connection, a gateway page, a cancelled start), then raises the error. A leave that fails is retried once, and again on aclose().

Troubleshooting

Nothing joins. Ask the box whether it can join LiveKit rooms at all:

curl https://avatar.zeligate.ai/livekit/readiness -H "X-Api-Key: $ZELI_API_KEY"

installed: false comes with a problem saying why. identity and topic show the identity the box expects by default and the audio topic it listens on.

The face joins but never talks. Your agent's audio is not reaching the avatar. Check that destination_identity matches the avatar's identity, and that avatar.start() ran before session.start().

The face talks but looks neutral. That is the fallback when no tone is sent and the audio reads as neutral. Send a tone, or check the avatar has takes recorded for the tones you use.

409 on every join. Another call holds the box. GET /livekit/status answers free, busy or draining, and the room name if it is yours.

Zeli Avatar · real-time avatars over WebRTC · self-hostable · AU data residency · source