Update memory: document v1.5.0 mic chunk size fix
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -43,20 +43,30 @@
|
||||
|
||||
## Plugin Status
|
||||
- **PS_AI_Agent_ElevenLabs**: compiles cleanly on UE 5.5 Win64 (verified 2026-02-19)
|
||||
- v1.1.0 — all 3 protocol bugs fixed (transcript fields, pong format, client turn mode)
|
||||
- v1.5.0 — mic audio chunk size fixed: WASAPI 5ms callbacks accumulated to 100ms before sending
|
||||
- v1.4.0 — push-to-talk fully fixed: bAutoStartListening now ignored in Client turn mode
|
||||
- Binary WS frame handling implemented (ElevenLabs sends ALL frames as binary, not text)
|
||||
- First-byte discrimination: `{` = JSON control message, else = raw PCM audio
|
||||
- `SendTextMessage()` added to both WebSocketProxy and ConversationalAgentComponent
|
||||
- Connection confirmed working end-to-end; audio receive path functional
|
||||
- `conversation_initiation_client_data` now sent immediately on WS connect (required for mic + latency)
|
||||
|
||||
## Audio Chunk Size — CRITICAL
|
||||
- WASAPI fires mic callbacks every ~5ms → **158 bytes** at 16kHz 16-bit mono
|
||||
- ElevenLabs VAD/STT requires **≥3200 bytes (100ms)** per chunk; smaller chunks are silently ignored
|
||||
- Fix: `MicAccumulationBuffer` in component accumulates chunks; sends only when `>= MicChunkMinBytes` (3200)
|
||||
- `StopListening()` flushes remainder so final partial chunk is never dropped before end-of-turn
|
||||
|
||||
## ElevenLabs WebSocket Protocol Notes
|
||||
- **ALL frames are binary** — `OnRawMessage` handles everything; `OnMessage` (text) never fires
|
||||
- **ALL frames are binary** — bind ONLY `OnRawMessage`; NEVER bind `OnMessage` (text) — UE fires both for same frame → double audio bug
|
||||
- Binary frame discrimination: peek byte[0] → `'{'` (0x7B) = JSON, else = raw PCM audio
|
||||
- Fragment reassembly: accumulate into `BinaryFrameBuffer` until `BytesRemaining == 0`
|
||||
- Pong: `{"type":"pong","event_id":N}` — `event_id` is **top-level**, NOT nested
|
||||
- Transcript: type=`user_transcript`, key=`user_transcription_event`, field=`user_transcript`
|
||||
- Client turn mode: `{"type":"user_activity"}` to signal speaking; no explicit end message
|
||||
- Client turn mode (`client_vad`): send `user_activity` **with every audio chunk** (not just once) — server needs continuous signal to know user is speaking; stopping chunks = silence detected = agent responds
|
||||
- Text input: `{"type":"user_message","text":"..."}` — agent replies with audio + text
|
||||
- **MUST send `conversation_initiation_client_data` immediately after WS connect** — without it, server won't process client audio (mic appears dead)
|
||||
- `conversation_initiation_client_data` payload: `conversation_config_override.agent.turn.mode`, `conversation_config_override.tts.optimize_streaming_latency`, `custom_llm_extra_body.enable_intermediate_response`
|
||||
- `enable_intermediate_response: true` in `custom_llm_extra_body` reduces time-to-first-audio latency
|
||||
|
||||
## API Keys / Secrets
|
||||
- ElevenLabs API key is set in **Project Settings → Plugins → ElevenLabs AI Agent** in the Editor
|
||||
|
||||
Reference in New Issue
Block a user