LiveKit Agent Simulations 如何測試語音智能體
Testing voice agents with Agent Simulations
LiveKit 的 Agent Simulations 以任務層級模擬測試語音智能體,現已向所有 LiveKit Cloud 專案開放。文字模擬只測試對話邏輯,執行成本較低,適合每次 pull request;音訊模擬涵蓋背景噪音、打斷及聽者附和聲,較適合夜間或發布前測試。
Testing a voice agent is hard. Your agent will take calls from someone on a bad connection in a parking lot, someone who speaks softly, someone who gives their account number in three pieces with a question in the middle, and someone who may try to get it to say something it shouldn't. The happy path is the easy part, and it's the only part most agents get tested on. The long tail is where they fail in production.
The LLM at the core compounds it, since the same input won't always produce the same output. And testing any of this by hand means calling your own agent, over and over, which doesn't scale with the number of tasks it performs or the number of languages your users speak.
Agent Simulations is a repeatable way to test voice agents and identify failures before your customers do. Serve AI used simulations to validate a full agent rewrite against a hundred real production calls before cutting any traffic over. Agent Simulations is now generally available for all LiveKit Cloud projects, and will work in concert with Agent Evaluations, coming soon after.
Why conversations need different tests#
The LiveKit Agents framework already supports turn-by-turn testing. But turn-by-turn testing can't replicate the unpredictability of a real conversation. A caller may have to provide basic information before the agent can reschedule their appointment, and the order in which they get there varies.
Simulations work at the task level. You state the user's task, and the agent takes whatever path it needs to accomplish it. Each scenario you write stays in the suite. When you add your tenth feature, the checks for your first three run automatically, so you find out if the new work broke the old.
These types of checks matter more with agents than with most software. With a web app you can visually see a regression: you ship a new screen, review the webpage, and notice a missing link. A voice agent has no interface to glance at, so the only way to know a capability still works is to actually try it.
How simulations work#
Agent Simulations uses an LLM-driven user to play out a scenario from start to finish. For each scenario, you can define the simulated user's personality, goals and what their behavior will be. An LLM-based judge grades the transcript against freeform expectations you set. Once a simulated conversation reaches a natural conclusion, the judge assigns a pass/fail verdict and the reason behind it. Scenarios are run in parallel and the results, transcripts and traces land in your LiveKit Cloud dashboard.
Generating scenarios#
With Agent Simulations you don't have to write scenarios from scratch. Point the CLI at your project and LiveKit's scenario generation tool reads your agent's code and produces scenarios grounded in its actual prompts, tools, and conversation logic rather than a generic set of conversations.
The CLI runs the scenarios against your agent and reports back with a link to the run in the dashboard. Within minutes you have a pass rate on paths you never explored, without making a single call yourself.
Generated scenarios are a starting point, not a finished suite. Keep the ones that surface something, adjust the instructions on the ones that don't, and write your own for the cases you already know matter. Your checked-in scenarios.yaml becomes the source of truth, versioned alongside the agent it tests.
Text and audio runs#
Text simulations bypass the speech models in your agent, focusing only on testing the conversation logic. They're fast and cost-effective to run at scale, and they're the default. Text-mode requests are scheduled at low priority and paced into spare capacity against your inference quota, so they don't compete with your project's live traffic. That makes them cheap enough to run on every pull request, and when a scenario fails the CLI returns a non-zero exit code, which fails the build.
Audio simulations cover the entire audio pipeline, including background noise, interruptions, and backchanneling, the sounds a listener makes while the other person is still talking. They test your cascaded or speech-to-speech/full duplex agent. They run in real time and are metered at a higher rate than text, so they're better suited to a nightly or pre-release schedule.
A realistic simulated user#
A simulation is only as good as the caller driving it. A simulated user that waits its turn and asks exactly what you expected will pass every scenario you write and tell you nothing.
Since real callers make small sounds while the agent is still talking, the simulator does too, placed where a listener would actually make them rather than on a timer. Near the end of a turn it uses only "mm-hmm" and "uh-huh," and drops "okay," so a backchannel doesn't accidentally answer the agent's question.
Text mode has no audio and no pauses, so it runs strictly one turn at a time. In audio, the simulated user can end up overlapping the agent through a backchannel or by replying during a long pause.
All of this adds up to increased confidence, which in practice means you can change your agent code without wondering what you inadvertently broke.
Getting started#
Agent Simulations is free through October. Note that simulations run your real agent, so it calls your LLM as usual, and in audio mode your STT and TTS too. Those are billed at your normal rates.
You will need the LiveKit CLI, LiveKit Observability enabled, and a LiveKit Cloud project to get started. Full setup, options, and the CI configuration are in the docs.
來源:LiveKit 產品與工程博客 · livekit.com