Test and deploy your agent
Running Simulations and reading Quality reports
Running simulations and reading quality reports are essential practices for maintaining and improving the performance of your Customer Agent. These tools help you catch issues early, validate that your agent responds correctly, and track its progress over time with measurable data. Understanding how to use each effectively will ensure your agent delivers consistent, accurate answers to customers.
running simulations
Simulations are scripted conversations designed to automatically test your Customer Agent by mimicking real customer interactions. Although this feature is currently in beta and not yet available, it is planned for future release. Once available, simulations will allow you to:
- Write detailed conversation scripts that include customer messages, follow-ups, and procedure clicks.
- Define conditions such as audience, context, topic, and surface to simulate specific scenarios.
- Set expected outcomes to score the agent’s responses.
- Group simulations into suites for scheduled or on-demand runs.
- Compare results across runs to detect regressions.
- Trace failures back to the exact source, procedure, or context causing the issue.
Until simulations are released, you can rely on the Preview feature for ad-hoc checks, Quality reports for scoring the agent against control questions, and Conversations to monitor live interactions for potential problems.
understanding quality reports
Quality reports provide a quantitative way to measure your agent’s performance by running a fixed set of control questions and scoring the answers. This feature is currently available and helps you track whether your agent is improving or declining over time.
key components of quality reports
- Control questions: A curated list of common questions customers ask, with expected answers defined once and reused for every report.
- Control answers: The benchmark answers you expect the agent to provide.
- Report runs: Each time you run a report, the agent answers all control questions, and the system scores each response by comparing it to the control answer using semantic similarity.
- Report score: An average score across all questions, giving a quick overview of the agent’s health.
- Per-question score: Detailed scores for each question, highlighting specific areas that need improvement.
how to use quality reports
You can add control questions manually or pull them from the top questions identified in the Analyze section. After setting the expected answers, run a report to see how the agent performs. Low scores indicate answers that differ significantly from the expected response, guiding you to the exact source that requires fixing or updating.
Quality reports are best used:
- After any change to Living Knowledge to verify that performance has not degraded.
- On a weekly basis as a regular health check.
- Before presenting agent results to leadership or auditors.
interpreting and acting on report results
When you see a low score for a question, open it to compare the agent’s answer, the control answer, and the sources used. This helps you decide whether to fix the source content, add new sources, or update the control answer if the agent’s response is actually correct.
conclusion
While simulations will soon provide a powerful way to automate scenario testing and catch regressions, quality reports already offer a robust method to measure your agent’s accuracy and track improvements over time. Using these tools together—simulations when available, and quality reports now—ensures your Customer Agent remains reliable, accurate, and ready to meet customer needs.