Skip to main content
Agent Observability

Agent evaluations

Written By Stanislas

Last updated 16 days ago

Overview

Agent evaluations provide a dedicated testing interface to assess the accuracy and consistency of your AI agent's responses. By comparing generated outputs against defined reference answers, you can systematically score how well your agent follows instructions and knowledge sources.

Use evaluations when refining agent prompts, updating connected documentation, or validating an agent before production rollout to catch regressions and maintain response quality.


Prerequisites

  • Role permissions: Only workspace owners and workspace administrators have access to configure and execute evaluations.

  • Active custom agent: At least one custom AI agent created in your workspace.

  • Test criteria: A set of test questions paired with expected ground-truth answers.


Step-by-step guide

1. Access the evaluations dashboard

  1. In the left sidebar of your workspace, click Agents and select the agent you want to evaluate.

  2. In the agent configuration navigation panel, scroll down to the Observability section and click Evaluations.

  3. The Evaluations dashboard opens, displaying the Questions and answers table, search bar, and action controls.

2. Add an evaluation manually

  1. Click the + Add dropdown button located above the table.

  2. Click Add an evaluation.

  3. In the modal, fill in the evaluation details:

  • Question: Enter the input prompt or query to send to the AI agent.

  • Expected response: Define the target response against which the agent will be measured.

  • Add comment (optional): Enter internal context, guidance, or testing criteria for reviewers.

  1. Click Save to add the test case to the questions list.

3. Import evaluations from a CSV file

  1. Click the + Add dropdown and select Import evaluations from csv file.

  2. Click the sample.csv link to download the formatted template.

  3. Prepare your spreadsheet containing question, expected response, and optional comment columns.

  4. Drag and drop your file into the Drop your file here box, or click the box to browse your computer.

  5. Click Save to upload and populate the evaluation table.

4. Run evaluations and review scores

  1. In the Questions and answers table, check the box next to each evaluation you want to run.

  2. Click the Run evaluation button in the top toolbar.

  3. Once the run finishes, the Semantic similarity column displays the similarity percentage for each evaluated item.

  4. Check the Global Score card beneath the table to view the overall semantic similarity rating and performance bracket across all evaluated items.

5. Inspect evaluation details

  1. Click the eye icon in the Action column of any completed evaluation row.

  2. Review the slide-over drawer showing the original Question, Expected response, the live Agent response after evaluation, and the calculated Semantic Similarity score.

  3. Compare the generated output with your expected text to identify missing details, phrasing issues, or incorrect knowledge retrieval.


Practical use cases

  • Pre-deployment benchmarking: Execute a suite of 25 reference customer questions against a new bot to guarantee compliance and accuracy before rolling it out company-wide.

  • Prompt iteration and regression testing: Run evaluations after modifying system instructions or switching LLM models to verify existing response quality did not degrade.

  • Knowledge update validation: Evaluate agent answers before and after uploading new documentation to confirm the agent correctly retrieves updated internal policies.


Tips & best practices

  • Run evaluations on targeted subsets of questions during quick iteration cycles to optimize credit usage.

  • Write clear, unambiguous expected responses that focus on essential facts rather than exact phrasing.

  • Remove outdated test questions using the red delete icon in the table toolbar whenever your business workflows change.

  • Review cases scoring below 80% to detect whether instructions need refinement or whether knowledge documents are missing key context.


Troubleshooting

  • Run evaluation button is disabled:

  • Cause: No evaluation rows are currently selected.

  • Fix: Select at least one checkbox in the evaluation table to activate the button.

  • Evaluations menu item is not visible:

  • Cause: Your account has a standard member role rather than workspace owner or admin permissions.

  • Fix: Ask a workspace administrator to grant your account administrator access.

  • Low semantic similarity score:

  • Cause: The agent omitted key facts, misunderstood instructions, or retrieved inaccurate context.

  • Fix: Click the eye icon to view the agent's full response, then update the system prompt or provide more detailed knowledge sources.


Additional resources

  • Agent analytics – Monitor agent run counts, completion rates, and credit consumption over time.

  • Monitoring agent sessions and inbox traces – Inspect individual conversation transcripts and technical error logs.

  • Agent settings – Configure system instructions, temperature, and model selection.