Skip to main content
Open Analytics → Oversight to manage the call-analysis signals available to your team. Oversight is available only when it is enabled and your role can view it. Oversight has five team-level pages:
  • Dashboard: monitor annotation volume and eval failures
  • Annotations: extract structured signals from call transcripts
  • Evals: test whether an applicable call met a pass/fail standard
  • Playground: try temporary definitions against a small call sample
  • Backfill: recalculate selected persisted definitions on historical calls

Read the Dashboard

The Dashboard defaults to the last 30 days. Change the date range to review:
  • Analyzed calls: calls with analysis available in the period
  • Calls with annotations: the count and share with at least one annotation value
  • Eval failures: failed eval results in the period
  • Most common annotations: values, counts, and shares by annotation
  • Top eval failures: failed count, failure rate, and applicable-call count
An eval’s failure rate uses applicable calls, not every analyzed call. Compare the Applicable count before ranking two evals by failure rate.

Configure annotations

An annotation extracts a reusable value from a transcript. In Annotations, choose an available template or create a custom annotation, then configure:
  • Name and Description
  • Type: Simple, Multiple choice, or Text
  • Options for a Multiple choice annotation
  • Prompt that describes the extraction logic
  • Voice agent scope when that control is available
  • Active and Visible
Use Simple for a yes/no signal, Multiple choice for a controlled set of categories, and Text only when a free-form answer is necessary.

Configure evals and alerts

An eval returns pass or fail only when the call is applicable. In Evals, configure:
  • A clear name and description
  • An Applicability prompt that says when the standard should be tested
  • An Evaluation prompt with the pass/fail criteria
  • A voice-agent scope when available
  • Active and Visible settings
Keep applicability narrow. A call that should not be judged must be excluded by the applicability prompt rather than treated as a failure. To receive email notifications for an eval, open its alert control and set:
  • A failure-rate threshold from 1% to 100%
  • A 1-, 6-, 12-, or 24-hour window
  • The minimum number of applicable calls required before alerting
  • One or more recipient email addresses
Use a minimum volume large enough to prevent a single call from creating a noisy alert. You can disable an alert without removing it, or remove the alert rule when it is no longer needed.

Understand Active, Visible, and scope

Active and Visible do different jobs:
  • An inactive definition remains saved but does not run on calls.
  • A hidden definition can still run, but its result does not appear in the dashboard call UI.
New custom definitions start inactive and hidden. Validate them before enabling either control. Custom and code-backed definitions can offer All voice agents or one voice agent. Shared templates apply across assistants and do not expose a voice-agent scope. Some definitions are managed at the enterprise level; those are read-only on the team page and must be changed by the appropriate enterprise owner.

Test in Playground first

Playground creates a temporary working session from the current annotation and eval library. Changes made there do not update the persisted definitions or production call results.
  1. Open Playground.
  2. Choose Last 5, Last 10, or Last 25, or paste up to 25 call UUIDs.
  3. Add the annotations and evals you want to test.
  4. Refine their temporary prompts or options.
  5. Choose Run playground.
  6. Review evaluated and skipped calls, token usage, transcripts, extracted values, explanations, quotes, pass/fail results, and evidence.
Use examples that include expected positives, expected negatives, and calls where an eval should be not applicable. Copy a proven change back into the saved definition only after reviewing the individual-call evidence.

Backfill historical calls

Backfill queues background work for eligible analyzed calls. It changes historical Oversight results, so test the definition in Playground first.
  1. Open Backfill.
  2. Select the first and last calendar day. The range is interpreted in the team’s timezone.
  3. Select only the persisted annotations and active evals that need recalculation.
  4. Review the day count and selected definitions.
  5. Choose Run backfill, then confirm.
The dashboard queues one job per eligible call. Successful results replace or create the historical value for each selected definition. If a selected definition fails to evaluate, its existing value is preserved; changing a voice-agent scope can also clear a stale selected result from calls that no longer match that scope.
A queued backfill has no cancel control on the team page. Start with a narrow date range, do not submit the same backfill twice, and wait for processing before judging the Dashboard totals.

Make changes safely

  • Set a definition inactive before retiring it; hiding it alone does not stop evaluation.
  • Prefer inactive and hidden over deletion while you confirm that no workflow depends on the result.
  • Deleting a team-created annotation or eval cannot be undone.
  • Backfill only the definitions whose logic or scope changed.
  • Do not use Backfill merely to refresh the Dashboard; change the date range or reload the page instead.
  • Review alert recipients and thresholds after changing an eval’s applicability.
Dashboard totals can lag while live analysis or a backfill is still processing. Inspect a few underlying calls before concluding that a definition or historical run is incorrect.