Open Analytics → A/B Test to create and run experiments across voice-assistant configurations. This page appears only for teams and roles with configuration access.
Starting an experiment sends live calls to every configured variant. Validate each assistant configuration and the assistant’s default configuration before you start.
Plan the experiment
Define these before creating anything:
- One behavior you want to change
- A hypothesis with an expected direction and business reason
- A primary success metric and any guardrail metrics
- The voice assistant that will receive the experiment
- A saved assistant configuration for every variant
- A traffic allocation and planned duration
- The default assistant configuration to use if the experiment is stopped
Use variants that differ only in the behavior being tested. If a prompt, voice, transfer rule, and booking policy all change at once, the result will not explain which change mattered.
Create an experiment
- Choose New Experiment.
- Enter a descriptive name and hypothesis.
- Select the voice assistant. If the team has one assistant, it is selected automatically.
- Enter a planned duration in days.
- Configure at least two variants. Give each one a distinct name, traffic percentage, and assistant configuration belonging to the selected assistant.
- If assistant overrides appear, optionally set a first-message or TTS-voice override for a supported variant.
- Select up to 10 primary metrics. The available metric catalog depends on your team’s features.
- Confirm that all traffic percentages add up to 100% and choose Create Experiment.
The new experiment is Not Started and does not receive traffic yet. The selected voice assistant is fixed as soon as the experiment is created.
Common outcome metrics can include booking, contained-booking, transfer, lead-booking, and average-call-duration measures. Use each metric’s description in the dashboard; a similar name does not guarantee the same population or preferred direction.
Edit before starting
While an experiment is Not Started, you can edit its name, hypothesis, planned duration, variant names, traffic splits, assistant configurations, supported overrides, and metrics. You cannot change its voice assistant.
Before choosing Start Experiment:
- Open every referenced assistant configuration and test a representative call.
- Confirm the intended control configuration.
- Confirm that splits total 100% and every variant has a configuration.
- Recheck which configuration is currently the voice assistant’s default.
- Verify the metric set and record the launch decision outside the experiment if your team requires an approval trail.
Starting begins live allocation immediately. Once live, the experiment structure is locked.
Monitor a live experiment
The detail page shows status, processed exposures, total exposures, days running, and results for each group. Results update once daily, so calls can appear before their metric results do.
For each metric, review:
- The control baseline and each variant’s mean
- Actual exposures by group, not only the configured split
- Absolute change from control
- The p-value and Significant or Not Yet Significant label
- Whether higher, lower, or neither direction is considered better
- The linked calls for each variant
The star marks the best observed value in that row; it does not by itself prove that the variant is a winner. Average Call Duration is treated as neutral because longer or shorter is not universally better.
The dashboard gives rough call-volume guidance for booking-rate changes: a smaller percentage-point improvement generally needs far more calls per variant. Avoid decisions based on the first few exposures, an unbalanced sample, or a positive metric with a harmful guardrail.
You can add metrics to a live experiment, up to the limit, but cannot remove a metric that was already attached while the experiment is live. Variant configurations, names, and traffic splits are also locked.
Stop and roll back safely
Choose Stop Experiment when a guardrail degrades, a variant is defective, or the planned decision point is reached. Enter a required stop reason and confirm Stop Serving Experiment.
Stopping does two things:
- It stops the experiment from serving new traffic.
- New calls fall back to the voice assistant’s current default configuration.
Stopping does not deploy the control or the best-looking variant. Verify the default configuration before stopping, especially if someone changed it after the experiment began. A stopped experiment cannot be resumed from the dashboard.
If the page displays a synchronization warning after stopping, preserve the warning and contact your Avoca support channel. Confirm the experiment status and test a new call against the default configuration before considering the rollback complete.
Decide what to ship
- Wait for the daily result update and enough exposures for the expected effect size.
- Check exposure balance and open calls from each variant.
- Review every primary and guardrail metric, including directionality.
- Treat Not Yet Significant as inconclusive, not as a loss.
- Record the decision and stop reason.
- Stop the experiment.
- Separately make the chosen assistant configuration the default, then test it.
Keeping the stop and deployment steps separate prevents an experiment result from silently changing the production default.