EVERYONE.
Start with restricted settings
Keep rollout, handoff, and follow-up disabled while establishing the test baseline. Restrict the audience before adding test recipients. API reference: GET Get Settings · PUT Replace Settings
Read the effective settings back after the update. Change one behavior at a time when testing handoff or follow-up, then restore the restricted baseline before moving to another scenario.
Add test recipients
Keepai_audience set to ALLOWLISTED_ONLY, then add each tester as an E.164 phone number. Store the returned allowlist entry id so you can remove the entry later.
API reference: GET List Allowlist · POST Create Allowlist Entry · DELETE Delete Allowlist Entry
Run a single-turn test
API reference: POST Test Agent
The current REST response does not include
estimated_token_usage.
Preserve context for multi-turn tests
Send the returnedconversation_id with the next customer message.
API reference: POST Test Agent
Test knowledge, skills, connectors, and tools
Cover the configured resources instead of testing only happy-path questions. API reference: POST Run Connector Tool · GET List Connector Logs- Ask questions answered by business information, FAQs, websites, and files.
- Check missing facts, contradictory sources, stale content, and unsupported requests.
- Verify when each skill should and should not run.
- Exercise every connector tool with valid, invalid, and incomplete input.
- Confirm that secret values never appear in responses or logs.
- Include cases that should trigger human handoff or produce no response.
/connectors/{connectorId}/tools/{toolId}/runs with representative input before allowing the agent to select it in a conversation:
Test through WhatsApp with allowlisted recipients
Validate the same scenarios from each allowlisted test phone number. This confirms live channel behavior that the test endpoint cannot reproduce completely.Run evaluations
The evaluation API provides these resources:
List the available cases, submit a run using the required
eval_case_ids value, retain the returned job_id, and poll the job endpoint.
scenario, categories, max_turns, and success_criteria. Detailed results include scores, turn labels, reasons, transcripts, and timestamps. Summaries aggregate scores, highlights, and failure categories.
Go-live checklist
- Required business facts are correct and non-contradictory.
- Multi-turn conversations preserve the intended context.
- Missing information produces a safe response instead of a fabricated answer.
- Connector tools succeed and fail safely with representative input.
- Handoff scenarios behave as expected.
- Evaluation failures are reviewed and either fixed or explicitly accepted.
- A rollback owner knows how to disable rollout.
Next: Roll out safely
Enable the agent for allowlisted recipients first, then expand to the full audience.

