Skip to main content
Test the agent against expected answers and failure cases before changing the audience to EVERYONE.

Start with restricted settings

Keep rollout, handoff, and follow-up disabled while establishing the test baseline. Restrict the audience before adding test recipients. API reference: GET Get Settings · PUT Replace Settings
GET Get Settings returns an array, even when only one settings object exists. PUT Replace Settings returns the updated settings object. YCloud omits null fields before calling Meta, so settings that are not supplied remain unchanged. Read the effective settings back after the update. Change one behavior at a time when testing handoff or follow-up, then restore the restricted baseline before moving to another scenario.

Add test recipients

Keep ai_audience set to ALLOWLISTED_ONLY, then add each tester as an E.164 phone number. Store the returned allowlist entry id so you can remove the entry later. API reference: GET List Allowlist · POST Create Allowlist Entry · DELETE Delete Allowlist Entry
Use GET List Allowlist to review the current test audience. Use DELETE Delete Allowlist Entry to remove a tester by the returned entry ID.
Do not use an end-user number with the +86 country calling code. Meta Business Agent does not currently reply to messages from +86 end users, even when the number is formatted as valid E.164 and added to the allowlist.

Run a single-turn test

API reference: POST Test Agent
A successful response can contain:
The current REST response does not include estimated_token_usage.

Preserve context for multi-turn tests

Send the returned conversation_id with the next customer message. API reference: POST Test Agent
Use a new conversation for scenarios that must not inherit previous context.

Test knowledge, skills, connectors, and tools

Cover the configured resources instead of testing only happy-path questions. API reference: POST Run Connector Tool · GET List Connector Logs
  • Ask questions answered by business information, FAQs, websites, and files.
  • Check missing facts, contradictory sources, stale content, and unsupported requests.
  • Verify when each skill should and should not run.
  • Exercise every connector tool with valid, invalid, and incomplete input.
  • Confirm that secret values never appear in responses or logs.
  • Include cases that should trigger human handoff or produce no response.
If a connector call fails, inspect the connector’s logs and verify credentials, certificate state, parameter bindings, and tool request definitions before changing the skill. Run each configured tool directly through /connectors/{connectorId}/tools/{toolId}/runs with representative input before allowing the agent to select it in a conversation:

Test through WhatsApp with allowlisted recipients

Validate the same scenarios from each allowlisted test phone number. This confirms live channel behavior that the test endpoint cannot reproduce completely.
During early access, the test endpoint can return an empty response or ELIGIBILITY_CHECK_FAILED while the audience is ALLOWLISTED_ONLY. Check no_response_reason, eligibility, rollout, audience, and allowlist settings. If API testing requires EVERYONE, use it only in a controlled environment and restore the restricted setting immediately after the test.

Run evaluations

The evaluation API provides these resources: List the available cases, submit a run using the required eval_case_ids value, retain the returned job_id, and poll the job endpoint.
The status field is returned as a string and is not currently constrained to a documented enum. Stop polling when the response provides a completed result or error instead of assuming undocumented status names. Evaluation cases describe a scenario, categories, max_turns, and success_criteria. Detailed results include scores, turn labels, reasons, transcripts, and timestamps. Summaries aggregate scores, highlights, and failure categories.

Go-live checklist

  • Required business facts are correct and non-contradictory.
  • Multi-turn conversations preserve the intended context.
  • Missing information produces a safe response instead of a fabricated answer.
  • Connector tools succeed and fail safely with representative input.
  • Handoff scenarios behave as expected.
  • Evaluation failures are reviewed and either fixed or explicitly accepted.
  • A rollback owner knows how to disable rollout.

Next: Roll out safely

Enable the agent for allowlisted recipients first, then expand to the full audience.