Hi all,
requests is the SDK library we are currently leveraging to orchestrate this workflow, and I want to walk through the exact sequence of events causing the async training job to fail. When we initially submit the TRAINING CONFIG payload, the initial handshake completes without issue, but the system immediately begins evaluating the distribution metrics. On the second polling interval, the backend responds with a 422 UNPROCESSABLE ENTITY. We have already verified that the UTTERANCE EXAMPLES and INTENT LABELS match the schema perfectly, which means the rejection is strictly driven by the CLASS IMBALANCE CONSTRAINT rejecting the current distribution.
Current specs include:
- Python 3.9 requests library
- Intent health metrics endpoint
- Snapshot comparison enabled for MODEL VERSIONING
- Audit logging routed to S3
To break down the failure path step by step, we first observe that the classImbalanceThreshold parameter is set to 0.15 in our configuration. Next, the API interprets this as a strict boundary, and once the distribution crosses it, the training state locks. Following that, the status polling loop continues executing, but after the third check, it returns a 503. Finally, because the job never reaches a terminal success state, the DEPLOYMENT PIPELINE reference never updates, leaving the infrastructure in a pending loop.
import requests
headers = {"Authorization": f"Bearer {token}", "Content-Type": "application/json"}
# Note: The specific training submission endpoint is not listed in the provided API spec.
# The following code demonstrates polling the intent health metrics to check status/metrics.
# Replace flowId, versionId, intentId, and language with your actual values.
url = f"{base_url}/api/v2/flows/{flowId}/versions/{versionId}/intents/{intentId}/health?language=en-US"
resp = requests.get(url, headers=headers)
print(resp.status_code, resp.json())