Cognigy.AI intent training job returns 422 then 503 during status polling

Hi all,

requests is the SDK library we are currently leveraging to orchestrate this workflow, and I want to walk through the exact sequence of events causing the async training job to fail. When we initially submit the TRAINING CONFIG payload, the initial handshake completes without issue, but the system immediately begins evaluating the distribution metrics. On the second polling interval, the backend responds with a 422 UNPROCESSABLE ENTITY. We have already verified that the UTTERANCE EXAMPLES and INTENT LABELS match the schema perfectly, which means the rejection is strictly driven by the CLASS IMBALANCE CONSTRAINT rejecting the current distribution.

Current specs include:

  • Python 3.9 requests library
  • Intent health metrics endpoint
  • Snapshot comparison enabled for MODEL VERSIONING
  • Audit logging routed to S3

To break down the failure path step by step, we first observe that the classImbalanceThreshold parameter is set to 0.15 in our configuration. Next, the API interprets this as a strict boundary, and once the distribution crosses it, the training state locks. Following that, the status polling loop continues executing, but after the third check, it returns a 503. Finally, because the job never reaches a terminal success state, the DEPLOYMENT PIPELINE reference never updates, leaving the infrastructure in a pending loop.

import requests
headers = {"Authorization": f"Bearer {token}", "Content-Type": "application/json"}
# Note: The specific training submission endpoint is not listed in the provided API spec.
# The following code demonstrates polling the intent health metrics to check status/metrics.
# Replace flowId, versionId, intentId, and language with your actual values.
url = f"{base_url}/api/v2/flows/{flowId}/versions/{versionId}/intents/{intentId}/health?language=en-US"
resp = requests.get(url, headers=headers)
print(resp.status_code, resp.json())

Documentation states: “Async training jobs require a valid access token with nlu:intent:edit and nlu:training:edit scopes throughout the entire lifecycle of the job.”
Why does it not work? You’re seeing a 422 turn into a 503. That’s usually the platform rejecting the poll because the token context is stale. The snapshot comparison needs higher privileges than a basic view scope.

Cause:
Your token rotation is lagging. The initial POST works because the token is fresh. By the second poll, the token service hasn’t refreshed the bearer token, or the scope doesn’t cover the audit logging payload size. The 503 is the platform giving up on the invalid auth context.

Solution:
Check the token scopes. You need to ensure the grant includes nlu:training:edit. Run this to verify the current token’s claims before hitting the poll endpoint.

import requests

# Verify token scopes first
token_response = requests.get('https://api.mypurecloud.com/api/v2/oauth/tokeninfo', headers={'Authorization': f'Bearer {access_token}'})
print(token_response.json())

# If nlu:training:edit is missing, you'll get 422 on poll
poll_url = f"https://api.mypurecloud.com/api/v2/nlu/intent-models/train/{job_id}/status"
headers = {
 'Authorization': f'Bearer {access_token}',
 'Content-Type': 'application/json'
}
# Ensure you handle the refresh token immediately upon 401, don't wait for 503
response = requests.get(poll_url, headers=headers)

The docs also mention: “If the training job payload exceeds the default size limit, the server returns 422 UNPROCESSABLE ENTITY.” You might be hitting that with the snapshot data. Chunk the utterances.

Weird how the 503 pops up. Usually means the backend worker crashed on the auth check.

1 Like