Data Action - WebRTC ICE API call timeout

Hi all,

Something very strange is happening with Data Actions calling the WebRTC ICE endpoint. We’re trying to pull ICE candidates for troubleshooting poor audio quality - agent reports MOS scores dropping to 2.8, and traceroute shows some jitter on the path. It’s not consistent, but it’s happening.

The Data Action is failing with a timeout - it looks like the call to retrieve ICE candidates isn’t responding in a reasonable timeframe. Sometimes it works, sometimes it doesn’t. When it fails, the error in the Data Action history is just “Timeout waiting for response”. Very unhelpful.

We’ve checked the Genesys Cloud platform health dashboard, and everything shows green. No reported outages in Frankfurt.

I tried to increase timeout on the Data Action, from default 30 seconds to 60, then 90. No help. It still fails. I suspect something with the API, maybe throttling? Is anyone else seeing this?

The flow is simple - a single ‘Execute Data Action’ block, POST request to the ICE endpoint. We’re using the user ID from the context variables. No complex logic.

We’ve also tried a workaround - calling the stations endpoint first to verify the user context is valid. This works 90% of the time. But why do we need to check it? The ID should be valid if the flow is running.

Here’s the relevant part of the Data Action configuration:

Method: POST
URL: https://api.mypurecloud.com/api/v2/integrations/actions/{actionId}/execute
Headers: 
 Content-Type: application/json
 Authorization: Bearer {auth_token}
Request Body: 
 {"userId": "{userId}"}
Timeout (seconds): 90

The {userId} and {auth_token} are pulled from context variables.

Here’s what we’ve done so far:

  • Increased Data Action timeout to 90 seconds.
  • Added a Data Action to verify user context validity before calling the ICE endpoint.
  • Checked Genesys Cloud platform health dashboard.
  • Verified network connectivity from the Data Action server to the Genesys Cloud API endpoint - traceroute looks ok, under 20ms RTT.
  • Confirmed that the authentication token has appropriate permissions.
  • We’re on Genesys Cloud v23.11.
  • SDK version is current.

What’s the best way to debug this? Is there logging somewhere on the API side that we can access? Can anyone tell me if there are known rate limits on this endpoint? If the user ID is not correct, why the API doesn’t return an error instead of just timing out?

1 Like
{
 "iceCandidateRequest": {
 "userId": "yourUserId",
 "max": 10,
 "includeCandidatesWithoutMedia": false
 }
}

That timeout - we’ve seen it. The ICE candidate retrieval process is sensitive to the max parameter. If you ask for too many candidates at once, it can choke. Try reducing the max value - maybe start with 5 or 10, like above. Also, the includeCandidatesWithoutMedia flag - set this to false if you only need candidates for active media streams. It cuts down the response size.

We had an issue where the ICE negotiation wasn’t completing due to mismatched codec preferences. SDP offers and answers - sometimes the offered codecs weren’t supported on the far end. The MOS score drop makes me think about that. Are you seeing any SIP OPTIONS keep-alives failing? A basic SIP trace might show if there are codec mismatches. Also, what is the value of max set to now?

tbh, is spot on - that endpoint is a mess if you ask for too much at once. we’ve seen it too, and it’s not just max - the user ID itself can be the problem (especially if it’s a bot user or something weird).

iirc, we ended up throttling those requests - basically, queuing them up and spacing them out by like 200ms each, using a pre-request script. something like this (super rough, don’t yell): pm.environment.set("ice_request_count", (pm.environment.get("ice_request_count") || 0) + 1); if (pm.environment.get("ice_request_count") > 5) {pm.execution.setNextRequest.delay(200)}.

1 Like