DogStatsD metrics dropping on /api/v2/employeeperformance/externalmetrics/data

Hey everyone,

Running into a weird gap in our dashboards. We’ve got a custom pipeline pushing external metrics into Genesys Cloud via POST /api/v2/employeeperformance/externalmetrics/data to track some custom KPIs in Datadog. Most of the time it’s smooth, but we’re seeing a 15% drop in data points for specific intervals. No errors in the logs, just missing data!!

Is it a rate limit issue? Probably not, since we aren’t seeing any 429s on the API side.

Tried a couple of ways to fix this:

Option A: Increase the batch size in the DogStatsD collector before pushing to the API.
Pros: Reduces the number of total calls.
Cons: If a larger batch fails, we lose more data at once and it’s harder to pinpoint the exact record that caused the hiccup.

Option B: Implement a local queue with a retry mechanism for any response that isn’t a 200 or 202.
Pros: Should stop the data loss for intermittent network blips.
Cons: Adds latency to the metrics and we’re already pushing the limits of our memory buffer on the collector.

The environment is v2 API and we’re using a custom Node.js middleware to bridge the two. The payloads are standard JSON and usually fly right through.

{
 "metrics": [
 {
 "metricName": "custom.queue.latency",
 "value": 124.5,
 "timestamp": "2023-10-27T10:00:00Z"
 }
 ]
}

Still seeing gaps in the external metrics data despite the collector saying the POST was successful.

The promotion of these metrics is likely failing due to a rate limit issue. I would recommend checking the aggregates to see if the threshold is being hit.

curl -X POST "https://api.mypurecloud.com/api/v2/analytics/ratelimits/aggregates/query" \
-H "Content-Type: application/json" \
-d '{
 "interval": "PT1H",
 "filter": {
 "type": "and",
 "clauses": [
 {
 "type": "or",
 "clauses": [
 {
 "type": "eq",
 "propertyName": "apiName",
 "value": "POST /api/v2/employeeperformance/externalmetrics/data"
 }
 ]
 }
 ]
 }
}'
1 Like