Hey everyone,
Running into a weird gap in our dashboards. We’ve got a custom pipeline pushing external metrics into Genesys Cloud via POST /api/v2/employeeperformance/externalmetrics/data to track some custom KPIs in Datadog. Most of the time it’s smooth, but we’re seeing a 15% drop in data points for specific intervals. No errors in the logs, just missing data!!
Is it a rate limit issue? Probably not, since we aren’t seeing any 429s on the API side.
Tried a couple of ways to fix this:
Option A: Increase the batch size in the DogStatsD collector before pushing to the API.
Pros: Reduces the number of total calls.
Cons: If a larger batch fails, we lose more data at once and it’s harder to pinpoint the exact record that caused the hiccup.
Option B: Implement a local queue with a retry mechanism for any response that isn’t a 200 or 202.
Pros: Should stop the data loss for intermittent network blips.
Cons: Adds latency to the metrics and we’re already pushing the limits of our memory buffer on the collector.
The environment is v2 API and we’re using a custom Node.js middleware to bridge the two. The payloads are standard JSON and usually fly right through.
{
"metrics": [
{
"metricName": "custom.queue.latency",
"value": 124.5,
"timestamp": "2023-10-27T10:00:00Z"
}
]
}
Still seeing gaps in the external metrics data despite the collector saying the POST was successful.