Analytics Alert - Event to PagerDuty - Duplicate Incidents

Fun one today. Think of the Genesys Cloud Analytics API as a leaky faucet - it drips events, and we’re trying to catch every drop with PagerDuty. The issue isn’t missing events, it’s getting the same event triggering multiple PagerDuty incidents, almost immediately after each other. It’s like trying to fill a bucket with that leaky faucet, and the overflow protection is way too sensitive.

We’ve got an Analytics alert set up monitoring queue abandon rate. When it crosses 5%, it triggers a webhook to our internal service, which then formats a PagerDuty v2 event. The alert itself seems accurate - GC Analytics shows the rate jumping, and the webhook fires. The problem is PagerDuty shows two, sometimes three, incidents for the same alert evaluation.

The service is written in Node.js, using the PagerDuty Events API v2 SDK (version 2.4.3). Here’s the relevant code snippet for creating the event:

const pagerDuty = require('pagerduty-events-api');
const client = new pagerDuty.Client();

async function createPagerDutyIncident(abandonRate, queueName) {
 const event = {
 routing_key: 'YOUR_ROUTING_KEY',
 event_action: 'trigger',
 payload: {
 summary: `High Abandon Rate - ${queueName}`,
 detail: `Abandon rate exceeded 5% (${abandonRate}%)`,
 source: 'Genesys Cloud Analytics'
 },
 dedup_key: 'abandon-rate-' + queueName //trying to dedup on queue name
 };

 try {
 await client.trigger(event);
 console.log('PagerDuty incident created.');
 } catch (error) {
 console.error('Error creating PagerDuty incident:', error);
 }
}

We’ve tried setting dedup_key - initially, just a static value, then the queue name. Doesn’t help. The PagerDuty UI shows no correlation ID - it’s as if the requests are just arriving independently. Network traces show the webhook firing, and the POST requests to PagerDuty are successful (202 responses). I also checked the GC alert history - it’s only triggering once per rate crossing, so it’s not a double-fire from the source. The GC environment is prod, using the Miami region. The webhook endpoint is TLS 1.2 secured.

It feels like a race condition somewhere, or maybe PagerDuty is having issues with handling similar events rapidly. The mic stays hot, just need to stop the flood.

Hey everyone,

It sounds like the event deduplication is not working as you expect. The documentation for the Analytics API states that “event data is sent on a best-effort basis, and duplicate events may occur”. It’s not ideal, but it explains why you’re getting the duplicates.

Here’s a few things you can check:

  • The PureCloudPlatformClientV2 doesn’t have a built-in deduplication feature for Analytics events, so you’ll need to handle it in your PagerDuty integration logic. You’ll need to store event IDs, and ignore duplicates within a short time window.
  • Check the alert condition definition - maybe the alert is too sensitive and triggers on slight fluctuations.
  • Consider increasing the threshold for the alert. 5% might be too low, and any minor fluctuation will trigger it.
  • You can use the metrics.aggregation.interval parameter when creating the alert to reduce the frequency of event triggers. The documentation says the minimum is 60 seconds.

It’s a bit of work to add the deduplication logic, but it should solve the duplicate incident problem.

2 Likes

the docs are right - the api doesn’t auto-dedupe. that’s a perf tradeoff, i guess. but blindly retrying on a 429 won’t fix it, it’ll just worsen the storm. the earlier post’s points are good, but incomplete. you need idempotency keys in your payload.

cause: the alert is firing repeatedly because gc is re-transmitting the same event data, potentially due to network hiccups or internal retries before it hits your webhook. without an idempotency key, pagerduty treats each transmission as unique.

solution: add idempotency_key to your alert config. generate a unique uuid for each event before you send it to gc. the gc api will then drop duplicates based on this key. example payload snippet:

{
 "event_type": "queue.abandon_rate",
 "abandon_rate": 0.06,
 "idempotency_key": "a1b2c3d4-e5f6-7890-1234-567890abcdef",
 "queue_id": "some-queue-id"
}

just a hunch, but gc’s internal retry logic likely includes a window for deduplication that relies on the key. without it, you’re fighting a losing battle. it’s not a guaranteed fix, but it’s the only reliable method i’ve found.

That fix worked - added a client-side deduplication layer using the event UUID as a key, and the duplicate incidents stopped. Was simpler than I thought.

resource "genesyscloud_dataaction" "event_dedupe_filter" {
 name = "Event-Dedupe-Filter"
 description = "filters duplicate analytics events"
 api_version = "1.0"
 status = "PUBLISHED"
 schema {
 type = "object"
 properties {
 event_uuid = {
 type = "string"
 }
 }
 }
}

so yeah, saw the fix worked with the uuid - good stuff!! might be wrong but i reckon you could also look at using a data action before pagerduty to filter events. like, if you’ve already got a data action set up for analytics stuff, just add a little logic to check if that uuid’s been seen in, like, the last 5 mins or something. took me like 2 hrs to get a similar thing working for webhooks, but it avoids doing it on the pagerduty side - could be cleaner?

worth a shot anyway. for what it’s worth, we usually add a ‘seen_events’ redis cache in front of the dataaction to make it really snappy.