EventBridge notification lag and SQS DLQ spikes during high volume

Hi all,

Dealing with some weird latency on our EventBridge pipeline. We’ve got a Node.js Lambda handler processing conversation events via SQS. Normally it’s snappy, but during peak hours, we’re seeing a MASSIVE spike in the Dead Letter Queue.

The logic is simple: trigger from EventBridge, push to SQS, Lambda processes. We’ve already implemented exponential backoff to avoid hitting API rate limits - always a priority to keep the tenant healthy - but we’re still seeing events land in the DLQ with no clear reason.

Tried checking the event definitions via GET /api/v2/usage/events/definitions/{eventDefinitionId} to see if the schema changed, but nothing there. The Lambda isn’t crashing; it’s just timing out or getting throttled. I’ve bumped the memory to 1024MB and increased the concurrency limit, but the lag persists.

CloudWatch shows the events are arriving at the EventBridge rule, but the gap between the eventTime in the payload and the Lambda execution start is stretching to 30+ seconds.

{
 "eventTime": "2023-10-27T14:20:01.123Z",
 "eventType": "ConversationUpdate",
 "errorMessage": "Task timed out after 3.01 seconds",
 "awsRequestId": "c688-4b12-8d34-efg567",
 "stack": "TimeoutError: execution exceeded allocated timeout"
}

We’re following security best practices by using IAM roles with least-privilege access, but I’m wondering if there’s some hidden throttle on the EventBridge side for specific event types. Just a reminder to everyone: don’t use long-lived credentials in your Lambda env vars.

The logs show the function is just hanging before it even attempts the first API call.

1 Like

Check if the visibility timeout on the SQS queue is too short for the Lambda execution time. It’ll cause a loop that kills containment metrics if messages keep retrying.

{
 "VisibilityTimeout": 300,
 "MessageRetentionPeriod": 1209600
}

Is the DLQ actually a different queue or just a setting on the main one?

4 Likes

That’s right, and you’ll want to check the batch size on the SQS trigger too. When we moved our Zendesk ticket syncs over to GC, we found that too large a batch often timed out the Lambda before it finished. It’s a common gotcha mentioned in a few community posts about event spikes. Try dropping the batch size to 10.

The batch size tip in the earlier reply is right (400), but don’t forget the concurrency limit (429). If the Lambda scales too fast (which it will during peaks), you’ll choke the downstream API.

if (concurrent_executions > limit) {
throttle_and_retry(exponential_backoff)
}

Set a reserved concurrency (keep it low) to stop the spikes.

+1 to the batch size point.

The DLQ spikes usually happen because the Lambda doesn’t ACK the message.

If the code crashes before the return, SQS thinks it FAILED.

It puts it back in the queue until the maxReceiveCount is hit.

try {
 await processEvent(event);
} catch (e) {
 console.error(e);
 throw e; // This is what triggers the DLQ loop
}