Notification API WebSocket stream - inconsistent event delivery, intermittent disconnects

The notification API stream - specifically events for voice.call - is dropping events. Not all calls trigger events. It’s intermittent, which makes debugging… challenging. We’re seeing this across multiple orgs now, so it’s not isolated. -1 to assuming it’s a localized config issue.

Data flow looks like this:

[Genesys Cloud] <---WebSocket---> [Rust Client]
 |
 v
[Event Processing Pipeline]

Client connects to wss://events.gc.genesys.cloud/events/v2/stream/ - standard setup. Subscription filters are correct. We’ve verified event types and object types are matched. The issue isn’t what events are coming through, it’s that they aren’t all coming through. We’ve confirmed the calls are definitely happening in Genesys Cloud.

The Rust client is using the tokio-tungstenite crate, version 0.17.5. Serde is version 1.0.152. The event deserialization is straightforward - using serde_json::from_str. No custom deserialization logic that could be the source.

Error codes are unhelpful. When the connection drops, the WebSocket simply closes without a specific error. The client attempts a reconnect - the reconnect logic is working, but the missed events are not retransmitted. The API doesn’t appear to have any provision for event replay.

We’ve tried:

  • Increasing the WebSocket keepalive interval. Worth a shot, but didn’t fix it.
  • Adding more solid error handling to the reconnect logic. Doesn’t address the root cause.
  • Checked platform health status. No reported outages.
  • Verified subscription filters against the documentation. Multiple times.
  • Increased logging to capture more detail around the connection lifecycle. Mostly just shows disconnects.
  • Different regions. No change.
  • Client libraries are updated to latest.

Here’s a snippet of the reconnect logic - it’s pretty standard:

use futures_util::stream::StreamExt;
use tokio_tungstenite::tungstenite::Message;

async fn reconnect(url: &str) -> Result<(), Box<dyn std::error::Error>> {
 loop {
 match connect(url).await {
 Ok((mut ws_stream, _)) => {
 // Process events
 while let Some(message) = ws_stream.next().await {
 // Deserialize and process
 }
 }
 Err(e) => {
 eprintln!("Connection error: {}", e);
 tokio::time::sleep(std::time::Duration::from_secs(5)).await;
 }
 }
 }
}

This isn’t a simple timing issue. It’s inconsistent. Some calls have events, others do not. Something is clearly dropping events at the platform level. +1 to thinking the event stream isn’t reliable.