So, we’ve got this Architect flow pulling data from Athena via the API gateway - basically trying to list all historical conversation summaries. From what I’ve seen, the standard way of querying historical conversation data doesn’t play nice with huge result sets. It just… times out.
Tried adding ?$top=1000 and $skip=0 to the query, hoping for pagination, but it looks like Athena’s not respecting those parameters through the gateway. It returns the full dataset (or errors, usually a 504 - but that’s intermittent) instead of chunking. I was expecting to see nextPageToken in the response, but no luck.
We’re on Genesys Cloud, notification_api, Architect flow version 3.1. EmbeddedClientAppSdk is at 4.2.1. The WebSocket connection is stable, but these baked-in reconnection quirks still cause headaches. Just yeet this into prod is never an option, obviously.
Here’s the basic WebSocket handshake code we’re using for channel subscriptions - not directly related, but it’s always good to have:
Okay - so you’re hitting the limits of the historical analytics endpoint, which - surprise - isn’t exactly designed for pulling everything at once. We ran into this last quarter when trying to backfill data for a compliance audit. Athena itself does support pagination - but the gateway’s implementation is…spotty, to say the least.
Try this - build the query in the Architect flow to call the historical analytics endpoint. Then, use a data action to parse the nextPage link out of the response body. It’s a JSON object, so it’s fairly straightforward. The nextPage link will contain the $skip value you need for the next request.
It’s clunky, and you’ll need to build a loop to iterate through all the pages - which means some extra scripting in the flow. We’ve got a training module on data actions and JSON parsing, though I’m guessing it’s already on the backlog for Q4.
Stakeholders won’t be thrilled with the performance, but it’s better than a timeout. Worth a shot.
It’s a good discussion so far - the limitations of the historical analytics endpoint are definitely a frequent source of trouble. I’d like to add a few thoughts, expanding on what’s already been shared.
The assumption that $top and $skip will function identically to their behavior in other API contexts is a common misstep. While those parameters are present in the documentation, their implementation with Athena through the gateway isn’t entirely consistent. Specifically, Athena itself will paginate, but the gateway doesn’t always pass those requests through cleanly.
Regarding the granularity setting - setting it to ‘hourly’ is a sensible approach, but it’s worth considering the potential impact on query performance. While it reduces the volume of data returned per call, a very granular setting over a long date range can still create a substantial workload for Athena. We’ve found that experimenting with daily or even weekly granularity - where the business requirements allow - can yield more stable results.
One gotcha to watch for is the maximum allowed date range. Athena has an upper limit on the time period a single query can cover, typically around 90 days. If you’re trying to pull data for a longer period, you’ll need to break it down into smaller, overlapping segments and stitch the results together on the application side.
The earlier reply mentioned building queries in Architect. That’s a perfectly valid approach, but it’s worth noting that Architect flows have their own execution limits and timeout settings. Complex queries, even when paginated, can exceed those limits, resulting in failed requests. Monitor the flow execution logs carefully.
Finally, don’t overlook the possibility of data sampling. Depending on your data volume and Athena’s configuration, the results may be based on a sample rather than the full dataset. If accurate counts are crucial, confirm Athena isn’t applying any sampling. This is less of a pagination issue, but it can lead to discrepancies in the data you retrieve.
So that whole Athena pagination thing is a mess, honestly - we’ve been banging our heads against it for months trying to pull historical data for reporting, and it’s just…slow, even when it works. The earlier reply about dropping the top to 100 is right - it’s more reliable, but you’ll need to build a proper loop in your Architect flow to handle the pagination yourself, and it’s going to be painful because the analytics API’s pagination limits are criminally low, like, seriously. Here’s a quick Python snippet for fetching all pages - it’s dirty, assumes you’ve got the API key/auth sorted, and it’s not handling errors properly because move fast, ship it, but it’ll give you the idea; you’ll need to adapt it to hit the Athena gateway from within Architect, which is… a whole other level of fun.
import requests
import json
base_url = "https://api.mypurecloud.com/api/v2/analytics/botflows/{botFlowId}/sessions"
# replace with your actual auth info
api_key = "YOUR_API_KEY"
auth_token = "YOUR_AUTH_TOKEN"
all_data = []
page_size = 100
page_number = 0
while True:
url = f"{base_url}?granularity=hourly&pageSize={page_size}&skip={page_number}"
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
}
response = requests.get(url, headers=headers)
if response.status_code == 200:
data = response.json()
if not data['results']:
break
all_data.extend(data['results'])
page_number += page_size
else:
print(f"Error {response.status_code}: {response.text}")
break
print(json.dumps(all_data, indent=2))
You’ll probably want to cache each page locally - or even better, write it straight to Redis or something - because re-querying for the same data is just…sad.
Okay, so the Athena pagination through the Genesys Cloud API is…rough. You’re right to suspect it’s not a straight $top/$skip situation - it’s more complex than that.
Here’s what we’ve done - and it’s a bit of a workaround, admittedly. Instead of wrestling with the API directly in Architect, we pull the data in batches using a Data Action, then iterate.
The Approach
Data Action Setup: Create a Data Action that requests historical conversation analytics with a small $top value (like 50 or 100). IMPORTANT: set the granularity to daily.
Looping Logic: In Architect, loop through the results from the Data Action. Each loop iteration fetches the next “page” of results.
Tracking Offset: Maintain a variable (using a Session variable in Architect) to track the current $skip value. Increment this variable after each Data Action call.
You’ll need to increment ${session.offset} by 50 (or whatever $top value you chose) on each loop. The loop continues until the number of returned results is less than $top, indicating the end of the result set.
It’s not ideal, but it’s more reliable than trying to force Athena’s pagination through the gateway. Worth a shot.