Constructing an Analytics API aggregation query that groups by queue and media type

Constructing an Analytics API aggregation query that groups by queue and media type

What You Will Build

You will build a Python script that queries the Genesys Cloud Analytics API to retrieve conversation volume metrics, aggregated by Queue and Media Type.
This tutorial utilizes the GET /api/v2/analytics/conversations/details/query endpoint via the official Genesys Cloud Python SDK.
The implementation covers authentication, query construction, pagination handling, and result parsing.

Prerequisites

Before executing the code, ensure you have the following resources and configurations:

  • OAuth Client Credentials: A Genesys Cloud OAuth client with the scope analytics:conversation:read.
  • Genesys Cloud Python SDK: Version 3.0.0 or higher. Install via pip install genesys-cloud-sdk-python.
  • Python Runtime: Version 3.8 or higher.
  • Environment Variables: You must configure the following in your shell or environment file:
    • GENESYS_CLOUD_CLIENT_ID
    • GENESYS_CLOUD_CLIENT_SECRET
    • GENESYS_CLOUD_REGION (e.g., us-east-1, eu-west-1)

Authentication Setup

Genesys Cloud uses OAuth 2.0 for API authentication. The Python SDK handles the token acquisition and refresh logic automatically when initialized with client credentials. This eliminates the need to manually manage token expiry in your business logic.

The following code initializes the platform client. Note that the SDK caches the access token and refreshes it transparently when it expires.

import os
from purecloudplatformclientv2 import (
    Configuration,
    ApiClient,
    AnalyticsApi
)

def get_analytics_api_client():
    """
    Initializes and returns an authenticated AnalyticsApi client.
    """
    # Load credentials from environment variables
    client_id = os.getenv("GENESYS_CLOUD_CLIENT_ID")
    client_secret = os.getenv("GENESYS_CLOUD_CLIENT_SECRET")
    region = os.getenv("GENESYS_CLOUD_REGION", "us-east-1")

    if not client_id or not client_secret:
        raise ValueError("GENESYS_CLOUD_CLIENT_ID and GENESYS_CLOUD_CLIENT_SECRET must be set.")

    # Configure the client
    config = Configuration(
        client_id=client_id,
        client_secret=client_secret,
        region=region
    )

    # Initialize the API client
    api_client = ApiClient(configuration=config)

    # Return the Analytics API instance
    return AnalyticsApi(api_client)

Implementation

Step 1: Constructing the Aggregation Query Object

The core of this tutorial is the QueryConversationDetailsRequest object. This object defines what data you want, the time window, and how you want it grouped.

To group by Queue and Media Type, you must set the group_by field to a list containing the strings "queue" and "mediaType". You must also define the interval to determine the time granularity of the results.

The following code constructs the request body. It uses the last_24_hours preset for simplicity, but you can replace this with custom start/end ISO 8601 strings.

from purecloudplatformv2.models import (
    QueryConversationDetailsRequest,
    Interval,
    QueryTime
)

def build_queue_media_query(start_time: str, end_time: str):
    """
    Builds the query object for conversations grouped by queue and media type.
    
    Args:
        start_time: ISO 8601 start time string.
        end_time: ISO 8601 end time string.
        
    Returns:
        QueryConversationDetailsRequest object.
    """
    
    # Define the time interval for the query
    time_query = QueryTime(
        start_time=start_time,
        end_time=end_time,
        # Using 'PT1H' for 1-hour intervals. 
        # Options include: PT1H, PT30M, PT15M, PT1M, P1D, P1W, P1M, P1Y
        interval="PT1H" 
    )

    # Define the query request
    query_request = QueryConversationDetailsRequest(
        query_time=time_query,
        # Group by Queue and Media Type
        group_by=["queue", "mediaType"],
        # Define the view. 'default' is standard for most metrics.
        view="default",
        # Optional: Filter by specific queues if needed. 
        # If None, it queries all queues accessible by the client.
        # filters=[{"type": "queue", "selector": "id", "values": ["queue-id-1"]}]
        filters=None
    )

    return query_request

Critical Parameter Explanation:

  • group_by: This is an array of strings. The order matters for the structure of the returned JSON. If you use ["queue", "mediaType"], the response will be a list of queues, each containing a list of media types. If you reverse the order, the structure changes accordingly.
  • interval: This defines the time buckets. If you set PT1H, you get one data point per hour. If you set P1D, you get one data point per day. Ensure your start_time and end_time span aligns with this interval to avoid partial buckets at the edges.
  • view: The default view provides standard metrics like conversations (total count), abandoned, and handled. Other views like agent or skill provide different metric sets.

Step 2: Executing the Query and Handling Pagination

The Analytics API returns paginated results. You must check the next_page_id in the response to fetch subsequent pages. Failing to handle pagination will result in truncated data, especially for large organizations with high conversation volumes.

The following function executes the query and iterates through all available pages.

from purecloudplatformclientv2.rest import ApiException

def fetch_aggregated_data(api_instance: AnalyticsApi, query_request: QueryConversationDetailsRequest):
    """
    Executes the analytics query and handles pagination.
    
    Args:
        api_instance: Authenticated AnalyticsApi client.
        query_request: The constructed query object.
        
    Returns:
        A list of ConversationDetailsResponse objects.
    """
    all_responses = []
    page_id = None
    max_pages = 100 # Safety limit to prevent infinite loops in case of API errors

    print(f"Starting query execution...")

    for _ in range(max_pages):
        try:
            # Execute the query
            # The API expects the request body to be passed directly
            response = api_instance.post_analytics_conversations_details_query(
                body=query_request,
                page_id=page_id # Pass page_id for subsequent requests
            )
            
            # Append the response data
            all_responses.append(response)
            
            # Check if there is a next page
            if response.next_page_id:
                page_id = response.next_page_id
                print(f"Fetched page. Next page ID: {page_id}")
            else:
                print("No more pages. Query complete.")
                break
                
        except ApiException as e:
            # Handle specific HTTP errors
            if e.status == 429:
                print("Rate limit exceeded. Implement exponential backoff in production.")
                import time
                time.sleep(10) # Simple wait for demonstration
                continue
            elif e.status == 400:
                print(f"Bad Request. Check your query parameters. Error: {e.body}")
                break
            else:
                print(f"Unexpected API Error: {e.status} - {e.reason}")
                raise e

    return all_responses

Step 3: Processing and Flattening Results

The raw response from the Analytics API is nested. When grouping by queue and mediaType, the structure looks like this:

{
  "entities": [
    {
      "id": "queue-id-1",
      "name": "Support Queue",
      "entities": [
        {
          "id": "voice",
          "name": "Voice",
          "metrics": { ... }
        },
        {
          "id": "web",
          "name": "Web Chat",
          "metrics": { ... }
        }
      ]
    }
  ]
}

You must flatten this structure to make it usable for reporting or database insertion. The following function extracts the metrics into a flat list of dictionaries.

def flatten_results(responses: list) -> list:
    """
    Flattens the nested analytics response into a list of flat records.
    
    Args:
        responses: List of ConversationDetailsResponse objects.
        
    Returns:
        List of dictionaries with keys: queue_id, queue_name, media_type, metric_name, metric_value.
    """
    flat_data = []
    
    for response in responses:
        if not response.entities:
            continue
            
        # Iterate over the first level: Queues
        for queue_entity in response.entities:
            queue_id = queue_entity.id
            queue_name = queue_entity.name
            
            if not queue_entity.entities:
                continue
                
            # Iterate over the second level: Media Types
            for media_entity in queue_entity.entities:
                media_type = media_entity.name # e.g., "Voice", "Web Chat"
                
                # Extract metrics
                if media_entity.metrics:
                    for metric_name, metric_value in media_entity.metrics.items():
                        # metric_value is often a dict with 'value' and 'count'
                        if isinstance(metric_value, dict):
                            value = metric_value.get('value', 0)
                            count = metric_value.get('count', 0)
                        else:
                            value = metric_value
                            count = 0
                            
                        flat_data.append({
                            "queue_id": queue_id,
                            "queue_name": queue_name,
                            "media_type": media_type,
                            "metric_name": metric_name,
                            "metric_value": value,
                            "metric_count": count
                        })
                        
    return flat_data

Complete Working Example

The following script combines all previous steps into a single, runnable module. It fetches conversation data for the last 24 hours, grouped by queue and media type, and prints the total conversations per queue per media type.

import os
import datetime
import json
from purecloudplatformclientv2 import (
    Configuration,
    ApiClient,
    AnalyticsApi
)
from purecloudplatformv2.models import (
    QueryConversationDetailsRequest,
    QueryTime
)
from purecloudplatformclientv2.rest import ApiException

def get_analytics_api_client():
    client_id = os.getenv("GENESYS_CLOUD_CLIENT_ID")
    client_secret = os.getenv("GENESYS_CLOUD_CLIENT_SECRET")
    region = os.getenv("GENESYS_CLOUD_REGION", "us-east-1")

    if not client_id or not client_secret:
        raise ValueError("Missing GENESYS_CLOUD_CLIENT_ID or GENESYS_CLOUD_CLIENT_SECRET")

    config = Configuration(
        client_id=client_id,
        client_secret=client_secret,
        region=region
    )
    api_client = ApiClient(configuration=config)
    return AnalyticsApi(api_client)

def build_query():
    # Calculate time window: Last 24 hours
    end_time = datetime.datetime.utcnow()
    start_time = end_time - datetime.timedelta(hours=24)
    
    # Format as ISO 8601 with Z for UTC
    start_iso = start_time.strftime("%Y-%m-%dT%H:%M:%SZ")
    end_iso = end_time.strftime("%Y-%m-%dT%H:%M:%SZ")
    
    time_query = QueryTime(
        start_time=start_iso,
        end_time=end_iso,
        interval="PT1H" # 1-hour granularity
    )
    
    return QueryConversationDetailsRequest(
        query_time=time_query,
        group_by=["queue", "mediaType"],
        view="default"
    )

def fetch_and_process():
    api_instance = get_analytics_api_client()
    query_request = build_query()
    
    all_responses = []
    page_id = None
    
    print("Fetching analytics data...")
    
    while True:
        try:
            response = api_instance.post_analytics_conversations_details_query(
                body=query_request,
                page_id=page_id
            )
            all_responses.append(response)
            
            if response.next_page_id:
                page_id = response.next_page_id
            else:
                break
        except ApiException as e:
            print(f"API Error {e.status}: {e.reason}")
            break
            
    # Flatten the data
    flat_data = []
    for res in all_responses:
        if not res.entities:
            continue
        for queue in res.entities:
            if not queue.entities:
                continue
            for media in queue.entities:
                if media.metrics and 'conversations' in media.metrics:
                    conv_value = media.metrics['conversations'].get('value', 0)
                    flat_data.append({
                        "queue": queue.name,
                        "media_type": media.name,
                        "conversations": conv_value
                    })
                    
    # Sort by queue name and then by conversation count descending
    flat_data.sort(key=lambda x: (x['queue'], -x['conversations']))
    
    # Print results
    print("\n--- Analytics Results ---")
    print(f"{'Queue':<30} | {'Media Type':<15} | {'Conversations':<15}")
    print("-" * 65)
    
    for row in flat_data:
        print(f"{row['queue']:<30} | {row['media_type']:<15} | {row['conversations']:<15}")
        
    return flat_data

if __name__ == "__main__":
    try:
        results = fetch_and_process()
        print(f"\nProcessed {len(results)} records.")
    except Exception as e:
        print(f"Fatal error: {e}")

Common Errors & Debugging

Error: 403 Forbidden (Insufficient Scopes)

Cause: The OAuth client used for authentication does not have the analytics:conversation:read scope.
Fix: Go to the Genesys Cloud Admin Console, navigate to Security > OAuth 2.0, edit your client, and ensure analytics:conversation:read is added to the Scopes list. Re-authorize the client if necessary.

Error: 400 Bad Request (Invalid Interval)

Cause: The interval string in the QueryTime object is invalid or inconsistent with the start_time and end_time. For example, using PT1H (1 hour) but querying a range of only 30 minutes.
Fix: Ensure the time range is at least as long as the interval. If you want hourly data, query at least one hour. Use standard ISO 8601 duration formats: PT1H, PT30M, P1D.

Error: 429 Too Many Requests

Cause: You have exceeded the API rate limit for the Analytics endpoint. The Analytics API has stricter rate limits than other APIs because queries are computationally expensive.
Fix: Implement exponential backoff. The SDK does not automatically retry 429s. You must catch the ApiException with status 429 and wait before retrying. Start with a 1-second delay and double it with each subsequent 429 response.

Error: Empty Entities List

Cause: The query returned successfully, but entities is empty.
Fix: This usually means no conversations matched the criteria. Check the following:

  1. Did you filter by a specific queue that had no activity in the selected time window?
  2. Is the time window in the future? Ensure start_time and end_time are in the past.
  3. Are you querying for a media type that your organization does not use? (e.g., querying for “Fax” if you only have Voice and Web Chat).

Official References