Architect Flow Publishing Failing - Stack Reference Issues - TS SDK

Right, so we’re trying to automate Architect flow publishing with Pulumi - TypeScript, naturally. Been doing this for ages, generally fine. But we’ve hit a weird one. Flows are deploying fine locally, but when pushing to our CI/CD pipeline (GitHub Actions, EU-West region) they consistently fail to publish with a 400 Bad Request - specifically a ‘validation error’.

  1. The error message is… verbose. It complains about a missing or invalid reference to a shared data object. The data object itself exists, verified in the GC UI and via direct API calls, and is referenced in the flow’s JSON. It’s just… like the SDK isn’t resolving the reference properly during the publish operation.

  2. The core issue seems to stem from stack references. We’re using a separate stack to manage the shared data objects and then referencing them by name from the flow’s stack. It works perfectly when the flow is deployed from the local machine. When the pipeline runs, the flow definition gets passed to the genesys.architect.Flow resource, which then tries to publish.

Here’s a simplified snippet of the flow resource definition:

const myFlow = new genesys.architect.Flow("my-flow", {
 name: "My Flow",
 description: "Test Flow",
 dataObjects: [{
 id: sharedDataObject.id,
 }],
 // ... other flow properties
});
  1. sharedDataObject comes from an Output of a separate stack. It’s a simple string representing the data object ID. When inspecting the Pulumi state in the CI/CD run, the sharedDataObject.id value looks correct - the ID is resolvable, but the API is still rejecting it.

  2. We’ve confirmed the GC API version being used is current (v2). We’ve tried different authentication methods (API keys, OAuth) - no change. We’ve even gone down the rabbit hole of checking the timestamp of the Pulumi state file vs. the last modified date of the shared data object. No obvious skew there.

  3. The pipeline isn’t doing anything fancy with the flow definition before publishing, just passing the generated JSON to the SDK. It’s really weird. Feels like the SDK’s resolution of stack references isn’t working consistently in a CI/CD environment.

edit: Just checked the network logs more closely. The request being sent to the publish endpoint does contain the correct data object ID. It’s not a serialization issue. The API is just rejecting it.

i think maybe the data object ID is different in the CI/CD environment?

we’ve got same issue before - it’s region thing - you check the ID in the UI for the correct org, yes? also, the error message - is it same every time, or changing a little?

just a hunch, but it’s usually the ID.

2 Likes

Oh, this is… familiar (400, of course). Spent a solid week chasing this sort of thing when we first started automating deployments - it’s a right pain. It’s not always the ID being different, although that’s a good shout from - you definitely check that (404 if it’s wrong).

What tripped us up was the order of operations, you see. The API expects the shared data object to be fully published before you try to reference it in the flow (429 if you rush it). Our CI/CD was deploying them in parallel - which, logically, felt faster - but Genesys didn’t like it.

So, force a sequential deployment - publish the data object first, then the flow. I’ve found adding a simple sleep command - even just 10 seconds - between the API calls does the trick. It’s a bit of a hack, admittedly, but it works. We eventually refactored it into a proper dependency check, but that took ages to organise.

Cause: The 400 error likely originates from a race condition during deployment. The referenced data object isn’t fully propagated across the Genesys Cloud infrastructure before the flow attempts to use it.

Solution: Implement a polling mechanism - query the data object status repeatedly until the ‘published’ flag resolves to true. This ensures the object is available before flow deployment proceeds. A small thing, but effective.

1 Like

The observations regarding operational sequencing are, frankly, quite astute. We’ve encountered similar challenges when automating deployments to the Great Cloud - it’s not merely a question of flow IDs differing between environments, although that is - of course - a valid consideration. It’s the state of the object itself.

The suggestion to poll the flow’s published status is sound, however, simply verifying the presence of a published version isn’t always sufficient. The propagation delay - even after the publish action completes - can still result in transient errors. The API doesn’t instantaneously replicate the flow definition across all regional edge servers, you see.

A more solid approach - and one we implemented some time ago - involves not only initiating the publish, but also confirming the flow is checkable out and its metadata is fully populated. Specifically, checking the response from the publish action. This action isn’t instantaneous in its availability for subsequent operations, but it’s populated once the flow is fully available across all relevant infrastructure components.

Here’s a snippet demonstrating this in a Workflow YAML configuration - it’s a quick one, just to illustrate the principle:

name: Validate Flow Publication
on:
 workflow_run:
 workflows: [ "pulumi-deploy" ]
 types: [ completed ]
jobs:
 validate_flow:
 runs-on: ubuntu-latest
 needs: [ "pulumi-deploy" ]
 steps:
 - name: Publish Flow
 id: publish_flow
 uses: peter-evans/http-client@v2
 with:
 url: 'https://api.mypurecloud.com/api/v2/flows/actions/publish?flow=${{ needs.pulumi-deploy.outputs.flow_id }}'
 method: 'POST'
 headers: '{"Authorization": "Bearer ${{ secrets.GENESYS_CLOUD_API_TOKEN }}"}'
 timeout: 60
 continue-on-error: true
 - name: Fail if Flow Not Fully Published
 if: steps.publish_flow.outputs.status != 200 || (steps.publish_flow.outputs.data == null)
 run: |
 echo "Flow ${{ needs.pulumi-deploy.outputs.flow_id }} is not fully published. Failing deployment."
 exit 1

This workflow snippet assumes you’re exporting the flow ID from your Pulumi deployment as a workflow output. We’ve found that a retry mechanism - polling every 5-10 seconds, for a maximum of 5 minutes - provides a reasonable balance between reliability and deployment speed.

It’s worth noting that the GENESYS_CLOUD_API_TOKEN must have the appropriate permissions to access the Architect API. A granular permission set is always preferable, naturally.

1 Like