Topic Detection - Precision Drop After Keyword Update

Hi all,

We’re seeing a significant precision drop in our topic detection model after updating the keywords for ‘payment disputes’. It’s odd, because the keywords themselves seem relevant - we added “refund request”, “chargeback”, and “billing error” to the existing list of “dispute”, “incorrect charge”, and “payment issue”. Before the update, precision was around 82%. Now, it’s hovering around 65% - it’s impacting our reporting.

The flow is pretty standard - Zoom Contact Center routing based on detected topic, then tagging the interaction in our CRM. We’re using the Zoom Contact Center speech analytics engine directly - no custom integrations there. The topic model uses phrase spotting, not sentiment analysis, so I initially ruled out calibration issues. But it feels…similar to what was discussed in the community post about the SIP trunk jitter affecting precision, only this isn’t the trunk.

The weird part is the recall actually increased slightly, from 78% to 80%. This is making it hard to pinpoint the issue. I checked the keyword spotting logs in the analytics dashboard, and it appears the new keywords are being triggered correctly - they aren’t just missing entirely. It’s like the engine is correctly identifying these phrases, but misclassifying the overall topic.

I’ve also reviewed the phrase match thresholds - they’re set to the default 0.75. Has anyone else experienced this kind of precision drop after adding keywords? Is there a hidden weighting system for keywords that we should be aware of? Or maybe a cache invalidation issue after updating the model? I’m wondering if there’s a limit to the number of keywords a topic can have without impacting performance.

I tried recreating the topic model from scratch, but it didn’t improve things. Any thoughts? I’m open to exploring different approaches.

2 Likes

That precision drop sounds like a classic over-eagerness problem with the keyword list- the model’s matching too broadly after you added those terms, so it’s flagging interactions that are tangentially related to payment disputes. Try weighting the original keywords higher, and the new ones lower- the Zoom Contact Center topic detection API lets you do that via the keyword_weight parameter when you’re updating the model; it’s a float between 0 and 1, where 1 is full weight and 0 is effectively ignored. For example, to give “dispute” a weight of 1 and “refund request” a weight of 0.2, you’d use a PATCH request like this:

{
 "keywords": [
 {"keyword": "dispute", "weight": 1.0},
 {"keyword": "incorrect charge", "weight": 1.0},
 {"keyword": "payment issue", "weight": 1.0},
 {"keyword": "refund request", "weight": 0.2},
 {"keyword": "chargeback", "weight": 0.2},
 {"keyword": "billing error", "weight": 0.2}
 ]
}

We’ve seen it before- the default weight is 0.5, and often lowering the weight on recently added terms will improve precision. Re-train the model after the change and see if that tightens things up.

Thanks that’s a really solid point about keyword weighting. It’s easy to forget that the model isn’t just looking for the presence of keywords, but also how strongly they indicate a specific topic.

What I’m wondering, though, is if the issue isn’t solely the weight - it’s about how the keyword_weight interacts with the existing training data. Zoom Contact Center’s topic detection uses a machine learning model, and those models are, well, sensitive. Adding keywords without considering the existing data distribution can sometimes throw things off.

Let’s talk a bit about how the model evaluates interactions. Each interaction gets a score for each topic. This score is calculated based on the keywords present, their weights, and - crucially - the model’s prior understanding of those keywords. If the model was previously trained on a lot of interactions where “dispute” and “incorrect charge” appeared, it will naturally assign a higher score to those terms.

Then you introduce “refund request” and “billing error” - terms that sound related but might appear in totally different contexts. The model hasn’t seen enough examples to properly assess their relevance yet. It’s essentially saying, “Okay, these words are here, but I don’t really know what to do with them.” This can lead to a lower overall precision because the model is misclassifying interactions.

To address this, you could try a multi-stage approach. First, instead of directly changing weights, test with a smaller subset of interactions. Run a set of recent interactions through the model with the new keywords, but with the original weights. Look at the results. You can use the Zoom Contact Center reporting API to pull interaction data and predicted topics - it’s a bit tedious, but you get a clear picture.

GET /api/v2/analytics/interactions?topic_detected=true&date_from=2023-10-26&date_to=2023-10-27

Examine the interactions where the model incorrectly flagged a payment dispute. Are there common themes? Are certain keywords appearing in contexts where they shouldn’t be? This will help you refine your keyword list and weights.

Then, when you do update the model, increase the keyword_weight gradually. Don’t jump from 1.0 to 0.5 immediately. Start with something like 0.8 and monitor the impact.

You’ll also want to look at the confidence_score returned by the API. The confidence score tells you how sure the model is about its prediction. Lower confidence scores suggest the model is uncertain and might be misclassifying interactions. If you’re seeing a lot of low confidence scores, that’s a sign you need to re-evaluate your keywords and weights.

1 Like

so that keyword weighting thing seems… obvious, right? but we ran into something similar last quarter - inc-4471 was the ticket, if anyone’s curious. the problem wasn’t just the weights, it was the data itself.

basically, the zoom contact center topic detection api looks for exact matches unless you tell it otherwise. so “billing error” is different than “billingerrors” or “a billing error”. we found a lot of variation in the call transcripts - mostly just typos, but enough to throw things off.

you might try adding some fuzzy matching to the keywords. it’s a bit of a pain to set up - you need to use regex in the keyword_pattern field instead of just a string. it’s clunky but it helps. something like billing\s?error would catch both “billing error” and “billingerror”. it’s a bit fiddly but worth a look.