Kount, an Equifax company · 2024

Trust score explainability

Rebuilt the score as an argument a reviewer can defend, not a number they have to take on faith. Analysts were approving and declining more than a million transactions a day on a score they could not explain.

In codeReview consoleGitHub


Role
Lead product designer
Team
1 PM · 1 solutions architect · 2 engineers · 3 data scientists
Timeframe
Q2–Q4 2024
Platform
Web · internal review console

Shipped

Transaction Trust Score modal showing score 85.6, waterfall chart of contributing factors, and two-column reason codes
The trust score explainer. Score bands render at equal width even though their ranges are unequal. Kount's published threshold guidance puts roughly 5% of volume below 61, so the vast majority of traffic sits in the top bands — proportional widths would compress the range where nearly every real decision happens.
Reconstructed payment review queue with ten transactions, trust scores, and decision statuses
Ten reconstructed payments. Open a row to work the payments-fraud case. This is the surface analysts lived in, not a component gallery.

Constraint

Kount scores over a million transactions a day. An analyst sees a number between 0 and 99.9 and decides: approve, hold, or decline. The score was accurate. It was also opaque. When a merchant challenged a decline, the answer was the number. Disputes escalated to engineering. Manual review felt safer than a decision they could not defend.

“I trust the model. I just can’t tell a merchant why we blocked their best customer.”

Fraud analyst, internal interview

Two limits were fixed from the start. The model is proprietary, so contributing factors could surface and weights could not. Reason codes were pre-scripted by the risk team, not generated per transaction. The design problem was how to make that limited set feel sufficient for a decision an analyst has to defend.

What one screen has to answer

QuestionSurface
What is the score?The number and its band, on the same screen as the decision.
How was it built?A waterfall that starts from a stable visual baseline.
What drove it?Active and inactive reason codes, side by side.
Research synthesis card showing the four questions analysts repeatedly asked about the trust score
Queue observation notes. Analysts re-derived the same four questions on every transaction.

Decisive moves

Separate composition from drivers

Queue observation kept returning two questions: how the score was built, and what specifically drove it. Composition became the waterfall. Drivers became a two-column list, increased and decreased. Data Science controlled which reason codes could surface and how many.

Pin the visual baseline

The mathematical start often sat in the 70s and shifted with every transaction. A waterfall only reads as an argument if it begins from a stable number. We pinned the visual baseline at 80.0 so the movements stayed legible. The underlying math stayed more fluid. That decision made the rest of the interface coherent.

Diagram comparing the shifting mathematical starting point of the score with the fixed visual baseline of 80.0 used in the waterfall
The mathematical starting point shifted per transaction. Pinning the visual baseline at 80.0 made the waterfall readable.

Keep absence visible

Inactive reason codes stay on the screen, dimmed rather than hidden, so analysts learn the full vocabulary and can see what did not fire. A feedback control in the footer flags explanations that do not hold up. That routes to the risk team as reason-code tuning, not a dead end.

Two-column reason code list with active codes at full opacity and inactive codes dimmed
Inactive reason codes stay visible at reduced opacity. Absence is information.

In code

The original Figma file is Equifax IP. The public kit is the same slice rebuilt: a ten-row review queue that opens the Payments Fraud case. The explainer lives on that case. Weights stay hidden. Reasons stay pre-scripted. The visual baseline stays pinned at 80.0. Approve, hold, and decline write back to the queue. The score itself does not change.

Reconstructed trust score explainer for order 48291-C showing score 85.6, a composition waterfall from the 80.0 baseline, two-column reason codes, and approve, hold, and decline actions
The same modal, now a React composition. Approve, hold, and decline write back to the queue. Feedback routes to reason-code tuning. The score itself does not change.

Open Northwind or Harbor in the live console. Source is on GitHub. The component kit is still in Storybook.

Evidence

The decision did not change. The defensibility of it did.

The feature shipped as Explainable AI for Omniscore. Equifax describes it as giving users factor-by-factor risk assessments that let them identify key risk drivers quickly, improve accuracy, and reduce time spent on manual investigation — and as reducing the need for overly complex rules, because analysts who trust the score stop compensating for it in policy.

Those are Equifax's published descriptions of the shipped product, not outcomes I measured. I designed the interface; I did not run the study that would let me claim a number, and I would rather say so than round one up.


NextRules engine