An AI risk metric/score of 87 tells an approver almost nothing.
It says the model is worried. It doesn't say why, how fresh the number is, or whether the data behind it was complete. If the pipeline failed 17 times this month, the 87 looks exactly the same.
That's a design problem, not a model problem. The fix is mostly disclosure: say where the number came from, show the signals behind it, put partial runs on the decision screen, and let the approver re-run it before signing off on $248,600.
The person clicking Approve owns the decision. The interface owes them the evidence.
Chris, this is spot on. On a payments agent we built, approvers only started trusting the flags once each one showed the reason behind it and when it last ran. A bare number just made people click Approve faster, which is the opposite of what you want.