Your model works. The problem is that people do not trust it, do not understand what it just did, or cannot correct it when it gets something wrong. That is a design problem, and it is the one that decides whether an AI product gets adopted or quietly abandoned after the demo.
I designed the human-in-the-loop review system for a claims platform that cut processing time by 48% and human error by 64%, and the conversational onboarding at Learntech that took course completion from 31% to 64%. In both cases the win came from designing what happens when the AI is uncertain or wrong - not from the happy path.