Valuate the Model, Don’t Just Evaluate It

Episode Summary

A data scientist finishes a model and reports an area under the curve of 0.83. Eric Siegel points out that those numbers are, in his words, entirely arcane to the business, that the person responsible for the operation cares about profit, savings and KPIs, and that the translation between the two is generally not done. The missing step is small: take the same test data already used to evaluate the model, and valuate it as well, expressing its performance in business terms. That pays off twice: it steers development toward value, and it gives the data scientist a way afterwards to say what the model is worth, which is what gets an organization to deploy it. This, he says, is why predictive AI is failing. Not the mathematics, but a no man’s land between business and technology where both sides point at the other and the hose never connects to the faucet.

Key takeaways

  • Predictive AI is most of what AI meant until about three years ago. Fraud detection, credit risk scoring, marketing targeting and predictive maintenance: learning from data to put odds on a per-case outcome, then using those odds to drive the operational decision

  • A data scientist is trained to report the wrong thing, and their tools support it. Area under the curve, precision, recall and accuracy are what the training and the tooling produce, and he calls them entirely arcane to the business stakeholder, who cares about profit, savings and KPIs

  • The fix is one more step on data you already have. Take the same test data used to evaluate the model and valuate it as well, expressing its performance in business metrics rather than technical ones

  • Valuating pays off twice. During development it lets the data scientist navigate the train and test iteration toward actual value, and afterwards it gives them a way to convey the model’s potential value in business terms, which he calls the carrot at the end of the stick that gets an organization to deploy

  • Both sides point at the other. He describes a no man’s land between business and technology in which responsibility is always somebody else’s, so the hose never connects to the faucet

  • Most predictive AI projects stall before deployment. He says the field is failing like crazy and calls it a crisis, while adding that this does not mean the value is unproven: if only 15 percent succeed, that 15 percent of many projects is a lot of success, but overall the track record is dismal

  • He is careful that this is not just his opinion. He has run Machine Learning Week since 2009, has been involved in rounds of industry research, and cites surveys of data scientists and an IBM study of executives, before saying plainly that the models are not deploying

  • If you do not measure business value, you cannot be pursuing business value. He states it as a principle rather than a checklist, and it is why the first of the two problems he names is a problem in the development of the model itself

  • Predictive AI can be a guardrail on a generative one. A system that performs at 95 percent cannot simply be unleashed, but a predictive layer can learn which cases are most likely to go wrong

  • Route the risky slice to a human and keep most of the promise. His worked example sends the top 15 percent of riskiest cases for human review, which is more expensive per case, and still realizes 85 percent of the autonomy, which he sets against the zero percent you get from a system that never deploys

  • The executive has to go one level in, not become technical. His guidance is that a business leader assess the model’s potential value in business terms before accepting it, because the alternative is having the model handed to you and having no basis to act on it. His framing is to dive in a little, which he says is not super technical

About Eric Siegel

Eric Siegel is CEO of Gooder AI, a product built to close the gap between what a predictive model does technically and what it is worth to a business, which he describes as a de facto business console for predictive AI projects. He has run the Machine Learning Week conference since 2009 and has been involved in rounds of industry research, including surveys of data scientists. He is the author of two books written for both sides but, as he puts it, first and foremost for the business reader: The AI Playbook, which sets out the six-step practice he calls BizML, and Predictive Analytics, his first. On air he is direct that predictive AI is failing to deploy at scale, that this is a crisis rather than a quibble, and that the cause is a biz-tech gap nobody is incentivized to bridge.

 

In this episode

00:00 Welcome
00:21 Guest introduction, Gooder AI and the two books
00:47 What predictive analytics is, for an executive who has not met it
01:03 Predictive AI was most of what AI meant until three years ago
01:20 Fraud detection, credit risk, marketing targeting, predictive maintenance
01:39 Playing the odds better, and how it relates to generative AI
02:12 Odds on the outcome, driving the per-case decision
02:32 Not generally used together yet, and why they need each other
02:49 Why a language model is harder to use well for concrete value
03:34 No magic crystal ball, only probabilities
03:48 Decades of positive track record, and how little is realized
04:11 Less magic-seeming, and the semi-technical understanding it needs
04:30 How each resolves the other’s weaknesses
05:28 A specialized chatbot built on Claude
06:52 Autonomy as the ideal, and why it is out of reach
07:13 A system that is right 95 percent of the time
07:31 Using predictive AI to find the cases most likely to go wrong
08:12 Sending the riskiest 15 percent to a person
08:30 85 percent of the promise, against zero for a system that never ships
08:44 The executive considering an agentic solution
09:19 Justifying the build with an expected outcome
11:06 Deployment means operations change
11:30 Area under the curve, precision, recall, accuracy
12:00 What the business side cares about instead
12:51 Where the financial support usually already is
13:04 Where the project stalls
13:24 Post-POC, pre-deployment
13:40 Predictive AI as a field is failing like crazy
14:04 The biz-tech gap, still not widely bridged
14:10 Adil Ajmal at Fandom, and the same combination in use
16:06 No man’s land between biz and tech
16:32 What pure predictive performance does and does not tell you
16:58 Performance in business terms instead
17:26 If you do not measure business value you cannot pursue it
17:50 Operations do not improve unless they change
18:26 Why he wrote them and who they are for
18:32 For both sides, but business first
19:32 BizML, and six steps across six chapters
20:13 Why deployment is the whole point
20:35 Running Machine Learning Week since 2009
21:06 They just do not deploy, and the failures get swept under the rug
21:24 Viable models that could be delivering value
21:50 More excited about the rocket science than the launch
21:55 Adoption lagging enthusiasm, and 3 percent at the first summit
22:14 The BARC research, and 19 percent today
22:32 Friction on both sides, and what still needs solving
23:33 His writing and his course at machinelearning.courses
23:53 James Taylor, from the business vantage
24:34 What Gooder AI was built to do
24:37 What executives are delegating without realizing
25:01 Diving in one level, without becoming technical
25:35 The leadership skill that matters most
26:42 Making the haystack smaller
27:03 The arithmetic that turns performance into a financial win
27:50 Hands dirty, or feet cold
28:17 Assessing potential value before you use it
28:36 What the business side has to be given
29:06 Wrap-up

In Guest’s words

“Let’s take this same test data that we’re using to evaluate the model, but not just evaluate the model, let’s valuate the model.”

— Eric Siegel   (16:32)

“There’s this sort of no man’s land between biz and tech, and both sides point to the other to say it’s their responsibility. So the hose isn’t connecting to the faucet.”

— Eric Siegel   (16:06)

“The translation from technical performance, pure predictive performance, to potential business value is generally not done.”

— Eric Siegel   (12:00)

“Predictive AI as a field is failing like crazy. It’s really a crisis.”

— Eric Siegel   (13:40)

“Because if you don’t measure business value, you can’t be pursuing business value.”

— Eric Siegel   (17:26)

“It’s like being more excited about the rocket science than the launch of a rocket.”

— Eric Siegel   (21:50)

“If you don’t get your hands dirty, then your feet will get cold.”

— Eric Siegel   (27:50)

“Operations don’t improve unless they change, and in this case we’re talking about changing them with probabilities.”

— Eric Siegel   (17:50)

 

Resources

Eric Siegel, Gooder AI and the books

Ideas and frameworks discussed

  • Predictive AI: His term for what most people called AI until about three years ago. Learning from data to put odds on a per-case outcome, then using those odds to drive an operational decision. Fraud detection, credit risk scoring, marketing targeting and predictive maintenance are his examples

  • BizML: The six-step business practice for running machine learning projects that The AI Playbook is built around, described across six chapters and aimed at both sides but at the business reader first

  • Evaluate versus valuate: His wordplay, and the practical core of the episode. Evaluating a model asks whether it beats guessing; valuating it asks what it is worth, on the same test data, expressed in business metrics. He puts it as one more step right after evaluation

  • The biz-tech gap: His diagnosis of why projects stall. Technical metrics on one side, business metrics on the other, a no man’s land in between where both sides point at the other, and no incentive system for anyone to bridge it

  • Predictive AI as a guardrail: Layering a predictive model over a generative or agentic system to score which cases are most likely to produce a bad outcome, so the riskiest slice can be routed to a person while the rest runs

  • The 15 and 85 arithmetic: His worked illustration. Send the riskiest 15 percent for human review, keep 85 percent of the promise of autonomy, and compare that to the zero percent delivered by a system nobody is willing to deploy

  • Making the haystack smaller: How he describes what a model does for fraud detection. Not finding the needle, narrowing where you look, and then doing the arithmetic that turns that into a financial win

Named on air

  • Claude, Anthropic: The model behind the specialized chatbot Gooder AI built into the product, which answers questions about the predictive AI project, the product and what the levers and sliders do. The what-if scenarios are the interface itself rather than the chatbot

  • IBM: Named for a study of executives which he says found these projects usually come out even on ROI. The clause is garbled in the transcript, so nothing beyond that is claimed here

  • James Taylor: Named as someone else working on the same gap, coming at it from the business and operationalization side

  • BARC: The research Christina cites for a 19 percent figure on generative AI deployment, which she says three or four other studies support

  • AI Business Value Is Not an Oxymoron: The title of the talk he says he will give at AI Realized Summit, which took place on 5 November 2025

Related AI Realized episodes and events

 

Frequently Asked Questions

 
 
 
 
 
 
 
 
 
 
 
Previous
Previous

Zero to Campaign With Everyday AI, in Four Steps

Next
Next

More Podcast Episodes