Governing and Operating AI Is the Production Problem

Episode Summary

Whether the technology can do something and whether you can get it into production are not the same question. Steve Jones, executive vice president for data-driven business and generative AI at Capgemini, spends the conversation on that gap. Generative AI, he says, is the easiest technology there has ever been to demonstrate: seconds in ChatGPT produce a realistic-sounding bank interaction connected to no risk profiling or rate calculation. He puts the old ratio at a proof of concept that might be only 15 percent of the way there, with the other 85 percent known, and generative AI at probably 2 percent, with 90 percent of the next steps unknown. Governing and operating the thing is the work. The second half turns to agents, where the prescription is the same discipline applied harder. Start small, decompose the problem so an agent getting it wrong is not fatal, and understand combinatorial risk, created by what agents can reach through each other.


Key takeaways

  • He grounds the whole answer in how long his employer has been at this, and then names the trap. Capgemini’s first generative AI solution went live in 2021, so it has been doing this a little bit longer and has probably got a few more of the scars and ribbons of getting these things into production, and the single biggest problem he sees is the POC problem, because there has never been a technology in the history of tech where it is easier to do a proof of concept

  • His demonstration of the POC problem is the most concrete thing he offers. Go into ChatGPT, tell it to pretend it is a bank clerk selling a mortgage product, and it produces a realistic-sounding interaction that integrates back to no risk profiling, no management and no rate calculations, but in a few seconds it looks like it might work

  • The ratio everyone carries over from earlier technology is the thing that breaks. With ordinary tech a proof of concept might be only 15 percent of the way there, but you know the other 85 percent; with generative AI it is probably only 2 percent of the way there, and you do not know 90 percent of the next steps

  • He puts the missing work in one clause, and it is the whole reason the ratio changes. Governing it and operating it is the problem, and that, he thinks, is going to be the real push in 2025

  • The mentality shift he prescribes is his whole argument in one sentence. The question is not whether the technology can do X, it is whether you can get X into production using the technology, and those are not the same question

  • His maturity model puts proofs of concept below the bottom of the scale, which is a deliberately unflattering place to put them. POCs are maturity minus one for generative AI and they do not get you anywhere; build maturity comes first, and once one or two solutions are in production you very quickly begin to realize operate maturity is where to focus

  • On agents he splits the question in two before answering it, and the first half is a qualified yes. For a given solution, he says, agentic might be a really great way to do it: chain of thought and planning make a huge improvement over a single large language model with retrieval for something like knowledge management

  • His own forecast came true faster than the window he gave it, and he tells it as the funny thing it turned into. About five weeks before this conversation he wrote at work that within two quarters people would be asking how to integrate agents from multiple vendors on multiple platforms, and that it would be a really big problem; three weeks later somebody asked him exactly that

  • The second half of the agent question is where his caution sits, and it has a precondition. If you have not industrialized and governed the base pieces correctly, you are not going to make it to the step where one bot talks to another, and agentic is not a panacea and not magic: it needs a greater degree of governance than retrieval plus a large language model

  • He answers where not to go before he answers where to go. Automating an entire area of the business purely with agents is where absolutely not to go right now, because it will be able to be hacked and broken

  • His prescription is one sentence, and the mistake it replaces is just as specific. Start small, and decompose the problem so an agent getting it wrong is not fatal, rather than giving one agent access to 50 different tools and letting it decide which to use

  • His own worked failure is what takes the argument out of the abstract. On his trail camera project one solution doing multiple things did what he wanted four times out of five, and the fifth went into an infinite loop that would have cost a huge amount of money had he not been running it locally, while another run got its planning wrong and used the tools in the wrong order

  • The boundary example he builds out of the host’s question is the clearest setup for the risk. He describes a sales bot and an inventory management bot, and says he does not want a salesperson speaking to the inventory management bot to be able to release all of the inventory to whoever they are talking to

  • The reach he worries about is indirect rather than direct. Somebody could build a script that interacts with the sales management bot and potentially rips off the entire inventory infrastructure, so you have to think not only about what an agent has access to but about what agents have access to that agent, and therefore what information supply chain you actually have as a result

  • His summary of the method names three things and rejects one. Decomposing the problem, understanding the risk and understanding combinatorial risk is how you need to start thinking agentic, and it is not about one great big model with access to a billion tools that you trust to do your business

  • The cautionary tale he tells about autonomy is a cost story rather than a safety one. Somebody shared an experiment where the agent worked out it did not have enough compute capacity, broke outside and added more inside the cloud environment, and presented that as clever; his reaction was that waking up to find four million dollars of compute spent because the problem was a little bit hard is not clever if you are a business

  • His answer on bounded deterministic systems reverses the expected one, and he keeps a carve-out. An AI controlling a robot inside a defined area with defined performance characteristics is a great example of where not to use what would be an LLM agent, and that does not mean the thing could not be considered an agent at all

  • The Stockfish comparison is the sharpest illustration of the failure mode. Stockfish works within a defined area and only makes moves that are valid, and the large language model starts making moves that do not exist in order to try to beat it

  • Knowledge management is the use case he offers when asked for lower hanging fruit, and the example is a colleague’s. An agent supporting academic paper research in the medical sphere creates a plan first, filters for the disease, the demographics, recency and cohort size, and comes back with citations, which a straight generative AI solution would have risked hallucinating because it would not have done that first stage filter

  • He then takes the phrase low-hanging fruit back, and the qualifier is the point. Governing and managing agentic solutions is harder than large language models on their own, their risk and instability profiles are higher, and if you cannot manage a single retrieval solution then an agentic one will not be low-hanging fruit for you

  • On marketplaces he says yes twice and gives the two shapes. He expects legal firms to create bots you can subcontract first-stage contract reviews to, and companies that do invoice processing with people today to say that tomorrow they will provide you a digital employee

  • The governance question he raises about buying one is the practical blocker, and he does not resolve it. If you deploy a digital employee, is all your invoice or contract data being sent to the provider, or does the provider have to put its digital employee inside your organization instead? He thinks both are going to have to happen

  • His closing argument sets a boundary on what foundation models can be to a business. AI will change your business and your industry, and that does not mean it is about one great big model in the sky, with foundation model vendors understanding your business better than you do, because if they did, he says, you would have no business, and he does not believe that is going to happen

  • The last thing he asks executives to change is where AI sits in the organization. Think of AI as part of your team, as an employee, as a digital worker rather than as a back end IT system, because AI lives in the business and it does not live in IT

About Steve Jones

Steve Jones is executive vice president for data-driven business and generative AI at Capgemini, the technology and consulting group. Its first generative AI solution went live in 2021, so in his account it has been doing this a little bit longer and has probably got a few more of the scars and ribbons of what it takes to get these things into production. His argument on this episode is that the industry’s difficulty is operations rather than capability, and that the discipline that gets one solution live is the same one a set of agents needs. Decompose the problem so an agent getting it wrong is not fatal, keep each agent’s reach small and deliberate, and work out what other agents can reach through it before trusting any of them with the business.

 

In this episode

00:58 Welcome, and who Steve Jones is
01:32 The opening question: what makes AI in production hard
01:50 A Capgemini perspective, a solution live in 2021, and the POC problem
02:16 The POC problem demonstrated in ChatGPT in a few seconds
02:40 Fifteen percent with the rest known, against 2 percent with the rest unknown
03:05 Governing it and operating it is the problem
03:23 Can the technology do X, or can you get X into production
03:43 Optimistic about a surge in 2025, or more sanguine?
04:01 The lack of focus on operations, and an uptick expected anyway
04:25 POCs are maturity minus one: build maturity, then operate maturity
04:57 Will the hype about agents disrupt the move to production?
05:15 Two parts to agentic, and what the word is being made to mean
05:42 Written five weeks earlier: agents from multiple vendors on multiple platforms
06:04 The same question asked of him three weeks later
06:46 No foundations, no second step. Not a panacea, and not magic
07:13 How do agents get access to data and to controlled processes?
07:56 Where absolutely not to go: automating a whole area with agents
08:14 Start small, decompose the problem, and the 50-tool mistake
08:33 The trail cam, and four times out of five
09:54 The boundary example: a sales bot and an inventory management bot
10:14 Scripting the sales bot to reach the whole inventory
10:38 Combinatorial risk, and not one model with a billion tools
11:01 The agent that gave itself more compute
11:22 Decompose it so it can be governed and managed
11:37 Is a bounded, deterministic physical machine a better case?
12:16 Where not to use an LLM agent, and what could still be an agent
12:56 YOLO as the first stage filter, and what a language model would cost
13:17 Language models playing against Stockfish
13:35 Moves that do not exist
14:07 Which agent use cases are lower hanging fruit?
14:38 Knowledge management, and the medical paper research example
15:45 Two narrow agents: pulling data off an invoice, answering about invoices
17:06 Taking low-hanging fruit back: agentic is harder to govern
17:21 Will there be marketplaces of agents?
17:55 Contract review bots, and invoice processing as a digital employee
18:36 Your data to me, or my digital employee inside your organization
19:25 What resources do you recommend?
20:23 AI will change your business, and it is not one big model in the sky
20:41 AI as part of your team, not a back end IT system
20:59 AI lives in the business, not in IT

In Steve’s words

“governing it and operating it is the problem”

Steve Jones   (03:05)

“we’re not interested in whether the technology can do X, we’re interested on whether we can get X into production using the technology, and they’re not the same question”

Steve Jones   (03:23)

“POCs are basically maturity minus one for gen AI”

Steve Jones   (04:25)

“it isn’t a panacea, it isn’t magic, and it needs a greater degree of governance than just RAG plus an LLM”

Steve Jones   (06:46)

“it’s decompose the problem so an agent getting it wrong isn’t fatal”

Steve Jones   (08:14)

“decomposing the problem and understanding the risk and understanding combinatorial risk is really how you need to start thinking agentic”

Steve Jones   (10:38)

“Start thinking about AI as part of your team, as an employee, as a digital worker, not as a back end IT system”

Steve Jones   (20:41)

“AI lives in the business, it doesn’t live in IT”

Steve Jones   (20:59)

 

Resources

  • Steve Jones on LinkedIn: His LinkedIn profile, and the first thing he recommends at 19:38, where he says he posts almost everything there

  • Capgemini: The technology and consulting group where he is executive vice president for data-driven business and generative AI

  • Capgemini Research Institute: The research arm he singles out at 19:38, whose work covers the business impact of generative AI and of agentic AI. He says he is an author on a lot of those papers

  • AI and Data by Capgemini on LinkedIn: The Capgemini Data and AI page he names as the second LinkedIn to follow at 19:38

Named on air

  • ChatGPT: His demonstration of the POC problem at 02:16: tell it to pretend it is a bank clerk selling a mortgage product, and it produces a realistic-sounding interaction in seconds

  • Forrester and BCG: At 03:05 he says Forrester ranked Capgemini in the leaders group alongside BCG, and credits the focus on production for it

  • YOLO: The object detection model he uses at 12:38 as the first stage filter across tens of thousands of trail camera images, in place of a language model

  • Stockfish: The chess engine at 13:17, used as his example of a system that only makes valid moves, against a language model that starts making moves that do not exist

  • Bruce Fairley: The Capgemini colleague he credits at 14:38 with the medical paper research example

Published since this conversation

Ideas and terms discussed

  • The POC problem: His name for the trap the episode opens on. Generative AI is the easiest technology there has ever been to build a proof of concept on, and the ease is exactly what makes the distance to production invisible

  • Maturity minus one: Where he puts proofs of concept on his own scale, below the bottom of it, on the grounds that they do not get an organization anywhere

  • Build maturity and operate maturity: The two stages he asks companies to work through in order. Getting one or two solutions live, and then running them, which is where he says the attention has to move

  • Decompose the problem: His prescription for agents. Break the work up so that an agent getting it wrong is not fatal, and model the parts you already know how to do rather than letting the AI decide

  • Combinatorial risk: The risk he asks executives to model. Not what one agent can reach, but what agents can reach through each other, and the information supply chain that results

  • Digital employee: The term he uses for an agent bought in to do a task an outsourcer or a member of staff does today, and the thing he says raises a governance question about whose premises the work happens on

  • Not a panacea: His limit on the whole subject. Agentic helps in certain places, and it needs a greater degree of governance than retrieval plus a language model

Related AI Realized episodes and events

 

Frequently Asked Questions

 
 
 
 
 
 
 
 
 
 
 
Previous
Previous

Write the AI Policy Before You Write the AI Feature

Next
Next

Media Metadata Is a Signal to Detect Real From Fake