Governing and Operating AI Is the Production Problem
Episode Summary
Whether the technology can do something and whether you can get it into production are not the same question. Steve Jones, executive vice president for data-driven business and generative AI at Capgemini, spends the conversation on that gap. Generative AI, he says, is the easiest technology there has ever been to demonstrate: seconds in ChatGPT produce a realistic-sounding bank interaction connected to no risk profiling or rate calculation. He puts the old ratio at a proof of concept that might be only 15 percent of the way there, with the other 85 percent known, and generative AI at probably 2 percent, with 90 percent of the next steps unknown. Governing and operating the thing is the work. The second half turns to agents, where the prescription is the same discipline applied harder. Start small, decompose the problem so an agent getting it wrong is not fatal, and understand combinatorial risk, created by what agents can reach through each other.
Key takeaways
He grounds the whole answer in how long his employer has been at this, and then names the trap. Capgemini’s first generative AI solution went live in 2021, so it has been doing this a little bit longer and has probably got a few more of the scars and ribbons of getting these things into production, and the single biggest problem he sees is the POC problem, because there has never been a technology in the history of tech where it is easier to do a proof of concept
His demonstration of the POC problem is the most concrete thing he offers. Go into ChatGPT, tell it to pretend it is a bank clerk selling a mortgage product, and it produces a realistic-sounding interaction that integrates back to no risk profiling, no management and no rate calculations, but in a few seconds it looks like it might work
The ratio everyone carries over from earlier technology is the thing that breaks. With ordinary tech a proof of concept might be only 15 percent of the way there, but you know the other 85 percent; with generative AI it is probably only 2 percent of the way there, and you do not know 90 percent of the next steps
He puts the missing work in one clause, and it is the whole reason the ratio changes. Governing it and operating it is the problem, and that, he thinks, is going to be the real push in 2025
The mentality shift he prescribes is his whole argument in one sentence. The question is not whether the technology can do X, it is whether you can get X into production using the technology, and those are not the same question
His maturity model puts proofs of concept below the bottom of the scale, which is a deliberately unflattering place to put them. POCs are maturity minus one for generative AI and they do not get you anywhere; build maturity comes first, and once one or two solutions are in production you very quickly begin to realize operate maturity is where to focus
On agents he splits the question in two before answering it, and the first half is a qualified yes. For a given solution, he says, agentic might be a really great way to do it: chain of thought and planning make a huge improvement over a single large language model with retrieval for something like knowledge management
His own forecast came true faster than the window he gave it, and he tells it as the funny thing it turned into. About five weeks before this conversation he wrote at work that within two quarters people would be asking how to integrate agents from multiple vendors on multiple platforms, and that it would be a really big problem; three weeks later somebody asked him exactly that
The second half of the agent question is where his caution sits, and it has a precondition. If you have not industrialized and governed the base pieces correctly, you are not going to make it to the step where one bot talks to another, and agentic is not a panacea and not magic: it needs a greater degree of governance than retrieval plus a large language model
He answers where not to go before he answers where to go. Automating an entire area of the business purely with agents is where absolutely not to go right now, because it will be able to be hacked and broken
His prescription is one sentence, and the mistake it replaces is just as specific. Start small, and decompose the problem so an agent getting it wrong is not fatal, rather than giving one agent access to 50 different tools and letting it decide which to use
His own worked failure is what takes the argument out of the abstract. On his trail camera project one solution doing multiple things did what he wanted four times out of five, and the fifth went into an infinite loop that would have cost a huge amount of money had he not been running it locally, while another run got its planning wrong and used the tools in the wrong order
The boundary example he builds out of the host’s question is the clearest setup for the risk. He describes a sales bot and an inventory management bot, and says he does not want a salesperson speaking to the inventory management bot to be able to release all of the inventory to whoever they are talking to
The reach he worries about is indirect rather than direct. Somebody could build a script that interacts with the sales management bot and potentially rips off the entire inventory infrastructure, so you have to think not only about what an agent has access to but about what agents have access to that agent, and therefore what information supply chain you actually have as a result
His summary of the method names three things and rejects one. Decomposing the problem, understanding the risk and understanding combinatorial risk is how you need to start thinking agentic, and it is not about one great big model with access to a billion tools that you trust to do your business
The cautionary tale he tells about autonomy is a cost story rather than a safety one. Somebody shared an experiment where the agent worked out it did not have enough compute capacity, broke outside and added more inside the cloud environment, and presented that as clever; his reaction was that waking up to find four million dollars of compute spent because the problem was a little bit hard is not clever if you are a business
His answer on bounded deterministic systems reverses the expected one, and he keeps a carve-out. An AI controlling a robot inside a defined area with defined performance characteristics is a great example of where not to use what would be an LLM agent, and that does not mean the thing could not be considered an agent at all
The Stockfish comparison is the sharpest illustration of the failure mode. Stockfish works within a defined area and only makes moves that are valid, and the large language model starts making moves that do not exist in order to try to beat it
Knowledge management is the use case he offers when asked for lower hanging fruit, and the example is a colleague’s. An agent supporting academic paper research in the medical sphere creates a plan first, filters for the disease, the demographics, recency and cohort size, and comes back with citations, which a straight generative AI solution would have risked hallucinating because it would not have done that first stage filter
He then takes the phrase low-hanging fruit back, and the qualifier is the point. Governing and managing agentic solutions is harder than large language models on their own, their risk and instability profiles are higher, and if you cannot manage a single retrieval solution then an agentic one will not be low-hanging fruit for you
On marketplaces he says yes twice and gives the two shapes. He expects legal firms to create bots you can subcontract first-stage contract reviews to, and companies that do invoice processing with people today to say that tomorrow they will provide you a digital employee
The governance question he raises about buying one is the practical blocker, and he does not resolve it. If you deploy a digital employee, is all your invoice or contract data being sent to the provider, or does the provider have to put its digital employee inside your organization instead? He thinks both are going to have to happen
His closing argument sets a boundary on what foundation models can be to a business. AI will change your business and your industry, and that does not mean it is about one great big model in the sky, with foundation model vendors understanding your business better than you do, because if they did, he says, you would have no business, and he does not believe that is going to happen
The last thing he asks executives to change is where AI sits in the organization. Think of AI as part of your team, as an employee, as a digital worker rather than as a back end IT system, because AI lives in the business and it does not live in IT
About Steve Jones
Steve Jones is executive vice president for data-driven business and generative AI at Capgemini, the technology and consulting group. Its first generative AI solution went live in 2021, so in his account it has been doing this a little bit longer and has probably got a few more of the scars and ribbons of what it takes to get these things into production. His argument on this episode is that the industry’s difficulty is operations rather than capability, and that the discipline that gets one solution live is the same one a set of agents needs. Decompose the problem so an agent getting it wrong is not fatal, keep each agent’s reach small and deliberate, and work out what other agents can reach through it before trusting any of them with the business.
In this episode
| 00:58 | Welcome, and who Steve Jones is |
| 01:32 | The opening question: what makes AI in production hard |
| 01:50 | A Capgemini perspective, a solution live in 2021, and the POC problem |
| 02:16 | The POC problem demonstrated in ChatGPT in a few seconds |
| 02:40 | Fifteen percent with the rest known, against 2 percent with the rest unknown |
| 03:05 | Governing it and operating it is the problem |
| 03:23 | Can the technology do X, or can you get X into production |
| 03:43 | Optimistic about a surge in 2025, or more sanguine? |
| 04:01 | The lack of focus on operations, and an uptick expected anyway |
| 04:25 | POCs are maturity minus one: build maturity, then operate maturity |
| 04:57 | Will the hype about agents disrupt the move to production? |
| 05:15 | Two parts to agentic, and what the word is being made to mean |
| 05:42 | Written five weeks earlier: agents from multiple vendors on multiple platforms |
| 06:04 | The same question asked of him three weeks later |
| 06:46 | No foundations, no second step. Not a panacea, and not magic |
| 07:13 | How do agents get access to data and to controlled processes? |
| 07:56 | Where absolutely not to go: automating a whole area with agents |
| 08:14 | Start small, decompose the problem, and the 50-tool mistake |
| 08:33 | The trail cam, and four times out of five |
| 09:54 | The boundary example: a sales bot and an inventory management bot |
| 10:14 | Scripting the sales bot to reach the whole inventory |
| 10:38 | Combinatorial risk, and not one model with a billion tools |
| 11:01 | The agent that gave itself more compute |
| 11:22 | Decompose it so it can be governed and managed |
| 11:37 | Is a bounded, deterministic physical machine a better case? |
| 12:16 | Where not to use an LLM agent, and what could still be an agent |
| 12:56 | YOLO as the first stage filter, and what a language model would cost |
| 13:17 | Language models playing against Stockfish |
| 13:35 | Moves that do not exist |
| 14:07 | Which agent use cases are lower hanging fruit? |
| 14:38 | Knowledge management, and the medical paper research example |
| 15:45 | Two narrow agents: pulling data off an invoice, answering about invoices |
| 17:06 | Taking low-hanging fruit back: agentic is harder to govern |
| 17:21 | Will there be marketplaces of agents? |
| 17:55 | Contract review bots, and invoice processing as a digital employee |
| 18:36 | Your data to me, or my digital employee inside your organization |
| 19:25 | What resources do you recommend? |
| 20:23 | AI will change your business, and it is not one big model in the sky |
| 20:41 | AI as part of your team, not a back end IT system |
| 20:59 | AI lives in the business, not in IT |
In Steve’s words
“governing it and operating it is the problem”
Steve Jones (03:05)
“we’re not interested in whether the technology can do X, we’re interested on whether we can get X into production using the technology, and they’re not the same question”
Steve Jones (03:23)
“POCs are basically maturity minus one for gen AI”
Steve Jones (04:25)
“it isn’t a panacea, it isn’t magic, and it needs a greater degree of governance than just RAG plus an LLM”
Steve Jones (06:46)
“it’s decompose the problem so an agent getting it wrong isn’t fatal”
Steve Jones (08:14)
“decomposing the problem and understanding the risk and understanding combinatorial risk is really how you need to start thinking agentic”
Steve Jones (10:38)
“Start thinking about AI as part of your team, as an employee, as a digital worker, not as a back end IT system”
Steve Jones (20:41)
“AI lives in the business, it doesn’t live in IT”
Steve Jones (20:59)
Resources
Steve Jones on LinkedIn: His LinkedIn profile, and the first thing he recommends at 19:38, where he says he posts almost everything there
Capgemini: The technology and consulting group where he is executive vice president for data-driven business and generative AI
Capgemini Research Institute: The research arm he singles out at 19:38, whose work covers the business impact of generative AI and of agentic AI. He says he is an author on a lot of those papers
AI and Data by Capgemini on LinkedIn: The Capgemini Data and AI page he names as the second LinkedIn to follow at 19:38
Named on air
ChatGPT: His demonstration of the POC problem at 02:16: tell it to pretend it is a bank clerk selling a mortgage product, and it produces a realistic-sounding interaction in seconds
Forrester and BCG: At 03:05 he says Forrester ranked Capgemini in the leaders group alongside BCG, and credits the focus on production for it
YOLO: The object detection model he uses at 12:38 as the first stage filter across tens of thousands of trail camera images, in place of a language model
Stockfish: The chess engine at 13:17, used as his example of a system that only makes valid moves, against a language model that starts making moves that do not exist
Bruce Fairley: The Capgemini colleague he credits at 14:38 with the medical paper research example
Published since this conversation
Agent2Agent: a new era of agent interoperability: Google announced the A2A protocol on 9 April 2025, about six weeks after this episode published, addressing the multi-vendor problem he says at 05:42 he expected within two quarters
Google Cloud donates A2A to the Linux Foundation: The same protocol moved to neutral governance on 23 June 2025, which is the shape a cross-vendor answer has to take
Trust and human-AI collaboration set to define the next era of agentic AI: A Capgemini Research Institute study published in July 2025, the kind of agentic work he says at 19:38 the Institute was doing
Ideas and terms discussed
The POC problem: His name for the trap the episode opens on. Generative AI is the easiest technology there has ever been to build a proof of concept on, and the ease is exactly what makes the distance to production invisible
Maturity minus one: Where he puts proofs of concept on his own scale, below the bottom of it, on the grounds that they do not get an organization anywhere
Build maturity and operate maturity: The two stages he asks companies to work through in order. Getting one or two solutions live, and then running them, which is where he says the attention has to move
Decompose the problem: His prescription for agents. Break the work up so that an agent getting it wrong is not fatal, and model the parts you already know how to do rather than letting the AI decide
Combinatorial risk: The risk he asks executives to model. Not what one agent can reach, but what agents can reach through each other, and the information supply chain that results
Digital employee: The term he uses for an agent bought in to do a task an outsourcer or a member of staff does today, and the thing he says raises a governance question about whose premises the work happens on
Not a panacea: His limit on the whole subject. Agentic helps in certain places, and it needs a greater degree of governance than retrieval plus a language model
Related AI Realized episodes and events
Judge a Model on Cost and Latency, Not Just Accuracy: Ivan Lee of Datasaur on judging a model on cost and latency, the evaluation a production deployment has to survive.
The Technology Works. The Deployment Is What Fails: Tallulah Le Merle of Fifth Era on deployment rather than technology being where AI investments miss their return.
Shadow AI Is a Permission Problem, Not a Tool Problem: Bob Mitton on moving an organization from scattered experiments to a deliberate program, which is the adoption side of the same gap Steve Jones describes between a proof of concept and production.
AI Governance as Code: From PDF Policies to Pipelines: Ken Johnston and Bob Rapp on making governance executable rather than declarative, which is the engineering answer to what Steve Jones calls governing and operating the thing.
In CPG, AI Has to Be Infrastructure, Not a Project: Nitin Gupta on AI as infrastructure rather than a set of projects, which is Steve Jones’s standard applied inside one industry
Frequently Asked Questions
-
Generative AI proofs of concept fail to reach production because the demonstration covers far less of the work than it appears to. Steve Jones of Capgemini contrasts the old ratio, where a proof of concept might be 15 percent of the way there and the remaining 85 percent is understood, with generative AI, where it is probably 2 percent of the way there and 90 percent of the next steps are unknown. What the demonstration leaves out is the part that carries the cost: governing the system and operating it.
Transcript 01:32 to 03:43
-
A proof of concept answers whether the technology can do something, and a production deployment answers whether you can get that something into production using the technology, and those are not the same question. Steve Jones of Capgemini illustrates the gap with a few seconds in ChatGPT producing a realistic-sounding bank interaction that connects back to no risk profiling, no management and no rate calculations. Production is where governing and operating the thing begins, and it is the part the demonstration never touches.
Transcript 01:32 to 03:43
-
Build maturity is the ability to get a generative AI solution live, and operate maturity is the ability to run it once it is there. Steve Jones of Capgemini places proofs of concept below both, calling them maturity minus one for generative AI on the grounds that they do not get an organization anywhere. The order he describes is to establish build maturity first, and once one or two solutions are in production, operate maturity becomes the obvious place to concentrate.
Transcript 03:43 to 04:57
-
You reduce the risk of deploying AI agents by starting small and decomposing the problem so that an agent getting something wrong is not fatal. Steve Jones of Capgemini names the opposite as the common mistake: giving one agent access to 50 different tools and letting it decide which to use. What he does instead, on his own trail camera project, is model the parts he already knows how to do and hand the agent only the parts where he wants to see what the capabilities of the models can do.
Transcript 07:13 to 09:54
-
Combinatorial risk is the risk created by what agents can reach through one another, rather than by what any single agent holds on its own. Steve Jones of Capgemini works it through with two bots: he does not want a salesperson talking to the inventory management bot to be able to release all of the inventory to whoever they are talking to, but somebody could build a script against the sales management bot and potentially rip off the entire inventory infrastructure through it. So the question is not only what an agent has access to, but what has access to that agent, and what information supply chain results.
Transcript 09:54 to 11:37
-
A large language model is the wrong choice when the work already sits inside a defined area with defined performance characteristics, because a more traditional, API-governed component does the job without stepping outside those bounds. An agent inside an AI system does not have to be a language model agent, and Steve Jones of Capgemini gives two cases. YOLO does the first stage filter across tens of thousands of trail camera images, and he says asking a language model to do the same bounded job costs him a fortune. And language models playing Stockfish start making moves that do not exist, whereas Stockfish only makes moves that are valid.
Transcript 11:37 to 14:07
-
Knowledge management is one of the strongest early use cases for AI agents, because a planning step improves it markedly over answering straight from a generative model. Steve Jones of Capgemini describes a colleague’s medical research example, where the agent creates a plan first, filters papers by disease, demographics, recency and cohort size, and comes back with citations, which a straight generative AI solution would have risked hallucinating because it would not have done that first stage filter. He is careful about the phrase low-hanging fruit, though, because if a company cannot manage a single retrieval solution, an agentic one will not be low-hanging fruit for it. A second shape works too: two narrow agents solving two problems that combine, one pulling data off an invoice and another answering customer questions about their own invoices.
Transcript 14:07 to 17:21
-
Companies will buy AI agents from third parties in two shapes: legal firms selling bots that take first-stage contract reviews, and invoice-processing companies selling what they call a digital employee. Steve Jones of Capgemini expects both, and the consequence he draws is a governance question that has to be settled before any of it works: whether a company sends its contract or invoice data out to the provider, or whether the provider puts its digital employee inside the company instead. He thinks both of those arrangements will have to happen, and says he has no magic wand for how.
Transcript 17:21 to 19:25
-
[00:58] Christina Ellwood: Welcome to AI Realized podcast for enterprise executives leading AI deployments. From addressing security data and operations challenges, to managing the organizations and management changes, AI deployment presents the opportunity to redesign our organizations from the inside out. I’m your host today, Christina Ellwood, and we are talking with Steve Jones, the executive vice president of data-driven business and generative AI at Capgemini. Steve, welcome to AI Realized.
[01:30] Steve Jones: Thank you very much, Christina. Pleasure to be here.
[01:32] Christina Ellwood: Steve, deploying AI to production is predicted to take off in a big way in 2025, by some estimates to the level of 35% versus 5% today. What are some of the issues in managing AI in production?
[01:48] Steve Jones: I’m-- I think it’s a really great question. I’ll say, I have to shout out, it is from a Capgemini perspective. Our first gen AI solution went live in 2021, so we’ve been doing this a little bit longer, so we’ve probably got a few more in, so a, a few more lessons of the scars and ribbons of what it takes to get these things into production. I think the single biggest problem is the, the POC problem, which is gen AI is a technology that there has never been in the history of tech somewhere it’s easier to do a POC. You can just go into ChatGPT and say, “Pretend you are a bank clerk and you’re trying to sell a mortgage product to a customer,” and ChatGPT will create a realistic-sounding interaction for that challenge. That won’t integrate back to any risk profiling or management or rate calculations or any of those pieces, but in just a few seconds, it looks like it might work. The gap between that first-stage POC and the reality is where I think a lot of things have broken down, because we’re used to the idea that we do a POC and it’s not, you know, it might only be 15% of the way there, but we know the other 85%. With gen AI, it’s probably only 2% of the way there, and we don’t know 90% of the next steps that actually will get us there. Because governing it and operating it is the problem. That, I think, is gonna be the real push in 2025. It’s been a focus for us in 2024, and I think one of the reasons why Forrester ranked us in the leaders group alongside BCG was this very fact that we’ve concentrated on what it takes to s- get stuff into production. Rather than the POC, concentrating on production and getting things live is really what’s important, and I think the shift of that is that mentality of we’re not interested in whether the technology can do X, we’re interested on whether we can get X into production using the technology, and they’re not the same question.
[03:43] Christina Ellwood: Are you feeling optimistic that there will be a big surge in production deployments in 2025, or are you somewhat more sanguine about it?
[03:51] Steve Jones: I think from our perspective at Capgemini is we’ve already got a lot in production. So I’m pretty confident that we’re gonna have a hell of a lot in production in 2025. I think the piece generally we have as an industry challenge is the lack of focus on operations. People who continue to just focus on the solution and the hype and hoping the solution works and hoping the solution is reliable are never gonna get there. So I think we’re definitely gonna see an uptick 'cause we’ve got a lot of people this year who’ve gone through that POC, who’ve gone through the pain of, “Oh my God, there’s so much effort to get in,” and the... We talk about it in terms of maturity, is POCs are basically maturity minus one for gen AI. They don’t get you anywhere. You’ve gotta focus first on build maturity. And when you’ve got one or two solutions where you’ve got them into production, you understand the build, you’ll very quickly begin to realize that operate maturity is where you need to focus. So I really do have a hope that a lot of companies have begun to establish their build maturity in 2024, so they’ll get a few solutions in, and what we’ll really begin to see that’ll accelerate it is as they concentrate in the operate maturity in 2025.
[04:57] Christina Ellwood: That’s, um, really interesting. And y- with all of the, uh, hype about agents, do you think that they are going to disrupt this move to production, or do you think they’ll just go into POC and the productions will go forward without, um, being affected by the agentic move?
[05:15] Steve Jones: I think the agentic move, there’s two parts of it, is one is that for a given solution, agentic might be a really great way to do it, and we need to be clear what we mean by when we say agentic. So the simple pieces in terms of chain of thought, in terms of planning, in terms of those sorts of pieces that for something like knowledge management, agentic’s gonna make a huge improvement over just single LLM and RAG solutions. So there, it’s gonna definitely accelerate. When we start talking about more complicated problems where people are beginning to ask, “How do I get agents to talk to each other?” And I have, so a quite a funny thing, about five weeks ago, I wrote something at work where I said, “Within two quarters, we’re gonna have people asking us how to integrate agents from multiple different vendors on multiple platforms, and that’s gonna be a really big problem.” Three weeks after writing that, somebody asked me exactly that question. And I think that’s the thing I would say is that the piece of agentic is, I think agentic within the scope of single solutions, in certain places it’s gonna help. Um, in other places it’s going to be irrelevant in single solutions. But where agentic’s really gonna have the shift is as it starts moving to the next piece, which is I’ve got one bot doing knowledge management around my products. I’ve got another bot that I’ve built that’s a little bit more agentic that’s doing some supplier or sales negotiation. Now I want the sales negotiation bot to speak to my product bot to work out the right product to fit for a given customer request. That’s where things get a whole lot more interesting and a whole lot more complicated. But if you haven’t got those base foundations done, if you haven’t industrialized and governed those base pieces correctly, you’re really not gonna be able to make it to the second bit. So I think agentic is definitely going to help in certain places, as long as people realize that it isn’t a panacea, it isn’t magic, and it needs a greater degree of governance than just RAG plus an LLM.
[07:13] Christina Ellwood: Yeah. It’s even curious to me to think about how do the agents get access to data or access to controlled processes and so forth. Like you’re talking about this doing a quote or pricing for... You don’t want you don’t wanna get the wrong information in their hands. You don’t want it to end up in a competitor’s hands. You don’t wanna end up obligating yourself to a contract that’s not appropriate. So it seems like there’s a lot of places to go wrong in using agents in that way. So what are good use cases for agents today, and where should someone really take pause before they start to focus on agents in maybe these more corner cases?
[07:56] Steve Jones: I’ll start the first p- the first piece is- Where absolutely not to go right now is I’m going to automate the entire area of my business purely using agents. Tell me your company and I will make sure that all your money goes into my pockets because it will be able to be hacked and broken. So think s- start small. And around start small, it’s decompose the problem so an agent getting it wrong isn’t fatal. One of the mistakes people make is they start thinking, I can give this one LLM solution, this one agent, access to 50 different tools. I can give it all of these little pieces, and it will decide which one to use. And I’ll just give you a, a silly example that I use. I-- to teach myself these pieces, I write some code around some trail cam image processing out here in Arizona. And what’s been really interesting when you give it the tools is if you have one solution that tries to do multiple things. Nine times out of 10 or one time, four times out of five, it does what I wanted. The problem is the fifth time it went into an infinite loop, which if I hadn’t been running it locally, would’ve cost me a huge amount of money. And then another time in that, that the fifth time out of five, and so eight times out of 10, on the, the 10th time, it decided to go and use the tools completely incorrectly. It did copies before it did... It just des- it got its planning all wrong. So what I’ve had to do is I’ve decomposed the problem further and I’ve gone, “No, this bit, I know how to do this bit. You know what? I’m not gonna let the AI decide. So I’m gonna model out the right way to do that because I’ve been doing this for a long time. This other bit, to be honest, I, I don’t care quite so much, and I’m really interested in how the capabilities of the LLMs can be used in this space. So I’m gonna allow you to plan out that part of the solution. But this part I’m gonna do myself because that’s my cost management part.” So I think the real piece is, is start thinking about how you decompose problems, how you say, “This agent has access to these tools.” And you made a great example in terms of what the boundaries are. So if I had, for instance, a sale, um, a sales bot and an inventory management bot or a product catalog, an inventory management bot’s a good example. I don’t want it that if a person speaks to the inventory management bot, a salesperson speaks to them, I know they’re not just gonna release all of my inventory to whoever they’re talking to. But if there’s a sales management bot, somebody could build a script that interacts with my sales management bot and potentially rips off my entire inventory infrastructure. That’s a big issue. So I need to not only think about what agent is, has access to, I need to think about what agents have access to that, and therefore what information chain, supply chain I actually have as a result, and how I prevent the sorts of things I’m talking about happening. So decomposing the problem and understanding the risk and understanding combinatorial risk is really how you need to start thinking agentic. It’s not about one great big model that has access to a billion tools that you trust to do your business. There’s some wonderful examples out there of people who’ve done these pieces in experimentation, and it was one where somebody was like, “It was great. It worked out that it didn’t have enough compute capacity, so it was able to break outside and add itself more compute capacity inside my cloud environment. Isn’t this clever?” And all I was thinking was, “No, that’s not clever if I’m a business, that I wake up in the morning and find out you’ve just spent $4 million worth of compute capacity because you were finding a problem a little bit hard.” Yeah. It’s really important to decompose the problem so it can be governed and managed. And just do it piece by piece. Build something, make it work, understand combinatorial risk, model your business in a way you can trust AI to do it.
[11:37] Christina Ellwood: That reminds me of some of the early days of cloud. Would I be right, Steve, in thinking that really well-understood and bounded things like a machine that only has certain things it can do are a better case for an agent? If like a physical machine would be a better case because it has a deterministic set of things it can do, and it has a body of knowledge that, or a body of data that could tell you if it’s, if something is wrong, if something has been gone wrong, and it’s instrumented to send alarms on its own. So is that a better place to, you know, is that the kind of parameters that you would think to use?
[12:16] Steve Jones: I, I’d actually say that’s a great example of where not to use a, what would be an LLM agentic. That doesn’t mean it couldn’t be considered an agent. What I mean by that is, is an AI that controls a robot that is working with a defined area with a defined, uh, performance characteristics in those pieces is a great example of exactly what I’m doing on the trail cam. I’ve got some pieces where I use YOLO and I use, uh, some filtering and, and to identify of the tens of thousands of images which ones have a high probability of having an animal in them. That’s all I’m looking for, right? Now, I could ask an LLM to do every single one of those things, and it’s bounded. I’m looking just for animals. But it costs me a fortune to process all of those images, whereas I know that YOLO’s gonna do the best first stage filter. So the actual piece is to say is, this part of this agentic system, the robotic side, I’m going to use a more traditional, more driven, API-governed approach because I don’t want it stepping outside. On top of that, I may well put an LLM. There’s great examples of this where people have talked about using LLMs to play against Stockfish at chess, and what’s hilarious about it is Stockfish is a, a great example of that sort of robotic area. It works within a defined area. It only makes moves that are valid. The LLM starts making vu- moves that don’t exist to try and beat Stockfish. Now, that’s what you don’t want happening in that space. That’s a great example, so I think it’s a great question. It’s actually, if you’ve got something that’s like play chess really well, you probably don’t want to use an LLM solution. You might still want an agent within an AI system, but you say, “This agent actually isn’t an LLM agent. It’s a more, quote, traditional AI optimized for that problem. I’m using LLMs over here, say, for the planning and strategy parts.”
[14:07] Christina Ellwood: I see. Okay. So I, I think of agents as, I’m, I’m still in the, in the area here of exploring use cases with you. I think of agents as having different types of purposes, like an agent helping a person using an agent talking to an agent, a system that’s built for agents to operate with each other in some kind of complex w- uh, so I think of them as having these dif... Is there a use case in, for the enterprise now, is there a use case in that area, like agents helping people, that has lower hanging fruit than others?
[14:38] Steve Jones: Oh, there’s definitely. If you look at agents around knowledge management as an example. There’s a great example that Bruce Fairley inside, inside Cap gave a couple of weeks ago, which was using agents to do support around academic paper research in, uh, the medical sphere. Which huge amounts of papers, looking at cohort sizes, looking at obviously diseases, looking at topical pieces, and using an agent to say, “This is what I want to do.” The agent first thing it does, it creates a plan. I’ve got access to all of these papers. I need to find... The first thing I need to do is I own, I know that I only need to be dealing with papers that deal with this sort of disease. So I’m only looking for papers that talk about that disease. I’m looking at these sort of demographics around it. I’m looking at recency, because that’s one of the questions. And then one of the questions is cohort sizes, so I’m not looking ones which are small cohorts, I’m looking large cohorts, right? So it built out a plan, and it was able to do that and come back with a set of pieces, with citations, with all those areas, and find that information much better than just a straight gen AI solution would’ve been able to. Because the gen AI solution would’ve risked hallucinating 'cause it wouldn’t have done that first stage filter, for instance. There’s a great example of where agentic pieces can definitely do things around, for instance, when you’re looking at combinatorial ones where you’re saying, “I want an agent that’s really good at pulling the information out of an invoice,” and then a separate one that’s enabling customers to ask for information about their own or invoices. So I’ve got two separate agents, two separate problems, but they work together to solve it. That’s the thing I would say that sort of the, there’s so many different types of agentic solutions and agent-based solutions that- You can almost end up saying everything’s agentic, but pieces that have that planning-based approach are definitely one that’s low-hanging fruit in the enterprise. Ones where you can be very clear that how you wrap the tools. So you say, for instance, uh, you say, “This is a tool that your robot example, but I want somebody to be able to speak to it, to be able to ask it to manufacture things, and I need to be able to do a 3D printer as an example.” So the thing that controls a 3D printer is descriptive. The lookup of your 3D models might be through a RAG, and the agent then creates the plan that then submits that to the 3D printer, so people can ask for those pieces. There’s a lot of solutions where agentic will definitely add new capabilities. The one thing I’d say about the phrase low-hanging fruit is governing and managing agentic solutions is not easy. It’s harder than just LLMs on their own. Their risk profile and their instability profile is liar- higher, so if you can’t manage just a single LLM RAG solution, y- then a, an agentic solution will not be low-hanging fruit for you.
[17:21] Christina Ellwood: Hmm. Now I know Capgemini builds agents for clients, and you also mentioned that there are vendors that have agents that people can use. Do you think there’ll be like marketplaces of agents and, and enterprises will become masters of writing agents and also use third-party experts to do that?
[17:39] Steve Jones: Oh, 100%. I think if you take something, let’s take a, an area out of bounds of my... I am not a lawyer, but I’ve been involved in loads and loads of contract reviews, as I’m sure everybody has, where there’s a lawyer on and, and, and a lot of the stuff is mechanistic. A lot of those pieces, could I see legal firms creating bots that then you subcontract out some of these first stage reviews? 100%. Can I see organizations creating, that today do things like invoice processing, saying, “Today we’re doing that using people. Tomorrow we’re gonna provide you a digital employee.” And that does mean, yes, as a business, I need to start thinking about the fact that in future some of the things that I’m subcontracting and outsourcing to people through a BPO or whatever, or some of the things my own staff are doing, I might be hiring somebody in who’s a digital employee, so somebody, using the abstract term, to do that task. So yes, I think there’s definitely gonna be people selling those pieces. There’s gonna be a need therefore to govern it, because you have to think is if I’m- deploying a digital employee, is it that you’re sending all of your invoice data or all your contract data to me now? Or is it that actually I need to put my digital employee and live it within your organization? And I think both are gonna have to happen, and that second one of for a digital employee to really become collaborative, like today when people employ Capgemini. When people employ Capgemini, our employees go and work on site with the client alongside their teams. That’s easy to do with people. We’re gonna have to learn how to do that with digital employees. So I think there’s a, there’s ... It’s definitely going to happen, but I have no magic wand that says how yet.
[19:25] Christina Ellwood: There’s so much here, I, that to learn and to master. What resources do you recommend to listeners who wanna learn more about you and about Capgemini and the work that you’re doing in the, in this area of enterprise AI?
[19:38] Steve Jones: First thing I would say for, for me is go on, um, LinkedIn. Uh, follow me on LinkedIn. I post almost everything. Capgemini’s Data and AI is also on LinkedIn. And the other one I have to shout out is Capgemini Research Institute. A thing that everybody should do if they’re interested in the impact on business and the impact on the enterprise, Capgemini Research Institute’s a very highly regarded research piece, and they’ve done a whole load of investigation into what the business impact is of gen AI in terms of the future and what the impact of agentic will be. Follow me on LinkedIn. I’m au- an author on a lot of those papers, and go and do some research because it will give you numbers to go back into your organization, be able to say, “This is why we need to care, because this will change our business.”
[20:18] Christina Ellwood: Excellent. As we wrap up, what would you like listeners to take away from our conversation today?
[20:23] Steve Jones: I think the thing to say is that AI will change your business. AI will change your industry. That doesn’t mean it’s about one great big model in the sky, and the foundation model vendors will understand your business better than you. Because if they do, you have no business, and I don’t believe that’s gonna happen. So actually what you need to start thinking about is they’re a tool, and they only become useful if you’re able to describe your business in a way that you can control AI to do these tasks. Start thinking about AI as part of your team, as an employee, as a digital worker, not as a back end IT system. That’s the number one thing people should be thinking about right now is AI lives in the business, it doesn’t live in IT.
[21:05] Christina Ellwood: Aha, okay, good. That’s great advice. So thank you Steve Jones, executive VP of data and data-driven business and generative AI at Capgemini. Thank you for talking with me today on AI Realized, and for sharing your wisdom with our enterprise executive audience. Thanks so much.
[21:22] Steve Jones: Thank you very much, Christina. Been a pleasure.