Only Content That Clears Every Agent Gets Monetized
Episode Summary
Fandom carries 50 million pages of content across 250,000 communities for 350 million monthly users, and Adil Ajmal, its chief technology officer, says that scale changes how you decide what to build. Agents do two different jobs. Helix, a targeting platform, segments fans by what they are feeling rather than who they are. A separate orchestrated set of narrow agents checks every edit for policy safety and then for brand suitability, because a page can be perfectly safe and still not be the best one for a particular brand. Then the conversation turns to money. Asked how the calculation changes now that tokens are a cost, he says the framework itself is not different; it is just a question of figuring out the right cost, which means a proof of concept every time and a recalculation whenever a new model arrives. What the agents decide is what can be sold: only content that clears all of them is monetized.
Key takeaways
Publishers are losing search traffic to AI answers, and he credits authoritative content for where Fandom sits. He says a lot of publishers are seeing their traffic go down because people now ask a question of AI instead of doing a regular search, and that Fandom has been in a very interesting position from that perspective because it has authoritative content in its space
Fandom is the fourth most referenced site in Google’s AI search results. He cites a report he says came out the month before the conversation, on Google’s new search AI experience, and says Fandom is referenced by the AI search engines way more than most other places because its content is so deep and so structured
Being a branded property helps to an extent, and depth is what earns the citations. Asked whether owning words like Star Wars is what lifts the authority, he says it is the depth of content and actually being able to get accurate information for those questions that really helps elevate it
Helix targets on what a fan is feeling rather than on who they are. He describes their targeting platform creating audience insights and segments for advertisers that are way deeper than demographic or social data: the emotions people may be experiencing at that point based on the shows or the storylines or where they are in a particular story, combined with viewing habits, consumption habits, and what game they are playing and where they are in it
Brand safety is somewhat solved and brand suitability is not. Detecting policy violations, hate speech, violence and discrimination he calls somewhat of a solved problem, because it is somewhat easier to detect. Suitability is very different and starts becoming complicated, because most brands do not want their ads showing up in content that may be related to terrorism or murder or crime
Entertainment breaks the category rule that suitability filtering runs on. In the entertainment and gaming space, he says, you are going to have legitimate content sitting in exactly that space, and he takes a James Bond movie as the example
Suitability decides what gets monetized, not just what gets published. He describes tiered monetization: a page can be perfectly safe from a policy perspective and still not be the best page to show a particular brand’s ad on, and that judgment is made by the agentic platform they built for content moderation and brand suitability analysis
Helix targeting moved four brand metrics, on the numbers he reports: a 50 percent lift in brand awareness, a 16 percent increase in brand consideration, about a 22 percent boost in preference, and a 72 percent surge in purchase intent
The platform is many narrow agents behind an orchestration layer, and the simpler decision runs first. Asked whether they use lots of very small agents each doing narrow things, he says that is correct, and starts with the policy check on an edit, which he calls the simpler task and which stops the page being saved at all
Only content that clears every agent gets monetized. He describes the sequence ending with a last agentic piece that checks brand suitability, and says content that does not pass ends up going for human review. That share keeps shrinking, he says, because the agents get more training data as humans look at the small segments that get flagged
Translation is a pretty solved problem, except for invented worlds. He says translation is pretty solved if you think about it, and then that translating content focused on virtual worlds is a much more complicated problem, because content built on real world or historical things translates a lot more accurately
If the translation is not what users call the thing, the opportunity is gone. He says that if your translated content is not exactly what the users are looking for or what the users are calling it, then you have missed the opportunity for the user and the user is not actually going to get it
Depending on the type of feature, there is a success metric and a cost-at-scale number before it is built. He says they are always going to have a success metric and an idea of what it is going to cost at scale, because given their scale they cannot just randomly scale stuff and it would end up being very expensive
AI does not change the business case, it changes one input. Asked how the calculation differs now that tokens and consumption are a new variable, he says in all honesty the framework itself is not different, and that it is just a question of figuring out the right cost
Nothing gets built to scale without a proof of concept, and the trial itself is costed first. He says they do not necessarily build anything to scale right off the bat, that there is always going to be a proof of concept and that is what gives them the cost, and that even the trial is going to cost something so they work that out before running it
Cost they are pretty good at predicting. Engagement and revenue are the harder half. He credits the data science team for the cost side and says the budgeting part is easier, and that what has a much bigger variable is what the engagement and the revenue are going to be, because too many other variables come into play
Every new model means recalculating the whole thing. He says they are very thorough when they run those proofs of concept, checking whether the assumptions are actually holding, and that as new models come in they reanalyze and recalculate all of it rather than sticking with what it was before
The expensive models are for starting something, not for running it. They reclassify the Helix data on an ongoing basis, and he says that doing that on commercial models would be ridiculously expensive for them. They still use commercial models for certain aspects, including starting it out, and run older models they have trained for explicit purposes very cheaply on the regular tasks
The feature that surprised him hallucinated once it was scaled, and the users caught it. He says it started hallucinating and generating content which was not outright incorrect but had nuanced things in it that could be very offensive, and that it got caught by their users very quickly, because at their scale so many eyeballs are on everything that you find out about things immediately
The fix was letting the community edit what the AI wrote. They stopped the feature, then changed the process so communities and power users could edit the AI-generated content. Users were then really excited about it because they still felt in control and could correct mistakes, and their corrections ended up teaching the model
Telling people to use AI does not go anywhere. He says it is not easy unless you have a plan for how to do it, how to encourage adoption and how to measure the results, and that if you just tell people to use AI a few will, but you are not measuring results and not seeing how it is actually transforming anything
Their AI policy classifies the data, not the tools. The secure instinct is to allow nothing corporate security has not approved, he says, but that is very restrictive because a small team cannot evaluate every single tool. So they wrote data classifications instead: named authorized tools for particular types of corporate data, under contracts stating the data will not be used to train the vendors’ models
The adoption target is 80 percent utilization, and it is tied to named outcome metrics. He says the technology organization had a goal the previous year and set another this year for 80 percent utilization, with metrics naming the needles that utilization is meant to move
His line to the creator community was enablement, not replacement. A lot of people were afraid Fandom would start generating content with AI and replace its creators, he says, and at a user conference two years before this recording his tagline was enablement, not replacement: the point is to let creators make better content, faster, and different types of content, not to put them out of the picture
His advice to an executive earlier in the journey is to empower rather than prescribe. He says that if you basically just prescribe tools to people it is usually not the best recipe for success, because it is very hard to prescribe a tool for each function and each work process across an org
About Adil Ajmal
Adil Ajmal is chief technology officer at Fandom, the fan platform for entertainment and gaming that carries 50 million pages of content across 250,000 communities for 350 million monthly users. He leads AI strategy for the company and runs its board working group on AI, formed after the first ChatGPT model reached the market because the board wanted to understand it. His career runs from startups through big tech into media and entertainment: earlier in it he was at Homestead, which he says was the 18th largest site on the internet at the time, and his last company before Fandom was in consumer lending. He started programming in sixth grade, and describes his own philosophy as simplifying a problem down to the one fundamental thing that has to be solved. He is a sci-fi and fantasy fan, which is part of why he took the job.
In this episode
| 00:42 | Welcome, and the guest introduction: chief technology officer at Fandom, 350 million monthly users |
| 01:26 | The through line from startups to big tech to Fandom |
| 04:07 | Simplify the problem down to its core |
| 06:08 | Changing user behavior at scale, and where the users come from |
| 08:58 | Publishers losing traffic to AI answers, and what authoritative content changes |
| 09:18 | Fourth most referenced site in Google’s AI search results |
| 09:47 | Above sites larger than Fandom, because the content is structured |
| 10:21 | Depth of content and accurate answers, not the brand |
| 10:35 | Hence the word authoritative, and on to the agentic use |
| 10:56 | Helix, the targeting platform, and the launches he cannot discuss yet |
| 11:22 | Targeting on emotion, storyline and position in a game |
| 12:10 | Segments like heroes and survival, very different from demos |
| 12:44 | Content safety first: policy terms, hate speech, violence, discrimination |
| 13:09 | Brand safety is somewhat solved, brand suitability is not |
| 13:28 | Legitimate entertainment content that sits in the risky category |
| 13:36 | A villain, murders, bomb blasts, and brands that advertise anyway |
| 13:49 | 50 million pages of unstructured content, converted to structured |
| 14:40 | Tiered monetization, and pages that are safe but not suitable |
| 15:06 | Human review for a very small percentage of pages |
| 15:59 | Measuring the impact, and the predictive AI engine |
| 16:22 | The numbers: awareness, consideration, preference, purchase intent |
| 16:47 | Launched in the market the year before, and still evolving |
| 17:53 | A tiered approach, transparency, and Grand Theft Auto |
| 18:36 | Advertisers picking categories, and insights on every bid request |
| 19:11 | Lots of small agents with narrow jobs, coordinated |
| 19:22 | The orchestration layer, and the policy check first |
| 19:38 | A policy violation stops the page being saved at all |
| 19:58 | Then the text of the whole page, because one edit can change its context |
| 20:32 | Only content that clears every agent is monetized |
| 21:03 | The whole system evolving, behind a single orchestration layer |
| 21:42 | Translation is a pretty solved thing, except for virtual worlds |
| 22:35 | Shogun, and why the translation worked |
| 22:54 | Mechanically spot on, and still wrong locally |
| 23:14 | If it is not what users call it, the opportunity is gone |
| 23:41 | Custom glossaries for virtual worlds |
| 23:53 | Why a glossary per world per language does not scale |
| 24:10 | Sounds like a job for an agent |
| 24:15 | An agentic system, and where the human part still comes in |
| 24:30 | Power users and native speakers training the agents and models |
| 25:04 | Privacy, and what they explicitly do not do today |
| 26:57 | A success metric and a cost at scale, before building anything |
| 27:31 | Engagement is not directly tied to revenue |
| 28:12 | The cost of translating content into another language |
| 29:20 | The token and consumption cost as the new variable |
| 29:58 | The framework itself is not different |
| 30:03 | Always a proof of concept, and what it tells you |
| 30:22 | What to tokenize, and what each API call costs |
| 31:10 | Cost they predict well, engagement and revenue they do not |
| 32:12 | The price of advertising moving while the cost stays put |
| 32:31 | Jeff at Amazon on inputs and outputs |
| 33:20 | Discipline, and never assuming the next model costs the same |
| 34:04 | Recalculating as new models arrive |
| 34:24 | The work of staying on top of it |
| 34:34 | Gemini models on the Vertex platform for translation |
| 34:46 | OpenAI elsewhere, and open source models trained on their own data |
| 35:05 | Reclassifying the Helix data on an ongoing basis |
| 35:23 | Commercial models would be ridiculously expensive at that volume |
| 35:50 | Disciplined operators, and a private equity-backed company |
| 36:39 | Optimizing for the outcome first, then for cost |
| 37:01 | Time to market, and being truly hybrid |
| 37:58 | The one compliance case, and where the data can sit |
| 38:42 | Extracting the common questions and their answers |
| 39:08 | Roughly ten wikis, then about a hundred communities, with human review |
| 39:26 | 250,000 communities, and why human review stops scaling |
| 39:49 | AI hallucinates a lot, and what verification cost at the time |
| 40:09 | Not outright incorrect, but nuance that could be very offensive |
| 40:36 | The implied nuances were absolutely incorrect, and it was taken down immediately |
| 40:55 | Letting communities and power users edit what the AI wrote |
| 41:22 | Users felt in control, and their corrections taught the model |
| 41:38 | Found on a Friday morning, pulled back immediately |
| 42:18 | The person who pulled the feature informed the exec team |
| 42:43 | The cultural through line since the first AI Realized Summit |
| 43:38 | Deliberate and thoughtful about adoption |
| 43:49 | Telling people to use AI does not go anywhere |
| 44:24 | Why approving every tool is too restrictive |
| 44:43 | Data classifications, authorized tools, and the training clause |
| 45:03 | A culture of bringing tools in rather than keeping them out |
| 45:23 | OpenAI without an enterprise license, and Gemini with one |
| 46:07 | Copilot, mixed results, and not letting one bad first try end it |
| 46:26 | Evangelists and champions per tool and per function |
| 46:55 | An 80 percent utilization goal |
| 47:15 | Product quality, the defect rate, and forming the AI committee |
| 48:00 | Leading the AI strategy, and a CEO who uses it |
| 48:24 | The board working group on AI |
| 48:45 | Enablement, not replacement |
| 49:13 | One day past the thousandth day of ChatGPT |
| 49:51 | Goals alone did not move anything |
| 50:37 | A structured bake-off between coding agents |
| 51:09 | People going in and out of the group |
| 51:24 | Cross-functional across technology and product |
| 52:17 | Embrace it, and be an evangelist yourself |
| 53:02 | Prescribing tools is not the recipe for success |
| 53:26 | Empower, encourage goals, then hold people accountable |
| 54:50 | Legal copy review, financial analysis, and coding |
| 55:11 | Where AI-generated code does not hold up at scale |
| 55:50 | Employees figuring out what value they add |
| 56:30 | Encourage training, but not everybody changes |
| 57:14 | Reasonably private, and the one writer he reads |
| 57:41 | The AI Realized podcast, newsletter and conferences |
| 58:39 | Curiosity and continuous learning as the leadership trait |
| 59:40 | Do not be afraid of AI, and learn to use it to your advantage |
| 59:52 | Wrap-up |
In Adil’s words
“And only content that passes all of these different agents and these different decision points is monetized on our platform. Otherwise, it actually ends up going for human review.”
— Adil Ajmal (20:32)
“Brand suitability, which is very different than just the safety of the content itself.”
— Adil Ajmal (13:09)
Bbecause our content is so deep and so structured, we actually get, uh, referenced by all the AI search engines and the searches way more than most other places.”
— Adil Ajmal (09:18)
“Given our scale, we can’t just randomly scale stuff. It would end up being very expensive.”
— Adil Ajmal (26:57)
“The framework itself is not different. It’s just a question of, you know, of figuring out the right cost.”
— Adil Ajmal (29:58)
“As new models come in, we reanalyze, uh, you know, and recalculate all of this. We, we don’t just stick with what it was before.”
— Adil Ajmal (34:04)
“It started hallucinating and generating some content which wasn’t, like, outright incorrect, but it had nuanced things in it that could be very offensive.”
— Adil Ajmal (40:09)
“If you just tell people to use AI, uh, it doesn’t really go anywhere.”
— Adil Ajmal (43:49)
“My tagline was enablement, not replacement.”
— Adil Ajmal (48:45)
“If you basically just prescribe tools, uh, to people, it’s usually not the best, uh, recipe for success.”
— Adil Ajmal (53:02)
Resources
Adil Ajmal and Fandom
Adil Ajmal on LinkedIn: His profile. Chief technology officer at Fandom, where he leads the company’s AI strategy
Fandom: The platform whose scale sets the terms of this conversation: 250,000 communities and 50 million pages of content, serving 350 million monthly users
FanDNA Helix: Fandom’s announcement of the targeting platform he calls Helix at 10:56 and 15:59. Dated 13 November 2024, which matches his description of launching it in the market the year before this recording
Fandom Debuts Helix to Unlock Unlikely Audiences: Adweek on the Helix launch. It independently reports the 50 million pages of content he cites, and quotes Fandom’s chief revenue officer on reaching audiences advertisers had not considered
Momentum: The predictive Helix offering Fandom announced on 11 August 2025, roughly two weeks before this conversation was recorded. He does not name it on air, where the product is only ever Helix
Ideas and terms discussed
Brand safety and brand suitability: Two different questions, and he separates them carefully. Safety is whether content violates policy: hate speech, violence, discrimination. He calls that somewhat of a solved problem because it is comparatively easy to detect. Suitability is whether a particular brand wants its advertising beside that content, and in entertainment it is the hard one, because a James Bond film is legitimately full of the things a category filter blocks and most brands will happily advertise with it anyway
Tiered monetization: What the suitability judgment feeds. Content is sorted into buckets, and a page can be perfectly safe from a policy perspective and still not be the best page for a particular brand’s ad. The decision is about what gets monetized rather than about what gets published
The orchestration layer: How the agents are arranged. The simpler check runs first: a policy violation on an edit stops the page being saved at all. An uploaded image or video takes longer and is decided separately. Then the text of the whole page, because one small edit can change the context of the page. Then a last agent for brand suitability. Only content that clears every one of those decision points is monetized, and what fails goes to human review
Helix: The targeting platform, and the reason older self-trained models do the ongoing reclassification. It builds audience insights and segments that go deeper than demographic or social data, reaching for the emotion a fan may be experiencing based on the show, the storyline or where they are in a game. Its full public name is FanDNA Helix; on air he says only Helix. The numbers he gives for it are a 50 percent lift in brand awareness, a 16 percent increase in consideration, about a 22 percent boost in preference and a 72 percent surge in purchase intent
The report behind the fourth-place ranking: What he cites at 09:18 when he says Fandom is the fourth most referenced site in Google’s AI search results. He dates it to the month before the conversation and gives neither its title nor its publisher
Custom glossaries for virtual worlds: Their answer to the translation problem, and one he says does not scale. A mechanically perfect translation of a game community can still fail, because the name that evolved in the local language is not the mechanical translation, and if the content is not what users are calling the thing the opportunity is gone. The AI has no training data for an invented world, so they hand-build a glossary per world, which he says is not easy or that scalable to do for every language
Costing the proof of concept: The step that carries his whole cost argument. Nothing is built to scale right off the bat; there is always a proof of concept, that is what produces the cost number, and the proof of concept itself is costed before it runs. As new models arrive the numbers are reanalyzed and recalculated rather than carried forward
Data classification in an AI use policy: Their alternative to approving tools one at a time, which he says is too restrictive for a small security team. The policy names authorized tools for particular classes of corporate data, under contracts stating the data will not be used to train the vendors’ models, and defines how much leeway people have to bring other tools in. If a tool someone brought in was working, the next step was asking security how to make it part of the enterprise offering
80 percent utilization: The adoption target the technology organization set, alongside metrics naming the outcomes that utilization is meant to move, including development velocity and the defect rate. He says they set three metrics around it, and that a deliberate change management process led up to it
The AI committee: Cross-functional across technology and product with somebody from content on it, formed once it was clear that setting goals alone did not move adoption. Some members stepped up because they were already doing more with AI and some were recruited per function. It reports to him as executive sponsor rather than as a hardline reporting line, membership is not fixed, and its current work is a structured bake-off between coding agents run with volunteers and a measurement window
Enablement, not replacement: His tagline at a user conference two years before this recording, given to a creator community that was afraid Fandom would generate content with AI and replace them. The point of using AI for customers, he says, is to let them create better content, faster, and different types of content
Named on air
Ethan Mollick, One Useful Thing: He names Mollick on air, not the newsletter. Asked at 57:06 what listeners should look at to learn more about him, he says he is reasonably private and names Mollick instead, as the writer whose pieces give you insight into how he thinks. Mollick is a professor at the Wharton School
Perkins Miller: The chief executive he credits at 48:00 as a huge user of AI from the first day and as the person who has pushed the company on it. He says only the first name on air
Google Gemini, on the Vertex platform: What they use for a lot of the translation work. He says Gemini has come a long way and that Fandom had an enterprise license for it from the start, because Google was pushing it
OpenAI: Used for other tasks, and the tool he still rates as one of the best. Fandom originally had no enterprise license for it and now does, which is what let them roll it out to a large amount of the company with their own data in it
Copilot: What he calls it on air, and what is publicly GitHub Copilot. His example of a tool that came out fast with mixed results, and the reason they keep re-checking whether something has improved since the last time it failed for a given use case
Shogun: The FX series, the older film and the books. His example of a translation that came out great, because so much written artifact exists for it
Grand Theft Auto: His worked example of content advertisers want and also find risky. He says it was going to be the largest game launch of 2025 and has been pushed to 2026
James Bond and Mission Impossible: The films he uses to show why category filtering fails in entertainment: a villain, murders, bomb blasts, and most brands happy to advertise alongside all of it
Jeff at Amazon: Quoted at 32:31: you can control your inputs but not all your outputs, so be very good about your inputs. He gives only the first name on air
Homestead: The company early in his career that he says was the 18th largest site on the internet at the time
Related AI Realized episodes and events
Detect Intent, Then Tailor Every Screen to the Person: Al Shanmugam of EchoStar on intent detection and personalization at consumer scale.
Valuate the Model, Don’t Just Evaluate It: Eric Siegel of Gooder AI on valuating a model rather than just evaluating it, which is the measurement question behind monetized content.
In CPG, AI Has to Be Infrastructure, Not a Project: Nitin Gupta of Mondelēz International on AI as infrastructure inside a consumer brand.
Stop Sending Employees to IT. Send Service to Them.: Lenin Gali of Atomicwork on handling employee service requests at volume, and on replacing mean time to resolution with ticket deflection.
Frequently Asked Questions
-
Content moderation at scale works by breaking the judgment into narrow agents behind an orchestration layer and running the simplest check first. Adil Ajmal, chief technology officer at Fandom, describes the platform doing exactly that: an edit is checked for a policy violation, which he calls the simpler task and which stops the page being saved at all; an uploaded image or video takes slightly longer and is decided separately; then the text of the whole page is read, because a small edit can change the page’s context; and a last agent judges brand suitability. Only content that passes every one of those decision points is monetized, and what fails goes to human review, which he says covers a very, very small percentage of pages, and that share keeps shrinking because the agents get more training data as humans look at the segments that were flagged.
-
Brand safety is whether content violates policy, and brand suitability is whether a particular brand wants its advertising next to that content. Adil Ajmal calls the safety question somewhat of a solved problem, because policy violations, hate speech, violence and discrimination are comparatively easy to detect. Suitability is the complicated one: most brands do not want their ads appearing beside content that may be related to terrorism or murder or crime, but in entertainment and gaming that content is legitimate. His example is a James Bond film, which has a villain, murders and bomb blasts, and which most brands will happily advertise with.
-
When an AI feature at Fandom started generating offensive content, Fandom’s own users found it and an employee pulled the feature the same day, informing the executive team afterwards rather than asking permission first. Adil Ajmal describes an AI feature that extracted commonly asked questions and their answers. It was proved out on roughly ten wikis and then about a hundred communities with heavy human review, and it began hallucinating once it was scaled with more confidence in the model: content that was not outright incorrect on its face, but whose implied nuances could be very offensive and which he says were absolutely incorrect. The community caught it very quickly, which he puts down to having so many eyeballs on everything at their scale. The fix was to change the process so communities and power users could edit the AI-generated content, and their corrections went on to teach the model.
-
Fandom is referenced heavily by AI search engines because its content is deep and structured, which its chief technology officer says counts for more than being a branded property, though the branded part does help to an extent. Adil Ajmal cites a report he says came out the month before the conversation, which put Fandom fourth among the sites Google references in its AI search results, above other sites larger than it, and says the article notes Fandom is hitting above its weight. Asked whether owning branded words like Star Wars is what does it, he answers that it is the depth of content and being able to get accurate information for those questions that really helps elevate it. He notes that a lot of publishers are seeing traffic fall as people ask a question of AI instead of running a regular search, and puts Fandom in a very interesting position from that perspective, because it has authoritative content in its space.
-
AI translation fails on games and fictional worlds because the model has no training data about the world it is translating. Adil Ajmal says translation is otherwise a pretty solved problem, and that content built on real world or historical subjects translates a lot more accurately: Fandom’s Shogun community came out great because so much written artifact exists for it. A game was different. The translation was spot on from a mechanical perspective, but the way that game is referred to in the local language had evolved differently, and if the translated content is not what users are actually calling the thing, he says, the opportunity for the user is gone. Their workaround is a hand-built glossary per virtual world, which he says is not easy or that scalable to do for every language.
-
You predict what AI will cost at scale by pricing a proof of concept before you build anything to scale, then recalculating when the model changes. Adil Ajmal says the framework for justifying an AI feature is not different from any other feature, and that it is just a question of figuring out the right cost. They do not necessarily build anything to scale right off the bat: there is always a proof of concept, that is what gives them the cost, and even the proof of concept is costed first. Because they handle large amounts of content constantly, his team has become very good at working out exactly what needs tokenizing, how much, and what each API call will cost. As new models come in they reanalyze and recalculate rather than sticking with the earlier numbers, and he is blunt that assuming a new model will cost the same as the last one puts you in a world of trouble.
-
An AI use policy stays workable by classifying the data rather than the tools, so people know what they can put where without every tool needing individual approval. Adil Ajmal says the secure instinct is to allow nothing corporate security has not approved, but that a small security team cannot evaluate every single tool, which makes that policy very restrictive. Fandom’s policy instead sets data classification categories: named authorized tools for particular types of corporate data, contracts stating that the data will not be used to train the vendors’ models, and defined leeway for bringing other tools in. The culture around it is deliberately inviting rather than restrictive, because they wanted people bringing in tools that made their work better, and a tool that was working became a conversation with security about making it part of the enterprise offering.
-
Employees adopt AI tools when there is a utilization target tied to outcome metrics and a champion inside each function, rather than a set of tools prescribed from the top. Adil Ajmal says that if you just tell people to use AI it does not really go anywhere, because a few will and nobody is measuring anything. Fandom’s technology organization set a goal of 80 percent utilization and three metrics naming what that utilization was meant to move, including development velocity and the defect rate. They deliberately found evangelists and champions for different tools across engineering, content, HR and accounting, so those people could train others in their teams, and formed a cross-functional AI committee that now runs structured bake-offs between tools with a set number of volunteers and a measurement window. His caution for anyone starting is that prescribing a tool for each function and each work process across an org is usually not the recipe for success.
-
[00:00] Christina Ellwood: Welcome to AI Realized, the podcast for enterprise executives leading AI adoption. From tackling security, data, and operational challenges to navigating organizational transformation, AI deployment offers a unique opportunity to redesign our organizations from the inside out. I'm Christina Ellwood, your host for today's episode, and today I have two guests with us who have spent decades deploying AI at companies like Microsoft, Ford, IBM, GE Healthcare, and General Motors. And they are now channeling that experience into something new. Ken Johnston is the VP of AI and data at Envorso and the founder of AiGovOps Foundation. Bob Rapp is a principal AI architect at General Motors AI Center and is the co-founder of the same organization. Together, they're making the case that AI governance needs to work like engineering: automated, testable, and embedded in the deployment pipeline, not filed away in a PDF. AI Realized recently co-hosted Governance After Hours in San Francisco with the AiGovOps Foundation, bringing together executives and startups working on AI governance. And the energy in that room made it clear this is a conversation the industry is ready for Ken, you spent 25 years at Microsoft, then ran a Ford subsidiary serving 20 million connected vehicles, and now you're writing the Lean AI Handbook. Bob, you're, you've architected AI systems at IBM Watson, GE Healthcare, Vodafone, and now GM. What was the specific moment each of you realized that AI governance just simply couldn't stay in the policy document anymore, that it had to become code?
[01:45] Ken Johnston: I guess I'll go ahead and go first on this. For me, it started actually with GDPR. And so heading into that, I actually had been working a lot with Microsoft data, was one of the leading people at generating insights off of data. But GDPR was one of the first standards that came out and said that you needed to ask permission to use data, that those permissions were granted sometimes for a limited time and for specific use. Prior to that, there was just a general sense of do no evil. When GDPR came out, I spent a lot of time with our lawyers at Microsoft, learning and sharing with them how data could be used and abused, and what we needed to do to lock it up inside of Microsoft. At that exact same moment is when Microsoft was working on responsible AI standards as we moved into this AI era. So I moved directly from privacy into data governance, into into AI governance. And for me, when I saw how we had to automate for privacy and for data governance, when I left Microsoft and I went to Ford Motor Company and I saw the lack of automation we had there and how many things were locked up in documents and spreadsheets, I realized that I needed to do a major push to bring to bring that company and autonomic up to speed on modern practices. Just like we did with privacy, but now with AI, the threat was much higher and the speed was much faster, and so the need was even greater than before, and that's what kind of pushed me in this direction. What about you, Bob?
[03:25] Christina Ellwood: Yeah, that makes sense. And of course, it's only gotten more compli- complex from a regulatory point of view, yeah? And Bob, how about you?
[03:32] Bob Rapp: First let me thank you for inviting us to this podcast. I really enjoyed meeting you and such a wonderful bunch of founders. Boy, the center of the known universe in AI governance is clearly San Francisco. That was a wonderful event. I think for me it was cancer. So as a two-time cancer recoverer both at Watson, where I came to try to fix some of the problems we were having with how we were doing data science in the cancer space with Watson, and also at GE Healthcare, where I'm proud to say we shipped some of the very first models to diagnose certain kinds of cancers. What I noticed is I stopped trusting the decks the day that the cancer models changed, but our controls didn't fire. So we had a really good policy that was stuck in a piece of paper, and it didn't change how the diagnosis and other models were working. So data drift and other things are very dangerous, and we've got to watch that in real time.
[04:27] Christina Ellwood: Yeah. Even when you automate those systems, which of course many people still have not automated, even the most basic thing like website compliance the automation can be out of sync as well if it's not provided the most up-to-date information. And of course, in the world of agents where we have agents talking to agents generating data that no one has ever seen before, it's a much more complicated problem. Bob, you've said compliance frameworks that exist only in PDFs are liabilities waiting to surface. Walk us through what AI governance as code actually looks like and what gets automated, what gets tested, and what changes in the deployment pipeline.
[05:06] Bob Rapp: That's a great question. L-let me give you a couple of short ones, some tests you can do right now. If your policies don't compile, you're not governing AI, you're just hoping. Or another way to say that is that governance as code means that the wonderful policies that you spent a long time building from compliance and legal and all the good people that are doing that, if they don't live in your software repositories right next to the models, if every change doesn't go through the same automated checks for lineage and fairness and risk, AI, CI/CD pipeline all the way, then you're not shipping policy as code. You're hoping and you're praying that you have a plan. But you're not... The policy has no relationship to the product you're shipping. That's the test.
[05:54] Christina Ellwood: Okay, so then does the governance as code, the code that is governing the use of the LLM, sit in the data pipeline or sit in the in the inter- the space between the prompt and the model? Or where does it r- where does it really live?
[06:10] Bob Rapp: It... and Ken's probably as one of my heroes of tests, Ken can talk a little more deeply about this. What I would say is every single gate has a yes or no answer, it has a contract, and it has a test. Just like we've always done with CI/CD pipelines so continuous integration, continuous development. If you don't pass the test, you don't pass the gate. None shall pass, a Monty Python moment there. If you pass the gate, you pass the contract, you get to go to the next place. So it, the policy lives everywhere, instantiated as code, and we check it at every single gate, and it doesn't make it out the door. And it, whether it makes it out the door, whether it's sustained, or whether we have a fail and we recover, we-- those things are immutable. They're rules that always occur, and we never, ever make a flexible movement there. We lock those in, right?
[07:01] Christina Ellwood: Gotcha. Okay. K-
[07:02] Bob Rapp: yeah, Ken- Yeah ... why don't you, why don't you
[07:04] Christina Ellwood: elaborate from a test perspective?
[07:06] Ken Johnston: Yeah, I just wanna add in that the- You, the CI/CD pipeline is really where we focus a lot of our conversation, but as we know, especially with agentic AI, and we're all gonna get together and talk about that in a couple of weeks on the 17th you need to actually have your gates running in runtime as well. And so we do this already in the cloud, and one of the things that we talk about in governance is that a lot of what we're trying to do with governance as code is the same thing we do with services today. We have we have watchdogs in production today for live services. We need those kinds of mechanisms, but they need to be there for your governance policies, and they need to be there at the edges between your agents. They need to be there in place to protect against prompt injection. All of these different things have to be in place. You'll hear Bob and me talk mostly about the CI/CD pipeline because that's where we really focus because most governance activities these days are processes people do after the model has left the barn. And so we're really emphasizing shift left, bring governance back in closer to the development cycle. But you really need to have it also there through your observability layer and at your runtime.
[08:21] Christina Ellwood: So are you envisioning that these governance policies as code will live inside your CI/CD tools? Is that where you think they're gonna belong? So if you're using Rollbar, they're gonna be in Rollbar or something like that? Are you looking to the vendors to develop the mechanism by which we pour our policies into that pipeline?
[08:44] Ken Johnston: Well- Yeah, let me- Take a- Yeah, let me jump in on that, Bob- Okay. Sure ... because that's kinda what we're seeing with the ecosystem and why we started the AiGovOps Foundation, is that there is a lack of technology in the space today. We know this much, it needs to live in the CI/CD pipeline. It needs to live in those tools. You need to have those protections, but honestly, the technology has not yet been developed, and that's one of the things that we're trying to promote. We know that it needs to be automated. There are literally a gap of solutions to allow this to happen. And you were gonna say something, Bob?
[09:20] Bob Rapp: Yeah. So I was gonna say that the purpose of the open source projects that we'll talk about a little bit later is to build the executable code that we can give to the community and they can improve that says, "We're listening for certain patterns. We know about regulations. We're keeping track. We're looking at risk models. We've got dashboards to see where we're doing, and we can get to a compliant state." It's all very early but it does work, and there's a ton of folks in the space that have very similar solutions. And I'm very impressed. We've got a group of founders actually tomorrow that are presenting some of their solutions. So there's a lot of growth, there's a lot of hope, but it's not completely a solved problem, which makes it a really exciting time to, to get in here and fix it.
[10:04] Christina Ellwood: I think it was really wise to use an open source project because it does... It not only marshals the talents of the open source community, but it means that people can adopt this tool without having to look to their vendors as the sole source, right? They can ask their vendors for it. They can explain what they're looking for. They can represent how they're doing it today. But they have something they can use without having to have vendor delivery or approval for additional budget in order to make it happen. It-- Is your vision that this same approach, the same project that you've built for the general question of AI governance in the CI/CD pipeline would be used by agents as well?
[10:48] Bob Rapp: Sure. Sure. We think and Joe at Glacius is one of my heroes, and he's actually presenting tomorrow. What he would say is, what we're looking at is the baseline, the bottom set of standards, and a framework that people can use. So policy people and code people can get together on a common understanding of what the gates and what the rules are. But there's a lot of folks above that line that are delivering more. Agents are certainly the place to deliver. That's the-- Those are the cool kids. Microsoft Build has been announcing all week some amazing agents that are terrifying and wonderful. Agents also, by the way, exist as swarms. So imagine instead of just having one or two agents, you have ten thousand agents, or as I like to call them, viruses with credit cards, unless you've got the right controls, right?
[11:33] Christina Ellwood: Yeah, for sure. That's a great way to think about that, isn't it? So Ken, at Ford, you delivered over-the-air updates to twenty million vehicles. So you have my sympathies for what that must have been like. Having some background in IoT, I can appreciate that. Where a bad deployment could be life-critical. And Bob, you at GM were building AI for connected vehicles in safety-critical environments. So give me one example from automotive specifically where automated governance caught something or where the lack of it caused a real problem Real world case, guys
[12:12] Ken Johnston: Yeah. Gotta be careful that we don't actually share confidential details here. What I can share with you is this, and it's actually true both for Ford and a case study that I read on the autonomous taxi company Zoox. One of the things we were doing at Ford is we were pulling logs from our vehicles whenever the autonomous vehicle engine triggered- Even for a millisecond, a question mark, a this doesn't make sense. And maybe it solved it because clearly an accident didn't happen, but for even a millisecond, the AI wasn't sure what was going on. And I had this wonderful opportunity to be with the team that was analyzing this data, and for a millisecond, And because at Ford we had both LIDAR and we had cameras, but for one millisecond, one of the cameras thought that the world-- that the car was upside down. Jeez. And the reason was that it was a sunny day, and it was driving, and there was a puddle, and in the puddle there was a perfect reflection of the trees on either side of a two-lane road that were going down into the water. Now, because we had multiple systems, of course the system as a whole worked. But it was something that we captured, and because we were always working on how do you prevent these kinds of risks, so that ended up being something that we put into our quality control pipeline so that we can then monitor that in the future to make sure that the model never drifted out of that, that there were always retries. That, just 'cause the camera sees something, LIDAR and motion sensors are more authoritative in the system than a camera that might get tricked by a puddle. So that was a fascinating lesson.
[14:06] Christina Ellwood: Yeah, just that's a really good one. Yeah. Bob, do you have one to add?
[14:10] Bob Rapp: I would say, and I have lots of friends at Ford, lots of friends at GM. I think the auto industry tries really hard to be very safe. One thing I've noticed that I really admire about General Motors is that we treat governance as a safety system. So just like we think about other safety systems, LIDAR, radar, brakes all of those things, not as a slide deck, because the most dangerous control is the one that's stuck in Word or a PDF. The only controls that matter are the things that happen in real time. 'Cause if you think about increasingly self-driving is becoming the thing for all automotive manufacturers, so we could have up to a million transactions in a few seconds, and all of those transactions could have some level of artificial intelligence in them, right? We don't know. But we do at GM, but we-- in the future that could be a thing that would sneak out. So I'm really pleased by how thoughtful the safety engineers are in, in that process of policy as code, right? In the car.
[15:14] Christina Ellwood: And have you seen a case where automated governance- Avoided a problem or the lack of automated governance caused a problem?
[15:23] Bob Rapp: Yeah, I-- bo-both, and I'll talk about the good cases. By using automated governance, the way that we do over-the-air updates has really changed significant in the auto industry. I would say at GM and Ford and Toyota, and we have lots of friends in the industry. But one of the things that we're able to do, and it's really more about how you use artificial intelligence, both deterministic and non-deterministic, to target the right platform. 'Cause if you realize that there are six or seven hundred thousand different variations of just a pickup truck. So that's where automated governance says, "Do we have the right package to the right vehicle at the right time in the right order?" And I've seen that across the industry get better and better, so I'm very hopeful about the future.
[16:07] Christina Ellwood: Gotcha. Okay. That seems like a really practical use case, too, of making sure you've got the right make and model for the push of the update. So at our AI Governance After Hours event in San Francisco that you both came down from San-Seattle for, which I was thrilled to have you there as representatives of AiGovOps Foundation. I saw executives and startup founders in the same room wrestling with the same governance questions that you re- you wrestle with at the enterprise level. So these are-- this is about the technology. It's not about the size of the organization, the type of organization, or even the scale of the deployment. Everyone is facing those same challenges. And you've launched AiGovOps Foundation, what was it? Six months ago. And you're already presenting at AICon USA. You're hosting meetups. You're building with the practitioner community. So Bob, what's the single biggest misconception About AI governance that you're encountering from these various enterprise leaders, and what do you tell them?
[17:07] Bob Rapp: It's the thing that makes me twitch, is that people think that AI governance is a break. What it actually is an accelerator. As some of my friends, regulators in the UK and in Australia say is, if your policy is stuck in a PDF, you can't ship software. But if your policy is shipped in the software as governance, policy as code, that actually accelerates your ability to ship AI. And so I'm actually seeing that at companies like Merck and John Deere and certainly GM, that when we get that we can go faster because we're safe enough to ship. And it reminds me s- sometimes I feel like enterprises could be that Ralph Nader slogan, unsafe at any speed. Because they're, they fill out forms and they've all had a retreat, and they've planned their next retreat, but they don't have any governance that they can execute. They just have governance in documents.
[18:00] Christina Ellwood: It also strikes me though that the degree of testing is really significant because obviously models can hallucinate any time. You don't get to choose whether or not they're hallucinating while you're doing your tests. So how do you know when you have had sufficient testing to be safe to ship your code?
[18:21] Bob Rapp: I'll defer that to Ken, my test hero here.
[18:24] Ken Johnston: So the state of the art in... That's a very good question. The state of the art in this continues to evolve and change rapidly. One of the techniques people are using is using different AI models to test AI models and drive them. We used to do things in testing called boundary value analysis. There's a practice within red teaming now to actually create these kinds of role-based adversarial approaches to try to un- in essence, unnerve your AI to get it off of its game. And so those are some techniques that help harden the system. There is no pat answer right now. And so one of the analogies I often compare this to and why we're so focused on the concept AiGovOps is Bob and I go back to the early days of security, when security, everybody was paying it lip service. We used to do pen testing after we shipped a product. Seriously. It's oh, ship a product, let's see if the bad guys can get inside the fence. That was brilliant. But, with security, we've done the whole shift left thing, and we've developed a lot of new technologies, so we help prevent those kinds of mistakes getting into production. And that literally is a gap we have today with governance and governance automation. We are saying you need to go have governance as code, but we're also saying the tools don't fully exist and the frameworks don't exist correctly to fully automate it in a way that actually protects people. But we're also focused on this because, as Bob already said, governance done right is an accelerator. I'm gonna steal one of Bob's lines. One of the problems you get sometimes with governance is you get all the way done with your project and you get to the maybe gate. You get to the person that's there that says maybe we can ship it, but I'm not ready to sign off." "Maybe this is a good idea, but I don't know that we've checked all the boxes." And you get stuck at maybe. Governance, when it's automated, has your- ha- has your standards embodied in code, and you know that it's done, and you can continue to invest to make it the best state-of-the-art standards as possible as we continue to accelerate the development of the technology. Governance is an accelerator. It gets you past the maybe gate and into production.
[20:33] Christina Ellwood: I wanna do one follow-up question on this issue of testing, because obviously you're testing against a model, and most organizations are now using more than one model, and they may be using more than one version of a given model. How many models do you include when you're doing challenge testing? And what is the magic configuration for te- for the challenging? Do you s- you start with the one you're gonna deploy and then challenge it with one you're not deploying? W- what is it? What's the logic that is used in designing the testing with the challenge models?
[21:05] Bob Rapp: Let me take this one because it's one of my pet peeves, and I talk to an awful lot of people that say we already know how to do all of the CICD and testing and dev test and moving th-" the one difference here is that in non-deterministic or generative AI, if you have a billion transactions, each transaction is by design a little bit different. And so all of the systems that we've had before, whether it's security or virus or whatever it is, starts with we ship some code and now we're looking for anomalies. If I have a billion generative AI transactions, I have a billion generative AI transactions, but I have no notion of what they mean unless I have context. So your question about what do you test it sounds strange to say we test everything, but what I would say is you make sure you listen to all the models that are on your network. You then have files for all of the regulations that you care about, which should be all of them probably, and then you do a determination of how you're sitting from a compliance point of view, and then you decide what N number, N equals 19, N equal 1,000, to take as a sample for immutable logging so that auditors and other folks that are in policy can get a handle on this. And that is what-- That is, by the way, I've just told you what Beacon is, which is our open source project, which is live now. And we also have a Lantern project that goes around Beacon to explain to auditors and folks that are not as technical how to go through that audit process because we want to keep everybody safe. We want humans to ship great AI. We're huge fans of AI- ... when it's good, and we'd like them to not ship bad AI because we don't wanna dig that bunker in the backyard.
[22:47] Christina Ellwood: I think this question of how many models and which models and how many tests and whatever is a kind of maybe gate, right? Because maybe I'm okay with three models, but maybe I'm not sure if three is enough. So I think there's still maybe gates that people are gonna run into and I think your approach of providing the s- the foundation and then layering on top of the foundation is very sound. So Ken, your book, "The Lean AI Handbook," comes out this summer from Pearson. For the executives listening who have 15 AI pilots but only two in production, which I'm sure is more than one, what is the single highest leverage operational change they should make tomorrow morning?
[23:31] Ken Johnston: Thanks for asking that question. And so the Lean AI Handbook is less about, governance and more about just building AI correctly and the good engineering practices. And the number one thing I found in the research is this: everybody in a rush to get to AI seems to have dropped many of the fundamentals of good software engineering, and that's what the book focuses on. And they don't do AI deployments, and we've even seen this with major large companies like xAI, where they don't have fully automated deployments with fully automated rollbacks. They don't have proper instrumentation for observability. And what's unique about AI that's a little bit different than normal services development is what I call the learning loop. Your whole goal in AI is to actually start the learning process. You're not trying to build the perfect model the first time out. You're trying to get a model, an agent, a tool into production with the correct instrumentation for observability, because we know... One thing we, a lot of us know about AI is that data is fuel, and when you get your models into production, if you've got them instru-instrumented correctly, you're now generating more fuel to make your models better. So I really focus on, and why the book is called The Lean AI Handbook, it's really about going back to agile development, incremental small changes, and managing that with blast radius control so that you can roll things back safely. And one thing you do with AI is accidents are going to happen. Do not even imagine for a second that you're gonna do an AI project that is not gonna have to have a rollback at some point, and you're kidding yourself. And that's really the fundamental problem is everybody seems to have just said, "Oh, it's AI, it's magic. Let's stop the fundamentals." And so my book is actually really simple. Fundamentals of engineering apply to AI as well.
[25:27] Christina Ellwood: I would argue fundamentals of data are also really im-important to be reminded of because, as my friend Steve Jones at Capgemini likes to say, people call data the new oil, and it's true in the following sense: that like oil, it has to be refined. You don't take it out of the ground and put it straight in the car. And he said the thing about AI, generative AI in particular, is that you're moving the consumption of the data to the wellhead. And that, of course, is so true, especially when we are using our systems of record as our primary source of data because the systems of records weren't built to generate good data. They were built to do a function and exhaust... The data is the exhaust of that process. So it's not designed to be clean, it's not designed to be normalized, it's not designed to be any of the things that we need it to be to use it in AI. So that is different than the things you're working on in the governance area, but it's related obviously because a big part of what you're learning is from the data that may or may not be well-suited to the process of training. So let's move into our wrap-up section here. What one resource would each of you recommend for listeners who wanna learn more about what you're building? Go ahead, Bob
[26:56] Bob Rapp: I would go to aigovops-foundation.org or aigovops-foundation.com. We've just put those set up too, and look at the two open source projects. The first one's called Beacon. That's really for folks that are writing software. And then I would look at Lantern or Umbrella, we haven't quite decided on that name yet, which is really the structure around how you would deliver that. And big thanks, by the way, to Glacius, Joe, and the team for being our first folks helping us with oversight on that open source project.
[27:31] Christina Ellwood: Gotcha. Okay, that's great. And Bob is there anything about your work that you would want to f- highlight for folks, that they could either read about something you've written or perhaps a presentation you've given?
[27:43] Bob Rapp: So I tend to put everything up either on LinkedIn or YouTube. Recently, it's been LinkedIn, but if you go to our foundation site, you'll see both Ken and I's posts on LinkedIn and other social media. We try to keep that up to date, and I'm excited we'll be recording our first virtual meeting with, meetup with about 200 folks, some in Seattle and then some around the world. We have a speaker from Singapore, and I think one from London tomorrow, so that's really exciting. So-
[28:10] Christina Ellwood: Fantastic ... come look. I will be on, I will be on that call. Ken- Excellent ... what about you? What res-- what one resource would you recommend for listeners who wanna learn more about what you're building?
[28:23] Ken Johnston: Again, LinkedIn and our website. I did create an e-book on AI governance fundamentals. I go through a lot of case studies. A link to that can be found on our website that Bob already mentioned, and that's free to download. Again, we're focused on building a community, not a product, so we're giving a lot of our learnings away. And keep the link fresh because honestly, the book is in its third version since I started working on it four months ago. I keep adding new sections here and there because I keep learning. One of the things I love about the foundation is this: Bob and I have worked in this space for a while, but because we're taking a helping hand approach, the number of people that are approaching us and sharing their knowledge and information, honestly the amount that I am learning about good governance has skyrocketed because of the approach that we're taking. And so we're gathering that knowledge, we're putting it into open source tools, and I'm the one that likes li- long form documents, so I'm putting our learnings into documents.
[29:26] Christina Ellwood: That's fantastic. So Ken, in the AI revolution, what leadership skill do you find most valuable in your work today?
[29:35] Ken Johnston: That's easy. It's clarity. And it's actually executive clarity. One of the challenges we're seeing, and we call it prototype theater or demo theater too often now especially with no code, low code, vibe coding, people can create really flashy and cool looking demos, and everybody gets excited and says, "Hey, let's go put that into production." But as we've discovered, sometimes those flashy demos are nothing more than a wire frame, and you've still got 90 plus percent of the work to do, and you get distracted. You're now thinking you're taking a shortcut, but really you've just derailed your progress. The best AI projects are the ones that align to executive objectives delivered by the board. When you get off track from that, like any other software project, you're gonna end up cutting it and not executing because there's not alignment. So clarity and alignment, those are the things that I really focus on.
[30:32] Christina Ellwood: Great. That's really good advice. And Bob, if listeners remember just one thing from our talk conversation today, what would you like it to be, and why?
[30:41] Bob Rapp: I think the best thing to take away is if you're just gonna remember one thing, let it be this: governance that lives in a PDF doesn't run in production. So if you want to be trustworthy with your AI at scale, your policies have to compile, your controls have to be testable, and your evidence has to be as automated as your deployments. All the great principles Ken just talked about are true, but it's even harder because it's a non-deterministic world. So if you can't ask the question, what do we do when this fails? How do we roll it back? How do we fix it? How do we get it to the previous gate? You're not doing software engineering, you're doing hope, prayer, and probably a lot of theater.
[31:26] Christina Ellwood: Wonderful. So Ken Johnston, the VP of AI and Data at Envorso and founder of AiGovOps Foundation, and Bob Rapp, the Principal AI Architect at General Motors and co-founder of AiGovOps Foundation, thank you both for sharing your experience with us today. It's been a pleasure.
[31:45] Bob Rapp: Thank you very much. Thank you. A pleasure. Let's continue the conversation. Appreciate