Visual AI Lets You Ask Business Questions of Your Video
Episode Summary
Companies have had huge amounts of visual data forever, Alex Gammelgard says, and it has stayed pretty inaccessible: unless someone was watching or tagging it by hand, there was not really a lot software could do with it. It is now becoming machine readable the way text already is, she argues, so the same cameras and the same footage answer business questions the company could not ask before. She works through where it applies: offshore pipe inspection where most of the footage shows nothing happening, a sports organization with eighty years of archive and no metadata, and work site security cameras that could also show whether helmets and vests are worn. The pattern she names is an existing human job where someone spends a lot of time watching, searching, inspecting or waiting. The second half is about what makes a pilot survive: defining good enough, having ground truth to measure against, and the cost levers that decide whether it scales.
Key takeaways
Start from an existing human job where someone is spending a lot of time watching, searching, inspecting or waiting for something to happen. Alex Gammelgard says the very first thing is defining a use case that is going to matter, and that pattern is the strongest one she saw
Do not pitch a visual AI project internally as a headcount cut. On the Mississippi Department of Transportation example she rejects that reading outright: her words are that it was not that too many people were doing the job and they were trying to get rid of someone, it was that someone was spending a lot of time looking at a wall and could not actually see everything and take action on it
Say what the system has to do before you test any technology. Every use case does not need the fastest model, she says, and you do not need to be monitoring 100 percent of the time: on the pipe corrosion example you can set up a really basic camera and have it monitor once a minute, and that really lowers the cost
Build a pipeline you can swap models into rather than buying cameras with the AI already in them. Her recommendation is a modular system where models are swapped out as the technology gets better, and she says that if you do that you can use even very basic cameras
Look first at the cameras you installed for something else. She calls that one of the biggest opportunity areas, because the video is already being captured, there are no new cameras to install, and the cost of getting started is going to be a lot lower
Check the camera before you commit it to a second job. Her nuance is that you cannot assume every camera is reusable for every purpose: a high, broad-mounted security camera may not have the resolution to tell whether a worker is wearing safety glasses
Put the people who need the outcome in the room with the people choosing the technology. In her diagnosis pilots usually stall because the evaluators are not well enough connected to the business people who need the tech to work for them, and the first symptom is that nobody ever defined what good enough means
Ask what doubles when the deployment doubles. Her framing is exactly that question: if you double your number of cameras, does the cost double? The levers she names for changing the answer are the sampling rate and how much you process at the edge before sending events upstream
Pick a first project where the highest value and the easiest lift meet. Asked where an executive should start next quarter, she says to look for places you already have cameras tied to potentially some kind of metadata you are not looking at today that would be useful for your business
About Alex Gammelgard
Alex Gammelgard is a technology marketing executive who has worked in visual AI across several industries. She was most recently VP of Marketing at Wowza, the video streaming and video infrastructure company, where she worked with enterprise customers deploying visual AI in transportation, industrial inspection, sports and healthcare, and where the team built the Video Intelligence Framework she describes in this episode. She is based in San Francisco and studied at California Polytechnic State University, San Luis Obispo.
In this episode
| 00:32 | Welcome, and why visual AI is not the same thing as generating video |
| 01:57 | What visual AI actually means now, and the side of it everyone talks about |
| 02:54 | Video becomes machine readable the way text already is |
| 04:07 | Industry examples: offshore pipe corrosion, oil rigs, eighty years of sports archive |
| 05:25 | Security cameras repurposed for worker safety, same feeds, new questions |
| 06:57 | The pattern that makes a use case worth doing |
| 07:34 | The Mississippi DOT video wall, and why this is not a headcount argument |
| 08:00 | Where pilots stall: testing the technology before defining the need |
| 08:26 | Not every use case needs the fastest model, or constant monitoring |
| 09:56 | Old cameras are not the barrier. A modular pipeline is the answer |
| 10:17 | The Video Intelligence Framework, and swapping models as they improve |
| 11:59 | Cameras installed for one purpose, reused for another |
| 12:27 | Why you cannot assume every camera is reusable for every job |
| 13:23 | The evaluators are not well enough connected to the people who need the outcome |
| 14:02 | Ground truth, and getting the evaluation away from being subjective |
| 14:40 | The environment is what breaks the pilot |
| 15:21 | Where air-gapped is usually necessary: defense, hospitals, courts |
| 18:19 | Why the economics change between ten cameras and a thousand |
| 18:49 | If you double the cameras, does the cost double? Sampling rate and edge processing |
| 21:13 | Video to action: traffic incidents, shelf compliance, and the NICU |
| 22:36 | Where to start: the highest value against the easiest lift |
| 23:11 | How long a return takes, and the two-day case |
| 25:18 | The resources she recommends |
| 26:24 | The leadership skill: resourceful, proactive, adaptable |
| 27:24 | Remember one thing: visual data, and the human judgment about it |
In Alex’s words
“Visual information is becoming machine readable in a way that text has been for a very long time. So instead of asking what happened in a transcript, you can ask what happened in this room, on this factory floor, in this game, in this inspection.”
Alex Gammelgard (02:54)
“Once they start thinking about visual AI, they realize they could use those same feeds for worker safety ... that’s all with the same cameras and the same footage, and what’s changed is now you can ask a number of business questions out of that data”
Alex Gammelgard (05:25)
“I think in a lot of cases where pilots stall, it is because you start testing the technology before you really figure out what is it you actually need.”
Alex Gammelgard (08:00)
“The important nuance is you can’t assume every camera is reusable for every purpose.”
Alex Gammelgard (12:27)
“The people that are evaluating the tech and making the choice aren’t well enough connected to the business people that need the tech to work for them.”
Alex Gammelgard (13:23)
“I’d say the other thing that we run into a lot, and this is something that we hear a lot about at Wowza, is the environment is what breaks the pilot.”
Alex Gammelgard (14:40)
“What we often heard from customers is that they didn’t necessarily care about the lowest price, they just wanted to understand what would scale and what would stay fixed.”
Alex Gammelgard (18:19)
“Visual data is probably the most powerful data that you have access to, no matter what your business is. Being able to replicate not just the eyes on to a situation, but also the human judgment about the situation, that is a huge unlock for almost every single business.”
Alex Gammelgard (27:24)
Resources
Alex Gammelgard and Wowza
Alex Gammelgard on LinkedIn: The first place she sends listeners. Asked what she would recommend, she says she is very approachable and always happy to chat about these topics
Wowza: The video streaming and video infrastructure company she describes throughout, and the second place she sends listeners. She credits its community and documentation as high quality
Video Intelligence Framework: The pipeline she names in this episode, built so that any model a customer needs can be connected to the cameras they already have instead of replacing the whole video pipeline
How MDOT Streamlines Statewide Traffic Video with Wowza Streaming Engine: Wowza’s write-up of the Mississippi Department of Transportation deployment she describes in this episode, where thousands of cameras were shown about thirty at a time on a video wall
Ideas and terms discussed
Eyes on video: Her term for the pattern she names as the strongest use case, a job where a person is watching, searching, inspecting or waiting for something important to happen, and the system replicates the human eyes on it
Good enough: The accuracy and the frequency the outcome actually requires, agreed between the business side and the technology side before a pilot starts. She contrasts 70 or 80 percent accuracy with 100 percent, and continuous monitoring with once a minute
Ground truth: Historical data a model can be run against so the result is measured rather than argued about. She names the absence of it as one of the reasons a pilot stalls
Air-gapped: Running the system with no network path off the site. She gives defense, hospitals and courts as the cases, and says that where the footage genuinely cannot leave the site, air-gapping is really the only safe way to do AI
Sampling rate: How many frames are pulled for inference. Her example of the cost difference is one frame every 30 seconds against 30 or 60 frames a second
Video to action: Christina Ellwood’s phrase for what happens after detection, with AI seeing a stopped vehicle and creating an incident routed into a traffic operations workflow, rather than only raising an alert
Also named on air
Sora and Veo: The video generation models she names as what the public conversation about visual AI has been dominated by, and the side of it this episode is not about
The AI Realized community: The third place she sends listeners, after her own LinkedIn and Wowza
Related AI Realized episodes and events
Agentic AI Business Strategy: Retrofit or Reimagine the Work: Blaine Mathieu on the strategic choice between retrofitting work you already do, which a competitor can probably copy, and reimagining it, which is harder to copy and more differentiating.
In CPG, AI Has to Be Infrastructure, Not a Project: Nitin Gupta of Mondelez International on computer vision reading what is actually on a store shelf, and on treating AI as infrastructure that sits across every function rather than as a project.
The Technology Works. The Deployment Is What Fails: Tallulah Le Merle on why AI programs fail in the deployment rather than in the technology, and on choosing use cases that survive contact with a business.
Smaller Models, Bigger Wins: Verify Before You Answer: Jason Williamson on getting useful work out of small models running at the edge, and on the efficiency case against reaching for the largest model available.
Frequently Asked Questions
-
Visual AI is software that understands visual information a business already has, such as live camera feeds, security footage, industrial inspection video, archived sports footage and screen recordings, rather than software that generates video. Alex Gammelgard draws that line deliberately: the public conversation has been dominated by the generation side, and she names Sora and Veo as what people get excited about, but the bigger unlock for an enterprise is understanding the visual information it is already creating. Her framing is that visual information is becoming machine readable in a way that text has been for a very long time, so instead of asking what happened in a transcript you can ask what happened in this room, on this factory floor, in this game, in this inspection.
-
Visual AI is being used for offshore pipe inspection, oil rig monitoring, sports archives, crowd safety in stadiums and worker safety on work sites. Alex Gammelgard describes a customer with hours of footage of infrastructure who wanted to know whether pipes were corroding, where most of the video is literally looking at nothing happening, so alerting someone when something is happening saved hundreds of eyeballs. She also describes a major sports organization holding eighty years of archived footage with no metadata on it, looking for what might indicate that an athlete is going to get injured, or whether a crowd of seven people near a fence means a fight is about to break out. One of her favorite examples is a mobile camera system installed at a work site for security and theft prevention, where the same feeds could also show whether people are wearing helmets, vests and goggles.
-
A good visual AI use case is an existing human job where someone spends a lot of time watching, searching, inspecting or waiting for something to happen. Alex Gammelgard calls that pattern eyes on video, and her example is a state transportation department where thousands of cameras were shown about thirty at a time on a video wall, so a stretch of highway might go unseen for a full minute. She is explicit that the argument is not about headcount: it was not that too many people were doing the job, it was that someone was spending a lot of time looking at a wall and could not see everything and take action on it. She also says that in a lot of cases pilots stall because teams start testing the technology before working out what they actually need, and that not every use case needs the fastest model or monitoring 100 percent of the time.
-
You do not need to buy new cameras to use AI on video, and Alex Gammelgard’s recommendation is to build a modular system instead, so models can be swapped out as the technology gets better. Her reason is that a camera bought with AI built into it carries today’s version and is out of date within two months, which makes an expensive upgrade a short-lived one. With a modular pipeline, she says, even very basic cameras work. At Wowza her team built one, the Video Intelligence Framework, so any model a customer needed could be connected to the cameras they already had, and she describes the prospect of replacing an entire video pipeline as a huge barrier for a lot of people.
-
Existing security cameras can often be reused for AI, and Alex Gammelgard calls that one of the biggest opportunity areas, because the video is already being captured and the cost of getting started is a lot lower. A mobile security trailer installed at a work site for theft prevention can also support PPE monitoring, incident review and unsafe zone detection, and retail cameras counting traffic past a store can answer other questions about the business. The caveat she flags is that you cannot assume every camera is reusable for every purpose: the angle, the resolution, the frame rate and the permissions all have to be good enough for the new job, and a high, broad-mounted security camera may not be sharp enough to tell whether a worker is wearing safety glasses.
-
Visual AI pilots usually stall because the people evaluating the technology are not well enough connected to the business people who need it to work for them. Alex Gammelgard names two consequences of that gap. The first is that nobody defined what good enough means, so a model produces outputs and everyone debates whether they serve the purpose, when the two sides should have agreed in advance whether 70 or 80 percent accuracy is enough and how often the inputs need to arrive. The second is that there is no ground truth to evaluate against, so you have to line up historical data you can run the models against, and compare their output to something real. She adds a cause that is not about people at all: the environment is what breaks the pilot, when the bandwidth, connectivity, power or privacy it needs is not there.
-
Visual AI needs to run on the edge or in an air-gapped environment wherever the footage cannot leave the site, and Alex Gammelgard’s rule of thumb is that any place you think the typical regulations apply, you are probably going to need to do it air-gapped. Defense, hospitals and courts are the three cases she gives. On the drivers, Christina Ellwood names privacy, the leakage of intellectual capital and cost. Alex Gammelgard agrees with all three and adds bandwidth and latency: putting compute close to the camera means at least part of the inference runs locally, whether on the camera itself, a nearby edge device, a local server or on-prem, and keeping the full video local while sampling certain frames for AI lowers the cost quite a bit. Where the video genuinely cannot leave, she says running it air-gapped is really the only safe way to do AI.
-
The cost of a visual AI deployment is driven by how much video goes through a model, because video is such a high volume input. Alex Gammelgard’s point is that a workload can look manageable in a pilot and get expensive very quickly once it goes from 10 cameras to 1,000, or from a few hours a day to 24-hour surveillance. More frames means more inference, and higher resolution means more compute and more bandwidth. The levers she names are the sampling rate, pulling one frame every 30 seconds rather than 30 or 60 a second, processing at the edge and sending only the important events upstream, and running event-triggered rather than continuously. What her team often heard from customers, she says, is that they did not necessarily care about the lowest price; they just wanted to understand what would scale and what would stay fixed.
-
How long a visual AI project takes to pay back depends on the use case and on how often the thing you are watching for actually happens. Alex Gammelgard’s fastest example is a transportation department, where catching a car stopped on the side of the road reduces the secondary accidents that follow it, and going from three or four of those to zero could pay off in two days. At the other end she describes a use case where the incident happens very infrequently but matters a great deal when it does, and you might not see it for a month. Christina Ellwood draws the consequence and Alex Gammelgard agrees with it: the experiment has to run longer than the interval between incidents, so a monthly incident needs a project longer than a month while a daily one can be proved out quickly.
-
[00:32] Christina Ellwood: Welcome to AI Realized, the podcast for enterprise executives leading AI adoption. From tackling security, data, and operational challenges to navigating organizational transformation, AI deployment offers a unique opportunity to redesign our organizations from the inside out. I'm Christina Ellwood, your host for today's episode, and we are talking today with Alex Gammelgard, a tech marketing executive who's worked in visual AI off and on for over 20 years. When most people hear visual AI right now, they think about generating video, and Alex works on the other side of it, machines understanding visual information that already exists, live camera feeds, archives, industrial footage, images, and the like, and turning what is happening inside all of that into something a business can actually use. She has most recently worked alongside enterprise visual AI evaluations across very different industries, so she has seen where companies found real value and where pilots quietly stalled out. Alex, welcome to the show.
[01:39] Alex Gammelgard: Thank you. Nice to be here.
[01:42] Christina Ellwood: Let's start with what we're actually talking about here because I think the phrase has been taken over, visual AI. So why don't you talk, tell us a little bit about what visual AI really means right now and what is really generating this video that needs to be analyzed?
[01:57] Alex Gammelgard: Yeah, of course. It's interesting, the public conversation around visual AI has been dominated on the generation side. People get very excited about Sora, Veo, um, and ob- obviously that's very interesting. But what has a much bigger unlock for enterprises is the ability to understand the visual information already in your business. So that could be live camera feeds, security footage, industrial inspection video, archive sports footage, screen recordings, really any visual information that your company is already creating. And what's interesting is companies have had this huge amount of visual data forever. It's, video is the closest thing to replicating real understanding for people. Video is way more telling than text, and it's been a part of businesses forever. But it's been pretty inaccessible, so if you didn't have someone watching it or manually tagging it, there wasn't really a lot that software could do with it. And now the opportunity is visual information is becoming machine readable in a way that text has been for a very long time. So instead of asking what happened in a transcript, you can ask what happened in this room, on this factory floor, in this game, in this inspection, and that's available to every enterprise, and almost every business probably has a use case that they're not thinking about.
[03:17] Christina Ellwood: We are very much focused on, in, in our community around sharing use cases and helping each other understand how those use cases can bring value in, in cross-industry situations. So take us through a couple of different industry examples of use cases of using a video in a business operation.
[03:38] Alex Gammelgard: Yeah. At a high level, it's where in the business are people using eyes as part of a process? So if someone is watching, inspecting, searching, categorizing, or verifying something, that is your first list of potential visual AI opportunities. At Wowza, I joined there a year and a half ago and spent a lot of time, it's a video streaming, video infrastructure company, and my early belief was that it would be a lot of media and entertainment, but actually it was a lot of enterprise companies. And as Wowza started moving into being more of an AI company, the use cases were varied and across almost every industry. But one example was, like, offshore inspection. So we get a customer with hours of road footage of infrastructure. Most, they were looking at pipes, and they wanted to understand if there was potential corrosion. Something looked a little different. Most of that video footage is literally looking at nothing happening, and so, you know, to be able to alert someone to when something was happening, that, that saved hundreds of eyeballs. A lot of the oil and gas companies have that same situation. Someone's out monitoring their oil rigs or onshore watching videos of what's happening out there. On the total flip side, there was a major sports organization. They had 80 years back of archived footage and no metadata around it. And so they were trying to understand trends, trying to understand what might indicate that an athlete's gonna get injured, what might indicate that someone's gonna do really well on a play. Even applying that further out to the stadium, what might indicate if there's a crowd of seven people near a fence, does that mean a fight's gonna break out? Being able to go back and understand and even predict things based on the historical video data, then being able to take action on live video data. One of my other favorite examples that I think shows the power of AI here is you think about someone with a mobile camera system in a work site, and the purpose of those cameras is security and theft prevention. Almost every organization has that, has some kind of security camera. But once they start thinking about visual AI, they realize they could use those same feeds for worker safety, being able to understand if people are wearing helmets, vests, goggles, are the safety procedures being followed. And that's all with the same cameras and the same footage, and what's changed is now you can ask a number of business questions out of that data
[06:04] Christina Ellwood: Yeah. You make me think of my work in scientific spaces where safety goggles and these types of things are really important, or in construction sites where you're concerned about whether equipment has been set up correctly to be safe for people to be working around and so forth. I could see that really being valuable. But tell me a little bit about what you found around what makes for a successful use case. Because while I might have a lot of video and I might have a concern about safety or I might have a concern about or want to understand how to optimize the performance of my athletes or something like that, it isn't free, and it is, there isn't necessarily an owner for that particular question. So how do, h-how, how do companies identify use cases that are sufficiently valuable to pursue them and get them into production?
[06:57] Alex Gammelgard: Yeah. I, that's a great question, and there's actually, there's a number of vectors that you look at there. I think the very first thing is defining a use case that's gonna matter, and the pattern that we saw was the strongest use case is an existing human job where someone is spending a lot of time watching, searching, inspecting, or waiting for something to happen. A couple more examples, like w- I think about our DOT customers just sitting watching a wall of cameras, like on a traffic wall, waiting for a, a, a car to be pointed the wrong way on the highway. We had a customer, the Mississippi DOT, that had, they had thousands of cameras, and it was all on a video wall. And they could show 30 at a time, let's say. And so someone would sit there and watch, and 30 c- 30 wall cameras would come up, and then it would flip through. So you might not see a certain area of the highway for a full minute, and being able to... It's not like there was some, there's too many people doing the job, and they were trying to get rid of someone, it was that someone was spending a lot of time looking at a wall and not being able to actually even see everything, and take action on it. In that, for that example, that was y- a very obvious there will be huge ROI here. But if you have someone in your business that's basically waiting for something important to happen, we call it eyes on video. There's always, it's replicating human eyes. That's going to be a good use case. I think in a lot of cases where pilots stall, it is because you start testing the technology before you really figure out what is it you actually need. Every use case does not need the fastest model. You don't need to be monitoring 100% of the time. If I go back to that oil rig example, they're looking for corrosion on pipes. That's not gonna happen where you need 90 frames a second to be watching and pulling AI out of that. That's something where you can set up a really basic camera, and you can have it monitor once a minute, and that really lowers the cost. Also thinking about, do you need this to be in the cloud, or can you have this happen on the edge where the camera happens, and be able to keep it on site, and not engage one of the big cloud vendors or one of the big, you know, frontier models. So those are the kinds of things you need to think through is what is it you actually need the system to do? I would say level of accuracy is another consideration. If you need it to be 100% accurate, that's gonna be a very different model than if you need it to be just good enough.
[09:22] Christina Ellwood: Okay. And i-i- in terms of the, I think beginning with the end in mind is always a good guide for any of these AI use cases, and we hear it over and over again that people jump to the technology before they've really defined their use case. And in, uh, Blaine Mathieu's model for strategy setting for agentic solutions, he has pointed out that one of the fundamental decisions is are you tr- are you, or observations is, are you retrofitting a process, or are you reimagining the process, or are you somewhere in between that? Sounds like that might be the case here, too.
[09:56] Alex Gammelgard: Yeah, that is really true. What we've noticed is most people don't think they can do AI because they have old cameras. They've had-- and it's a huge investment, and the problem is you buy a camera, it's got today's version of AI on it, and that's out of date within two months in this world that we're living in. So what we've really recommended is for people to create a modular system where you can swap models out as the technology gets better. And if you do that, you can use even very basic cameras. At Wowza we had developed a solution, the Video Intelligence Framework, that was a pipeline for people that wanted to connect really any model that they need to use to their existing cameras. And that became a huge unlock because for a lot of people thinking about needing to completely upgrade their entire video pipeline, their cameras, their pooling to be able to do AI was a huge barrier.
[10:53] Christina Ellwood: And so that's really no longer necessary. The other el-element is there are some of these video use cases that don't involve standalone cameras where you're installing infrastructure. You're using cameras that you have built into something you were already using, like the computers we're using right now. They all have cameras in them, and we record a lot of video in that way. Or we might be using robots, and the robots are capturing the video, and that's inherent in the work that they're doing. So my immediate thought, Alex, in listening to you talk about these use cases is that we're focused on cameras that were installed for, say, surveillance purposes as dedicated cameras. There are a lot of cameras that are just hanging around in our world now, right? We have them on our laptops. We have them in other environments, like we might be using robots that are using cameras for their navigation, and they're looking at the whole space, and that video could be used and repurposed for other reasons as well. So where are you finding, are you finding use cases where the video was initially captured for some o-other reason and then being repurposed?
[11:59] Alex Gammelgard: Yeah. Uh, that actually is one of the biggest opportunity areas because where you're already getting video, you're not having to install a bunch of new cameras, so the cost of getting started is gonna be a lot lower. If that's installed for one purpose, the same visual feed can be used for a bunch of other ones. If the angle, the resolution, frame rate, l- permissions and all that are good enough for the new job, that's the caveat, is that all has to work like that. You could frame it like, where do you already have cameras? What else could that visual data support? Because once it's machine readable, it can, it can do anything. And I had mentioned earlier the security trailers, that could be threat prevention and security, but also PPE monitoring, you know, incident review, unsafe zone detection. The important nuance is you can't assume every camera is reusable for every purpose. So if you have a really high, broad-mounted security camera, does it actually get the resolution you need to be able to understand, does a worker wear safety glasses? Is the existing camera good enough for the new visual task? So as long as that's there, really it's up to you. Retail was a great example, thinking about monitoring traffic coming by the store, but what else can that tell you about your business that's gonna be important?
[13:12] Christina Ellwood: Yeah, for sure. When people have picked a good use case and have the video already and they're running a pilot for that use case, where do the pilots stall?
[13:23] Alex Gammelgard: So the biggest place that they stall is usually because the people evaluating the technology, this is true for every single since software's ever been created, the people that are evaluating the tech and making the choice aren't well enough connected to the business people that need the tech to work for them. A couple things that become problems. So one is they never defined what good enough means. There's a pilot, a model has outputs, and everyone's debating, will this serve our purposes? So you have to get really clear, is 70 or 80% accuracy good enough for this? Does it need to be 100% accuracy? How often do you need the inputs to be happening? So you have to get really clear between the business and the tech folks and understand, for this outcome, what does, what, what requirements must be met? Another one is if you don't have any kind of ground truth to evaluate against. So it is gonna be important to have some historical data that you can run against the models. You're gonna wanna compare the model's output to something real. And again, this is really just about getting the evaluation away from being subjective. Because it is a risk to remove, to remove humans from any part of your business. There's always a risk there, and you wanna be really sure that you feel confident enough in what's happening. I'd say the other thing that we run into a lot, and this is something that we hear a lot about at Wowza, is the environment is what breaks the pilot. You're trying to do something with AI, and you don't have the bandwidth, the connectivity, the power, the privacy to make it happen. A lot of people run into issues where the footage can't leave the site. You wanna send it to a model, but unless you have a way to run that on the edge, you're not gonna be able to make that, that work.
[15:05] Christina Ellwood: We're seeing increasingly use cases where models need to run on the edge in an air-gapped environment for various reasons. Can you talk about some of the cases you've seen for that configuration for doing visual AI?
[15:21] Alex Gammelgard: Yeah. A lot of our customers at Wa- at Wowza actually that was the main reason that they chose Wowza, and Wowza was one of the few providers that allowed people to actually do thing on Edge. A lot of the clientele ended up moving more towards those use cases. I'll say d- defense is one, one use case. Another would be like hospitals. If you're thinking about running AI on visual data in a hospital, and you need to meet all the regulations there. Another would be like the courts. We are trying to get AI based on something that's happening in a courtroom. Any place where you think typical regulations apply, you're probably gonna need to be able to do this air-gapped.
[16:03] Christina Ellwood: So we are seeing more use cases for privately hosted models and even air-gapped privately hosted models, and part of that is about the leakage of intellectual capital into the models. And part of it is for privacy. Part of it is for cost. Are those the same drivers that you're seeing for air-gapped A- visual AI?
[16:33] Alex Gammelgard: Yeah. You are, are absolutely right. The way that we're seeing it work out in practice is that people want to put compute close to the camera or the video so they can run at least part of the inference locally, whether it's on the camera or a nearby Edge device or a local server or on-prem. And the drivers for that are the things you named, as well as it can also help with bandwidth and latency. So on the cost front, for example, if you wanna keep the full video local and sample certain frames for AI, that's gonna lower your costs quite a bit. If you're on somewhere where the video just simply cannot leave because it's sensitive and you have to do it air-gapped for regulations, and really that's the only safe way to do AI. I liked-- That was one of the things I felt like Wowza really had done a good job with, and part of the reason there was such high demand for its solutions was the fact that it had found a way to create a pipeline. They've always been an on-prem company, and being able to bring AI into that world versus forcing people to move to the cloud, that suddenly meant that people with all kinds of use cases that were highly sensitive and regulated could actually do AI when they couldn't before.
[17:51] Christina Ellwood: Interesting. The economics are really important in any of these use cases. I would imagine that in the case of this, these video or visual type use cases, the predictability is a factor. Do you wanna talk about use cases that have either predictable or unpredictable costs and how they differ in the justification for, for the, for the business case?
[18:19] Alex Gammelgard: Yeah, the economics matter a lot with visual AI because video is such a high volume input. So a workload might look manageable in a pilot, but then get expensive very quickly if you're going from 10 cameras to 1,000 or a few hours a day to 24-hour surveillance. Those are extremely different costs. And I, to your point, what we often heard from customers is that they didn't necessarily care about the lowest price, they just wanted to understand what would scale and what would stay fixed. If you double your number of cameras, does the cost double? This is where you can play around with things like sampling rate and processing things on the edge. If you reduce the sampling rate and you're only pulling one frame per every 30 seconds or something like that, that's gonna be a much lower cost than pulling 30, 60 frames a second. And then if you're able to process a lot of it at the edge and just send the important events upstream, that's a much more cost-effective model as well. You know, more frames means more inference, higher resolution means more compute, more bandwidth. And if you're gonna try to monitor something continuously, that's different than event-triggered. These are all levers that you can and should play with, and it has to map to what is your use case, how critical is it to have any of these things, and that'll really help you start understanding the costs.
[19:41] Christina Ellwood: Yeah, and in the case of the edge compute, you could have the edge compute be on a server tied to many cameras, or you could have the cameras themselves have some of the models in them and do some of the processing locally to the camera, right?
[19:55] Alex Gammelgard: Yeah,
[19:56] Christina Ellwood: yeah. Mm-hmm. Absolutely. And the cost has to be very different in those cases. Is there a kind of rule of thumb that if you do it in the camera, it's X times less expensive than doing-- Is there some rule of thumb, or is it all depends?
[20:08] Alex Gammelgard: I would say this is still pretty new, and one of the, one of the things that we were working on with some of the new customers coming on to do visual AI with Wowza was what is that repeatable, predictable cost model. And it is true, the answer is depends, depends. Video, people don't realize how, and I certainly didn't before I, I started working at Wowza, how, how technical video is. I think people don't really understand the process of what needs to happen from something getting filmed, getting transcoded, getting processed, and then being distributed to massive audiences. There's these-- The folks I was working with there are some of the most technical people I've ever met. Being able to bring AI into it, there's just in so many variables that, that tie into the cost.
[20:57] Christina Ellwood: Yeah. The shifting from the cost to the capabilities, as we move from video to text to video to action, what are the, what are the kinds of decision workflows that are being driven by video to action?
[21:13] Alex Gammelgard: Yeah. Some of the most common ones that we saw just out the gate is we did a lot with traffic and smart cities. So AI sees a stopped vehicle or obstruction, and rather than just being an alert, it can actually create an incident and route it to your traffic operations workflow. If in retail, AI would see something on a shelf or a display out of compliance, and can it actually create a task for the store manager? And, uh, I'm trying to think of another good example. Sports or media, if AI sees a goal or a player or an event, and it can make that vid- A great example would be if you see the teams going for a huddle and, okay, now it's time to go to a commercial break, and trigger that. Versus seeing somebody getting close to the goal and, okay, we need to have all the cameras directed at this person 'cause the action's about to happen. And as, and in healthcare, there's a ton of great examples as well. We had a customer that had a really great product where you could monitor, they did, they monitored NICU units, so they had cameras up for premature babies in a hospital. And being able to actually trigger, like a, a serious alert if something happens, those things can be lifesaving.
[22:27] Christina Ellwood: Absolutely. If an executive listening today wanted to run a small visual AI project in the next quarter, where would you tell them to start?
[22:36] Alex Gammelgard: I would tell them to start with defining where they think this could bring the most value, combined with where it would be the easiest lift. So places you already have cameras tied to potentially some kind of metadata that you're not looking at today that, that would be useful for your business. And of course, in video, the answer of what that could look like depends, but that would be the right way to think about getting started.
[23:02] Christina Ellwood: And how much time does it take to see a return even on a simple project for visual AI?
[23:11] Alex Gammelgard: Yeah, that depends on your project as well. For, I'm thinking about the DOT example. You get up and you catch one... Secondary car accidents are, are sec- sorry, let me start over. So that also depends on, on the use case and kind of the urgency. For example, when I think about our DOT customer, it's widely known that the biggest car, the biggest cause of car crashes are like secondary kind of accidents where someone's on the side of the road or a car's pointing the wrong way, or somebody's doing the looky loo thing, and that causes an accident. So the longer you have a car sitting on the side of the road that hasn't been caught, the more chances are that you're gonna have another accident, and the cost of a major accident on a highway is pretty high. For them, that could pay off in two days. If they're able to catch, say, there would normally be like three or four accidents, and they were able to bring it down to zero, that's a big deal. I don't know how that might be. That is measured in dollars, but it's also measured in a lot of other ways. Another use case where something happens very infrequently, but when it happens, it's really important. Maybe you don't have that incident happen for a month. So there is a little bit of, there's the dollar value, and then there's kind of the intangible value. And I think we're at a phase where a lot of companies doing this are doing very mission-critical tasks, and so they're, what I'm seeing is they're weighing both of those things.
[24:32] Christina Ellwood: So if we were picking up on the question about how could someone get started on a, a project right away, what, they would need to plan for the horizon of their experiment to be beyond the threshold of the incident that they're tying the use case to, it sounds like. So if your incidents are happening once a month, your project's gonna have to take longer than a month. But if your incident is happening daily, you might be able to run a pretty short project to see the value add.
[25:00] Alex Gammelgard: Yeah. Absolutely. That's exactly right.
[25:03] Christina Ellwood: Okay. So before we wrap up, let me ask a, a few closing questions. Okay? What resources would you recommend for listeners who wanna learn more about AI, visual AI use cases and more about you?
[25:18] Alex Gammelgard: Oh, gosh. I'm very approachable. You can find me on LinkedIn. I'm always happy to chat with people about these topics. Of course, I would send people to the Wowza website. They really are on the forefront of making these, making visual I- AI accessible, and they've got a ton of great resources. Their community and documentation are really high quality. Those would be two places I would start, and of course, the AI Realized community. It's always a great place to go.
[25:45] Christina Ellwood: So in the AI revolution, Alex, what is the leadership skill you find most valuable today? And I wanna preface this question with saying you're a marketing leader who has adopted AI with, but in a full-throated way. In fact, I think you just published the architecture that you used for deploying AI across marketing. So that leadership that you exerted there to, if you will, to AI-ify your marketing organization, I'm curious about the leadership skills you found most valuable in that work.
[26:24] Alex Gammelgard: I think being resourceful in this period of time and being proactive is really important. I think if you wanna solve a problem, there are an infinite number of ways you can go about solving it now, and there, no one's written the guidebook really on this. There's lots of AI, quote, "experts" out there. Anyone who's one day ahead of you is an expert, right? Like- ... and you're not gonna be that yourself until you get in and do it, so I think being willing to go look for answers and try a lot of things. Maybe that comes back to being resourceful, proactive, adaptable. I'm sure, like the thing I built yesterday that I'm really proud of might be really terrible in a month. Someone will find a way better way to do it.
[27:07] Christina Ellwood: Yeah, so you have to be willing to revisit.
[27:10] Alex Gammelgard: Yeah. We're, we're in such a, a crazy time where nothing is settled, and you have to be okay with that.
[27:18] Christina Ellwood: So if our listeners remember one thing from today's conversation, what should it be and why?
[27:24] Alex Gammelgard: I think it's understanding that visual data is probably the most powerful data that you have access to, no matter what your business is. Being able to replicate not just the eyes on to a situation, but also the human judgment about the situation, that is a huge unlock for almost every single business, and I think there's still use cases being written for it, and there's still opportunity for people to be kind of leaders in that area. So I would really encourage people to think about where they might need that in their own company.
[27:58] Christina Ellwood: Alex Gammelgard, thank you so much for sharing your experience and knowledge about visual AI with us today and being on the show.
[28:07] Alex Gammelgard: Thank you for having me.