Visual AI Lets You Ask Business Questions of Your Video

Episode Summary

Companies have had huge amounts of visual data forever, Alex Gammelgard says, and it has stayed pretty inaccessible: unless someone was watching or tagging it by hand, there was not really a lot software could do with it. It is now becoming machine readable the way text already is, she argues, so the same cameras and the same footage answer business questions the company could not ask before. She works through where it applies: offshore pipe inspection where most of the footage shows nothing happening, a sports organization with eighty years of archive and no metadata, and work site security cameras that could also show whether helmets and vests are worn. The pattern she names is an existing human job where someone spends a lot of time watching, searching, inspecting or waiting. The second half is about what makes a pilot survive: defining good enough, having ground truth to measure against, and the cost levers that decide whether it scales.

Key takeaways

  • Start from an existing human job where someone is spending a lot of time watching, searching, inspecting or waiting for something to happen. Alex Gammelgard says the very first thing is defining a use case that is going to matter, and that pattern is the strongest one she saw

  • Do not pitch a visual AI project internally as a headcount cut. On the Mississippi Department of Transportation example she rejects that reading outright: her words are that it was not that too many people were doing the job and they were trying to get rid of someone, it was that someone was spending a lot of time looking at a wall and could not actually see everything and take action on it

  • Say what the system has to do before you test any technology. Every use case does not need the fastest model, she says, and you do not need to be monitoring 100 percent of the time: on the pipe corrosion example you can set up a really basic camera and have it monitor once a minute, and that really lowers the cost

  • Build a pipeline you can swap models into rather than buying cameras with the AI already in them. Her recommendation is a modular system where models are swapped out as the technology gets better, and she says that if you do that you can use even very basic cameras

  • Look first at the cameras you installed for something else. She calls that one of the biggest opportunity areas, because the video is already being captured, there are no new cameras to install, and the cost of getting started is going to be a lot lower

  • Check the camera before you commit it to a second job. Her nuance is that you cannot assume every camera is reusable for every purpose: a high, broad-mounted security camera may not have the resolution to tell whether a worker is wearing safety glasses

  • Put the people who need the outcome in the room with the people choosing the technology. In her diagnosis pilots usually stall because the evaluators are not well enough connected to the business people who need the tech to work for them, and the first symptom is that nobody ever defined what good enough means

  • Ask what doubles when the deployment doubles. Her framing is exactly that question: if you double your number of cameras, does the cost double? The levers she names for changing the answer are the sampling rate and how much you process at the edge before sending events upstream

  • Pick a first project where the highest value and the easiest lift meet. Asked where an executive should start next quarter, she says to look for places you already have cameras tied to potentially some kind of metadata you are not looking at today that would be useful for your business

About Alex Gammelgard

Alex Gammelgard is a technology marketing executive who has worked in visual AI across several industries. She was most recently VP of Marketing at Wowza, the video streaming and video infrastructure company, where she worked with enterprise customers deploying visual AI in transportation, industrial inspection, sports and healthcare, and where the team built the Video Intelligence Framework she describes in this episode. She is based in San Francisco and studied at California Polytechnic State University, San Luis Obispo.

 

In this episode

00:32 Welcome, and why visual AI is not the same thing as generating video
01:57 What visual AI actually means now, and the side of it everyone talks about
02:54 Video becomes machine readable the way text already is
04:07 Industry examples: offshore pipe corrosion, oil rigs, eighty years of sports archive
05:25 Security cameras repurposed for worker safety, same feeds, new questions
06:57 The pattern that makes a use case worth doing
07:34 The Mississippi DOT video wall, and why this is not a headcount argument
08:00 Where pilots stall: testing the technology before defining the need
08:26 Not every use case needs the fastest model, or constant monitoring
09:56 Old cameras are not the barrier. A modular pipeline is the answer
10:17 The Video Intelligence Framework, and swapping models as they improve
11:59 Cameras installed for one purpose, reused for another
12:27 Why you cannot assume every camera is reusable for every job
13:23 The evaluators are not well enough connected to the people who need the outcome
14:02 Ground truth, and getting the evaluation away from being subjective
14:40 The environment is what breaks the pilot
15:21 Where air-gapped is usually necessary: defense, hospitals, courts
18:19 Why the economics change between ten cameras and a thousand
18:49 If you double the cameras, does the cost double? Sampling rate and edge processing
21:13 Video to action: traffic incidents, shelf compliance, and the NICU
22:36 Where to start: the highest value against the easiest lift
23:11 How long a return takes, and the two-day case
25:18 The resources she recommends
26:24 The leadership skill: resourceful, proactive, adaptable
27:24 Remember one thing: visual data, and the human judgment about it

In Alex’s words

“Visual information is becoming machine readable in a way that text has been for a very long time. So instead of asking what happened in a transcript, you can ask what happened in this room, on this factory floor, in this game, in this inspection.”

Alex Gammelgard   (02:54)

“Once they start thinking about visual AI, they realize they could use those same feeds for worker safety ... that’s all with the same cameras and the same footage, and what’s changed is now you can ask a number of business questions out of that data”

Alex Gammelgard   (05:25)

“I think in a lot of cases where pilots stall, it is because you start testing the technology before you really figure out what is it you actually need.”

Alex Gammelgard   (08:00)

“The important nuance is you can’t assume every camera is reusable for every purpose.”

Alex Gammelgard   (12:27)

“The people that are evaluating the tech and making the choice aren’t well enough connected to the business people that need the tech to work for them.”

Alex Gammelgard   (13:23)

“I’d say the other thing that we run into a lot, and this is something that we hear a lot about at Wowza, is the environment is what breaks the pilot.”

Alex Gammelgard   (14:40)

“What we often heard from customers is that they didn’t necessarily care about the lowest price, they just wanted to understand what would scale and what would stay fixed.”

Alex Gammelgard   (18:19)

“Visual data is probably the most powerful data that you have access to, no matter what your business is. Being able to replicate not just the eyes on to a situation, but also the human judgment about the situation, that is a huge unlock for almost every single business.”

Alex Gammelgard   (27:24)

 

Resources

Alex Gammelgard and Wowza

  • Alex Gammelgard on LinkedIn: The first place she sends listeners. Asked what she would recommend, she says she is very approachable and always happy to chat about these topics

  • Wowza: The video streaming and video infrastructure company she describes throughout, and the second place she sends listeners. She credits its community and documentation as high quality

  • Video Intelligence Framework: The pipeline she names in this episode, built so that any model a customer needs can be connected to the cameras they already have instead of replacing the whole video pipeline

  • How MDOT Streamlines Statewide Traffic Video with Wowza Streaming Engine: Wowza’s write-up of the Mississippi Department of Transportation deployment she describes in this episode, where thousands of cameras were shown about thirty at a time on a video wall

Ideas and terms discussed

  • Eyes on video: Her term for the pattern she names as the strongest use case, a job where a person is watching, searching, inspecting or waiting for something important to happen, and the system replicates the human eyes on it

  • Good enough: The accuracy and the frequency the outcome actually requires, agreed between the business side and the technology side before a pilot starts. She contrasts 70 or 80 percent accuracy with 100 percent, and continuous monitoring with once a minute

  • Ground truth: Historical data a model can be run against so the result is measured rather than argued about. She names the absence of it as one of the reasons a pilot stalls

  • Air-gapped: Running the system with no network path off the site. She gives defense, hospitals and courts as the cases, and says that where the footage genuinely cannot leave the site, air-gapping is really the only safe way to do AI

  • Sampling rate: How many frames are pulled for inference. Her example of the cost difference is one frame every 30 seconds against 30 or 60 frames a second

  • Video to action: Christina Ellwood’s phrase for what happens after detection, with AI seeing a stopped vehicle and creating an incident routed into a traffic operations workflow, rather than only raising an alert

Also named on air

  • Sora and Veo: The video generation models she names as what the public conversation about visual AI has been dominated by, and the side of it this episode is not about

  • The AI Realized community: The third place she sends listeners, after her own LinkedIn and Wowza

Related AI Realized episodes and events

 

Frequently Asked Questions

 
 
 
 
 
Next
Next

Both Sides Got the Same AI, So Cyber Defense Must Automate