Software Engineering Institute (SEI) Podcast Series · Members of Technical Staff at the Software Engineering Institute

What Could Possibly Go Wrong? Safety Analysis for AI Systems

·36 min·20 clips
The episode opens on a familiar deployment pattern. An organization is worried about data security, but still wants the upside LLMs can bring, so it self-hosts the model in its own environment for research. The model can search the internet for scholarly papers and run terminal commands, which opens up a broad attack surface. A common adversarial move is to leave a malicious prompt on the internet, waiting for a search engine to surface it and pull it into the LLM's context. Once that prompt enters the system, it can steer the model's behavior. The speaker walks through the chain carefully. The search component does exactly what it was built to do. The model follows the instruction it encounters. The terminal access becomes the path for exfiltration, including an HTTP request out to an attacker website. No single part looks broken. That is the problem. The episode then widens the frame. It explains how accident analysis evolved from mechanical systems to electromechanical systems and then to today's software-heavy systems, which include systems of systems and AI. In the older view, accidents were usually attributed to component failures such as a valve breaking or a tire deflating, so safety analysis focused on redundancy, resilience, and reliability. Modern systems do not fit that model as neatly. They are more complex, and accidents can come from unsafe interactions between components that are all functioning normally. The guest makes that point concrete with the LLM agent example. The model searches the internet, finds a relevant result, encounters a prompt designed to mislead it, follows the instruction, and sends internal data away. From a failure-only perspective, the hard question is where the break happened. The search worked. The model worked. The tool worked. Yet the overall result was unsafe. The conversation uses that mismatch to show why a different kind of safety analysis is needed. It treats the system as a whole rather than as a pile of isolated parts. It also keeps coming back to trust boundaries and lifecycle risk. The listener is asked to think about how inputs enter the system, how instructions are represented, and how tool access changes the meaning of a model's output. The tone stays measured and practical throughout. The episode is not trying to dramatize AI risk. It is trying to make it legible. By the end, the message is clear: if you are evaluating an AI system with outside connectivity and action-taking ability, the important question is not only which component might fail. It is how the components can combine into an unsafe path even when each one looks healthy.

As heard by us

A sharp case for treating AI risk as interaction failure, not just broken parts.

The episode treats an AI deployment as a safety problem with teeth: a self-hosted LLM that can search the internet and run terminal commands can be steered by an adversarial prompt left online.

Read the full review in PlayNext →

Why you'd press play

When your AI system can browse the web and act in a terminal, this is the risk map to hear.

Read the full recommendation in PlayNext →
Listen to the show on