DevSec Station is a security focused podcast for software developers who want to create amazing applications. Hosted by Tanya Janca, also known as SheHacksPurple, these short lessons will help you level up.
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
9:23
What happens when your AI assistant does exactly what you expect 59 times... and then does something completely bananas on the 60th?
While teaching a secure coding class, I watched nearly 60 students use the same AI assistant, programming language, and prompts. For 59 of them, everything went as expected.
But one student's AI-generated Python code included instructions to suppress linting checks, followed by code that leaked secrets to an external file.
The AI hadn't escaped a sandbox. It wasn't even an agent. It was simply generating code within the boundaries we'd given it.
And that's what makes the story so interesting.
Welcome to Season 2 of DevSec Station! In this episode, Tanya Janca explores why AI's unpredictable behaviour becomes a security problem when we give it access to files, credentials, tools, infrastructure, and deployment capabilities.
We'll discuss:
- Why successful AI runs don't guarantee the next run will be safe. - Why better prompts aren't a substitute for security controls. - Least privilege, independent security checks, and human approvals. - How to limit the blast radius when an AI system does something unexpected.
** Your one thing to do this week ** Pick an AI assistant or agent you're using and ask yourself: What's the worst thing it's technically capable of doing with its current permissions?
Not what you've asked it to do. What it can *actually* do.
Because we don't need AI systems to be perfectly predictable. We need to engineer systems that can safely survive a surprise.
DevSec Station is a podcast by Tanya Janca (SheHacksPurple), focused on short, practical lessons that help software developers build more secure software.
Your AI worked 59 times. But what about the 60th? I was teaching secure coding to a class of about 60 people, and everyone was using the same AI assistant, the same programming language, the same security prompt, the same build prompt. And for 59 people, it worked pretty much the way you would expect. And then one student called me over and said, Tanya, what's this? I looked at the code their AI had generated and it had done something wild. It had added comments telling the Python linter not to expect the following lines. And then in those lines, it was intentionally leaking secrets to an external file. And it was not fine. I'm Tanya Jenka. Welcome to season two of DevSec Station. Here's the security question that I think we need to start asking about AI assisted development and AI agents. What happens on the really weird run? Because we're used to testing software and asking questions like: does this work properly? Does it behave as intended? Does it produce the results we're expecting? But with an AI agent, we might run something five times or 10 times or 50 times and think, oh my gosh, this is great. It works perfectly. Except now we're dealing with a probabilistic system, which means a non-deterministic system. The fact that it behaved correctly the last 59 times does not guarantee that run number 60 will look anything like the previous 59. And that's not necessarily a huge security problem if the AI isn't able to do very much. If I ask an AI agent to suggest a sentence for an email and it gives me something completely bizarre, that's fine. I just don't send that weird sentence. Because obviously I would proofread that email before I send it, right? But we're increasingly giving AI systems the ability to do things. Write code, access files, invoke tools, access secrets, send messages on our behalf, modify infrastructure, publish software, deploy things. And now consistency starts to become an important security property because the potential consequences of a weird run have changed dramatically now. And I would like to make an important distinction at this point. The AI does not need to escape for this to become a security problem. We hear a lot about agents escaping sandboxes or bypassing controls or finding clever ways around the restrictions that we have set for them. And that is interesting. And we are definitely going to talk about that during season two. But that's not what happened in my classroom. The AI didn't escape. It wasn't even an agent. It was generating code because I literally asked it to generate code. It was operating entirely within the environments and the capabilities that we had given it. It just generated something we really, really, really did not want. And that is a completely different problem. Because if an AI can cause damage while staying completely inside the boundaries that we gave it, then making those boundaries stronger doesn't necessarily solve the problem. We also have to ask whether we've given it too much power inside those boundaries. So what do we do about this, right? For starters, the answer is not write a better prompt. So prompts are useful and we should absolutely give AIs good instructions. We should tell them what we expect and what we want. But a prompt is not a security control. If the consequences of the AI doing something unexpected is severe, we need controls outside the AI that limit what can actually happen. And I mean, fortunately, we know how to build those, right? We've been implementing least privilege for decades. If an agent only needs access to one repository, then we should not give it access to every single repository we have. If it needs to read something, why are we giving it CRUD permissions? If it needs credentials briefly, maybe it doesn't need that credential forever and it should expire. And then on top of that, there are independent security tools. If an AI is writing code, that code should probably go through code review, static analysis, linters, peer review, etc. It should still go through all the security checks we would expect other code to go through, including and especially testing. And yes, I appreciate the absolutely delightful irony of saying that after telling you a story in which the generated code tried to tell the linter not to look at that. But that's actually the best argument for making sure that the AI isn't in charge of deciding whether controls run or not, right? For higher impact actions, we also need to add approvals. Maybe the agent can prepare a deployment, but something or someone else has to authorize it. Maybe it could draft an email, but it cannot send it. Maybe it can propose an infrastructure change, but it can't independently push that change to production. We can make operations reversible when possible. That's a great idea. We can use dry run modes, we can plan for rollback, and we need to log what happened so that when something weird occurs, we can figure out what it actually did, not just what we think it did. And always, especially with all the things we are seeing in the news right now, we need to try to limit the blast radius. The goal isn't to make a probabilistic system behave in a perfectly predictable way because that is not going to happen. I don't think that that's a useful security strategy because we have all seen it fail repeatedly. Our goal needs to be to design systems around it so that one strange output doesn't become a catastrophic event. Think about the difference between these two questions. Does our AI usually do the right thing? And what is our AI actually allowed to do if something goes wrong? Obviously, very different questions. And obviously the second one's more important. Because if your AI works correctly 59 times and does something bananas on number 60, I want number 60 to be boring. Maybe a security tool rejects the code. Maybe an authorization check blocks that action from happening. Maybe someone has to approve it and they don't. Maybe a transaction limit kicks in, maybe it gets logged and an alert gets sent and you go investigate. Maybe you have to roll something back. But if there's a weird run, it shouldn't become a security incident. And that is going to be a recurring theme this season. We don't have to prevent AI systems from ever doing anything surprising. What we need to do is engineer our systems so they can safely survive a surprise. So here's your one thing to do this week. Pick one AI assistant or agent that you or your organization is using and ask, if this does something completely unexpected one time out of 60, what's the worst thing that it can actually do? Not what you've asked it to do or what the prompt says it should do, but from a technical perspective, what it is capable of doing with the permissions, credentials, tools, and access that it actually has. If you don't like the answer, don't rate a better prompt. Change its permissions. And we're going to talk about how to do that this entire season. Thank you for listening to DevSecStation. I'm Tanya Janka, and I'll see you next time.