Crazy Wisdom
Episode #448: From Prompt Injection to Reverse Shells: Navigating AI's Dark Alleyways with Naman Mishra
- AI & Agents
- Prompt injection
- indirect prompt injection
- red teaming
- layered security
- model layer
- infrastructure layer
- application layer
- reverse shell
This video isn't available here. Try YouTube, or listen to the audio below.
Watch on YouTubeAbout this episode
Timestamps
Key Insights
- AI security requires a layered approach. Naman emphasizes that GenAI applications have vulnerabilities across three primary layers: the model layer, infrastructure layer, and application layer. It's not enough to patch up just one—true security-by-design means thinking holistically about how these layers interact and where they can be exploited.
- Prompt injection is more dangerous than it sounds. Direct prompt injection is already risky, but indirect prompt injection—where an attacker hides malicious instructions in content that the model will process later, like an email or webpage—poses an even more insidious threat. Naman compares it to smuggling weapons past the castle gates by hiding them in the food.
- Red teaming should be continuous, not a one-off. One of the critical mistakes teams make is treating red teaming like a compliance checkbox. Naman argues that red teaming should be embedded into the development lifecycle, constantly testing edge cases and probing for failure modes, especially as models evolve or interact with new data sources.
- LLMs can unintentionally leak sensitive data. In one real-world case, a language model fine-tuned on internal documentation ended up leaking a Windows activation key when asked a completely unrelated question. This illustrates how even seemingly benign outputs can compromise system integrity when training data isn’t properly scoped or sanitized.
- Denial-of-wallet is an emerging threat vector. Unlike traditional denial-of-service attacks, LLMs are vulnerable to economic attacks where a bad actor can force the system to perform expensive computations, draining API credits or infrastructure budgets. This kind of vulnerability is particularly dangerous in scalable GenAI deployments with limited cost monitoring.
- Agents amplify security risks. While autonomous agents offer exciting capabilities, they also open the door to complex, compounded vulnerabilities. When agents start reading web content or calling tools on their own, indirect prompt injection can escalate into real-world consequences—like issuing financial transactions or triggering scripts—without human review.
- The Indian AI ecosystem needs to balance speed with sovereignty. Naman reflects on the Indian and global context, warning against simply importing models and infrastructure from abroad without understanding the security implications. There’s a need for sovereign control over critical layers of AI systems—not just for innovation’s sake, but for national resilience in an increasingly AI-mediated world.
Episode transcript
Welcome to the Crazy Wisdom Podcast. This podcast is for you. If you have an insane drive to find the truth of things, it's not the good answers that we seek, but the good questions. I interview a range of different guests from many different fields, all with the intention to uncover the simple truths that are hidden in plain sight. Most people don't want to go there. I go there, my guests go there, and you benefit. Please let me know if you enjoy these episodes and as always, subscribe on itunes, Spotify, or wherever you listen to the podcasts.
Welcome to the Crazy Wisdom Podcast. I've got Naman Mishra here and he is the CTO of Repel AI. Very interesting. It's so funny because I've just noticed that nobody talks about security. It's such an important thing. And I'm so glad that you're. You're here to. To kind of demystify what is going on with security. And welcome to the show.
So there's. It feels like there's two separate worlds of security within LLMs. There's one, the basic just the usage of LLMs, prompt injection, getting them to do things that they're not trained to do, and all these different things. So far, it doesn't seem like there's been a huge amount of. Maybe I just haven't heard of them, but people kind of getting screwed over by these types of things. Like, there's a lot in the beginning, but you don't hear that too much about prompt injection. I do know that Claude sort of did some things to help help their prompt injection. But then there's the other part of security, which seems like you guys might be doing with Repello, which is like, now you're building this machine learning app. There are all these sort of traditional cybersecurity things that have to do with computation that have been problems for. And it is also a mysterious world for me.
Your first full transcript is free. After that, an email opens every transcript in the index — a list of readers we can write to, not a guest book.