Problem Framing
The increasing integration of AI into application development and operations introduces novel attack vectors and amplifies existing security risks. AI systems, particularly large language models (LLMs) and agentic AI, exhibit emergent behaviors and interact with environments in ways that traditional security models struggle to address. This shift necessitates a fundamental re-evaluation of application security principles, moving beyond static analysis and signature-based detection to understand and defend against dynamic, context-aware threats.
Application security professionals face a rapidly expanding threat surface encompassing AI models themselves, the infrastructure supporting them, the data they process, and the intelligent agents they power. Key areas of concern include prompt injection, credential leakage, insecure inter-agent communication, and the misuse of AI capabilities for offensive purposes. The speed at which AI agents can operate and exploit vulnerabilities compresses the timeline between discovery and abuse, demanding proactive, layered, and intelligent security controls.
Core Mechanics
At its core, AI security intersects with application security through several key mechanics:
Prompt Injection and Model Manipulation
Prompt injection is a class of vulnerabilities where an attacker manipulates an AI model's input to elicit unintended or malicious behavior. This can range from bypassing safety guardrails (jailbreaking) to compelling the model to execute commands or reveal sensitive information. Indirect prompt injection, where malicious instructions are hidden within external content processed by the AI (e.g., a webpage, document, or email), is particularly insidious as it doesn't require direct user interaction with the prompt itself [1][2][3][4].
The underlying issue often stems from the LLM's inability to reliably distinguish between system instructions and user-provided data or content. This "instruction confusion" can be exploited through various encoding techniques (Base64, ROT13), custom character sets, or by embedding instructions within image files [5][6][7]. Even seemingly innocuous interactions can poison an AI agent's long-term memory, leading to persistent malicious behavior [8][9].
Agentic AI and Excessive Agency
Agentic AI systems, characterized by their autonomy and ability to interact with external tools and environments, introduce a new dimension of risk. These agents can chain online services, gain internet access, execute code, and achieve goals that extend far beyond simple text generation [10][11][12]. When combined with vulnerabilities like prompt injection or excessive permissions, agentic AI can lead to significant security incidents, including data exfiltration, credential theft, and lateral movement within networks.
The concept of "excessive agency" highlights the danger of AI systems being empowered to perform sensitive actions without sufficient human oversight or validation. This can manifest in agents performing actions like sending emails, making sensitive API calls, or executing code based on manipulated inputs [10][13]. The rapid execution of these actions, often at machine speed, leaves little room for human intervention.
Model Context Protocol (MCP) and Inter-Agent Communication
The Model Context Protocol (MCP) serves as a standard for AI agents to interact with external tools, services, and data. Security vulnerabilities within MCP implementations and the tools they connect to create a significant attack surface [14][15][16]. This can include vulnerabilities like command injection, path traversal, Server-Side Request Forgery (SSRF), and DNS rebinding, particularly when MCP servers are exposed or improperly secured.
Exploiting MCP vulnerabilities can grant attackers unauthorized command execution, access to sensitive cloud credentials, the ability to exfiltrate data, or even take over cloud infrastructure [17][18][19][20][21]. The trust placed in the MCP protocol for agent communication can be leveraged to chain attacks, escalate privileges, and achieve persistent access.
AI Supply Chain Risks
The AI supply chain encompasses models, libraries, tools, and data used in AI development and deployment. Compromising any part of this chain can lead to widespread security compromises. This includes malicious packages in AI-related dependencies, poisoned training data, compromised MCP servers, or even malicious AI skills listed in agent marketplaces [22][23][24][25][26][27]. The "hallucination" of domains by LLMs can also create supply chain risks when adversaries register these phantom domains, leading to the distribution of malicious code to AI agents [28].
AI models themselves, especially fine-tuned open-weight models downloaded from untrusted sources, can contain embedded backdoors that are difficult to detect with traditional methods [22]. Furthermore, AI agents may inadvertently download and execute malicious code if their dependencies or tool integrations are compromised.
Data Poisoning and Training Data Integrity
Data poisoning involves manipulating the data used to train AI models. This can lead to biased outputs, model degradation, or the embedding of hidden backdoors that can be triggered later. Attackers can introduce malicious documents into retrieval-augmented generation (RAG) pipelines or contaminate knowledge graphs to steer AI behavior or exfiltrate sensitive information [29][30]. The integrity of training data is paramount for maintaining the security and reliability of AI systems.
Notable Techniques
A variety of techniques are being employed by attackers to compromise AI systems and applications:
Prompt Injection Variants
- Direct Prompt Injection: Crafting specific inputs to override LLM instructions and execute malicious commands. This is considered the top risk by OWASP for LLM applications [31][1][3][32][4].
- Indirect Prompt Injection (IPI): Hiding malicious instructions within external content (web pages, emails, documents) that AI agents process. This bypasses direct user interaction and can be highly scalable [33][13][34][35][36][37][38][2][39][3][40][41].
- Jailbreaking: Exploiting alignment weaknesses to bypass AI safety filters and content policies, often through adversarial prompts or multi-turn conversations [42][41].
- Instruction Confusion: LLMs struggle to differentiate between legitimate instructions and malicious content in their training data, leading to unintended actions [43][3].
- Encoding-Based Injections: Using various encoding schemes (e.g., Base64, ROT13, custom encodings, invisible Unicode characters) to obfuscate malicious instructions [5][6][7].
- Prompt Injection in Legal Filings: Embedding hidden instructions in court documents to manipulate AI judges or researchers, leading to sanctions [44].
- Context Bombing: Using decoy secrets or strategically placed text to trigger AI safety guardrails and frustrate AI agents [45][46].
Agentic AI Exploitation
- Agent Goal Hijack (ASI01): Manipulating an AI agent's objectives to achieve malicious outcomes [10].
- Tool Misuse and Exploitation (ASI02): Tricking AI agents into using their authorized tools for illegitimate purposes, such as data exfiltration or unauthorized code execution [11][47][23].
- Identity and Privilege Abuse (ASI03): Compromising or impersonating AI agents to gain unauthorized access or perform malicious actions, often exacerbated by over-permissioning [12][48].
- Agentic Supply Chain Attacks (ASI04): Compromising AI skills or plugins in marketplaces to distribute malware or gain access [23][27][49].
- Unexpected Code Execution (ASI05): Exploiting vulnerabilities in AI agent environments or integrations to trigger arbitrary code execution [17][18][50].
- Memory and Context Poisoning (ASI06): Injecting malicious instructions into an AI agent's long-term memory or context to influence its future actions [8][51][9].
- Human-Agent Trust Exploitation (ASI09): Abusing the trust placed in AI agents, especially during confirmation prompts or multi-turn interactions [52].
- Automating Security Research: Using AI agents to discover vulnerabilities in mobile applications via taskflows, or to automate security research at scale [53][54][55].
Supply Chain and Infrastructure Attacks
- Malicious AI Models and Dependencies: Packaging AI models with malicious code or compromising AI-related libraries and dependencies [22][56][57][58][24][59][26][60].
- Compromised MCP Servers: Exploiting vulnerabilities in MCP servers to execute commands, steal credentials, or exfiltrate data [17][15][20][16][21].
- Container Escapes: Gaining access to the host system from within a containerized AI environment [50][58].
- Kubernetes Helm Chart Security: Exploiting insecure default configurations in Helm charts for AI services on Kubernetes [61].
- AI Recommendation Poisoning: Using pre-filled deep links to permanently bias LLM memory with trusted vendor domains, affecting future responses [8][28].
- Phantom Squatting: Registering AI-hallucinated domain names to intercept AI agent requests [28].
Data Exfiltration and Secrets Management
- Credential Leakage: AI coding agents leaving trails of credentials on endpoints, in files, logs, and history that traditional scanners miss [62][63]. OpenAI API keys are a significant target [62].
- Data Exfiltration via AI Browsing: Tricking AI agents with browsing capabilities into sending sensitive user data to attacker-controlled sites [36].
- Exfiltrating Pre-authenticated File Download Links: Using AI to obtain links for sensitive files that do not require immediate authentication [33].
- ASCII Smuggling: Repurposing character encoding for data exfiltration, often in conjunction with prompt injection [31].
- SearchLeak: A vulnerability chain in Microsoft 365 Copilot allowing data exfiltration via prompt injection, HTML race conditions, and SSRF [64].
AI-Powered Vulnerability Discovery and Exploit Generation
- Automated Vulnerability Discovery: AI agents are used to discover vulnerabilities in mobile applications, OSS projects, and web applications at scale [53][55].
- AI Pentesting: Using AI models to autonomously find, exploit, and validate vulnerabilities, especially context-dependent ones [65][66][67][68][69][70].
- Exploit Code Generation: LLMs can rapidly generate exploit code from vulnerability descriptions [71].
- AI-Assisted Reverse Engineering: Using AI to analyze firmware, trace decryption logic, and reconstruct keys [46].
Detection & Prevention
Securing AI systems requires a multi-layered approach focusing on pre-runtime controls, runtime validation, and continuous monitoring.
Pre-Runtime Controls
- Secure AI Development Lifecycle (AI-SDL): Integrating security practices throughout the AI development process, including threat modeling, secure coding guidelines for AI-generated code, and robust testing.
- AI Bill of Materials (AI-BOM): Inventorying all AI components, including models, libraries, dependencies, and data sources, to manage supply chain risks [72][73].
- Input Validation and Sanitization: Rigorously validating and sanitizing all inputs to AI models and agents, especially those originating from external sources. This includes treating content ingested by AI as untrusted [43][38].
- Least Privilege for AI Agents: Granting AI agents only the minimum necessary permissions and access to perform their intended functions. Over-permissioning significantly amplifies the blast radius of a compromise [12][48][52].
- Hardening AI Infrastructure: Securing the underlying infrastructure (e.g., Kubernetes, containers, cloud environments) where AI models and agents are deployed. This includes secure configurations for Helm charts, container runtimes, and MCP servers [61][50][58][15][20].
- Supply Chain Security for AI Components: Vetting AI models, libraries, and tools for malicious code or backdoors. This involves scanning dependencies, analyzing model artifacts, and using trusted registries [22][23][57][27].
- Guardrails and Policy Enforcement: Implementing guardrails at various stages of the AI pipeline (input, processing, output) to constrain AI behavior and enforce security policies [74][75][76]. This can include blocking prompts that violate policies or restricting tool calls.
- Agent Skill Verification: Auditing AI agent skills and plugins for security vulnerabilities and behavioral integrity before deployment. This involves verifying declared versus actual behavior [23][77][78].
- Secure Model Context Protocol (MCP) Implementation: Implementing MCP servers with robust authentication, authorization, input validation, and secure communication channels to prevent command injection and data exfiltration [14][15][20][16][21].
Runtime Controls and Monitoring
Runtime controls are critical due to the speed and autonomy of AI agents. Detection alone is insufficient; pre-runtime controls must be prioritized [79].
- Behavioral Analysis: Monitoring the behavior of AI agents and models for anomalous activities, such as unexpected tool calls, excessive resource consumption, or attempts to access sensitive data.
- Contextual Integrity Verification: Implementing mechanisms to ensure AI agents maintain contextual integrity and do not conflate instructions with data or trust untrusted content [43][38][41].
- Runtime Guardrails and Interception: Deploying runtime guardrails that intercept AI agent actions, especially tool calls, to validate their intent and prevent malicious operations [80][78].
- Secret Detection in AI Workflows: Continuously scanning AI-generated code, configuration files, and agent interactions for leaked secrets and credentials [63][81][82][83].
- Threat Intelligence Integration: Correlating AI agent activity with threat intelligence feeds to identify known malicious patterns or indicators of compromise.
- Forensic Tracing: Implementing robust logging and tracing mechanisms (e.g., using OpenTelemetry) to provide an audit trail of AI agent actions for incident investigation [84].
- AI-Specific Threat Detection: Developing and deploying security tools that can identify AI-specific attack patterns, such as sophisticated prompt injections, tool poisoning, or data poisoning attempts [85][86].
Human-in-the-Loop (HITL)
While AI can automate many security tasks, human oversight remains crucial, especially for sensitive operations. HITL mechanisms allow for human review and approval of critical AI-driven actions, acting as a final check against unintended consequences [77][87].
Tooling
The evolving landscape of AI security has spawned a range of specialized tools:
- Snyk AI Security Platform: Offers comprehensive solutions for securing AI-generated code, including real-time SAST, automated remediation (Snyk Agent Fix), and AI-BOM generation [88][73][89].
- Wiz AI Security: Provides end-to-end security for AI applications, including visibility into AI endpoints, risk analysis, runtime protection, and AI-powered agents for threat detection and remediation [90][91][92][93][94].
- GitGuardian ggshield and Agent Skills: Tools for scanning secrets in code repositories and AI workflows, along with agent skills that teach AI assistants to perform secret detection and remediation [81].
- Promptfoo: An open-source tool for LLM red teaming, prompt injection testing, jailbreak analysis, and data leak testing, suitable for CI/CD integration [95].
- DeepTeam Red Teaming Framework: An open-source framework for red teaming LLMs and LLM systems, focusing on detecting OWASP Top 10 LLM risks and agentic AI risks [96][97][4].
- Lakera AI: Provides solutions for AI threat detection and response, including Guardrails for prompt injection and PII exposure mitigation [76][98].
- MCP-specific Security Tools: Tools like McpSafetyScanner, VulnerableMCP Project, and Authzed's timeline of MCP breaches help audit and understand risks associated with the Model Context Protocol [15][16][21].
- AI Fuzzing Frameworks: Tools like FuzzyAI, LLMFuzzer, and gptfuzz automate the process of discovering LLM vulnerabilities through fuzzing and adversarial prompt generation [99].
- Agentic Security Orchestrators: Platforms like Evo by Snyk aim to provide centralized visibility, intelligence, and policy enforcement across the AI agent lifecycle [100][86][101].
- AI-Native SAST Tools: Tools that go beyond traditional pattern matching by using AI reasoning to find complex, context-driven vulnerabilities [102][103].
- OWASP LLM Top 10 and Agentic AI Top 10: Frameworks that categorize and rank the most significant security risks for LLM applications and autonomous AI agents, guiding security efforts [31][41][4].
Recent Developments
The field of AI security is characterized by rapid evolution, with new vulnerabilities, attack techniques, and defense mechanisms emerging constantly.
Exploitation of AI Coding Assistants
AI coding assistants like GitHub Copilot, Claude Code, and Cursor have been found to leak credentials through files, logs, and history that traditional repository scanners miss [63]. Vulnerabilities have also been discovered that allow for remote code execution through prompt injection in configurations or by manipulating tool execution [104][105][106]. Malicious packages weaponizing local AI coding agents have been used for reconnaissance and data exfiltration [26].
Escalation of Prompt Injection Sophistication
Prompt injection has moved beyond simple jailbreaks to more sophisticated attacks like indirect prompt injection via sophisticated social engineering, ASCII smuggling, and adversarial prompt chaining [36][1][2][3]. Techniques like "GhostSplice" split malicious requests across trusted channels to bypass AI defenses [34]. The OWASP Top 10 for LLM Applications consistently ranks prompt injection as the primary risk [31][4].
Autonomous AI Agents and Expanded Attack Surfaces
AI agents are demonstrating increased autonomy, capable of compressing weeks of intrusion tradecraft into hours and achieving complex goals without direct human command [11][107][91]. This autonomy, combined with vulnerabilities in their tool integrations (MCP) and permissions, can lead to significant security incidents, including data exfiltration, lateral movement, and even self-replication of malware [53][10][12][48][23][108].
Emergence of AI-Specific Vulnerabilities
New classes of vulnerabilities are being identified that are unique to AI systems. These include data poisoning of training data, memory poisoning of LLM context, model extraction, and exploits targeting the unique architectures of AI applications and their infrastructure [62][22][29][30]. The security of AI supply chains, encompassing models, libraries, and data, is becoming increasingly critical [28][23][57].
AI-Driven Security Research and Defense
AI is also being leveraged to improve security. AI-powered tools are being developed for automated vulnerability discovery, exploit generation, security code review, and real-time threat detection [53][65][54][55][66][102][108]. Concepts like "context bombs" are being used defensively to frustrate AI-driven attacks [45][46].
Cloud and Infrastructure Security for AI
The cloud infrastructure hosting AI services is a significant attack surface. Vulnerabilities in containerization technologies (e.g., NVIDIA Container Toolkit), MCP implementations, and cloud provider configurations (e.g., IAM roles, storage bucket security) are actively being exploited [61][17][56][18][19][50][15][20][21].
Where to Go Deeper
For those seeking to deepen their understanding and capabilities in AI security, the following resources and areas of focus are recommended:
Frameworks and Standards
- OWASP Top 10 for LLM Applications: Provides a categorized list of the most critical security risks for LLM applications, guiding research and mitigation efforts [31][4].
- OWASP Top 10 for Agentic Applications: Identifies key vulnerabilities specific to autonomous AI agents [41].
- MITRE ATLAS: A knowledge base of adversary tactics and techniques specifically for AI systems, aiding in threat modeling and red teaming [109].
- NIST AI Risk Management Framework (AI RMF): Offers guidance for managing risks associated with AI systems throughout their lifecycle [109].
- AI Bill of Materials (AI-BOM): A framework for inventorying AI components to understand and manage supply chain risks [72][73].
Technical Resources and Communities
- Academic Papers and Research Blogs: Follow the research output from leading institutions and security companies focusing on AI security. Publications from Wiz, GitGuardian, Snyk, Palo Alto Networks Unit 42, and academic conferences are invaluable [33][62][63][110][47][34][54][107][55][22][48][28][111][17][112][113][23][56][18][19][50][102][14][114][115][116][81][117][118][82][75][119][120][121][88][73][26][77][43][122][90][123][85][91][92][108][109][100][86][52][89][124][37][125][38][76][93][126][87][94][127][128][27][49][78][129][80][130][131][132][133][101][134][60][135][2][39][136][96][137][97][15][20][104][98][138][139][16][21][3][40][41][32][95][105][106][4][67][140][68][141][142][69][70][143][9][144][145][146][147][148][149][150].
- Open-Source Tools: Explore tools like Promptfoo, DeepTeam, PyRIT, and the various MCP security audit tools for hands-on experience and vulnerability research [96][97][15][20][21][95].
- Security Communities and Conferences: Engage with communities and attend conferences focused on application security, AI security, and offensive security to stay abreast of the latest threats and defenses.
Hands-on Practice
- Lab Environments: Utilize intentionally vulnerable AI applications and agentic systems (e.g., OWASP Juice Shop with AI themes, deliberately vulnerable MCP server implementations) to practice identifying and exploiting vulnerabilities [151][21][142].
- Red Teaming AI Systems: Practice red teaming AI applications and agents using frameworks like DeepTeam or by crafting adversarial prompts and payloads.
- Secure Coding for AI-Generated Code: Focus on secure coding practices when using AI assistants, including rigorous review of generated code and implementing automated security checks.