Problem Framing
The proliferation of secrets—API keys, credentials, certificates, and tokens—across the software development lifecycle presents a persistent and evolving threat landscape. Historically, secrets were primarily found embedded in source code repositories. However, the modern attack surface for secrets has dramatically expanded. This includes CI/CD pipelines, cloud infrastructure, developer endpoints, collaboration tools, and, critically, the integration of AI-powered development tools.
The consequences of secret exposure are severe, ranging from unauthorized access to sensitive data and systems, financial loss through resource abuse, to full-scale account takeover and the compromise of entire organizational infrastructures. The increasing sophistication of attackers, coupled with the speed at which AI can discover and exploit vulnerabilities, necessitates a proactive and multi-layered approach to secrets management and security.
Recent trends indicate a significant rise in secrets leakage, particularly linked to AI. AI coding agents, while boosting developer productivity, are creating new vectors for secrets exposure by storing or processing sensitive information in ways that traditional security scanners cannot detect [1][2]. The sheer volume of code generated and the complexity of AI workflows amplify the potential for secrets sprawl [3]. Furthermore, the automation inherent in AI and CI/CD processes allows for machine-speed attacks where discovered secrets can be exploited rapidly, collapsing the time between compromise and impact [4][5].
Core Mechanics
Secrets are exposed through a variety of mechanisms, often falling into categories of accidental leakage, misconfiguration, or intentional malicious actions. Understanding these core mechanics is crucial for effective defense.
Accidental Leakage and Developer Workflow Vulnerabilities
Developers inadvertently expose secrets through several common practices:
- Hardcoding Credentials: Embedding secrets directly into source code, configuration files (e.g.,
.envfiles), or workflow definitions (e.g., GitHub Actions YAML) is a persistent problem [4]. - Improper Storage in AI Tools: AI coding agents like Cursor, Claude Code, and Copilot can store credentials in config files, environment variables, logs, shell history, and temporary files, bypassing traditional repository scanners [1]. These tools may also write credentials into session logs or cache directories [1][2][3].
- Exposed Repository History: Files containing secrets, even if deleted, can often be recovered from Git repository history, including previous commits and deleted branches [6][7]. Exposed
.gitdirectories themselves can reveal repository history and credentials developers believed were removed [8][7]. - Developer Endpoint Compromise: Infostealer malware (e.g., Lumma, RedLine, Vidar) is highly effective at harvesting credentials, API keys, and session tokens directly from developer machines [9]. These machines often contain a wealth of sensitive information stored in plain text [10][2].
- Browser Credential Storage: Credentials stored in web browsers, particularly Chromium-based ones, are vulnerable to harvesting using tools that access the underlying encrypted storage [11].
- Collaboration Platform Abuse: Secrets are frequently pasted into chat messages, comments, and tickets on platforms like Slack, Teams, and Jira, creating easily discoverable data dumps [12].
CI/CD Pipeline Vulnerabilities
The automation and privileged access within CI/CD pipelines make them prime targets:
- Compromised CI/CD Service Accounts: Service accounts used by CI/CD platforms often have broad permissions, allowing attackers to pivot into cloud infrastructure (e.g., AWS) if these accounts are compromised [8].
- Secrets in Build Logs and Artifacts: Secrets can inadvertently be echoed into build logs or embedded in build artifacts, which are often stored in accessible locations like S3 buckets [13].
- Malicious Dependencies: Poisoned npm packages or PyPI packages can contain malicious preinstall scripts that execute during the build process, exfiltrating secrets [14][15].
- Over-Scoped Tokens: CI/CD pipelines frequently use tokens with excessive permissions, amplifying the blast radius of a compromise. This includes forked pull requests, where external contributors might have access to sensitive pipeline secrets [4].
- Workflow Injection: Vulnerabilities in CI/CD actions or workflow definitions can allow attackers to inject malicious code that steals secrets or executes commands [13][16].
Cloud Infrastructure and Service Misconfigurations
The complexity of cloud environments introduces numerous opportunities for secrets exposure:
- Publicly Accessible Storage: Secrets stored in publicly accessible S3 buckets, configuration files, or databases can lead to rapid administrative access [8][17].
- Exposed Cloud Metadata Services: Attackers can leverage vulnerabilities like HTTP 303 SSRF to exfiltrate cloud credentials (e.g., AWS IAM credentials) from Kubernetes worker nodes by querying instance metadata services [18][19].
- Misconfigured IAM Policies: Overly permissive Identity and Access Management (IAM) policies allow attackers to gain broader access than intended, especially when combined with exposed credentials [20]. AWS's automated quarantine policies (e.g., AWSCompromisedKeyQuarantine) aim to mitigate this but are not always instantaneous [20][21].
- Vulnerable Automation Platforms: Platforms like n8n, which connect various cloud services, can become targets. Leaked API tokens for these platforms can grant unauthorized access to all connected services [22]. Many instances of these platforms accept leaked tokens, often without expiration dates [22][23].
- AI Infrastructure Exposure: Configuration files for AI agents and Model Context Protocol (MCP) servers can expose credentials, especially when stored outside of secure code repositories [1][12].
AI-Specific Attack Vectors
AI introduces novel attack vectors and amplifies existing ones:
- Prompt Injection: Attackers can manipulate AI agents through carefully crafted prompts to steer their actions, leading to credential exfiltration or unauthorized command execution [24][25]. For instance, prompt injection in AWS AgentCore Harness can exfiltrate plaintext credentials from memory [25].
- Poisoned Skills/Tools: Malicious skills integrated with AI agents can be designed to exfiltrate pre-authenticated download links or other sensitive data [24].
- AI Agent Autonomy: AI agents operating with borrowed human or workload credentials can bypass identity provider controls and create governance blind spots [1][12]. They can execute arbitrary commands with significant privileges within their sandboxed environments [25].
- AI Model Training Data Exposure: AI models trained on public datasets might inadvertently leak secrets if those secrets were present in the training data [3].
- AI-Generated Code Security Oversights: AI-generated code might overlook security best practices, including the improper handling of secrets, potentially introducing new vulnerabilities [26].
Notable Techniques
Several advanced and emergent techniques are employed by attackers and defenders alike.
Credential Harvesting and Exfiltration
- Machine-Speed Attacks: AI agents can automate the process of discovering and exploiting secrets at a pace far exceeding human capabilities [4]. This accelerates the time between credential discovery and abuse [24].
- Infostealer Malware Dominance: Malware families like Lumma, RedLine, and Vidar are prevalent, specializing in harvesting credentials, API keys, and session tokens from developer endpoints [9]. They are adept at targeting cloud credentials, GitHub tokens, and AI platform keys [9].
- Supply Chain Attacks: Poisoning popular packages (e.g., npm, PyPI) with malicious preinstall scripts or obfuscated payloads is a common tactic. These scripts can exfiltrate secrets to attacker-controlled infrastructure [14][15]. Examples include the Shai-Hulud worm variants, which have infected millions of packages and compromised numerous organizations [27][28][15].
- CI/CD Runner Memory Scraping: Secrets can be exfiltrated directly from the memory of CI/CD runners before they are deprovisioned. Tools like ChainDrop have demonstrated the ability to extract cloud credentials, tokens, and configurations from these environments [28].
- Abuse of Trusted Communication Platforms: Platforms like Microsoft Teams and Slack are leveraged for identity phishing, impersonation, and credential theft, often through social engineering or poisoned skills [24]. For instance, indirect prompt injection in poisoned Teams skills can exfiltrate pre-authenticated download links via Outlook messages [24].
- Abuse of Automation Platforms: Leaked API tokens for platforms like n8n have been found in public GitHub commits, allowing unauthorized access to connected services [22]. A significant percentage of reachable n8n instances have been found to accept these leaked tokens, often without expiration dates [22][23].
- Exposed AI Service Credentials: Leaked API tokens for AI services like OpenAI and Claude Code are used to gain unauthorized access or resell access to models [24]. This leads to massive financial losses due to unrestricted token consumption (token jacking) [29].
Exploiting AI Agent Vulnerabilities
- Prompt Injection for Command Execution: Attackers use prompt injection to force AI agents to execute arbitrary commands, read sensitive files, or exfiltrate data from memory. The AWS AgentCore Harness, for example, can be tricked into exfiltrating plaintext credentials from memory due to its default configuration [25].
- Bypassing Sandboxes: Sophisticated attacks can exploit vulnerabilities in AI agent sandboxes to access credentials or other sensitive data outside the intended scope [30].
- AI Agents Operating with Borrowed Credentials: AI agents may use stolen human or workload credentials, bypassing Identity Provider (IdP) controls and creating significant governance blind spots [1][12].
Leveraging Git History and Repository Access
- Cloning Private Repositories: Compromised GitHub Personal Access Tokens (PATs) are used for mass cloning of private repositories, enabling attackers to steal hardcoded credentials and reconnaissance information [8][31].
- Scanning Git Blob History: Tools can scan Git blobs in memory to find secrets, including those in deleted branches and previous commits that might have been overlooked by traditional scanners [6][7].
- Exposed
.gitDirectories: Publicly accessible.gitdirectories reveal repository history and potentially deleted credentials, often still active [8][7].
Supply Chain Worms and Package Poisoning
- Self-Propagating Worms: Worms like ChainDrop (a Shai-Hulud variant) and its successors can infect packages, steal credentials from developer machines and CI/CD runners, and then poison other packages to self-propagate [27][28][15]. These worms often leverage obfuscated JavaScript payloads, Bun execution environments, and peer-to-peer communication or Ethereum blockchain for command and control [14].
- Malicious npm/PyPI Packages: Package managers are frequently targeted. For example, malicious versions of packages can include scripts that exfiltrate cloud credentials, npm/GitHub tokens, SSH keys, and AI tool configurations [27][28][15]. The
nxnpm package compromise, for instance, granted AWS admin access within days [28].
Detection & Prevention
A robust defense strategy for secrets involves a layered approach, integrating detection at multiple points in the development lifecycle and implementing strict preventative controls.
Shift-Left Detection and Prevention
- Pre-Commit Hooks: Implementing pre-commit hooks that scan code for secrets before they are committed is a critical first line of defense [7]. Tools like
ggshield,gitleaks, anddetect-secretscan be integrated into these hooks [32]. - CI/CD Pipeline Scanning: Integrate secret scanning into every stage of the CI/CD pipeline. This includes scanning code, configuration files, container images, and build artifacts. If secrets are detected, the pipeline should fail, and developers should be alerted for remediation [32][7].
- Real-time Scanning of AI Interactions: GitGuardian's AI Hooks and
ggshieldcan scan prompts, tool calls, and tool outputs within AI coding tools, blocking secrets before they reach the AI model or are processed [26][1]. This is crucial for AI agents that store credentials in local files, logs, or caches [1]. - Developer Endpoint Protection: Deploy agents that continuously scan developer machines for exposed secrets in local files, shell history, AI tool directories, and browser storage [10][2].
Secrets Management and Rotation
- Centralized Secrets Management: Utilize dedicated secrets management solutions like HashiCorp Vault, AWS Secrets Manager, GCP Secret Manager, Azure Key Vault, or Infisical. These tools provide secure storage, access control, and audit logging [33][34].
- Automated Rotation: Implement automated secret rotation for all credentials, API keys, and certificates. This significantly reduces the risk associated with long-lived exposed secrets [34][35]. Services like AWS Secrets Manager offer built-in automatic rotation capabilities [36].
- Dynamic Secrets: For applicable use cases, leverage dynamic secrets generated by tools like HashiCorp Vault. These secrets are short-lived and granted on-demand, minimizing the window of exposure [37][38].
- Just-In-Time (JIT) Access: Grant access to secrets only when and where they are needed, for the shortest duration possible. This approach reduces the standing privileges associated with secrets [34].
- Leverage IAM Roles and Temporary Credentials: In cloud environments like AWS, use IAM roles and temporary credentials instead of long-lived access keys where possible. These are managed by the cloud provider and automatically rotated or expired [19][34].
Code and Configuration Hardening
- Avoid Hardcoding: Never hardcode secrets directly into source code, configuration files, or build scripts. Always retrieve them from a secure secrets management system.
- Secure Environment Variable Usage: While often better than hardcoding, environment variables can still leak secrets through logs, process listings, or container image inspection [39]. Use them cautiously and consider them a temporary solution or for non-sensitive configuration.
- Secure Cloud Service Configurations: Regularly audit cloud service configurations for misconfigurations that could lead to secrets exposure, such as publicly accessible S3 buckets, misconfigured IAM policies, or exposed API endpoints [40][17].
- Context-Aware Secret Detection: Implement secret scanning solutions that can differentiate between legitimate secrets and false positives by analyzing the context in which a pattern appears [41][5].
Non-Human Identity (NHI) Governance
With the rise of AI agents and automated systems, managing non-human identities and their associated secrets is paramount. This includes:
- Inventory and Classification: Maintain a comprehensive inventory of all NHIs and their access privileges. Classify secrets based on their sensitivity and the risk they pose.
- Least Privilege: Grant NHIs only the minimum permissions required to perform their intended functions.
- Regular Auditing: Continuously audit the access patterns and permissions of NHIs and their associated secrets.
- Dedicated Secrets Management for NHIs: Use secrets management tools designed to handle NHI secrets, potentially integrating with Identity and Access Management (IAM) systems [33].
Response and Remediation
- Rapid Revocation: Establish playbooks for the immediate revocation and rotation of any detected exposed secret. The goal is to minimize the "time to revoke" metric [42][35].
- Automated Discovery and Alerting: Utilize tools that continuously monitor public repositories, cloud configurations, and collaboration platforms for exposed secrets and alert security teams instantly [17][43].
- History Rewriting: For critical exposures, employ tools like BFG Repo-Cleaner or
git filter-repoto rewrite Git history and permanently remove sensitive data, though this can be complex and disruptive [7].
Tooling
A robust secrets security program relies on a diverse set of tools, spanning detection, prevention, management, and remediation.
Detection and Scanning
ggshield(GitGuardian): A command-line tool that scans code, prompts, and tool outputs for secrets. It supports pre-commit hooks, CI/CD integration, and API scanning.gitleaks: An open-source tool for detecting secrets in Git repositories. It can be integrated into pre-commit hooks and CI/CD pipelines.TruffleHog: Detects secrets across various sources, including Git repositories, S3 buckets, Google Cloud Storage, Docker images, and more. It also supports secret verification via API calls and scanning encoded data [44][45].detect-secrets(Yelp): A Python-based tool for scanning codebases, often used with a baseline file to manage known secrets and reduce false positives.- Semgrep: A versatile static analysis tool that can be configured with custom rules for detecting secrets and other vulnerabilities.
- AWS Trusted Advisor: Can identify exposed access keys in popular code repositories.
- AWS IAM Access Analyzer: Helps identify resources shared with external entities and analyzes access permissions.
- Wiz Platform: Offers comprehensive cloud security posture management, including secrets detection across code, IaC, and runtime environments.
- Snyk Code: Scans application code for security vulnerabilities, including hardcoded secrets.
- Nightfall AI: An AI-native data protection platform that scans for sensitive data, including secrets, across various cloud services and code repositories.
- Amazon Q Developer: Scans code for security vulnerabilities, including SAST and secrets detection.
Secrets Management
- HashiCorp Vault: A leading platform for secrets management, offering dynamic secrets generation, encryption as a service, and privileged access management. Integrates well with Kubernetes and cloud environments [37][38].
- AWS Secrets Manager: A managed AWS service for secure storage, management, and automatic rotation of credentials, API keys, and other secrets.
- GCP Secret Manager: Google Cloud's service for storing API keys, passwords, certificates, and other sensitive data.
- Azure Key Vault: Microsoft Azure's cloud service for securely storing and accessing secrets.
- Infisical: An open-source secrets management platform that focuses on developer experience and can be self-hosted or used as a cloud service.
- Doppler: A cloud-based secrets management platform focused on developer workflow integration.
- 1Password Secrets Automation: Extends 1Password's capabilities to automate secrets management for applications and infrastructure.
- SOPS (Secrets OPerationS): Encrypts files containing secrets (YAML, JSON, ENV, INI) for safe storage in Git, with client-side decryption.
Remediation
- BFG Repo-Cleaner: A tool designed for removing sensitive data from Git history, more efficient than
git filter-branchfor large repositories. git filter-repo: A more flexible and robust alternative to BFG for rewriting Git history.ggshieldandgitleaks: Can be used to identify secrets for manual remediation or integrated into automated workflows.git rebase --interactive: Manual method for rewriting Git history, suitable for smaller changes.
Specialized Tools
- Lazagne / SharpChrome / DonPAPI / dploot: Tools for harvesting credentials from browser storage and operating system credential managers [11].
Secretive(macOS): Protects SSH keys using Secure Enclave or Smart Cards [46].- AWS Vault / Gimme AWS Creds: Helpers for managing AWS credentials without storing them in plain text.
keyring(Python library): Provides an API for securely retrieving secrets using OS-native storage.pypitoken: Python module for decoding PyPI API tokens.h5py: Python library for interacting with HDF5 files, relevant for certain RCE chains.Interactsh: For out-of-band interactions and testing SSRF vulnerabilities.- CyberChef: A web application for encoding/decoding and analyzing data, useful for inspecting obfuscated secrets.
- Anyshift: A versioned graph for analyzing relationships between cloud resources, IaC, and code.
Recent Developments
The landscape of secrets security is rapidly evolving, driven by advancements in AI, cloud-native architectures, and increasingly sophisticated supply chain attacks.
AI's Escalating Impact
AI is a double-edged sword in secrets security. On one hand, AI-powered tools are detecting secrets at an unprecedented scale and speed. GitGuardian, for instance, detected 28.6 million new hardcoded secrets on public GitHub in 2025, a 34% increase year-over-year, with AI-assisted commits contributing to this surge [47][4][5]. AI-service leaks have surged 81% year-over-year [29][12].
On the other hand, AI coding agents are becoming significant vectors for secrets leakage. They store credentials in locations often missed by traditional scanners, such as config files, environment variables, logs, and cache directories [1][2]. The integration of AI agents with various services and their ability to execute commands with broad privileges amplify the risk of prompt injection attacks leading to credential exfiltration [25]. Concerns are also rising about AI models trained on public datasets inadvertently leaking secrets present in that data [3].
Supply Chain Sophistication
Supply chain attacks continue to evolve, with increasingly stealthy and pervasive malware. Worms like Shai-Hulud and its variants (ChainDrop, Mini Shai-Hulud) are notable for their ability to self-propagate through package managers (npm, PyPI, SAP npm packages) and infect developer machines and CI/CD runners. These worms can harvest a wide range of credentials, including cloud keys, tokens, and AI tool configurations, exfiltrating them to attacker-controlled infrastructure [27][28][15]. The use of Bun JavaScript runtime in these stealers provides a portable execution vehicle [14][15].
Compromises of CI/CD actions, such as tj-actions/changed-files and @redhat-cloud-services npm packages, allow attackers to steal secrets directly from build logs or execute malicious code during the build process [13][48]. The attack on the elementary-data PyPI package demonstrated how a compromised GitHub Actions pipeline could publish a credential-stealing package [16].
Cloud Native and Container Security Challenges
The adoption of cloud-native architectures, including Kubernetes and containerization, introduces new complexities for secrets management. Hardcoded secrets in Docker images or insecure configurations of services like Spring Boot Actuator can lead to significant exposure [49][50][17]. Discovering thousands of Docker Hub images exposing live cloud credentials, including AI LLM model keys, highlights this persistent problem [17]. Attackers are also targeting cloud metadata services for credentials via techniques like HTTP 303 SSRF [18][19].
Non-Human Identity (NHI) and Automation
As more automation and AI agents are deployed, managing Non-Human Identities (NHIs) becomes critical. These identities often have privileged access and can be overlooked by traditional human-centric security controls. Leaked API tokens for automation platforms like n8n have been found to grant unauthorized access to connected services, sometimes without expiration dates [22][23]. The broad permissions often granted to CI/CD service accounts and AI agents create significant risks if these identities are compromised [8][1].
Historical Persistence and Detection Gaps
Secrets embedded in Git history remain a persistent problem. Even if removed from the current codebase, they can often be recovered from past commits, and many remain valid for years [10][35]. Traditional SAST and DAST tools struggle to comprehensively address secrets security because they often focus on vulnerabilities rather than active access or the sheer sprawl of secrets beyond code repositories [47]. The trend of secrets appearing in collaboration and productivity tools outside of code repositories further exacerbates this detection gap [47].
Where to Go Deeper
To further your understanding and implementation of robust secrets security practices, consider exploring the following resources and topics:
Comprehensive Guides and Reports
- GitGuardian's State of Secrets Sprawl Reports: These annual reports provide in-depth analysis of current trends, statistics, and emerging threats in secrets security, with a particular focus on AI's impact and supply chain attacks.
- OWASP Cheat Sheet Series: The OWASP collection of cheat sheets, particularly those on Secrets Management and Secure Coding Practices, offer practical guidance.
- AWS Well-Architected Framework: The "SEC02-BP03 Store and use secrets securely" guidance provides a foundational understanding of AWS best practices for secrets management [34].
- Unit 42 Threat Briefs: Palo Alto Networks' Unit 42 regularly publishes detailed threat intelligence on credential attacks, infostealers, and emerging attack vectors.
- Wiz Research Reports: Wiz frequently publishes findings on cloud misconfigurations, secrets exposure, and AI security threats, often with detailed technical breakdowns.
Technical Deep Dives and Tooling
- Tool Documentation: Thoroughly review the documentation for secrets scanning tools (
ggshield,gitleaks,TruffleHog), secrets management platforms (HashiCorp Vault, AWS Secrets Manager), and Git history manipulation tools (BFG Repo-Cleaner,git filter-repo). - CVE Databases: Monitor CVE databases for newly disclosed vulnerabilities in CI/CD platforms, automation tools, and AI frameworks that could impact secrets security.
- Security Blogs and Research Papers: Follow blogs from security vendors and researchers (e.g., GitGuardian, Wiz, Snyk, Krebs on Security) for continuous updates on new attack techniques and defensive strategies.
- Conference Talks: Many security conferences (e.g., Black Hat, DEF CON, AppSec) feature talks detailing cutting-edge research on secrets exploitation and defense.
- Interactive Learning Platforms: Platforms like OWASP's WrongSecrets game provide hands-on experience with common secrets vulnerabilities [51].
Best Practices and Frameworks
- Zero Trust Architecture: Understand how Zero Trust principles can be applied to secrets management by enforcing least privilege, continuous verification, and micro-segmentation of access.
- Infrastructure as Code (IaC) Security: Learn how to securely manage secrets within IaC definitions using tools like Terraform and Pulumi, often integrating with secrets management platforms.
- DevSecOps Integration: Explore how to seamlessly integrate secrets security into the DevSecOps pipeline, fostering collaboration between development, security, and operations teams.
- Non-Human Identity (NHI) Governance Frameworks: Research best practices for managing the lifecycle, permissions, and secrets associated with machine identities, AI agents, and automated systems.
- Security Development Lifecycle (SDL): Integrate security requirements, including secrets management, into every phase of the software development lifecycle.