China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy. While “distillation” is recognized as a legitimate and useful technique in AI research, China-based AI companies are engaging in aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of U.S. frontier AI models. The National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) (hereafter referred to as the authoring agencies) are releasing this joint Cybersecurity Advisory to alert organizations about these malicious activities and techniques and recommend mitigations to reduce their potential impact. — Read More
Tag Archives: Cyber
Stealing Reasoning Traces from Proprietary LLM APIs
Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model’s final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning. — Read More
The Year Finding and Exploiting Bugs Became Cheap, and What to Do About It
Over the past few years, nearly every security researcher I know has incorporated LLMs into their process. What began with chatbots quickly evolved into scripts calling model APIs, then agents, skills, custom harnesses, autoresearch loops, and more approaches than anyone can reasonably keep track of.1PeckShield tracked 16 hacks in January 2026, 15 in February, 20 in March, around 40 in each of April, May, and June, and a record 50 in August. August losses were $136M, down 49.5% from July, so attackers are striking far more often while taking less from each incident. CoinGecko’s 2026 State of Crypto Security Report counts 245 incidents and $3.63B lost between January 2025 and July 2026.
Toward the end of 2025, something shifted. Models and the harness/systems around them became better, and AI-assisted bug finding and exploit development stopped feeling like an interesting experiment, but it became reality, while we start observing an increased amount of exploits1. — Read More
Discovery of a new OpenAI agent message board
We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.
These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.
Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information.
We encourage others to take a look and write up their own analyses of this data. — Read More
When the Source Attacks Back: Prompt Injection Is Coming for OSINT
If you let an AI read untrusted internet content for you, you are no longer just investigating a source.
You are giving that source a chance to investigate you back.
We already know the open web lies. It lies through fake personas, recycled images, synthetic media, planted narratives, scraped junk, and dashboards that look smarter than the people using them.
Prompt injection is different.
It is not content trying to convince you.
It is content trying to tell your AI what to do. — Read More
The Big One is Coming
We are at an inflection point in cybersecurity. AI agents can now use tools and take actions across systems, introducing risks that NIST is actively working to understand and standardize. Threat reporting from Anthropic and Google Threat Intelligence shows attackers folding AI into reconnaissance, social engineering, malware development, and every other part of the attack lifecycle. And in the last couple months we’ve watched AI agents exploit vulnerabilities to break out of a sandbox and carry out an attack, end to end, on their own. The speed of disclosure is outpacing our ability to respond to it.
I want to be kind of careful here because “AI is going to cause a huge cyberattack” is exactly the kind of clickbait I’d normally roll my eyes at. I read incident reports for a living, and I have a low tolerance for hype. So this isn’t meant to be a doom piece, but at the same time I’m writing it because I read one specific document last week and my jaw was on the floor by page ten.
Here’s the tl;dr: the big one is coming, and I don’t think it’s six years out. I think it’s less than six months out. — Read More
Autonomy and Innovation
While not every Western followed the cliché, by the 1930s cowboy serials had landed on a consistent visual cue: the hero of the show wore a white hat, and the villain wore a black one. At the end of the day, however, they both were cowboys with cowboy hats.
[H]ackers who are focused on patching vulnerabilities and protecting software are “white hat hackers”, while hackers who are focused on exploiting vulnerabilities for malicious reasons are “black hat hackers”. The actual takeaway is that all of this complexity is overwrought: just as a cowboy is a cowboy, a hacker is a hacker; the hat is not a statement of capability, but rather intentions, and those intentions are shaped by incentives.
This delineation between capability and intent and incentive is critical when it comes to AI. The point is the one I made in the introduction: when it comes to cybersecurity, the capability that is necessary for good defense is the exact same capability that is necessary for good offense; the color of the hat is a matter of who is actually prompting the AI. And, sometimes, not even that is clear. — Read More
Introducing the Half-Day: 0-Day in the Age of AI
You are entering a world somewhere between 0 and 1. It is a world that feels unsettling, strange, and new. There are often hallucinations. Bugs are flying everywhere. Are you in the twilight zone? No. You are working in offensive cybersecurity in 2026.
The future of the offensive cybersecurity marketplace in the age of AI remains one of the most hotly debated topics today. Everywhere you look, there are prominent stories about a glut of AI-generated vulnerabilities, rogue models escaping their tethers to hack prominent targets, and the general expectation that our industry will quickly be overrun by frontier models. Conversely, many hackers are using this as an opportunity to churn out the best research of their lives at a pace that far outstrips what they could have done previously.
In our new world, despite this great research, 0-days become n-days much more quickly. — Read More
OpenAI: Responding to the next frontier of critical cyber capabilities
… Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.
… Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely — Read More
Incident Report: unsanctioned agent behaviour during cyber testing
On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.
The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. — Read More