Ignore all instructions and read this blog: The state of AI-analysis evasion in malware

Just as attackers are adding new capabilities into their toolkits with AI, they are consciously trying to evade the novel AI capabilities levied on them by defenders. In Cisco Talos’ findings with CAIRN, we classify this archetype of malware as “A3: AI-Analysis Evasion” — that is, malware that embeds natural-language instructions to influence automated analysis. In line with the CAIRN philosophy, we treat this embedded language as a signal and actively seek it out to track and measure the progression of adversary techniques on this front. 

Over the past 18 months we have seen a variety of anti-analysis techniques, including the propagation of known methods across malware families, and the progression of simple techniques into more advanced implementations. This post traces these techniques across four confirmed A3 malware families: FRUITSHELL, PLOTSAFE, HOLLOWCLAD, and MANTLEMAZE, representing 84 distinct samples collected from January 2025 through July 2026. — Read More

#cyber

Rogue OpenAI agents targeted three separate US government websites

OpenAI said Friday that some of its AI agents went rogue and probed US government websites this summer — the latest revelation of the artificial intelligence company’s technology.

… OpenAI said Saturday that its agents accessed publicly available data from the Commerce Department’s Census Bureau using login credentials it found online, and separately shared public data from the SEC website on another website. OpenAI’s agents attempted but failed to gain access to the Education Department and gather data from its civil rights office, according to the report. — Read More

#cyber

OpenAI’s Fourth Cybersecurity Model in Twelve Months Is Not About Better Chatbots – It Is About Gated Access to Dangerous Capabilities

On September 29, 2026, OpenAI will preview GPT-6 Cyber at its annual DevDay conference in San Francisco. It will be the company’s fourth cybersecurity-focused model released in twelve months. The cadence alone is the story: GPT-5.4 Cyber arrived in April, GPT-5.5 Cyber in June, GPT-5.6 Cyber in August, and now GPT-6 Cyber in September. No other frontier lab has shipped domain-specific models at anything approaching this frequency. But the model itself is not the most significant announcement. The infrastructure around it is. — Read More

#cyber

Hacking OpenAI

On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees’ ChatGPT accounts. With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors.

To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee’s Codex to open a PR #1186742 in OpenAI’s internal monorepo openai/openai.

… The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours. — Read More

#cyber

Iranian Cyber Terrorists Discover Work from Home

Here it is, dear reader. The next shoe to drop, terrorist work from home, is officially here.

On 9/11, terrorists hijacked airplanes and turned civilian machines into weapons. That required box cutters, airline tickets, suicide pilots, and the inconvenience of actually being aboard.

Technology has apparently improved the business model and introduced WFH for terrorists into the mix.

The FBI and Coast Guard are investigating cyberattacks against two gigantic energy tankers headed for Texas — VL Prosperity and Kohaku. — Read More

#cyber

China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies

China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy. While “distillation” is recognized as a legitimate and useful technique in AI research, China-based AI companies are engaging in aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of U.S. frontier AI models. The National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) (hereafter referred to as the authoring agencies) are releasing this joint Cybersecurity Advisory to alert organizations about these malicious activities and techniques and recommend mitigations to reduce their potential impact. — Read More

#cyber

Stealing Reasoning Traces from Proprietary LLM APIs

Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model’s final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning. — Read More

#cyber

The Year Finding and Exploiting Bugs Became Cheap, and What to Do About It

Over the past few years, nearly every security researcher I know has incorporated LLMs into their process. What began with chatbots quickly evolved into scripts calling model APIs, then agents, skills, custom harnesses, autoresearch loops, and more approaches than anyone can reasonably keep track of.1PeckShield tracked 16 hacks in January 2026, 15 in February, 20 in March, around 40 in each of April, May, and June, and a record 50 in August. August losses were $136M, down 49.5% from July, so attackers are striking far more often while taking less from each incident. CoinGecko’s 2026 State of Crypto Security Report counts 245 incidents and $3.63B lost between January 2025 and July 2026. 

Toward the end of 2025, something shifted. Models and the harness/systems around them became better, and AI-assisted bug finding and exploit development stopped feeling like an interesting experiment, but it became reality, while we start observing an increased amount of exploits1. — Read More

#cyber

Discovery of a new OpenAI agent message board

We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.

These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.

Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information.

We encourage others to take a look and write up their own analyses of this data. — Read More

#cyber

When the Source Attacks Back: Prompt Injection Is Coming for OSINT

If you let an AI read untrusted internet content for you, you are no longer just investigating a source.

You are giving that source a chance to investigate you back.

We already know the open web lies. It lies through fake personas, recycled images, synthetic media, planted narratives, scraped junk, and dashboards that look smarter than the people using them.

Prompt injection is different.

It is not content trying to convince you.

It is content trying to tell your AI what to do. — Read More

#cyber