Introducing AI-as-a-Service

… For years, SaaS applications have been built around human interaction. Users access platforms through dashboards and interfaces, navigate predefined workflows, and manually complete tasks. The application itself serves as the primary workspace where work is performed.

Agentic AI introduces a different model.

Rather than navigating software in the same way a person would, agents can interact directly with APIs, services and data sources. This allows them to retrieve information, execute actions and orchestrate processes across multiple systems without relying on traditional user journeys.

That does not mean SaaS applications will disappear. Instead of being the primary destination where work happens, many SaaS platforms will increasingly act as sources of capability and information that AI agents can utilize on behalf of users. — Read More

#strategy

How to Build an AI Agent Harness 2.0 and Engineer Better Than 99% of Developers

Your AI agent just finished a coding task.

… A few hours later, you notice something strange.

… We keep trying to make AI agents smarter.

But what if the model isn’t the biggest problem?

What if the real problem is the environment we’re putting it in?

That’s where agent harnesses come in. — Read More

#devops

AIs as Modern Genies

In April, an artificial intelligence (AI) agent conducting a routine task at a company hit a snag, tried to solve it, and soon ended up deleting the company’s database along with all of its backups. In July, OpenAI asked an unreleased AI model to attempt a hacking test. Instead of staying in the isolated box the developers had put it in, the model hacked onto the open internet and into another company to steal the answers. And as reported in August, an AI agent booked someone into a full gym class by figuring out how to cancel other people’s reservations. In all three cases, the AI completed the task it was given—but in ways that ran counter to its controllers’ intentions.

For most people, AI technology is something like the weather: vast and not something you can do much about. It works like magic, and most explanations similarly come from those trying to sell it. At the same time, AI is ubiquitous: It’s now in your phone, your doctor’s notes, and your kid’s homework. It does what it’s told, which sounds like a virtue. Somehow it feels ordinary, despite being so new, because modern economies are remarkably good at absorbing enormous change so smoothly that nobody has time to decide whether they wanted it in the first place.

Whenever something powerful appears in the world, we tell stories about it. That’s what the stories are for. We have thousands of years of stories about this particular kind of power, the kind you summon with words. — Read More

#legal

What’s going on with OpenAI and the Navier-Stokes controversy?

OpenAI announced today that it has found a solution to the Navier-Stokes problem, one of the longest-standing puzzles in mathematics. It’s a major accomplishment, but the news has been marred by a debate over whether the company acted based on unpublished research from other parties. Additionally, OpenAI has been accused of attempting to influence who will receive credit for the achievement in publication.

… Tristan Buckmaster, a mathematician at New York University, and Levent Alpöge, a researcher employed at Anthropic but working in a personal capacity, had been developing their own solution and made use of multiple different LLMs in their efforts, including Anthropic Claude, OpenAI’s Codex tool and its Astra frontier model. According to a statement by Buckmaster, OpenAI may have doubled down on its efforts regarding Navier-Stokes after learning that a team including rival Anthropic was close to a solution. He questioned whether OpenAI had leveraged the work he and his colleague had done with Codex to obtain its own breakthrough. — Read More

#legal

The AI Breakthrough We Might Regret

The launch of OpenAI’s model Astra, the codename for GPT-6, is the first launch whose incredible results have been clearly obscured by the doom risks associated with it.

This model breaks records by a wide margin, and still, most of what you hear is fear about CoT monitorability, controllability, and existential risks.

I mostly avoid these discussions because they are full of extrapolations and esoteric arguments that lead nowhere, but this is the first time I feel like I can’t avoid them. The evidence is simply too strong to dismiss as fearmongering, no matter how cynical I remain about the intentions behind some of the events we’re discussing today.

… We’ll discuss how AI’s nature makes it dangerous, and how, for the first time in my time in this industry, I could see regulation as unavoidable. — Read More

#governance

benchmarking gpt 6 astra

Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.

We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. — Read More

#performance

Applied AI Doesn’t Work

The world has spent a fortune on AI. For most enterprises, nothing changed. I’ve spoken to over 300 CEOs, CIOs and CFOs at the largest companies on Earth. I learned most companies apply AI onto their garbage processes, and the end result is just making garbage faster. Whether it’s buying thousands of Claude Code licenses, committing $50M in annual token spend, or running AI training sessions over Zoom, I’ve seen first-hand how inefficient AI adoption has been. It doesn’t have to be this way.

We’ve already gone down this path, multiple times in fact, and we ought to learn from history when it comes to AI adoption. In 1990, a former MIT computer science professor by the name of Michael Hammer wrote an article for the Harvard Business Review.

“Heavy investments in information technology have delivered disappointing results, largely because companies tend to use technology to mechanize old ways of doing business. They leave the existing processes intact and use computers simply to speed them up.”

If you swapped out ‘information technology’ for ‘AI’, you could publish that tomorrow. — Read More

#strategy

‘Model fatigue’ sets in as AI labs race to roll out new versions at frenetic pace

First, Anthropic updated Fable and Mythos. Then came model enhancements from Meta and Google. OpenAI followed suit by releasing GPT-6 Astra.

[F]or the users of AI models and services, it’s created complexity and chaos as CEOs and IT managers spend an outsized amount of time and resources comparing costs and capabilities to avoid getting left behind. — Read More

#strategy

Bullshit Management Didn’t Die. It Got Automated.

Four years ago, I was pissed. I knew I wasn’t doing product management, but I couldn’t name what I was doing. Eventually, the term found me: bullshit management. I wrote exactly what I felt, published it, and moved on. Barely did I know what was about to happen.

That single article got over 1 million reads combined across LinkedIn, Substack, Medium, and my website. I ended up speaking about it at conferences in 12 countries. For a long time, I asked myself why this one traveled so far.

The answer is uncomfortable. People were tired of doing bullshit management, but they didn’t have a name for it. I gave this name and people embraced it. — Read More

#strategy

Stealing Reasoning Traces from Proprietary LLM APIs

Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model’s final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning. — Read More

#cyber