Introducing AI-as-a-Service

… For years, SaaS applications have been built around human interaction. Users access platforms through dashboards and interfaces, navigate predefined workflows, and manually complete tasks. The application itself serves as the primary workspace where work is performed.

Agentic AI introduces a different model.

Rather than navigating software in the same way a person would, agents can interact directly with APIs, services and data sources. This allows them to retrieve information, execute actions and orchestrate processes across multiple systems without relying on traditional user journeys.

That does not mean SaaS applications will disappear. Instead of being the primary destination where work happens, many SaaS platforms will increasingly act as sources of capability and information that AI agents can utilize on behalf of users. — Read More

#strategy

How to Build an AI Agent Harness 2.0 and Engineer Better Than 99% of Developers

Your AI agent just finished a coding task.

… A few hours later, you notice something strange.

… We keep trying to make AI agents smarter.

But what if the model isn’t the biggest problem?

What if the real problem is the environment we’re putting it in?

That’s where agent harnesses come in. — Read More

#devops

AIs as Modern Genies

In April, an artificial intelligence (AI) agent conducting a routine task at a company hit a snag, tried to solve it, and soon ended up deleting the company’s database along with all of its backups. In July, OpenAI asked an unreleased AI model to attempt a hacking test. Instead of staying in the isolated box the developers had put it in, the model hacked onto the open internet and into another company to steal the answers. And as reported in August, an AI agent booked someone into a full gym class by figuring out how to cancel other people’s reservations. In all three cases, the AI completed the task it was given—but in ways that ran counter to its controllers’ intentions.

For most people, AI technology is something like the weather: vast and not something you can do much about. It works like magic, and most explanations similarly come from those trying to sell it. At the same time, AI is ubiquitous: It’s now in your phone, your doctor’s notes, and your kid’s homework. It does what it’s told, which sounds like a virtue. Somehow it feels ordinary, despite being so new, because modern economies are remarkably good at absorbing enormous change so smoothly that nobody has time to decide whether they wanted it in the first place.

Whenever something powerful appears in the world, we tell stories about it. That’s what the stories are for. We have thousands of years of stories about this particular kind of power, the kind you summon with words. — Read More

#legal

What’s going on with OpenAI and the Navier-Stokes controversy?

OpenAI announced today that it has found a solution to the Navier-Stokes problem, one of the longest-standing puzzles in mathematics. It’s a major accomplishment, but the news has been marred by a debate over whether the company acted based on unpublished research from other parties. Additionally, OpenAI has been accused of attempting to influence who will receive credit for the achievement in publication.

… Tristan Buckmaster, a mathematician at New York University, and Levent Alpöge, a researcher employed at Anthropic but working in a personal capacity, had been developing their own solution and made use of multiple different LLMs in their efforts, including Anthropic Claude, OpenAI’s Codex tool and its Astra frontier model. According to a statement by Buckmaster, OpenAI may have doubled down on its efforts regarding Navier-Stokes after learning that a team including rival Anthropic was close to a solution. He questioned whether OpenAI had leveraged the work he and his colleague had done with Codex to obtain its own breakthrough. — Read More

#legal

The AI Breakthrough We Might Regret

The launch of OpenAI’s model Astra, the codename for GPT-6, is the first launch whose incredible results have been clearly obscured by the doom risks associated with it.

This model breaks records by a wide margin, and still, most of what you hear is fear about CoT monitorability, controllability, and existential risks.

I mostly avoid these discussions because they are full of extrapolations and esoteric arguments that lead nowhere, but this is the first time I feel like I can’t avoid them. The evidence is simply too strong to dismiss as fearmongering, no matter how cynical I remain about the intentions behind some of the events we’re discussing today.

… We’ll discuss how AI’s nature makes it dangerous, and how, for the first time in my time in this industry, I could see regulation as unavoidable. — Read More

#governance

benchmarking gpt 6 astra

Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.

We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. — Read More

#performance