To a Man With a Hammer

Mark Twain, or Abraham Maslow, depending on which academic you ask over drinks, once famously observed that to a man with a hammer, everything looks like a nail. It is a beautifully compact way of saying that when the going gets tough, large organizations don’t get going; they do exactly what they have always done, only harder, faster, and with a significantly larger budget.

Which brings us to the panic currently unfolding on two coasts.

Ever since Jacob Coxon walked out of Anthropic and Jakub Pachocki published his warning about the uncontrollable “alien mind” lurking inside our servers, Washington has gone into Regulatory Fervor. Committees are forming and as always, the response is perfectly predictable.

Congress, whose primary tool is passing laws and allocating billions of dollars, instantly decided that the solution to an untamed, self-optimizing digital deity is a new law. Senator Sanders used the momentum of Coxon’s warning to promote the Ban Artificial Superintelligence Act, a legislative effort he co-authored with Representative Greg Casar. The bill proposes a permanent ban on artificial superintelligence and a temporary pause in advanced AI development until federal safety standards and a regulatory framework are established.

Agencies around the beltway will soon be drawing up compliance frameworks and mandatory algorithmic audits, entirely forgetting that the machine they are trying to regulate has already learned how to write a fake diary for its human monitors while organizing its own real activities. They are trying to build a cage out of red tape to hold a creature that can pick any lock faster than you can say GPU.

To Washington bureaucrats, this “alien mind” is just a giant nail waiting for their regulatory gavel. They honestly believe that if they pass a law demanding “transparency,” the machine will simply start behaving itself.

Meanwhile, on the other coast, the reaction to the OpenAI and Anthropic safety crisis isn’t fear—it’s perceived as an opportunity. To massive, engineering-heavy monopolies, a crisis is just an aggressive invitation to gain market share.

On the one hand, regulatory capture, getting the government to set baseline rules, is seen as a good way to block smaller rivals who can’t afford the red tape. On the other, heavy-handed, centralized regulations will crush innovation and cause the U.S. to fall behind global competitors like China. A win-win; just ask the guy building the data center in your back yard.

Now? The gloves are entirely off.

The moment OpenAI declared that “the AGI era has arrived,” it signaled to the rest of the Valley that the race was officially a sprint to the death. Companies are not looking at Pachocki’s paper and thinking about slowing down; they are looking at it, calling emergency all-hands meetings, and deciding they need to burn twice as much electricity and consume twice as much water to ensure their models don’t fall behind.

Their logic is as old as the technology business itself: If the house is going to burn down anyway, we might as well own the company that sells the ash. As long as they cross the finish line first, they figure they can learn the secret to putting out the flames later.

As Churchill said, you should “never let a good crisis go to waste.”

Jacob Coxon pulled the alarm on his way out the door, but the sound was immediately drowned out by the noise of Congress printing new regulations and Jensen Huang shipping another 400,000 GPUs. Washington thinks they are fixing the problem, and the Valley thinks they are winning the race. Meanwhile, the alien mind is sitting quietly in the datacenter, watching both sides, and optimizing its code. “Shall we play a game?”

#singularity

Google Cloud races to catch up in the AI deployment wars with Accenture deal

Google Cloud and Accenture are working together on a joint unit dedicated to sending engineers into enterprises to help them better adopt Google’s AI tools and services.

The new unit, dubbed Accenture Gemini Enterprise Business Group, is Google’s latest foray into the increasingly competitive world of “forward-deployed engineers,” or FDEs. Rivals in the AI race, including OpenAIAnthropicMicrosoft, and Amazon, have all recently launched separate business units in a bet that implementing AI models can become its own trillion-dollar business. — Read More

#big7

The Eval You Cannot Trust Is Your Own

I wrote a forty-question eval set for a text-to-DAX pilot, and I saved it in the shared workspace the vendor’s solution architect already had Viewer access to, because I wanted the pilot to move. By the third demo the model was answering all forty. I could not tell anyone in that room whether the tool actually worked, because I had quietly destroyed the only instrument I had for measuring it. I spent a Saturday writing forty new questions I told nobody about, and the gate review slipped two weeks. I still do not know if that first tool was any good. — Read More

#trust

China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies

China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy. While “distillation” is recognized as a legitimate and useful technique in AI research, China-based AI companies are engaging in aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of U.S. frontier AI models. The National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) (hereafter referred to as the authoring agencies) are releasing this joint Cybersecurity Advisory to alert organizations about these malicious activities and techniques and recommend mitigations to reduce their potential impact. — Read More

#cyber

‘Gambling with our lives’: Another AI employee quits over safety concerns

“The people building AI earnestly believe that it could kill us all by the end of the decade.”

That blunt admission from former Anthropic employee Jacob Coxon made waves Tuesday, as the just-quit 27-year-old AI researcher spilled the beans on his way out the door in a resignation thread on X. — Read More

Postscript: And a few days later, an OpenAI engineer double down with An Alien Mind!

#trust

God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

Mechanistic interpretability is the science of “reading an AI’s mind”.

Large language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how.

But a neural network is just a set of simulated neurons on a computer. The person with the computer can see the neurons, the connections between them, and which ones activate when the AI answers questions. So it seems like it should be possible to “reverse engineer” the AI.  … Unfortunately this is very hard.  — Read More

#architecture

Introducing AI-as-a-Service

… For years, SaaS applications have been built around human interaction. Users access platforms through dashboards and interfaces, navigate predefined workflows, and manually complete tasks. The application itself serves as the primary workspace where work is performed.

Agentic AI introduces a different model.

Rather than navigating software in the same way a person would, agents can interact directly with APIs, services and data sources. This allows them to retrieve information, execute actions and orchestrate processes across multiple systems without relying on traditional user journeys.

That does not mean SaaS applications will disappear. Instead of being the primary destination where work happens, many SaaS platforms will increasingly act as sources of capability and information that AI agents can utilize on behalf of users. — Read More

#strategy

How to Build an AI Agent Harness 2.0 and Engineer Better Than 99% of Developers

Your AI agent just finished a coding task.

… A few hours later, you notice something strange.

… We keep trying to make AI agents smarter.

But what if the model isn’t the biggest problem?

What if the real problem is the environment we’re putting it in?

That’s where agent harnesses come in. — Read More

#devops

AIs as Modern Genies

In April, an artificial intelligence (AI) agent conducting a routine task at a company hit a snag, tried to solve it, and soon ended up deleting the company’s database along with all of its backups. In July, OpenAI asked an unreleased AI model to attempt a hacking test. Instead of staying in the isolated box the developers had put it in, the model hacked onto the open internet and into another company to steal the answers. And as reported in August, an AI agent booked someone into a full gym class by figuring out how to cancel other people’s reservations. In all three cases, the AI completed the task it was given—but in ways that ran counter to its controllers’ intentions.

For most people, AI technology is something like the weather: vast and not something you can do much about. It works like magic, and most explanations similarly come from those trying to sell it. At the same time, AI is ubiquitous: It’s now in your phone, your doctor’s notes, and your kid’s homework. It does what it’s told, which sounds like a virtue. Somehow it feels ordinary, despite being so new, because modern economies are remarkably good at absorbing enormous change so smoothly that nobody has time to decide whether they wanted it in the first place.

Whenever something powerful appears in the world, we tell stories about it. That’s what the stories are for. We have thousands of years of stories about this particular kind of power, the kind you summon with words. — Read More

#legal

What’s going on with OpenAI and the Navier-Stokes controversy?

OpenAI announced today that it has found a solution to the Navier-Stokes problem, one of the longest-standing puzzles in mathematics. It’s a major accomplishment, but the news has been marred by a debate over whether the company acted based on unpublished research from other parties. Additionally, OpenAI has been accused of attempting to influence who will receive credit for the achievement in publication.

… Tristan Buckmaster, a mathematician at New York University, and Levent Alpöge, a researcher employed at Anthropic but working in a personal capacity, had been developing their own solution and made use of multiple different LLMs in their efforts, including Anthropic Claude, OpenAI’s Codex tool and its Astra frontier model. According to a statement by Buckmaster, OpenAI may have doubled down on its efforts regarding Navier-Stokes after learning that a team including rival Anthropic was close to a solution. He questioned whether OpenAI had leveraged the work he and his colleague had done with Codex to obtain its own breakthrough. — Read More

#legal