When an LLM answers a question, is it reasoning like humans, or just producing text that looks like reasoning? The distinction isn’t just philosophical, this determines what we can trust AI to do, how closely we need to supervise it, and ultimately what its real-world impact will turn out to be.
Melanie Mitchell at the Santa Fe Institute argues that we lack adequate methods for measuring machine cognition, and that AI is a form of “alien intelligence” that operates through non-human cognitive mechanisms. — Read More
Daily Archives: August 24, 2026
Measuring benchmark optimization in speech recognition
Public voice AI benchmarks increasingly suggest that models are performing at human levels. Yet those scores don’t always reflect how models work in the real-world. Since public benchmarks are open and widely used, models can also become optimized for the tests themselves. Their scores may improve because they have learned benchmark-specific patterns and not because they have become better at the underlying task.
… However, broader measurement alone does not solve the problem. This phenomenon, sometimes called benchmark optimization or “benchmaxxing,” is often discussed around machine learning, however, it has been difficult to measure in speech recognition.
Our latest research introduces three tests to help quantify it. — Read More
The AI-Native SDLC playbook
The traditional software development lifecycle (SDLC) is process-heavy to ensure accountability and control at each step. However, the traditional SDLC was designed to maximize efficiency in an era where the most time-consuming and expensive stage was writing and implementing code, which is no longer the case. PRDs, estimation rituals, and product security reviews all existed to force alignment during what could be weeks, months, or quarters of development work.
To better realize the productivity gains of and secure agentic AI, the traditional SDLC lifecycle requires the same level of transformation as the implementation phase has undergone. — Read More
This 1-Hour Andrej Karpathy Lecture Explains Modern AI Better Than Most Courses
For enterprises, the cautious AI era has begun
As companies mature in their AI deployment and double down on their previous investments into the technology, pressure on executives has reached a fever pitch
Early 2026 was the era of tokenmaxxing, or ramping up use of AI compute units as much as possible to appear productive. But momentum from AI providers to transition from flat-rate subscriptions to consumption-based pricing has ramped up in the last few months.
… The “use-AI-for-everything” mindset many companies adopted in 2025 and into 2026 now bears a much larger price tag. — Read More
Stanford researchers create viruses not found in nature using genomes designed by artificial intelligence
US researchers have for the first time successfully synthesised brand-new viruses not found in nature, based on designs generated by artificial intelligence.
The researchers’ paper, published in the journal Science on Thursday, details how scientists from Stanford University and the Arc Institute, a California-based non-profit dedicated to “high-risk, high-reward” research, were able to create the viruses using DNA sequences generated by two “genome language models” named Evo 1 and Evo 2. — Read More
Who’s behind the new ‘stealth model’ Ox Alpha?
A mysterious new AI model called Ox Alpha has driven certain corners of the internet into a frenzy of speculation about who actually built it.
The free model was released on OpenRouter on Thursday, where it was described as “a reasoning model designed for coding, sustained agentic work, and production workload.” On X, Stripe CEO Patrick Collison (whose company is acquiring OpenRouter) described Ox Alpha as “very impressive.” — Read More