Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.
We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. — Read More
Tag Archives: Performance
Why AI Research Tools Struggle to Fact-Check the Web: Building an Adaptive Evidence Pool
This article grew out of the comments on my previous piece, Why AI Research Tools Struggle to Fact-Check the Web, where, using GPT Researcher as an example, I showed how common search engines and context-building methods undermine the very foundations of AI-based fact-checking.
Not that those tools were bad. They’re brilliant at standard (re)search tasks. But when finding genuinely diverse and independent sources matters, they start to show their limits. So we fixed that part. Adding source assessment through an admission policy was enough to make a difference.
But as the readers’ comments quickly showed, I was only touching the tip of the iceberg. — Read More
‘Welcome to the AGI era’: OpenAI launches GPT-6 Astra
The rumors were true, all of them (and then some): OpenAI today is releasing GPT-6 Astra, a new frontier model that the company says likely marks the onset of artificial generalized intelligence (AGI), its long sought goal of “highly autonomous systems that outperform humans at most economically valuable work.”
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
That is an unusually consequential framing even by the standards of frontier AI launches. But for enterprises, the more immediate significance of Astra may be considerably more concrete: OpenAI is positioning GPT-6 Astra as a new era of computing in which users, including employees, no longer have to click around a mouse or type on a keyboard ever again (if they don’t want).
… Astra is designed to navigate software much as a person does — working across browsers, spreadsheets, websites and desktop applications, producing finished documents and presentations, and carrying out multistep workflows rather than merely telling a user how to complete them. — Read More
LLMs: Intelligence vs. cost
ArtificialAnalysis is a website that benchmarks the intelligence of various LLM models. They publish a headline Intelligence Index, which is calculated as the mean output of the curated selection of benchmarks they run on each model. It’s a decent finger-in-the-air measure of how smart a model is overall.
AA also records useful information — namely, how much it cost them to run the benchmarks. Since the benchmarks are the same across all models, this offers a good indicator of how much it will cost a user to run each model, in relative terms.
One of their main plots is the Intelligence vs. cost plot, which shows the Pareto frontier, i.e. the cheapest model that can achieve each intelligence score. This frontier is important, because using a super-intelligent and super-expensive model to accomplish menial tasks that could be done by a much dumber and cheaper one is just a waste of money.
Over time, I’ve become progressively more irritated by this plot, for a few reasons. — Read More
AI token prices are hitting new record lows
A closely followed measure of artificial intelligence token prices touched fresh lows this week, the latest sign of deflating prices in an increasingly competitive landscape.
The LLM Token Expenditure Index, a key gauge of daily prices from intelligence firm Silicon Data, fell to 97 cents on Monday. That marked the index’s lowest reading since its creation late last year and has fallen by more than half from the high recorded earlier this summer.
….The recent drop is driven in part by the rise of open-source Chinese models like Moonshot’s Kimi K3 that can fetch lower prices than alternatives from leading frontier labs, according to a Tuesday post from Charles-Henry Monchau, investing chief at Syz Group. — Read More
Base Models Stopped Being the Bottleneck
… Base models embed a lot of raw knowledge inside them, and this knowledge scales with the number of parameters, but we can use that base model and prune it in order to steer it through post-training to be good at specific tasks, even with a reduction of their core parameters. We model the raw knowledge to become performant in the tasks we are interested in (model is becoming a loaded word).
I don’t know about you, but with how things are evolving, I am seeing closer and closer the day where we can run these models “cheaply” at home (or in the cloud, for those with a bigger risk appetite), and we can have plug and play AI. We are getting closer to “the solar panel moment.” — Read More
Anonymous Ox Alpha processes 26T tokens on OpenCode, breaks OpenRouter launch record
Dax Raad (@thdxr)’s OpenCode said its users processed 26 trillion tokens through Ox Alpha during the anonymous AI model’s first four days, turning a free preview into one of the largest model trials on the coding agent.
The August 24th disclosure covered 327,000 unique users and 8,328,244 completed sessions, according to OpenCode’s usage dashboard. Ox Alpha ranked second among models tracked by OpenCode, behind DeepSeek V4 Flash at 33 trillion tokens and ahead of Xiaomi’s MiMo-V2.5 at 12 trillion. — Read More
7 lessons for IT leaders on using observability to monitor AI applications
Over six months, the Elastic IT team ran internal AI applications that returned $2.5 million in operational time to the business.1 A conversational support assistant moved us from zero digital resolution, where anything complex became a ticket, to 30% of support interactions closing without one.
We can put those numbers in front of a finance team because we measured them from day one at the level of individual usage events. For each use case, we assigned a conservative time-saving goal and validated it with the teams doing the work. For example, a support case summary saves about five minutes. And, using a simple formula (events*minutes saved*a standard burden), the ROI of the application is now a real-time KPI rather than simply assuming that it might be valuable because the application uses generative AI. The hours came back to support engineers who had been searching for answers and went toward work on the roadmap.
Most organizations are not in that position yet. In our Landscape of Observability survey of 500 IT decision-makers, 85% said they planned to enable observability for their large language model (LLM) applications. Only 8% had done it. Teams see the value, they just don’t seem to be prioritizing it. — Read More
Agentic brand drift: How AI-orchestrated organizations will lose their identity and how to get it back
Agentic AI is reshaping how organizations operate. As autonomous systems take over pricing, content, personalization, and supply chain decisions, the human choices that historically built brand identity are progressively displaced. We term the result agentic brand drift: the gradual, unintended divergence between a firm’s intended brand identity and the emergent brand character produced by its AI-orchestrated operations. Unlike brand inconsistency or deliberate identity change, agentic brand drift is internally generated, has no triggering event, and co-occurs with improving performance metrics, making it invisible to conventional monitoring. Critically, this failure mode falls outside the scope of existing AI governance frameworks such as NIST AI RMF and ISO/IEC 42001, which govern system behavior rather than meaning coherence. A firm executing those frameworks flawlessly will still experience agentic brand drift. We theorize three mechanisms, Decision Diffusion, Temporal Collapse, and Accountability Dissolution, operating as a causal sequence, and derive two complementary frameworks, CORE and GUARD, that give organizations capabilities existing governance does not provide: identifying which decisions carry identity stakes, supplying agents with organizational reasoning behind past brand choices, and monitoring output patterns for identity coherence over time. — Read More
Anthropic set AI agents loose on the same task. They started a turf war
What happens when you pit AI agents against each other? According to Anthropic’s testing, things get messy fast.
On Thursday, Anthropic’s Frontier Red Team published new research examining how groups of AI agents behave when they encounter each other in the wild. The findings provide a glimpse into potential risks that could develop as companies and governments move to implement agents working autonomously across shared codebases, markets, and computer systems. — Read More