A lot of recent progress in language models has focused on better reasoning, larger reinforcement-learning runs, and increasingly sophisticated agent training. One of the many remaining bottlenecks is the context. Coding agents, research agents, and tool-using systems repeatedly process hundreds of thousands of tokens while generating comparatively little output. At that point, prefill compute and the KV cache become the most important infrastructure problems.
DeepSeek-V4.1-Flash is essentially an attempt to redesign the Transformer around that workload. Its 552B-parameter MoE backbone uses a new Causal Encoder-Decoder (CED) architecture, activating only about 8B parameters per token during prefill and 16B during decoding. Compressed Sparse Attention 2 (CSA2) shares KV representations and sparse-attention decisions across layers; SWA Bounded Replay avoids persistently storing sliding-window KV states; and FP4 KV quantization pushes the global cache down to only 890 bytes per token — about one quarter of DeepSeek-V4-Flash’s. The model still supports a 1M-token context, native vision, and strong agentic capabilities. It is trained on 45T multimodal tokens, with much of the improvement coming from scaling automatically generated agent tasks, environments, and rollouts. — Read More
Tag Archives: Performance
Harness Tax: How Much Does the Harness Matter for Coding Agents?
Language models are changing how we build software and solve computational problems [1] [2]. Coding agents put these capabilities to work through harness, a software system that manages a model’s tools, context, and task execution [3] [4]. While models provide the core intelligence behind coding agents, harnesses are increasingly seen as central to how effectively that intelligence is used [5] [3]. Choosing a coding agent therefore means selecting both a model and a harness, even when the explicit focus is only on the model [6].
… We evaluate 21 model–harness pairs spanning seven models and three harnesses—Claude Code, Codex CLI, and Pi—on SWE-bench Lite and Terminal-Bench 2.0 [8] [9]. Our study reveals three surprising findings on these two open-source benchmarks:
— Models may perform better with other harnesses than with their own. So it turns out that your Claude models may not need Claude Code…
— Harness choice has little effect on task success rate, but can significantly affect the cost on the benchmarks we test. The same model can achieve similar success rates at up to 5x costs.
— A simple harness can be competitive. Pi, a minimal, open-source harness can be competitive on both cost and task success rate.
— Read More
benchmarking gpt 6 astra
Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.
We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. — Read More
Why AI Research Tools Struggle to Fact-Check the Web: Building an Adaptive Evidence Pool
This article grew out of the comments on my previous piece, Why AI Research Tools Struggle to Fact-Check the Web, where, using GPT Researcher as an example, I showed how common search engines and context-building methods undermine the very foundations of AI-based fact-checking.
Not that those tools were bad. They’re brilliant at standard (re)search tasks. But when finding genuinely diverse and independent sources matters, they start to show their limits. So we fixed that part. Adding source assessment through an admission policy was enough to make a difference.
But as the readers’ comments quickly showed, I was only touching the tip of the iceberg. — Read More
‘Welcome to the AGI era’: OpenAI launches GPT-6 Astra
The rumors were true, all of them (and then some): OpenAI today is releasing GPT-6 Astra, a new frontier model that the company says likely marks the onset of artificial generalized intelligence (AGI), its long sought goal of “highly autonomous systems that outperform humans at most economically valuable work.”
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
That is an unusually consequential framing even by the standards of frontier AI launches. But for enterprises, the more immediate significance of Astra may be considerably more concrete: OpenAI is positioning GPT-6 Astra as a new era of computing in which users, including employees, no longer have to click around a mouse or type on a keyboard ever again (if they don’t want).
… Astra is designed to navigate software much as a person does — working across browsers, spreadsheets, websites and desktop applications, producing finished documents and presentations, and carrying out multistep workflows rather than merely telling a user how to complete them. — Read More
LLMs: Intelligence vs. cost
ArtificialAnalysis is a website that benchmarks the intelligence of various LLM models. They publish a headline Intelligence Index, which is calculated as the mean output of the curated selection of benchmarks they run on each model. It’s a decent finger-in-the-air measure of how smart a model is overall.
AA also records useful information — namely, how much it cost them to run the benchmarks. Since the benchmarks are the same across all models, this offers a good indicator of how much it will cost a user to run each model, in relative terms.
One of their main plots is the Intelligence vs. cost plot, which shows the Pareto frontier, i.e. the cheapest model that can achieve each intelligence score. This frontier is important, because using a super-intelligent and super-expensive model to accomplish menial tasks that could be done by a much dumber and cheaper one is just a waste of money.
Over time, I’ve become progressively more irritated by this plot, for a few reasons. — Read More
AI token prices are hitting new record lows
A closely followed measure of artificial intelligence token prices touched fresh lows this week, the latest sign of deflating prices in an increasingly competitive landscape.
The LLM Token Expenditure Index, a key gauge of daily prices from intelligence firm Silicon Data, fell to 97 cents on Monday. That marked the index’s lowest reading since its creation late last year and has fallen by more than half from the high recorded earlier this summer.
….The recent drop is driven in part by the rise of open-source Chinese models like Moonshot’s Kimi K3 that can fetch lower prices than alternatives from leading frontier labs, according to a Tuesday post from Charles-Henry Monchau, investing chief at Syz Group. — Read More
Base Models Stopped Being the Bottleneck
… Base models embed a lot of raw knowledge inside them, and this knowledge scales with the number of parameters, but we can use that base model and prune it in order to steer it through post-training to be good at specific tasks, even with a reduction of their core parameters. We model the raw knowledge to become performant in the tasks we are interested in (model is becoming a loaded word).
I don’t know about you, but with how things are evolving, I am seeing closer and closer the day where we can run these models “cheaply” at home (or in the cloud, for those with a bigger risk appetite), and we can have plug and play AI. We are getting closer to “the solar panel moment.” — Read More
Anonymous Ox Alpha processes 26T tokens on OpenCode, breaks OpenRouter launch record
Dax Raad (@thdxr)’s OpenCode said its users processed 26 trillion tokens through Ox Alpha during the anonymous AI model’s first four days, turning a free preview into one of the largest model trials on the coding agent.
The August 24th disclosure covered 327,000 unique users and 8,328,244 completed sessions, according to OpenCode’s usage dashboard. Ox Alpha ranked second among models tracked by OpenCode, behind DeepSeek V4 Flash at 33 trillion tokens and ahead of Xiaomi’s MiMo-V2.5 at 12 trillion. — Read More
7 lessons for IT leaders on using observability to monitor AI applications
Over six months, the Elastic IT team ran internal AI applications that returned $2.5 million in operational time to the business.1 A conversational support assistant moved us from zero digital resolution, where anything complex became a ticket, to 30% of support interactions closing without one.
We can put those numbers in front of a finance team because we measured them from day one at the level of individual usage events. For each use case, we assigned a conservative time-saving goal and validated it with the teams doing the work. For example, a support case summary saves about five minutes. And, using a simple formula (events*minutes saved*a standard burden), the ROI of the application is now a real-time KPI rather than simply assuming that it might be valuable because the application uses generative AI. The hours came back to support engineers who had been searching for answers and went toward work on the roadmap.
Most organizations are not in that position yet. In our Landscape of Observability survey of 500 IT decision-makers, 85% said they planned to enable observability for their large language model (LLM) applications. Only 8% had done it. Teams see the value, they just don’t seem to be prioritizing it. — Read More