The thread this time is deployment, not raw capability: agents need to be judged on getting it right every time, small open models are taking over cited long-form writing, and enterprises are short on engineers who can wire models into the business.
Covers 507 stateful business workflows, each run 20 times per model, graded on terminal backend state and side effects rather than the agent's own text.
Among failed runs that ended cleanly with no tool error, executable checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%.
Kimi-K3 solves 93.89% of tasks at least once but only 13.41% on all 20 attempts; Claude Opus 5 solves 79.09% at least once and 47.53% on every attempt.
Framework (MIT) and benchmark data are on GitHub; by its cost-per-dependable-task metric, GPT-5.4 is lowest at an estimated $6.80.
Builder's takeSingle-attempt pass rates flatter agents. I'm applying the same idea to scoring in my AI Interview product: run the same answer sheet several times and check that what lands in the database agrees, not what the model says. When picking a model, put all-20-pass rate and cost per dependable task in your eval sheet; it tells you more than a leaderboard.
Built on Qwen3-8B; takes a research question plus retrieved literature excerpts and writes the full cited report in one pass instead of section by section.
In Asta, Fast mode averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, about 3.5x faster.
Post-trained on 47K filtered examples; Ai2 is releasing the model and the training data so others can reproduce and build on it.
Ai2 says open weights let institutions run it on their own infrastructure when research questions are sensitive.
Builder's takeRetrieve-then-write-a-cited-answer is the most common request in my AI Cloud Drive's document Q&A. An 8B model that does it in one call means private deployment and a much lower bill. I'll compare citation accuracy on my own documents before replacing the large-model call.
Copilot code review can be requested through the REST and GraphQL APIs, with an optional review effort level per request.
Balanced is now the default effort level for new and existing repositories and organizations; an explicit Lite setting is respected. The change took effect September 28, 2026.
Generally available on Copilot Pro, Pro+, Max, Business and Enterprise; effort can be set at enterprise, organization, repository or personal level, each overriding the one above.
Builder's takeWith an API, AI review can plug into your own release flow, for example raising the effort only when a change touches payments or permissions. The default effort changed, so review time and usage may shift too; check your bill and PR wait times this week. I wrote on the blog about how to split review work when AI writes most first drafts.
Backed by a $100 million commitment, with a goal of 10,000 Frontier Deployed Engineers (FDEs) by the end of 2027.
Engineers pass an in-person program and graded practical, then lead a real Claude use case at their own organization in a 12-week residency; the first FDE badges are expected in early 2027.
First cohorts run in San Francisco, New York and London, with engineers from Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley and Novo Nordisk.
Builds on the Claude Partner Network, where people across 46,000 firms hold more than 175,000 Claude certifications.
Builder's takeWhen a model vendor starts training the people who wire models into businesses, the enterprise bottleneck clearly isn't capability but integration and delivery. For independent developers and small teams, a track record of shipping real use cases is worth more than a badge, and that's where the work is.
Researched and drafted with AI assistance; editorial standards and views set by Darius.