AI agents
Putting agents to work: evaluation, reliability, tool use and cost. Every brief item about AI agents, newest first, each with its primary source.
Where it stands
Agents are moving from whether they can do the job to whether they can run reliably and safely. ThinkingBox on October 5 showed a wide gap between succeeding at least once and succeeding every time: 79.09% vs 47.53% for Claude Opus 5. October 7 brought progress on the runtime: Microsoft made MXC generally available on Windows so agents can run persistently in the background under OS control; GitHub Copilot's local sandboxing is built on MXC too, restricting which files, networks and credentials agent commands can touch; and NVIDIA's RTX Spark laptops bring a 125B-parameter model onto the device with up to 128GB of unified memory.
Agent isolation and governance are moving down into the operating system, and always-on agents that keep data on the device now have support from both the OS and the hardware.
What to watch next: which third-party agents get built on MXC, how good local models really are inside always-on agents, and whether more evaluations switch to repeated runs and end-state checks to measure reliability.
My advice: before shipping any agent that writes data, define its isolation policy and state checks and look at pass rates over repeated runs; if you're considering local deployment, validate local model quality on your own tasks before investing in hardware.
Darius · Updated Oct 8, 2026
Timeline 3 items
-
Oct 8 · Thu · 2 items
- Tools github.blog ↗Local sandboxing for GitHub Copilot is generally available, restricting agent commands’ file, network and credential access with enforceable enterprise policies
Take · The biggest worry about letting coding agents run freely has always been what they can touch on your machine. I'd set the team default to: write only to the current repo, no access to Git credentials, network via an allowlist, with local MCP servers included. That beats reviewing every command after the fact, and teams on other coding agents should hold their isolation to the same bar.
From the Oct 8 brief · item 03 → - Industry blogs.nvidia.com ↗Microsoft makes MXC agent containers generally available on Windows as NVIDIA opens RTX Spark laptop preorders with 1 petaflop of FP4 and up to 128GB of unified memory
Take · This is the same MXC that powers Copilot's local sandboxing above, so the isolation layer for running agents on your own machine now ships with the OS. With 128GB of unified memory holding a 100B-plus model, keeping private data on the device can be validated on a single machine, and products handling private files like my AI Cloud Drive could consider a local-inference edition. Don't rush the hardware purchase, though: first check whether a local model is good enough for your workload.
From the Oct 8 brief · item 05 →
-
-
Oct 5 · Mon · 1 item
- Research huggingface.co ↗Microsoft open-sources ThinkingBox: 507 business workflows, 20 runs each, graded on final database state
Take · Single-attempt pass rates flatter agents. I'm applying the same idea to scoring in my AI Interview product: run the same answer sheet several times and check that what lands in the database agrees, not what the model says. When picking a model, put all-20-pass rate and cost per dependable task in your eval sheet; it tells you more than a leaderboard.
From the Oct 5 brief · item 01 →
-