AI agents

Putting agents to work: evaluation, reliability, tool use and cost. Every brief item about AI agents, newest first, each with its primary source.

Subscribe via RSS Updated · Oct 8, 2026

Where it stands

Agents are moving from whether they can do the job to whether they can run reliably and safely. ThinkingBox on October 5 showed a wide gap between succeeding at least once and succeeding every time: 79.09% vs 47.53% for Claude Opus 5. October 7 brought progress on the runtime: Microsoft made MXC generally available on Windows so agents can run persistently in the background under OS control; GitHub Copilot's local sandboxing is built on MXC too, restricting which files, networks and credentials agent commands can touch; and NVIDIA's RTX Spark laptops bring a 125B-parameter model onto the device with up to 128GB of unified memory.

Agent isolation and governance are moving down into the operating system, and always-on agents that keep data on the device now have support from both the OS and the hardware.

What to watch next: which third-party agents get built on MXC, how good local models really are inside always-on agents, and whether more evaluations switch to repeated runs and end-state checks to measure reliability.

My advice: before shipping any agent that writes data, define its isolation policy and state checks and look at pass rates over repeated runs; if you're considering local deployment, validate local model quality on your own tasks before investing in hardware.

Darius · Updated Oct 8, 2026

Timeline 3 items

  1. Oct 8 · Thu · 2 items

  2. Oct 5 · Mon · 1 item