<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI agents · AI Brief · darius.wiki</title><link>https://www.darius.wiki/en/topics/ai-agents/</link><atom:link href="https://www.darius.wiki/en/topics/ai-agents/rss.xml" rel="self" type="application/rss+xml"/><description>Putting agents to work: evaluation, reliability, tool use and cost</description><language>en</language><lastBuildDate>Sat, 10 Oct 2026 00:00:00 GMT</lastBuildDate><item><title>Google open-sources AQuA, a quality agent that samples production sessions to diagnose a live agent, at $3.76 for a 32-session sweep</title><link>https://www.darius.wiki/en/daily/2026-10-10/google-aqua-ambient-quality-agent/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-10/google-aqua-ambient-quality-agent/</guid><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>An agent that checks a production agent from outside the request path, reading but never changing it. Take: The step I value most is verification: have the model find issues, then check them against the raw transcripts to drop false positives, the same reason I insist on primary sources. Conversational products like AI Interview can copy this: sample a few dozen real sessions after each release for a few dollars, far cheaper than waiting for complaints.</description></item><item><title>Hugging Face shows ML-Intern: an agent trains six models in a few days for about $103 in total</title><link>https://www.darius.wiki/en/daily/2026-10-10/hugging-face-ml-intern/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-10/hugging-face-ml-intern/</guid><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>A machine learning agent that asks for a budget, runs small tests and then trains and publishes, now available in HuggingChat. Take: Start at zero budget and ask before spending is a design every agent that touches money should copy. For small teams, a specialized small model now costs a few dollars per experiment; I'd try it on fixed tasks such as content moderation or tagging first, then compare cost against calling a large model.</description></item><item><title>Local sandboxing for GitHub Copilot is generally available, restricting agent commands’ file, network and credential access with enforceable enterprise policies</title><link>https://www.darius.wiki/en/daily/2026-10-08/github-copilot-local-sandboxing/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-08/github-copilot-local-sandboxing/</guid><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>Tools and commands Copilot runs on a developer’s own machine can now run inside a policy-restricted sandbox, at no extra cost. Take: The biggest worry about letting coding agents run freely has always been what they can touch on your machine. I'd set the team default to: write only to the current repo, no access to Git credentials, network via an allowlist, with local MCP servers included. That beats reviewing every command after the fact, and teams on other coding agents should hold their isolation to the same bar.</description></item><item><title>Microsoft makes MXC agent containers generally available on Windows as NVIDIA opens RTX Spark laptop preorders with 1 petaflop of FP4 and up to 128GB of unified memory</title><link>https://www.darius.wiki/en/daily/2026-10-08/nvidia-rtx-spark-windows-mxc/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-08/nvidia-rtx-spark-windows-mxc/</guid><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><category>Industry</category><description>At Microsoft’s Windows AI and Surface event, the two companies pushed always-on local agents forward at both the OS and hardware layers. Take: This is the same MXC that powers Copilot's local sandboxing above, so the isolation layer for running agents on your own machine now ships with the OS. With 128GB of unified memory holding a 100B-plus model, keeping private data on the device can be validated on a single machine, and products handling private files like my AI Cloud Drive could consider a local-inference edition. Don't rush the hardware purchase, though: first check whether a local model is good enough for your workload.</description></item><item><title>Microsoft open-sources ThinkingBox: 507 business workflows, 20 runs each, graded on final database state</title><link>https://www.darius.wiki/en/daily/2026-10-05/microsoft-thinkingbox/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-05/microsoft-thinkingbox/</guid><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><category>Research</category><description>An agent saying it's done doesn't count; the records it leaves do. How it measures, what it found, and what dependable completion actually costs. Take: Single-attempt pass rates flatter agents. I'm applying the same idea to scoring in my AI Interview product: run the same answer sheet several times and check that what lands in the database agrees, not what the model says. When picking a model, put all-20-pass rate and cost per dependable task in your eval sheet; it tells you more than a leaderboard.</description></item></channel></rss>