<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Coding agents · AI Brief · darius.wiki</title><link>https://www.darius.wiki/en/topics/coding-agents/</link><atom:link href="https://www.darius.wiki/en/topics/coding-agents/rss.xml" rel="self" type="application/rss+xml"/><description>What Claude Code, Codex, Copilot, Cursor and similar tools can and can't do</description><language>en</language><lastBuildDate>Fri, 09 Oct 2026 00:00:00 GMT</lastBuildDate><item><title>GitHub Copilot CLI 1.0.94-0 discovers local Ollama models in /model and switches to them in-session</title><link>https://www.darius.wiki/en/daily/2026-10-09/github-copilot-cli-local-models/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-09/github-copilot-cli-local-models/</guid><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>A coding agent can now switch between local and cloud models as needed without restarting the CLI. Take: The easy trap is the last point: switching to a local model doesn't mean your code stays on the machine. For teams with confidentiality requirements, I'd put COPILOT_OFFLINE=true in shared config instead of relying on everyone to remember to switch models.</description></item><item><title>Codex CLI 0.161.0 makes GPT-6.1 Sol the default and lets you pick a cyber access program per turn</title><link>https://www.darius.wiki/en/daily/2026-10-09/codex-cli-0-161/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-09/codex-cli-0-161/</guid><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>This release changes the default model and smooths permissions and MCP sign-in, worth a look before upgrading. Take: When the default model changes, the same scripts change in cost and output style. If you run codex exec in CI, I'd pin the model explicitly and switch only after comparing Sol's results and bill on non-critical jobs.</description></item><item><title>Local sandboxing for GitHub Copilot is generally available, restricting agent commands’ file, network and credential access with enforceable enterprise policies</title><link>https://www.darius.wiki/en/daily/2026-10-08/github-copilot-local-sandboxing/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-08/github-copilot-local-sandboxing/</guid><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>Tools and commands Copilot runs on a developer’s own machine can now run inside a policy-restricted sandbox, at no extra cost. Take: The biggest worry about letting coding agents run freely has always been what they can touch on your machine. I'd set the team default to: write only to the current repo, no access to Git credentials, network via an allowlist, with local MCP servers included. That beats reviewing every command after the fact, and teams on other coding agents should hold their isolation to the same bar.</description></item><item><title>GitHub ships a purpose-built model for leaked secrets that reads surrounding code to catch unformatted passwords, coming to push protection and Copilot /security-review</title><link>https://www.darius.wiki/en/daily/2026-10-08/github-secret-detection-model/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-08/github-secret-detection-model/</guid><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><category>Safety</category><description>GitHub is applying its own fine-tuned model to secret scanning across alerts, push protection and Copilot security reviews, and is publishing the billing model ahead of time. Take: With agents committing at every step, the odds of a secret slipping into history only go up, and blocking it at push time is far cheaper than rotating it later. I'd pilot push protection on repos that already have GHSP and set a budget cap on AI Credits, since checks bill even when nothing is blocked. GitHub's note that agents shouldn't enable credit-consuming features without explicit authorization is worth copying straight into a team's agent rules.</description></item><item><title>AWS releases the aws-ai-ml skill so Claude Code, Codex and Kiro can benchmark SageMaker inference endpoints and generate deployment code</title><link>https://www.darius.wiki/en/daily/2026-10-06/aws-sagemaker-inference-agent-skill/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-06/aws-sagemaker-inference-agent-skill/</guid><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>A cloud vendor packages which-instance-and-how-to-deploy know-how as a coding-agent skill: you describe the goal, it writes code you can review and run. Take: Skills are becoming how cloud vendors compete for the coding-agent entry point: whoever's deployment know-how lives inside the agent gets picked more often. I like that it produces code rather than opaque actions, but benchmarks spin up real endpoints and real bills; give the agent a separate account with a budget cap before letting it run.</description></item><item><title>GitHub Copilot code review gets REST and GraphQL APIs; Balanced becomes the default effort level</title><link>https://www.darius.wiki/en/daily/2026-10-05/copilot-code-review-api/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-05/copilot-code-review-api/</guid><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><category>Tools</category><description>Reviews can now be triggered from your own scripts and internal tools, with the effort level set per request. Take: With an API, AI review can plug into your own release flow, for example raising the effort only when a change touches payments or permissions. The default effort changed, so review time and usage may shift too; check your bill and PR wait times this week. I wrote on the blog about how to split review work when AI writes most first drafts.</description></item></channel></rss>