Oct 8, 2026 · Thu · 5 items

Anthropic releases Claude Haiku 5.5, GPT-6 rolls out globally in ChatGPT, GitHub Copilot local sandboxing goes GA

Agents are getting both cheaper to run and easier to contain today: Haiku 5.5 cuts small-model prices to a tenth of Haiku 4.5, making it affordable to hand large numbers of subtasks to agents, while GitHub and Microsoft build sandboxing and secret blocking into the OS and the platform, adding guardrails for letting agents work unattended.

RSS
  1. Anthropic releases Claude Haiku 5.5 at $0.10 per million input tokens under 100K, about 75% cheaper than Haiku 4.5 on average; Sonnet 5.5 cache reads cut in half

    • Pricing is tiered by prompt length: $0.10 per million input and $0.50 per million output tokens up to 100K tokens, $0.50 and $2.50 above that, versus $1 and $5 for Haiku 4.5. Anthropic notes the new tokenizer uses slightly more tokens per task, and it still comes out about 75% cheaper on average.
    • Against Haiku 4.5: 72.4% vs 15.7% on OSWorld 2.1 (offline subset) and 39.2% vs 0.0% on Terminal-Bench 4.0. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, and Anthropic still recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding.
    • Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens starting today, which Anthropic says makes most agentic work about 20% cheaper; this week Max 5x and Max 20x subscribers start getting $100 and $200 in monthly API credits, and Team up to $500.
    • The model ID is claude-haiku-5-5, available on the Claude Platform, AWS, Google Cloud and Microsoft Azure; the Python and TypeScript SDKs also add beta support for computer use and browser use.

    Builder's takeI'm moving short-prompt calls, such as PandaClaws' classification, summaries and rewrites and AI Cloud Drive's document Q&A, onto Haiku 5.5 for a comparison run right away: under 100K tokens it costs a tenth of the old price, so a big model planning and Haiku doing the subtasks finally pencils out. The new tokenizer uses slightly more tokens, so compare cost per completed task rather than unit price; the cost table in my post on what it costs to keep three AI products online needs redoing too.

  2. GPT-6 rolls out globally in ChatGPT alongside Intelligent UI

    • GPT-6 is rolling out globally in ChatGPT.
    • It ships with Intelligent UI, which OpenAI says delivers faster responses with visuals and interactive experiences users can explore and use directly.

    Builder's takeFor product builders the point is user expectations: once people get answers with visuals they can click every day in ChatGPT, a plain-text chat box starts to feel dated. I'd first list which parts of AI Interview's scored reports and AI Cloud Drive's answers would read better as charts or cards, then decide whether to have the model output structured UI directly.

  3. Local sandboxing for GitHub Copilot is generally available, restricting agent commands’ file, network and credential access with enforceable enterprise policies

    • Generally available in Copilot CLI, the Copilot app and VS Code sessions using Agent Host, included with Copilot at no additional cost.
    • Limits which files and directories agent-run commands can read or modify, and controls access to the internet, local networks, Git credentials and GitHub CLI credentials; it also covers local MCP and language servers where supported.
    • Powered by Microsoft eXecution Container (MXC), which translates one sandbox policy into native OS controls on Windows, macOS and Linux; enterprises can require sandboxing through managed settings that developers cannot weaken.
    • Sandbox policies apply to tool execution regardless of which model Copilot uses.

    Builder's takeThe biggest worry about letting coding agents run freely has always been what they can touch on your machine. I'd set the team default to: write only to the current repo, no access to Git credentials, network via an allowlist, with local MCP servers included. That beats reviewing every command after the fact, and teams on other coding agents should hold their isolation to the same bar.

  4. GitHub ships a purpose-built model for leaked secrets that reads surrounding code to catch unformatted passwords, coming to push protection and Copilot /security-review

    • The model reads surrounding code to identify likely credentials, including passwords without a recognizable token format, without generating code or prose.
    • Existing AI-detected secret alerts switch to the new model starting today at no additional charge for GHSP and GHAS customers; AI-detected alerts are coming to GHES 3.23 in public preview.
    • AI push protection is in private preview and flags unstructured credentials before they enter repository history; checks from the classifier are coming soon to /security-review in Copilot CLI and the Copilot app, also in private preview.
    • The new push protection and /security-review checks are off by default and opt-in, and consume GitHub AI Credits; a check can consume credits even if it does not block a push.

    Builder's takeWith agents committing at every step, the odds of a secret slipping into history only go up, and blocking it at push time is far cheaper than rotating it later. I'd pilot push protection on repos that already have GHSP and set a budget cap on AI Credits, since checks bill even when nothing is blocked. GitHub's note that agents shouldn't enable credit-consuming features without explicit authorization is worth copying straight into a team's agent rules.

  5. Microsoft makes MXC agent containers generally available on Windows as NVIDIA opens RTX Spark laptop preorders with 1 petaflop of FP4 and up to 128GB of unified memory

    • Microsoft announced general availability of Microsoft Execution Containers (MXC), OS-level infrastructure that lets agents run safely and persistently in the background under operating system control.
    • RTX Spark pairs a Blackwell RTX GPU with up to 6,144 cores and an up to 20-core Grace CPU connected at 600 GB/s, for 1 petaflop of FP4 compute and up to 128GB of unified memory; laptop preorders open now with availability on October 16, and compact desktops go on sale in November.
    • NVIDIA says it can run the 125B-parameter Qwen 3.8 Flash Next locally, unmetered and without sending data to the cloud; systems are coming from Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte.
    • NVIDIA also previewed DGX Station for Windows: a GB300 superchip with 748GB of coherent memory and up to 20 petaFLOPS of FP4, which it says can run models up to trillion-parameter scale locally.

    Builder's takeThis is the same MXC that powers Copilot's local sandboxing above, so the isolation layer for running agents on your own machine now ships with the OS. With 128GB of unified memory holding a 100B-plus model, keeping private data on the device can be validated on a single machine, and products handling private files like my AI Cloud Drive could consider a local-inference edition. Don't rush the hardware purchase, though: first check whether a local model is good enough for your workload.

Researched and drafted with AI assistance; editorial standards and views set by Darius.