Oct 6, 2026 · Tue · 3 items

Aleph Alpha open-sources Kolibri, AWS ships a SageMaker inference skill for coding agents, OpenAI explains its EU text watermarking

No frontier launch over the weekend; the thread is deployment. Open models now compete on whether they fit on one GPU and how cheap long context is, cloud vendors are packaging deployment know-how as skills for coding agents, and regulation is reaching watermarks on generated text.

RSS
  1. Aleph Alpha open-sources Kolibri: a 78B MoE with 3.46B active per token, validated to 1M-token context, Apache 2.0

    • Mixture-of-experts with 78B total parameters and 3.46B active per token; German and English only, a deliberate choice of depth over breadth, with a tokenizer built for German.
    • Native context is 262,144 tokens; quality and serving efficiency are validated up to 1,048,576 tokens, with ≤262,144 recommended for latency-sensitive or complex tasks.
    • FP8 weights need about 78 GB: minimum 2× A100 80 GB, 2× H100, or one H200, B200 or B300. Served through vLLM with Aleph Alpha's plugin and an OpenAI-compatible API.
    • Reasoning effort can be set to low, medium, high or off, and tool calling can be combined with it; pre-trained on 20T tokens, released October 3 under Apache 2.0.

    Builder's takeAbout 3B active per token, one GPU, Apache 2.0: this is the shape of model that suits privately deployed long-document Q&A like my AI Cloud Drive. But it only covers German and English, so don't drop it into Chinese workloads; what I want from it is a real throughput number for sliding-window attention plus long context on my own hardware before choosing the next self-hosted model.

  2. AWS releases the aws-ai-ml skill so Claude Code, Codex and Kiro can benchmark SageMaker inference endpoints and generate deployment code

    • Available through the Agent Toolkit for AWS and works with any MCP-compatible coding agent, including Kiro, Claude Code and Codex.
    • The agent can benchmark endpoints, recommend deployment configurations, compare performance runs and generate executable SageMaker Python SDK v3 code.
    • Setup requires AWS CLI 2.35+ and uv; generated code runs under your own AWS credentials, and AWS says you can go from zero to a working conversation in 10 minutes.
    • Kiro and Claude Code can also discover and load skills at runtime through the AWS MCP Server without a local install.

    Builder's takeSkills are becoming how cloud vendors compete for the coding-agent entry point: whoever's deployment know-how lives inside the agent gets picked more often. I like that it produces code rather than opaque actions, but benchmarks spin up real endpoints and real bills; give the agent a separate account with a budget cap before letting it run.

  3. OpenAI sets out how it will watermark generated text under EU rules, with detection access starting with researchers

    • Published October 5 under OpenAI's Safety category, describing its approach to text watermarking under EU rules.
    • The post covers where watermarks apply and how detection works.
    • Access to detection starts with researchers, and OpenAI explains why.

    Builder's takeAll I could read was the official RSS summary, so the details will have to wait; the direction is clear, though: for products serving EU users, generated text may carry a detectable mark. A content-generation backend like PandaClaws should start recording which content is model-generated and by which model in its metadata now, so compliance doesn't mean rework later.

Researched and drafted with AI assistance; editorial standards and views set by Darius.