Aleph Alpha open-sources Kolibri, AWS ships a SageMaker inference skill for coding agents, OpenAI explains its EU text watermarking
No frontier launch over the weekend; the thread is deployment. Open models now compete on whether they fit on one GPU and how cheap long context is, cloud vendors are packaging deployment know-how as skills for coding agents, and regulation is reaching watermarks on generated text.
Mixture-of-experts with 78B total parameters and 3.46B active per token; German and English only, a deliberate choice of depth over breadth, with a tokenizer built for German.
Native context is 262,144 tokens; quality and serving efficiency are validated up to 1,048,576 tokens, with ≤262,144 recommended for latency-sensitive or complex tasks.
FP8 weights need about 78 GB: minimum 2× A100 80 GB, 2× H100, or one H200, B200 or B300. Served through vLLM with Aleph Alpha's plugin and an OpenAI-compatible API.
Reasoning effort can be set to low, medium, high or off, and tool calling can be combined with it; pre-trained on 20T tokens, released October 3 under Apache 2.0.
Builder's takeAbout 3B active per token, one GPU, Apache 2.0: this is the shape of model that suits privately deployed long-document Q&A like my AI Cloud Drive. But it only covers German and English, so don't drop it into Chinese workloads; what I want from it is a real throughput number for sliding-window attention plus long context on my own hardware before choosing the next self-hosted model.
Available through the Agent Toolkit for AWS and works with any MCP-compatible coding agent, including Kiro, Claude Code and Codex.
The agent can benchmark endpoints, recommend deployment configurations, compare performance runs and generate executable SageMaker Python SDK v3 code.
Setup requires AWS CLI 2.35+ and uv; generated code runs under your own AWS credentials, and AWS says you can go from zero to a working conversation in 10 minutes.
Kiro and Claude Code can also discover and load skills at runtime through the AWS MCP Server without a local install.
Builder's takeSkills are becoming how cloud vendors compete for the coding-agent entry point: whoever's deployment know-how lives inside the agent gets picked more often. I like that it produces code rather than opaque actions, but benchmarks spin up real endpoints and real bills; give the agent a separate account with a budget cap before letting it run.
Published October 5 under OpenAI's Safety category, describing its approach to text watermarking under EU rules.
The post covers where watermarks apply and how detection works.
Access to detection starts with researchers, and OpenAI explains why.
Builder's takeAll I could read was the official RSS summary, so the details will have to wait; the direction is clear, though: for products serving EU users, generated text may carry a detectable mark. A content-generation backend like PandaClaws should start recording which content is model-generated and by which model in its metadata now, so compliance doesn't mean rework later.
Researched and drafted with AI assistance; editorial standards and views set by Darius.