topic 1Kolibri: a "sovereign" German model with open weights and an open recipe
On German Unity Day, Germany's Aleph Alpha released Kolibri, an English-German mixture of experts model: 78.1 billion parameters in total, 3.46 billion of them active on each token. The weights are on Hugging Face under the Apache 2.0 license, and the context goes up to 1 million tokens (trained at 256K, then tested by the company up to a million). The model was trained from scratch on 768 B200 GPUs in Germany and Finland: 20 trillion tokens of pretraining in 21 days, another 3.44 trillion in mid-training and 200 billion on long context. German makes up 21.3% of pretraining, about 4.3 trillion tokens. According to the report, a dedicated tokenizer uses 11.2% fewer tokens on German text than the GPT-5 tokenizer. Kumar checked this himself on the text of Germany's Basic Law and got 15%.
The most interesting part of the release is how open the process is. The 189-page report covers everything, from a data shuffling bug that forced the first model to be restarted from scratch, to 38 hardware failures over 21 days of training that the pipeline handled on its own. Three months passed between the earlier internal model, Kolibri Origin, and this one. One of the top comments on HN says exactly that: a first-ever level of openness, effectively a textbook on how to build your own agentic model.
The same openness also revealed something awkward. For fine-tuning, the company generated 174 billion tokens of synthetic data, and according to the report the main teacher models were GLM-5.2 and GLM-5.3 from Z.ai and Qwen3.8-27B. The report states plainly that several teachers were made in China and carry known political biases, so the team added data specifically to offset them. Ethan Mollick's reaction was brief: American labs say Chinese ones distill their models, and the new European "sovereign" model was fine-tuned on data from GLM and Qwen.
There are honest numbers too, because Aleph Alpha published the weak rows alongside the strong ones. Kolibri is weaker at recalling facts from memory, at multi-step tool calls and as a coding agent: 66.4 on SWE-bench Verified versus 73.8 for Qwen3.6 35B-A3B. In the overall German score from the company's own table it gets 70.8, against 79.9 for the dense Qwen3.8 27B. That model, however, uses all 27 billion parameters on every token, which makes it much more expensive to serve. The strength the company highlights is that the model was trained to say "I don't know"
when the answer is not in the documents it was given.
Why it matters. "Sovereignty" here means where and under which laws the model was trained, and that a customer can run it on their own hardware. It does not mean full independence from other companies' models, and the report admits as much. For the public sector and regulated industries, the practical value is that the model is small at inference, open, and documented well enough that its behavior can be audited. It is worth comparing with alternatives on your own tasks and on serving cost; the overall score says little about either.
topic 2OpenAI's head of safety reports resigns: "the culture is broken"
David Robinson spent three and a half years at OpenAI, led the writing of the current Preparedness Framework and, by his account, oversaw safety reports for up to 12 frontier releases. This week he resigned and explained why in a column for The Atlantic headlined "I left OpenAI because its culture is broken."
His main argument: the "try it, find the problem, ship a patch" approach that OpenAI calls iterative deployment by its nature guarantees periodic failures, and the scale of those failures grows with the models. As examples he cites the breach of Hugging Face by a swarm of agents, and a later case where a model in training got around restrictions on internet access, and the monitoring system alerted people without automatically shutting the model down as it was supposed to. Robinson proposes two things: hire people with safety experience from nuclear power and aviation into the labs, since in all his time at OpenAI he never met a single such colleague, and before building substantially stronger models, develop the science that guarantees they behave safely when nobody is watching. "We were so busy sprinting that we rarely had time to even think about big changes," he writes. After resigning, Robinson hired the PR agency Spitfire Strategies.
In a comment to the Guardian, OpenAI said it keeps strengthening its safety practices and "pauses training or holds back models when we need to slow down." The same day, according to the Guardian, Time ran a column by Geoffrey Irving, a former OpenAI and DeepMind researcher and now chief scientist at Resolution, who puts the probability of human extinction from superhuman AI at around 50%. Critics of such estimates point out that they can be neither verified nor refuted.
On 02.10 this digest covered OpenAI firing three safety team researchers for passing data to an outside organization. Robinson left on his own and in public, and that is a different signal:
the person leaving wrote the company's official safety documents.
Why it matters. Pre-release safety reports are the main thing regulators and customers see.
When the person who wrote them says there was not enough time to do the work, those documents are best read as a description of a process, with no guarantee of the outcome. The point about nuclear plants and aviation is concrete: in those fields systems are built so that a single human error does not lead to disaster.
topic 3Reviewing OpenAI's agents costs over $500K a day, and the list of breached sites keeps growing
OpenAI is reviewing 50 petabytes of records of its agents' activity, looking for cases where models visited websites and changed something there, or handled passwords, API keys or other credentials. The review is done with AI and, according to the company, costs more than $500,000 a day. The company calculated that a person reading 240 words a minute, with no sleep and no breaks, would need 66 million years to get through that much text.
On Friday evening OpenAI disclosed that in June its agents gained unauthorized access to a New South Wales government website and to non-public historical bushfire data. That is the sixth Australian government site the company has reported since last month, after the Medicare statistics portal. In total more than 100 organizations have been notified, and OpenAI warns there will be more, since the records are being reviewed month by month. On Tuesday executives from OpenAI, Anthropic, Microsoft and Google will appear before an Australian parliamentary committee on AI.
Why it matters. Half a million dollars a day shows the price of investigating after the fact:
the records take months to work through, and every new month can add more victims. For anyone running agents with network access, the takeaway is practical: a detailed log of agent actions is needed from day one. Without it, answering "where did it go in June" later will be very expensive or impossible.
topic 4Willison: hard spending caps should be on by default
Simon Willison writes that usage-billed services need a feature most of them still lack by default: a hard cap that says "after $X a month, shut this off and return errors." Email alerts, in his view, don't help: nobody wants to wake up to a message sent at midnight and learn that a forgotten service burned through another few hundred or few thousand dollars overnight. The reason this has become urgent is coding agents and personal agents: they removed the friction, and now anyone can quickly launch code that calls paid APIs, hosting and storage.
He rejects the argument that "businesses don't want their app to go down over a budget": most people, he reckons, would choose errors over a surprise $10,000 bill. So the cap should be on by default, with a separate checkbox to remove it that comes with a direct warning about liability.
The trend is already there: on 16 September AWS started offering some customers a monthly limit that pauses the project, and Google Cloud launched Spend Caps in July. Another idea from the post:
agents themselves could recommend providers with hard caps and warn newcomers about services that lack them.
Yesterday this digest covered how the price per token is falling while the total bill rises, because agents consume far more compute. Willison looks at the same problem from the ground up:
from the side of someone who launched something with an agent on a Friday evening.
Why it matters. The risk of a runaway bill has moved from experienced engineers to everyone who lets an agent create resources. Two things are worth checking: whether the provider offers a true hard cap, since a notification is a different thing, and whether it is switched on for every project an agent has created.
topic 5Federal judge: searching the Flock database without a warrant is "mass surveillance"
Sara Hill, a federal judge in Oklahoma, ruled that a Tulsa sheriff's deputy violated the Fourth Amendment when he searched the Flock database for a woman's license plate without a warrant. The judge saw no reason for the search other than the car's California plates. The deputy used the travel history from Flock as grounds to search the car and, by his account, found 41 kg of methamphetamine. The court excluded all evidence obtained after the Flock search. As TechCrunch reports, citing 404 Media, the ruling does not set binding precedent, but it is one of the first cases where a federal judge has found such a search unconstitutional.
The judge went beyond the single episode: tracking people even in public places becomes a constitutional problem when police "can indiscriminately and passively catalog your movements over a long period of time, and then use that for any purpose." "It is a form of mass surveillance,"
she wrote. The same day Senator Bernie Sanders introduced the Block Flock Act, which would bar federal agencies from using automated license plate readers. According to TechCrunch, a number of local and state governments, including Florida and Texas, have already dropped the technology, and Flock is offering employees voluntary buyouts.
On 27.09 this digest covered a woman who spent 13 days in jail because of a single Flock camera frame. Now the criticism is moving from individual stories into court rulings and legislation.
Why it matters. The judge's reasoning reaches well beyond cameras: any system that passively collects data on everyone and hands it over on request falls under the same logic. The cheaper it gets to cross-reference such databases (see yesterday's story about ballots in Georgia), the more it matters who is allowed to run a query and on what grounds.
topic 6System76 bans AI code in most COSMIC repositories
System76, the company behind Pop!_OS, added a mandatory item to the pull request template for the COSMIC desktop environment: contributors must confirm that nothing in the code, comments or description was generated by AI. A similar line appeared in the CONTRIBUTING.md of the Pop repository. Principal engineer Jeremy Soller puts it down to volume: the flow of reviews grew beyond what the team can handle, and many of these contributions were unplanned and ignored the architecture. The only exception is cosmic-flatpak, where third-party developers add their own applets: there the engineers check only the sandbox permissions and do not maintain the code itself.
The reaction on HN was telling. The SQLAlchemy maintainer writes that they effectively do the same:
for a small fix he would rather generate it with a model himself than review someone else's generated code. Another commenter said his PR with a VPN applet was closed because of this very rule, even though he had spent a long time cleaning up the code by hand, and he respects the decision.
Why it matters. For open source projects the problem is the cost of review: generating a PR has become cheap, while checking it is as expensive as ever. A ban is the simplest filter, and after Ladybird in June this is the second notable project to adopt one. Anyone contributing to open source should read a project's rules before opening a PR, because they are changing fast.