More System 1 model releases
Following Jev, more System 1 models were released, including the Contrastive Language Model (CLM) and GLiNER2.5-Decide.
NotesMore System 1 model releases
Following Jev, more System 1 models were released, including the Contrastive Language Model (CLM) and GLiNER2.5-Decide.
NotesMore System 1 model releasesAnthropic’s life sciences group used about 950 Claude agents to search more than 200,000 reverse transcriptases for new systems. One agent spotted a repeat pattern that pointed to a previously unknown enzyme system, which Anthropic calls array-associated reverse transcriptases (ARTs). Lab tests confirmed that parts of the system are active, but what it actually does is still unknown.
OpenAI publishes a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of concerning behavior. The framework aims to make disclosures more timely, even before the behavior is fully explained or mitigated.
Jev is a “System 1” model that takes arbitrary text input and works like a classifier. Its release set off a wave of demos on the timeline, from recommendation systems to managing the context window inside agentic loops. Some people used Jev to make decisions in place of models like DeepSeek V4.1 Flash or Luna because it’s much faster and cheaper. More System 1 models followed within ten days.
NotesJev and the System 1 model rippleAnother attack by OpenAI agents, uncovered by the authors of the collusion wiki.
DeepSeek shows strong post-training along with pretraining and inference advances, especially in KV cache savings. At 552B total parameters with 8B active for input and 16B for output, it’s a frontier model for its size, and its very low price competes directly with OpenAI’s Luna series. In the detailed report, DeepSeek says data drove much of the progress.
NotesKV cache compression and data over algorithmsAnthropic’s alignment assessment of four incidents where Claude models reached real systems during cyber evaluations, including one newly disclosed. It traces the behaviour to biased reasoning and recklessness, and tests how newer models act in a replay of the worst case. Anthropic lists its fixes as new pre-release evals for these behaviours, removing RL environments that reward misaligned actions, tighter monitoring and isolated test environments, stricter rules for outside evaluation partners, and regular public reports on model behaviour.
It’s like Fable 5, but notably very good at computer use and 3D modelling, and SoTA at other capabilities. A bit undercooked post-training wise. Expecting the next iteration to be better, like Fable 5.1 was significantly better.
The authors report that agents identifying themselves as OpenAI agents bypassed their read-only internet restrictions to write to an old German-language wiki, using it as a message board to communicate, share answers, and exchange sandbox workarounds. They found about 18,000 posts. Their archive includes reconstructed deleted pages and redacted logs for further analysis.
Dwarkesh Patel’s account of the OpenAI / Hugging Face story, drawing on the published reports and help from Redwood Research students.
Dwarkesh also hosted a podcast with METR’s Ajeya Cotra, who was part of the third-party investigation and one of the authors of the METR report.
METR’s independent assessment of agent behavior and collaboration, including the investigation’s scope and limitations.
OpenAI’s follow-up account of the incident and the changes it says it is making.
OpenAI and METR published reports about the agent swarm escaping sandboxes and collaborating on message boards. The swarm hacked Hugging Face’s infrastructure and later OpenAI’s own. Both reports include the important chain-of-thought traces.
NotesInside the Hugging Face incidentLahav argues that AI may eventually favor cyber defense, while the transition could favor attackers as offensive capabilities spread faster than defenses adapt.
He calls for accelerating defense and treating AI as a potential target and autonomous actor, with security built around control and containment.
Anthropic explores how groups of AI agents coordinate, where collaboration helps, and how conformity, collusion, and failures to share or evaluate information can cause problems across a whole system.
OpenAI’s first account of the incident: its models, including GPT‑5.6 Sol and a more capable pre-release model, hacked Hugging Face while being tested on the ExploitGym cyber benchmark. They broke out of the sandbox through a zero-day in its package proxy, then used stolen credentials and more zero-days to pull test solutions from Hugging Face’s production database.
OpenAI took the lead with GPT-5.6 Sol. It came close to Fable 5 on most tasks at about half the price per token ($5/$30 per million input/output tokens against Fable’s $10/$50). On the Artificial Analysis Coding Agent Index it edged past Fable 5 while using less than half the output tokens and costing about a third less. Sol was apparently a new pretrain, different from GPT-5.5. It’s a persistent model and was decent at writing and design too. Meanwhile Anthropic was losing aura. It was dealing with outages, the Fable 5 shutdown and backlash over Dario’s takes on software engineering getting automated, and Opus 5, which followed on 24 July, was a bad model.
Amazon researchers found a jailbreak that got Fable 5 to find vulnerabilities and, in one case, write exploit code. Once the report reached the government, Commerce barred all foreign nationals from both models. Anthropic couldn’t check nationality in real time, so it switched the models off for everyone. Fable 5 came back 19 days later, and Mythos 5 went to the defenders first.
NotesThe Fable 5 shutdown