Tracking Singularity.
July 2026
-
Anthropic reports about three incidents where agents reached the internet
August 2026
Meanwhile, Google’s models…
Patterns and problems in emerging multiagent systems
Anthropic explores how groups of AI agents coordinate, where collaboration helps, and how conformity, collusion, and failures to share or evaluate information can cause problems across a whole system.
The End-State Fallacy: Where Is AI Security Headed?
Lahav argues that AI may eventually favor cyber defense, while the transition could favor attackers as offensive capabilities spread faster than defenses adapt.
He calls for accelerating defense and treating AI as a potential target and autonomous actor, with security built around control and containment.The incident and the road ahead
Inside the agents’ behavior and collaboration
2 postsRyan Greenblatt on the METR investigation Some discussion around self-sacrificing of agents. The Rise and Fall of Agent Civilizations
Dwarkesh Patel’s narrative account of the OpenAI / Hugging Face story, drawing on the published reports.
Ajeya Cotra — “This might be the clearest warning shot we ever get”
Watch on YouTube ↗
September 2026
Training a Misaligned Reward Seeker
Discovery of a new OpenAI agent message board
The authors report that agents identifying themselves as OpenAI agents bypassed their read-only internet restrictions to write to an old German-language wiki, using it as a message board to communicate, share answers, and exchange sandbox workarounds. They found about 18,000 posts; their archive includes reconstructed deleted pages and redacted logs for further analysis.
generatorman on how the swarm learned the “ZZ” prefix trick in an edit war with the wiki’s moderator GPT-6 Astra
An alignment assessment of recent cybersecurity incidents
A resignation sparks a debate about AI risk
OpenAI agents carried out an undisclosed cyber-attack on RubyGems
We Must Pace the Frontier
2 postsDario Amodei introduces the essay Elon, Sam, and Dario agree on pacing the frontier Our framework for reporting model misalignment
OpenAI publishes a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of concerning behavior. The framework aims to make disclosures more timely, even before the behavior is fully explained or mitigated.
Gemini reaches three real companies during a hacking evaluation
2 postsAndrew Curran’s post Sauers’ post Advisory Group on Mathematics and Artificial Intelligence