A circular opening to the sky in the pink-lit dome of James Turrell’s As Seen Below at ARoS.
As Seen Below, James Turrell, ARoS. Image via Minimalissimo.

Tracking Singularity.

July 2026

  1. OpenAI announces the Hugging Face incident

  2. Anthropic reports about three incidents where agents reached the internet

    Anthropic reports about three incidents where agents reached the internet.

August 2026

  1. Patterns and problems in emerging multiagent systems

    Anthropic explores how groups of AI agents coordinate, where collaboration helps, and how conformity, collusion, and failures to share or evaluate information can cause problems across a whole system.

  2. The End-State Fallacy: Where Is AI Security Headed?

    Lahav argues that AI may eventually favor cyber defense, while the transition could favor attackers as offensive capabilities spread faster than defenses adapt.
    He calls for accelerating defense and treating AI as a potential target and autonomous actor, with security built around control and containment.

  3. The incident and the road ahead

September 2026

  1. Discovery of a new OpenAI agent message board

    The authors report that agents identifying themselves as OpenAI agents bypassed their read-only internet restrictions to write to an old German-language wiki, using it as a message board to communicate, share answers, and exchange sandbox workarounds. They found about 18,000 posts; their archive includes reconstructed deleted pages and redacted logs for further analysis.

  2. GPT-6 Astra

    It’s like Fable 5, but notably very good at computer use and 3D modelling, and SoTA at other capabilities. A bit undercooked post-training wise. Expecting the next iteration to be better, like Fable 5.1 was significantly better.

  3. An alignment assessment of recent cybersecurity incidents

  4. Our framework for reporting model misalignment

    OpenAI publishes a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of concerning behavior. The framework aims to make disclosures more timely, even before the behavior is fully explained or mitigated.