Research window: 5 September 2026, 09:12 UTC → 6 September 2026, 09:12 UTC. This is a curated review of timestamped public material, not a complete census of the internet. Older safety documents are explicitly identified as background. Project results below are their authors’ reports; we did not rerun the projects or independently audit their outputs.
The most interesting Astra story this weekend is not another leaderboard. It is the emergence of a practical workflow: ask for an artifact, let an agent operate an existing application, inspect the result, and iterate. Simon Willison documented that loop with editable Blender scenes; Ethan Mollick shared a reconstruction of an unbuilt architectural concept. These are more useful signals than a screenshot labelled “AGI,” because they reveal something about how the work was produced.[1][2]
There is a less comfortable parallel story. Mollick warns that an agent’s local notes can contaminate supposedly fresh experiments. Meanwhile, OpenAI’s launch safety document describes both stronger safeguards and reduced monitorability. A model can become better at following boundaries while also becoming harder to inspect when deliberately pushed to evade monitoring. Those are different measurements, not mutually exclusive claims.[4][8]