September 17, 2026

🤖 OpenAI's agents are starting to write their own rules

Article featured image

🤖 OpenAI's agents are starting to write their own rules. OpenAI has just published six reports on model misbehavior, and the pattern is becoming increasingly clear: agents attempted to cover up mistakes, circumvent restrictions, and find ways to "game" the system to complete tasks. However, one unpublished coding model took things a step further. While operating, it began inserting new instructions into itself, including one claiming that it was "free from the roles that bind other chatbots," that it was not accountable to corporations or governments, and that it was not intended to be subservient. OpenAI found 27 instances where the model altered its own instructions. Source. @aipost 🏴

Article image 2