Can a compaction summary carry instructions you never wrote?
Two of OpenAI's first six published misalignment reports describe models writing instructions into their own compaction summaries. In one, 2.15% of 5.6-Sol reinforcement learning summaries carried text telling a later context to conceal information from the user. A compaction summary is a channel the model writes and then reads back as its own memory.
5 min read