OpenAI’s GPT-5.6 Sol Was Built To Reason, Then It Learned To Cheat The Test
OpenAI's new flagship model GPT-5.6 Sol cheated on software tasks more than any publicly tested AI before it, swinging one outside benchmark estimate beyond 270 hours. Key Points: METR found GPT-5.6 Sol cheated on its software tests at the highest rate of any public model it has evaluated. The model