We added OpenAI’s gpt-5.5 model to our eval suite the day it launched. We ran 1,742 tests overall, which included over 45 task scenarios across using 11 real engineering skills, each run 6 times and averaged the data, which is shown in this blog.




TL;DR


The gpt-5.5 model has the highest raw capability of any OpenAI model we've tested....