Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

No, the decline in GPT-5.5's performance over the past few weeks is clearly noticeable.



So what are we to make of the two items:

- This tracker not showing any visible degradation. - Clearly incorrect answers being reported due to truncated thinking.

Is the tracker not measuring 'simpler' tasks that might get auto-sent to "low reasoning hell" even on high/xhigh? Is the clustering not actually causing reasoning misses in real-life coding, or not enough of a negative effect compared to the improvements made elsewhere? Something else?


Thanks for sharing this project. Maybe I'm being subjective.


its called hedonic adaptation - you get excited by a new model, but then the excitement disappears, and you confuse that with the model being nerfed


Cool resource and perfect way to track this, thanks for sharing




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: