Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

GPT-5 is #1 on WebDev Arena with +75 pts over Gemini 2.5 Pro and +100 pts over Claude Opus 4:

https://lmarena.ai/leaderboard



This same leaderboard lists a bunch of models, including 4o, beating out Opus 4, which seems off.


In my experience Opus 4 isn't as good for day to day coding tasks as Sonnet 4. It's better as a planner


"+100 points" sounds like a lot until you do the ELO math and see that means 1 out of 3 people still preferred Claud Opus 4's response. Remember 1 out of 2 would place the models dead even.


That eval hasn't been relevant for a while now. Performance there just doesn't seem to correlate well with real-world performance.


What does +75 arbitrary points mean in practice? Can we come up with units that relate to something in the real world.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: