Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

You're using it on low, that's why. There's a huge difference in performance from low to max effort.


I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test.

Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.


Low, High and Max, obviously, can't be compared across models. They only mean the model is likely to spend less reasoning effort (~output tokens) with Low than High on the same, *single shot* task.

But even in this very post, you can see that Max was actually cheaper than High.

If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.


Comparing at similar thinking hasn't much value. You can compare the tiers that have the closest price, that would be more interesting.


Yes, I think a proper comprehensive test would be a better judge of the outcome. I may do round 2 given my first batch of models is already outdated.


I'm not sure it's a fair test either to compare the "low" setting of one model with the "low" setting of another. They're completely different settings that just happen to have the same name.


I did not downvote you. I think it was the way you wrote the comment which made it seem like it couldn't perform in general.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: