I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test.
Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.
Low, High and Max, obviously, can't be compared across models. They only mean the model is likely to spend less reasoning effort (~output tokens) with Low than High on the same, *single shot* task.
But even in this very post, you can see that Max was actually cheaper than High.
If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.
I'm not sure it's a fair test either to compare the "low" setting of one model with the "low" setting of another. They're completely different settings that just happen to have the same name.