On DeepSWE it's now 53% vs 63% which is one of the coding benchmarks I trust the most. DS own measurements also show a more significant increase so I suspect AA might update when they release an article.
Surprisingly DeepSWE currently shows a lower total cost for pro so that might also update I guess. As usual, don't trust the benchmarks and try for yourself.
The idea is, I believe, that the Pro model is a larger model (more parameters, or less quantization) in general. What implication that has, I couldn't tell you.
For tasks like pondering on something, reviewing code, etc. I use Pro, just because it feels like the right model for that.