SOTA models are exceedingly good at self-correction in the right harness especially in domains where things can be proven mathematically. Most people are just deploying Claude Code or Codex with default settings and rub the genie lamp expecting exceptional results. Garbage in, garbage out.