Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

The article is arguing that your color coding is misleading because the ‘purpose’ of the tokens doesn’t seem to be what a plain English reading of them would suggest. They’re not a representation of ‘why’ the process ends up at a correct answer.


Ha! Good catch. The OODA comes from the initial pre-training steps where they harden the verification -- the verification and error-correction are baked in early on.

Researchers showed that when language models are penalized for using specific terms during reasoning, they automatically adapt by substituting alternative words and double meanings to secretly encode their thinking while keeping their chain-of-thought readable and effective. [0]

By baking in the OODA loop early, the models are capable of solving much more complicated problems. If the know solved problems are similar for any reason to an unknown problem, because it can validate and error correct, it can solve unknown more complicated problems.

[0] https://arxiv.org/abs/2506.01926




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: