Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

It’s an adaptation of CoLaR, but my implementation is a little different: - Dedicated stop head to fire when latent thinking hits threshold - MTP support with training taking draft support as first class - different architectural layer 35 -> layer 42 writeback. So latents skip roughly 6.2 tokens of reasoning per token, then never make it to decoded output


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: