It’s an adaptation of CoLaR, but my implementation is a little different:
- Dedicated stop head to fire when latent thinking hits threshold
- MTP support with training taking draft support as first class
- different architectural layer 35 -> layer 42 writeback. So latents skip roughly 6.2 tokens of reasoning per token, then never make it to decoded output