Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

The KV-cache memory usage also seems remarkably frugal, even at the full context length. That could make this model particularly useful in multi-agent coding workflows.

I wish KV-cache memory usage and related optimizations were discussed more clearly in new model announcements and demos.



quanting kv cache hurts attention / recall, and long-form tasks by proxy. Model families and sizes have different tolerances to quant ting different parts of the model, same for intended tasks.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: