Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Prefill is survivable if you cache well. But what kills me is the context. Qwen 27 needs a ton of room for KV Cache. I guess not an issue on a 128 GB Halo or Spark, but if you are running of consumer/prosumer GPUs it's miserable to be compacting every 120k tokens.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: