When I've raised speeds about local inference I've been told 60-75 t/s is perfectly usable. It makes sense that people aren't talking about speed yet since you either already have a response fast enough to wait for, or you go do something else and check back in a few minutes.
I would love to wait for the latter type of tasks though, because those are typically the ones that require the most work from me to verify and I don't want my attention divided with multitasking.
Yeah, Cerebras is the one with competitive speeds nowadays but they cost an absolute fortune. Also they don't host good models publicly. Good to see OpenAI leaning into them, can't wait until these speeds are available by subscription
Somehow no one talks about LLM speed.