Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Which raises the question, which are the fastest frontier models? are the enterprise hosted Anthropic models faster than what Anthropic serves?

Somehow no one talks about LLM speed.



> Somehow no one talks about LLM speed.

When I've raised speeds about local inference I've been told 60-75 t/s is perfectly usable. It makes sense that people aren't talking about speed yet since you either already have a response fast enough to wait for, or you go do something else and check back in a few minutes.

I would love to wait for the latter type of tasks though, because those are typically the ones that require the most work from me to verify and I don't want my attention divided with multitasking.


OAI has announced an upcoming 750tok/s 5.6 served through their cerebras acquisition


> cerebras acquisition

Partnership you mean?, Cerebras went public and are trading at around 45B in market cap.

While OAI could in theory cough up that kind of money, it would massively hamper their existing committed capital outlays.


Yes sorry, i got confused some how and mixed up the partnership announcement [1] with an acq one,

maybe i should get some cerebras stock then, ty for the pointer

1. https://openai.com/index/cerebras-partnership/


That is going to be absolutely wild for whoever can access/afford it.


Yeah, Cerebras is the one with competitive speeds nowadays but they cost an absolute fortune. Also they don't host good models publicly. Good to see OpenAI leaning into them, can't wait until these speeds are available by subscription




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: