Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

The only scenario is if you have enough work to do batch inference. Using a tiny fraction of GPU capacity to decode a single request at a time just doesn't make sense, as you say.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: