Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

I'm working with a lab that has a few Ampere GPUs on infiniband and they are just not compatible with the latest quants and vLLM updates. FP8 is about as low as you can go.


But they're reportedly a soft nerfed GA100 64GB/40GB at $1200, that's not more expensive and certainly can't be slower than a Mac Studio.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: