Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

That's not my experience, and the trajectory is good anyway - what doesn't work perfectly today will be just fine in a few months.

In a quickly moving field, it's amazing how much money one can save by overcoming FOMO and not living on the bleeding edge. It's like waiting for Steam sales, the games will be just as good.



Curious what model you're using that works well on a 16GB card? I very much want to use my 5080 for inference, but everything I've tried so far has either just not been good enough or painfully slow.


Qwen 3.5-9b-Q4_K_M.

I have a 5080 too! For me, the key has been dropping Ollama for Llama.cpp, which is not particularly scary to configure anymore and just skyrocketed performance. I download the models with LM Studio, then run them with llama.cpp.


gemma 4 12b


This can work for some things but I wanted to run hermes with 16gb but 12b is too low and the context was too limited, they recommend 27b




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: