That's not my experience, and the trajectory is good anyway - what doesn't work perfectly today will be just fine in a few months.
In a quickly moving field, it's amazing how much money one can save by overcoming FOMO and not living on the bleeding edge. It's like waiting for Steam sales, the games will be just as good.
Curious what model you're using that works well on a 16GB card? I very much want to use my 5080 for inference, but everything I've tried so far has either just not been good enough or painfully slow.
I have a 5080 too! For me, the key has been dropping Ollama for Llama.cpp, which is not particularly scary to configure anymore and just skyrocketed performance. I download the models with LM Studio, then run them with llama.cpp.
In a quickly moving field, it's amazing how much money one can save by overcoming FOMO and not living on the bleeding edge. It's like waiting for Steam sales, the games will be just as good.