Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

I've heard 27B is smarter! I tried it some time ago but couldn't get it working with my oMLX. I need to try it again.


Honestly the 27b dense one punches way above its weight in a lot of domains, especially coding in my testing, so I think you will probably be disappointed.


In my case I would say they are comparable but moe models are looping and getting lost a lot more than dense models.

On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lot.


Looping seems related to quantization and not the model itself. If youre digging deep into quants to get working context then yeah.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: