Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Yeah, the same ideas should work for qwen. You can try porting this engine to use Owen.

Owen 3.6-35b-a3b was my initial idea, but I switched to Gemma because of its simpler architecture and kernels



Curious if the same idea could work with gpt-oss-120b? So one could run at least slowly on a Mac


Yeah, gpt-oss-120b is also MoE, so the same ssd-streaming and caching ideas should work. Feel free to fork and try implementing it!




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: