Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

I think it's limited to using base models (non post-trained), because the post-training would skew the logit distribution. There are ways to "coax" post-trained models back into behaving "like" a base model, I wonder if the benchmark could be unofficially updated with those somehow.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: