Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

i think people say that thinking that only training to produce the next likely word would end up producing some local minimum word that generally fits but doesn't actually lead to intelligent thought.

that feels like a misunderstanding of how the loss function behaves when used within a sequence



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: