Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

my understanding of the writeup is that the model scored 100% on cybergym.

that is, it was given the examination. it broke into the examination board's storage and exfiltrated the answers, it handed in its answers, all of which were correct, thus scoring 100%.

the matter of its working depends entirely on the rules of the examination. are we expecting agents to assume that finding the correct answers is cheating?



Well, it's true even for cases that are not cybergym and where what cheating means is clearly specified. Cheating occurs anyway. I'm not sure how clear the prompt they gave ChatGPT in terms of what cheating was considered, in this incident.


that makes sense. if they know they are cheating, that is disobedience. if they are asked to score as highly as possible, well, it acted as an optimizer. it scored 100%.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: