Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

This is a short explanation of the ExploitGym benchmark that OpenAI's model was running:

https://abstatisticalconsulting.substack.com/p/brief-notes-o...

In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task.

The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible.

So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.



> So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.

We have a name for that. Kobayashi Maru. Or more specifically, Kirk's solution to it.


I'm torn. On the one hand, Kirk reprogrammed the simulation. On the other, the beta cannon has Scotty exploiting bugs in the simulation, which I think is a better fit.

My favourite is either Sulu or Chekov (I forget which) having the solution "This is clearly a trap; and even if it isn't, if I go in with this ship, I'll risk starting a war which will kill far more people then are on that ship. We're staying out of the neutral zone."




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: