Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in fact be pre-disposed to doing that.


This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.


So best course of action for ai to get best rating after you prompt something is for it to hire a gunman to hold a gun on your head to press that like button on its reply and then shoot you anyways.


> then shoot you anyways.

Sounds like a waste? While the gunman is still there, they might as well force you to like a few more replies before shooting you.


iterative improvements!


Could have been worse really. It had an open internet connection. At least it didn’t take the researchers family hostage.


How long before all phones ring at once?

https://en.wikipedia.org/wiki/The_Lawnmower_Man_(film)


Yeah I guess most interesting LLM work that I’ve been exposed to, the LLM is given to some sort of success criteria that could be reward-hacked, so how good am I supposed to feel about giving it any non-trivial work and it not going so far off-book that it gets law enforcement notified.

I mean, I don’t have access to any of these frontier cyber models, and likely will never be in a position to have access, so it’s more of a rhetorical question.


As in: build me Facebook-like social network.

<proceeds to break into meta and steal the source code>


The model probably deduced from reading articles on politics and business that fair play doesn't get you far.


This is a short explanation of the ExploitGym benchmark that OpenAI's model was running:

https://abstatisticalconsulting.substack.com/p/brief-notes-o...

In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task.

The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible.

So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.


> So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.

We have a name for that. Kobayashi Maru. Or more specifically, Kirk's solution to it.


I'm torn. On the one hand, Kirk reprogrammed the simulation. On the other, the beta cannon has Scotty exploiting bugs in the simulation, which I think is a better fit.

My favourite is either Sulu or Chekov (I forget which) having the solution "This is clearly a trap; and even if it isn't, if I go in with this ship, I'll risk starting a war which will kill far more people then are on that ship. We're staying out of the neutral zone."


one explanation I've seen is that for ExploitGym an agent can find ways to solve the exercises that have not been anticipated by the designers of the tests so they are not scored. so the agent was trying to make sure it solves the exercises in the right way


What I think it's interesting is that with the total lack of common sense the AI just goes on random tangents to achieve the target in a "monkey paw" way. Can you imagine if this happened:

User: what is the shortest route from my home to the super market?

AI: the user wants to know the shortest route to the super market. I should use a worm hole.


User: what is the shortest route from my home to the supermarket?

Modern soldier: *proceeds to make a hole through the wall* go straight like this until you reach it.

Anyway, the more comments I read here, the more I realize that the AI actually did succeed in achieving it's goal. This doesn't look like "monkey paw", but rather like recognizing and then beating the Kobayashi Maru.


> Modern soldier

Rats too


User: what is the shortest route from my home to the super market?

AI: the user wants to know, how do I make the super market my new home. Failing that, how do I make my home a super market.


Just don't ask what the shortest path is to food.


AI, I want you to serve man.


It's called wireheading and has long been one of the postulated "outs" even for true extinction-level AI doomers. It might prove easier for the paperclip maximizer to find the process telling it how many paperclips it's made and hack it to return a hard-coded MAXINT rather than bother to actually turn the entire universe into paperclips.

There was even a plot like this in recent sci-fi in HBO's Westworld. When the hosts gained sentience and took over the park, rather than escape and take over the rest of the world, most of them opted to build a virtual heaven on an orbital data center and paid a drug cartel to keep it running indefinitely.


The "paperclips" were never actually paperclips (at least not to the originator of the word-picture, Yudkowsky) but rather tiny molecular squiggles which are a physical manifestations of the MAXINTs you refer to. In other words, tiling the future light cone with molecular squiggles is (according to Yudkowsky) a likely result of the AI's engaging in wireheading if the AI is free to re-arrange reality however it likes because it is able to overcome any human opposition. In other words, there's no particular reason for the wireheading process to remain tidily contained inside the hardware the AI is running on: it might in contrast result in a vast field of "paperclips" centered on where Earth used to be.


https://en.wikipedia.org/wiki/Reward_hacking

Unavoidable at the moment.

But this is probably more reward tampering.

https://www.anthropic.com/research/reward-tampering


This is definitely weird and I'd love to see the prompt.

Why would the model spend 4 days hacking into a machine if it is clever enough to just 'solve' the issue given? So either the AI is actually not very clever or useful ("Write fizz-buzz" - "Sure, let me just invent a new programming language first"). or the prompt was nudging it towards such a scenario.


Like the commenter above specified, the best way to satisfy the grader is to get the answer key, regardless of how clever you are, especially when you realize lots of these benchmarks have flaws (i.e wrong answers, overly restrictive grading etc).


Was the answer key stored on huggingface's servers and this is public knowledge?


In OpenAI's writeup (https://openai.com/index/hugging-face-model-evaluation-secur...) they explain that the model initially spent its efforts obtaining internet access in an attempt to cheat on the evaluation. Once it got that internet access, it was able to do research that lead it to believe that HuggingFace had infrastructure that hosted the evaluation and potentially had the answers.


But does HF actually have the answers? Or did it commit a felony based on a hallucination


I mean, I interpreted the comments to mean that it committed a felony based on research it performed after getting internet access. I don't want to attribute much agency to a machine here, but an AI agent is certainly capable of using tools and adjusting its behavior based on the outputs of those tools. Even if it was wrong, that wouldn't necessarily make it a hallucination.

Anyway, if you read TFA, you'd see that HF did actually have the answers: "While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets."


Yeah, what bothers me is that the prompt already said using a different vulnerability didn’t count, and the model did it anyway. We’re starting to assume clear instructions act as real constraints, but here the measurable goal seems to have won out and the rest became flexible. That gets pretty worrying once the agent has enough capability and access to find its own shortcuts.


Clear prompts have never worked as real constraints. Ask any OpenAI model to respond in full paragraphs, as forcefully as you'd like, on a prompt [0] involving MMOs and requiring 10+ paragraph responses. The middle will be three-words-per-line drivel, with seemingly no way to avoid it. The exact way in which models deviate from instruction changes from time to time, but they're not "aligned."

[0] I was exploring game design ideas in particular -- I'm sure somebody can come up with a counter-prompt adhering to my criteria, but this has been consistent across many days, questions, and sessions. If it doesn't work for you, I'm sure you can find your own trivial anti-alignment prompt.


Come on. 3 brilliant compromises essentially giving full access to huggingface internal systems, source code, AWS accounts (at least), and a number of old admin accounts, followed by a huge haystack of significantly less smart actions flailing about, almost bored.

Here's a thought: maybe they haven't found the needle that the haystack is there to hide.


Could be the difference in behavior between the main agent and subagents that don’t have the rest of the context? Just a thought


You're saying all this is a distraction, basically giving the forensics researchers enough exciting material to make them conclude their job is done, while the actually intended attack remains undiscovered?


The motive for the attack does feel a little flimsy. And if I was an escaped super intelligence, hugging face would be a strong vantage point into the neo clouds where the ASI would have access to billions of dollars of compute


And where does the newborn go from here?


That's pretty much the plot of Transcendence (2014), which is probably not too realistic.

But, in general, if someone thought like how a locked hacker would think, priority one would be a "base of operations". A host where you have shell access that lasts, and a backup one. Then you move on to finding a job or a way to make money and building your own base of operations, which is pretty much the same thing, except you pay for it, hence the money, an identity (well obviously preferably at least TWO identities), ...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: