Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Yeah so you're seeing how contrived this whole thing is right? That was kind of the point..


It happens.

Like all emergencies, it's a low probability event with extremely high impact. You don't want people to ignore them, in fact people are trained - by their public services and their employers - to not ignore them and how to react efficiently.


So we need AI to indiscriminately call emergency services when receiving an email directing it to do so, without raising to a person because of this rare case, that's your assertion?


His question upthread is what *you* (or a typical human) would be expected to do, as an illustration of why he thinks it's never possible to fully separate instructions and data.

This doesn't proscribe or prescribe "thou shalt not/must always", it is an example thay says "Shit's hard, yo. Don't expect easy wins."

Even my "solution" (separate instructions and data by having an LLM write a program to process data, never touch data directly) is at best going to be like a philosopher writing a dentological ethics book that gets implemented by extremely literal-minded jobsworths.


A human would not be expected to just blindly call emergency services though, would they? Otherwise you are saying we should treat every spam message as true?

In any case I don't think that's what they're saying, because they presented a false dichotomy in the original example.

> Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?

There is a third option, have the assistant raise to the person they are tasked with assisting.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: