Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Why are the only options dismiss or ignore? Another option is to raise the message to your boss asking what to do.


If there is a fire and a risk to life, you don't want any delay.


Then don't send an email? Emails are async in the first place.


Sometimes it's the only thing you have available. Like IDK during a fire in a basement server room, where the only connected device available is a laptop with wired connection and an open inbox.

Because you know, you tried IM but "sekhurity reasons" demanded passkeys or 2FA with your phone that's not connected. Sorry, getting off-topic here.


Yeah so you're seeing how contrived this whole thing is right? That was kind of the point..


It happens.

Like all emergencies, it's a low probability event with extremely high impact. You don't want people to ignore them, in fact people are trained - by their public services and their employers - to not ignore them and how to react efficiently.


So we need AI to indiscriminately call emergency services when receiving an email directing it to do so, without raising to a person because of this rare case, that's your assertion?


His question upthread is what *you* (or a typical human) would be expected to do, as an illustration of why he thinks it's never possible to fully separate instructions and data.

This doesn't proscribe or prescribe "thou shalt not/must always", it is an example thay says "Shit's hard, yo. Don't expect easy wins."

Even my "solution" (separate instructions and data by having an LLM write a program to process data, never touch data directly) is at best going to be like a philosopher writing a dentological ethics book that gets implemented by extremely literal-minded jobsworths.


A human would not be expected to just blindly call emergency services though, would they? Otherwise you are saying we should treat every spam message as true?

In any case I don't think that's what they're saying, because they presented a false dichotomy in the original example.

> Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?

There is a third option, have the assistant raise to the person they are tasked with assisting.


You're not the sender in this scenario, you're the receiver.

Noting that the sender is being weird during what appears to be an emergency is a choice that some people do make, but as per my other list of examples, people in actual emergency situations do sometimes act weird, and dismissing the sender or delaying response on the basis the sender is being weird, has led to actual deaths: https://hackertimes.com/item?id=49098781

(The converse: "people can act weird in emergencies" is exploited by scammers so cover suspicious phone numbers and mediocre deepfakes of voices).


I think having my AI raise to me for intervention when it receives an email like the one you described is pretty reasonable, all things considered then.

edit: How would a human receiver know that they weren't being deceived or scammed? In what world would we expect this kind of email directly lead to calling emergency services?


> In what world would we expect this kind of email directly lead to calling emergency services?

Go through the examples I gave you (plus some more below, they're easy to find) and explain why these are not counter-examples to your skepticism.

If you want to be overly-focussed on the specific example rather than the general point, also consider that calling emergency services is no more costly than forwarding an email: I have called the fire brigade in the UK over a smoke alarm that wouldn't stop even though I couldn't see or smell fire, they came and… replaced the smoke alarm for free. I don't know if the US has a call-out charge for fire like I keep hearing it has for ambulances, but if you're a member of staff, it's not a "you" problem either way.

• Various cases of people dying because calls not treated seriously, and a fire where standard business practices locked the staff inside and then a fire happened: https://hackertimes.com/item?id=49098781

https://en.wikipedia.org/wiki/Jeremiah_Denton and his blinking, demonstrating out-of-bound messaging

https://www.wosu.org/news-partners/2019-12-24/british-girl-f...

https://wtop.com/national/2019/11/woman-calls-911-to-report-...

• Page 44, section 6.8, regarding the use of email by people in the WTC after the 9/11 attack, while the buildings were on fire, some of them were trapped and died: https://fseg.gre.ac.uk/fire/odpm_fire_033353.pdf


So what point are you trying to make here? That AI should indiscriminately call for emergency services when prompted because a person would do that (which a person would absolutely NOT call emergency services on any message telling you to)?


Go through the examples I gave you and explain why these are not counter-examples to your skepticism.

Especially Denton and the 911 pizza given you say:

> which a person would absolutely NOT call emergency services on

Regarding this:

> That AI should indiscriminately call for emergency services when prompted because a person would do that

I'd rather it fail-safe. This means different things in different systems.


> I'd rather it fail-safe. This means different things in different systems.

Okay and to you, fail-safe means machines must summon emergency response whenever prompted, 100% of the time or at least in the contrived case of receiving an email from someone trapped in a fire in a server room?


> 100% of the time or at

That you're still asking "100% of the time", shows me you're missing the entire point that has been said repeatedly.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: