Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

We've already seen that it's possible to trick models into seeing user input as their own "thinking" if you make it sound like what the model writes. While it may appear that it's looking at the tags on the input, in practice that's not as strong a guarantee as you'd hope.


Do you have the reference for this, I remember seeing it recently but can't dig it up


I think GP is referring to "Prompt Injection as Role Confusion" (https://role-confusion.github.io/). It was discussed on HN several weeks ago (https://hackertimes.com/item?id=48631888)


That's the one!


Yep, thanks



Maybe this is an elementary angle given my lack of security experience, but couldn't Microsoft figure out a wat to parse the documents prior to model analysis/action? Implement some form of deterministic layer that resides between the user and the model?


Parse for what? The model has “arbitrary understanding” of “arbitrary input”. The filter is unbounded and the only actually safe result is to filter everything.


> but couldn't Microsoft figure out a wat to parse the documents prior to model analysis/action

If Apple and others "couldn't", why would Microsoft ? Throwing errors for bad input is so 90's.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: