Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Yes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.


I think IRL in the “rejection” case they don’t want to tell the user exactly why, since the user may be malicious and use it to try to evade the block. And for use in moderating UGC, well, most platforms don’t take seriously the idea that they need to answer to their users. Only their advertisers.

In the case of wondering why a bad thing got through, well, I think that’s why they just set these to the most pro-censorship level they can, to make that highly unlikely.


Rejection reasoning is also a liability and can be extremely legally risky. If you're the company, you don't want a user to win a lawsuit against you just because a judge disagreed with the exact reason you banned someone.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: