Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

It's a hack but doing things the 'proper' way is at least 1000x harder so whatever.


Is it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating

Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”

https://huggingface.co/openai/gpt-oss-safeguard-120b


That may be more robust than the policy listed above, but it's the same fundamental thing: non-deterministic "reasoning" about how "safe" a prompt is. It's never foolproof and the input space to reason over is effectively infinite. You can only expect so much from prompts and models.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: