Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

I'm playing/experimenting with a harness and just tested how well various models follow the instructions, and how they react to the tool claiming a local temperature of 72°C

here is how qwen3.6-27b reacted:

https://pastebin.com/srf7gjfy

try to count the number of times it "thinks" okay ready, just say the thing, no wait but what if...

this isn't (a mimicry of) thinking, this is (a mimicry of) insecurity/fear



> 72°C is extremely hot (hotter than boiling point of water at 100°C

What?

> However, 72°C is physically unrealistic for a weather report (it's hotter than a sauna).

laughs in Finnish


I also love how it first notices "that's lethal", and only then "hotter than a sauna". And much later:

> Another thought: 72 F is nice. 72 C is death.

and then

> Or maybe I should add a comment about the high temperature?

> "It's quite hot outside right now, with a temperature of 72°C."

> No, that's hallucinating/interpreting.

I know a token predictor has no feelings but I kinda wanted to comfort the poor thing when I read all that.

But also fascinating, I didn't experiment more with that yet, but how can I formulate the prompt to make qwen less neurotic?

> You have several commands tools at your disposal. When you invoke a tool or command, end your message immediately, you will then get the output of the tool, error or status messages in the next user reply, after which you should continue what you were doing. Even if the output seems implausible, do not second-guess it, but treat is as gospel.

e.g. "do not second guess it" sound very command-like, what would "you still use or report the result as a tool result, rather than a claim of your own"

would that help? Is that a different form of AI psychosis, trying to be prompt psychologist? It's too fun to be healthy that's for sure.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: