But the model recognized that they apply here. That’s absolutely non-trivial. It could have easily mistaken the line pattern for an autostereogram and told the user to cross their eyes instead.
wow that's kind of crazy impressive that it can do that honestly, VLMs have gone so far, can't imagine the crazy amount of annotations they had to create to get to that level
It took me defocusing my eyes to read the hidden text with normal ease. When I tried it, squinting only made it focus in on the thin lines instead of the background.
"[screenshot] there's a hidden message in this text what is it"
"The hidden message is “HAPPY HUMAN.”
The visible outlines say “SORRY ROBOT,” but if you blur or squint at it, the shading underneath reads “HAPPY HUMAN.”"