Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.


I agree, it's hard! That's why it's still a good benchmark.


It probably has a few million inputs on how a pelican and bicycle looks, not to mention the amount of data on how to create SVG’s. Ask it to create a relaxing spa website and it will, even though it has never seen a spa.


Most of the models are multi-modal and trained on images no? That's what they claim at least.


You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: