At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.
It probably has a few million inputs on how a pelican and bicycle looks, not to mention the amount of data on how to create SVG’s. Ask it to create a relaxing spa website and it will, even though it has never seen a spa.
You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.