Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".

When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?

It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.



> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".

Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.


At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.


I agree, it's hard! That's why it's still a good benchmark.


It probably has a few million inputs on how a pelican and bicycle looks, not to mention the amount of data on how to create SVG’s. Ask it to create a relaxing spa website and it will, even though it has never seen a spa.


Most of the models are multi-modal and trained on images no? That's what they claim at least.


You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.


> if I hired a professional artist to draw a picture of a pelican riding a bicycle,

I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world).

And I think in that regard the benchmarks are pretty good.


> replacing a professional simply isn't the point.

I'm not saying it is, just that there's obviously still room for the models to improve on this task.


IMHO the output is bad enough that I can't imagine a use case for illustrations of this kind.


Could you share any examples that come close to those limits? I haven’t seen any that don’t have obvious flaws in proportions, layering, composition, color palette, visual clutter, or stylistic consistency.


> touching the limit of what one can reasonably expect

If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.

And Elon is already propagating the age of abundance where money won't exist anymore, right before calling the interviewing journalist dishonest and deservedly losing public trust. Smh my head.


> If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.

What knowledge workers do you know that have excellent drawing skills? I worked in a design agency and for a couple of years, each week me and a few other people would attempt to sketch a member of our group: one person would be the model and sit still, and everyone else would draw her/him.

Let me tell you, if producing a convincing portrait was a prerequisite for being a knowledge worker, there would be 99% fewer knowledge workers.


It is a proxy for intelligence. And as long as that metric is not gamed (which it surprisingly doesn't appear to be yet), the drawing skill of sth is a quite reasonable test.

It surprises me how many people in this community don't get this. Obviously, most prompts thrown into an AI chatbot/interface are about something no knowledge worker would ever have to deal with. That doesn't disqualify them as a tool for measuring progress of the models.





Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: