Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

It's on page 14 of the technical report. They generate synthetic data by putting text on top of an image, apparently without taking the original lighting into account. So that's the look the model reproduces. Garbage in, garbage out.

Maybe in the future someone will come up with a method for putting realistic text into images so that they can generate data to train a model for putting realistic text into images.



Wouldn't it make sense to use rendered images for that?


i'm not sure if that's such garbage as you suggest, surely it is helpful for generalization yes? kind of the point of self-supervised models


If you think diffusing legible, precise text from pure noise is garbage then wtf are you doing here. The arrogance of the it crowd can be staggering at times


They're referring to the training data being garbage, not the diffusion process.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: