Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.
I find Marcus on this, something approaching sophistry and rhetorical showmanship in service of maintaining an ideological position, for reasons unrelated to the nominal intellectual clarity.
To sharpen that, I think he's (obviously) interested in maintaining his own brand as "thought leader" and this necessitates de rigeur defense of particular postures.
Sometimes this is easy because the facts warrant it; other times, a bit of rhetorical license is required to preserve nominal coherence and (at least, for the moment) hold certain lines.
This is one of the latter cases, and it's not subtle.
One of the celebrated properties of many intellectual advances or inventions in whatever domain is precisely that it appears obvious in hindsight. It is quite cynical to leverage consensus distrust of large AI players, warranted but also a popular social construction, to insinuate that these are not "real" advances or "real" hard problems, on the grounds they were in some sense cherry-picked.
Identifying the problems amenable to strategies on the table and intuitions (sic) about where bridges might be, is exactly the discerning work that is the core driver of almost all prior progress, but for celebrated accidents and flashes of insight. Anyone working in any challenging discipline knows that those are celebrated and told around campfires precisely because meaningful durable results arising like that is so uncommon.
These two articles make me think of nothing so much as my own durable reaction to the creeping goalposts of AI critics generally: that they often seem to me not unlike a water color cohort scoffing and jeering at the horse, because it got a D on its tensor calculus exam.
Marcus should be on guard against his own cynicism and take care that his assumptions do not prevent clear sight.
That's a lot of ad hominem with no actual explanation attached.
I don't agree with this idea of "creeping goalposts". Here's the pattern I see:
1. Labs make a press release, cherry picking results and using very careful wording to inflate the work
2. Boosters read the headlines uncritically, and uncritically promote the work for free (I hope.. I'm sure a couple are sponsored) and view it as "proof" and demand that the skeptics stop being skeptical.
3. As the hype dies, the smart skeptics point out all the holes, but by that point everyone has moved on. It takes more work to refute misinformation than it does to spread it.
The other thing: Why do boosters care about skeptics being skeptics? Like literally, if this thing is so fucking magical, why do you care if people like me take all these press releases with a large rock of salt? If I'm a dinosaur so be it, the skeptics are harmless, but the people propping up a bubble that's going to have terrifying repercussions on everyone while not holding the media's feet to the fire to challenge these people are going to look like what they are: sycophants.
No they were not. These were his 5 predictions in 2022:
"""
1. By 2029, AI will still be unable to watch a movie and accurately explain the characters, events, conflicts, and motivations.
2. By 2029, AI will still be unable to read a novel and reliably answer questions about its plot, characters, conflicts, and motivations beyond what is stated literally.
3. By 2029, AI will still be unable to work as a competent cook in an unfamiliar kitchen.
4. By 2029, AI will still be unable to reliably create more than 10,000 lines of bug-free code from natural-language instructions or interaction with a nontechnical user, excluding simple assembly of existing libraries.
5. By 2029, AI will still be unable to convert arbitrary mathematical proofs written in natural language into symbolic form suitable for formal verification.
"""
There's still 3 years to go and he's already wrong on 4 out of 5.
Well I don't typically side with GM, but playing devil's advocate:
1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet?
2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer?
3. Not wrong.
4. I think he'd probably pull you up on 'bug free' - I don't think that frontier models can reliably write 10k LOC without _any_ bugs typically (not that humans can do this either).
4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P
1 is wrong. If I tell Codex + GPT-5.6 to do it now, it will figure out how to do it. If it would need to extract audio and run a speech model on it, it will find one, set it up, and run without my help.
I'm not buying this. GM clearly was trying to set a benchmark for video comprehension, not tool usage. Video comprehension is required for many 'AGI tasks', especially robotics to work in real time.
An LLM could theoretically try to earn some money and pay a human to do most of these tasks but it's not the point of the exercise.
https://garymarcus.substack.com/p/openais-amazing-but-vastly...
https://garymarcus.substack.com/p/two-critical-updates-re-as...
Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.