Yeah it's a trick question, the human error rate for it was about 30% (higher depending on the country).
The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.
That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)
The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.
That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)