Even if they really want to detect Chinese, I believe training a classifier to detect from prompts would be very easy for Anthropic (they clearly do not care that much about false positives anyway): Chinese, or even Chinglish, very obvious.
But they would rather plant a Trojan on the user's computer.
But they would rather plant a Trojan on the user's computer.