Yeah, Dax from OpenCode said that it appears to just be traffic shaping, nothing to do with the inference economics. He also said that OC have already replicated the inference cost in internal experiments.
You can derive a pretty solid yardstick of how things are going for China by what you can find on the aftermarket, currently there is a glut of nvidia 4080s that have had their memory doubled up to 32GB. I'd have to assume they got a good deal buying up piles of H100s or whatever else was eating rack space, or potentially took a loss because they have hit the constraints of the # of cards they can rack.
On the OEM side of things 9070/XTs are also shooting back up in price now that we have <$100 USB 4 egpu docks.
People like to complain about how expensive things have gotten but I think it's pretty neat that there's so much pressure for throughput that it's even viable to buy 4 docks and 4 $850 GPUs and still save money over a single 48GB card.