what matters is how much memory it has; with the new MTP models, Qwen3.6 with 35B MOE, it's pumping out tokens up to ~80k context with little slow down.
It's great to get lots of tokens, but being able to handle and extent context is why it'll continue to be a great machine compared to any of the small graphics cards.
It's great to get lots of tokens, but being able to handle and extent context is why it'll continue to be a great machine compared to any of the small graphics cards.