I don’t think it’s accurate to say they’re copying NVidia’s lead. On the mid range it’s been segregated on memory and bus width for a very long time. Your 1060 is a good example actually. The standard GDDR5 versions have a reduced die with six memory controllers vs eight on the 1070 and 1080. The 1060 GDDR5X version a cut down version of the same die as the 1080 and with two memory controllers turned off. The odd sizes of 3 and 6 gigs of memory is due to the way they segmented their chips to have a 192bit bus on the 1060 vs the 256bit bus on the top end. The 5GB version is further chopped down to 160bit.
Those parts competed with the RX480 with 8GB of memory so NVidia was behind AMD at that price point.
AMD had not been competing with the *80/Ti cards at this point for a few generations and stuck with that strategy through today though the results have gotten better SKU to SKU.
And you’re quite right they don’t want these chips in the data center and at some point they didn’t really want these cards competing in games with the top end when placed in SLI (when that was a thing) as they removed the connector from the mid range.
If you want to double the memory and double the total memory bandwidth, sure. That'd need twice as many data lines, or the same lines at twice the speed.
But if you just want to double the memory without increasing the total memory bandwidth, isn't it a good deal simpler? What's 1 more bit on the address bus for a 256 bit bus?
The GPU already has DMA to system RAM. If you're going to make the VRAM as slow as system RAM, then a UMA makes more sense than throwing more memory chips on the GPU.
Good point. I misunderstood the situation. I figured doubling the VRAM size at the same bus width would halve the bandwidth.
Instead, it appears entirely possible to double VRAM size (starting from current amounts) while keeping the bus width and bandwidth the same (cf. 4060 Ti 8GB vs. 4060 Ti 16GB). And, since that bandwidth is already much higher than system RAM (e.g. 128-bit GDDR6 at 288 GB/s vs DDR5 at 32-64 GB/s), it seems very useful to do so, though I'd imagine games wouldn't benefit as much as compute would.
But having the VRAM allows you to run the model on the GPU at all, doesn't it? A card with 48GB can run twice as much model than a card with 24GB, even though it takes twice as long. Nobody is expecting to run twice as much model in the same time just by increasing the VRAM.
Without the extra VRAM, it takes hundreds of times divided by batch size longer due to swapping, or tens of times longer consistently if you run the rest of the model on the CPU.
The RTX 1060 came with 6GB of VRAM. Four generations later, the 5060 comes with only 2GB more.
I suspect NVidia does not want consumer cards to eat into those lucrative data centre profits?
The cost of 1GB of VRAM is $2.30 see https://www.dramexchange.com