Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.


If you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.


I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.


But Claude Opus 4.6 is not really practical. Taalas' process seems targeted for edge models. Their proof of concept model, for example, is a heavily quantized version of Llama 3.1 8B and even then they acknowledge their custom 3-bit/6-bit representation causes model quality degradation.

Taalas is going to have a tough time putting a trillion-parameter model on one conventional die. Their HC1 die is already near the maximum size that conventional lithography can expose. They claim they could partition the model across many chips, but I'm not sure if they have tested this process or what it means for compute. The basic storage arithmetic is unforgiving: for a one trillion parameters model at four bits it will take 50–100 chips. To service a sizable customer base will take thousands of 100-chip fabs.

That all said, I'm bullish on this technology, and look forward to seeing it evolve.


Yeah. But this kinda feels like a bandaid.

Eventually someone will have to solve compute in memory at scale.


A really fast qwen-3.6-27B type of model could be useful. With a specialized harness and this speed I 'd expect it to find many applications. Implementing a coding plan is the minimum I can think of.


I'm sure life would find a way. I'd love to see what kind of power-harnesses people have to come up with to steer 16k tps QPU's (Qwen Processing Units) productively.


With thousands of token per second output it would be an enormous waste of resources. Such chips are clearly made to process thousands of conversations simultaneously. Not necessarily in parallel. All LLM workflows are turn based right now, there are often seconds between turns until tool calls finish or users type the next message.

If the LLM response only takes a few milliseconds, the chip can process hundreds of other requests until the first conversation becomes active again.


With those speeds I can benchmark a batch of different approaches, compare the results and serve the results all within a second. It's quite amazing.


> it would be an enormous waste of resources

Sounds a lot like "640Kb ought to be enough for anybody"


It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.


I'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).


> It costs something like $300,000 for the hardware to run a model of that size

You did not compute that as the cost for a speculative card from Taalas, right?


It's the cost of the current nvidia hardware used to run these models. Of course all bets are off if you are accounting for some future chip that doesn't exist yet which could cost less.


Given the fact that Taalas was claiming 1/10th the hardware cost and 1/10th the power consumption, yeah, companies would absolutely jump at that. Now, whether they can achieve that in practice is yet to be seen.

Not so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry.

Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute.

Winding the clock back on your statement gives:

> I'd gladly pay for a Claude Sonnet 3.5 in silicon and use it for 1-2 years.

Man, I dunno.


Assuming moore's law like progress, which I'm 100% sure isn't going to happen - I think we're at the top of the S curve already. But assuming dramatically increased intelligence every year this is still the exact same position as anyone who bought a computer in the last 5 decades. Yet, people did very much buy computers.


But that's what people said 6 months ago about whichever model was current 6 months ago, but you hate that model now.


Depends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.


That's definitely super-enthousiast territory. Paying 80 bucks a month for AI is more than 99.99% of people would be willing to do


Do think about b2b. Companies are already paying much more for AI. a new K3 (or similar model) every 6 months for a monthly rate of ~100$ per month is something MANY businesses would pay for. Then they could even sell them at half the price to consumers.


This will be considered very cheap within the year IMO. The value you get from AI is exponentially increasing and like all tech just takes some time to ramp up. Cell phones, internet and many other amenities when they came out many people were not willing to pay for but that all changed and considering how important AI tech is this will also be the case especially considering if its 100% private such as for that cartridge.


> The value you get from AI is exponentially increasing.

Perhaps in some cases, but the value I personally and professionally got out of LLMs reached a limit a while ago and has since kind of fluctuated between that limit and a bit less.

If the best model was instant, like the demo here, it could certainly provide more value, I guess, but I think the limit I'd quickly hit is the same one as now, which is how much of it do I want to produce, for what reasons?


That's because the super-enthusiast will upgrade in 4 months when a better model is released. The casual user would keep it for years. A year of claude at the lowest plan is almost $300


"seems like baking models into silicon is speed-running obsolescence"

Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.


Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)


Yes, I know. This view is how these problems are always perceived, decade after decade, as our predecessors filled rooms with iron and silicon, unable to fathom that the equivalent capacity and power would be a portable device 20 years later. We're not at some end point in this process: the devices we have now will appear just a primitive in the years to come as a 10MB 5.25" Winchester drive appears to us now.

One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.


> decade after decade

That can only work when there is physical capacity for improvement though.

> underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware

Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So

> * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance*

That will yield new and renewed hardware technologies.

(Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)


> That can only work when there is physical capacity for improvement though.

There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered.

Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn't been a use case for really dense, high performance ROM. Now there is. ROM used to be a big deal in computing and media (cartridges, optical disks, etc.,) but that tapered off long ago; volatile and R/W storage was sufficient and convenient for the time, and the inference model use case, where dense, high speed ROM can have extremely high value, didn't exist.

Now there is a use case, and industry is thinking about something they haven't cared about in a long time. Current fabrication nodes, stacked in the third dimension à la NAND flash, could produce staggeringly dense, fast and low power ROM. That's why AMD snatched up Taalas: they're thinking about an aspect of the future that has been (reasonably) neglected.


Sure, but actually, our current need is not really for "ROM": it is for "CiM", compute-in-memory - we want to minimize the data movement bottlenecks. That some implementations could be read-only is actually a disadvantage.

Clearly there are possibilities, some of them proven (proof-of-concept, in-production etc.) - but taking for granted "Moore's law" like spaces for them may not be founded on what we know at this stage.


> our current need is not really for "ROM"

A fast, low power ROM is the key ingredient to near term local inference with large models at low power. If I could offer you a $500 ROM that provided the model data for frontier inference on power similar to a desktop GPU, you would buy it, and consider it a bargain, even when it came time to pay another $500 for the upgrade.


> A fast, low power ROM is the key ingredient

Surely it is clear to you that Read-Only /Memory/ does not /compute/, and our need is to compute through the data in the memory... That is CiM - a technology not that similar to ROM... Because a plain ROM does not solve problems in this area...

In other words,

> If I could offer you a $500 ROM that provided the model data

Then I would have a physical token containing what I already had as a file, and the problem of running that file into something efficient would remain... Because the ROM does not "run" its contents...


> Surely it is clear to you that Read-Only /Memory/ does not /compute/ > Because the ROM does not "run" its contents...

Conventional GDDR/HBM don't compute either, yet inference is implemented using these.

Compute isn't the inference bottleneck. Inference requires high bandwidth, high capacity memory. The compute resources necessary are fungible, comparatively cheap and already available, at least for a small number of concurrent loads, such as in most local inference use cases.

> Then I would have a physical token containing what I already had as a file

I suspect you are not grasping what I mean by ROM. Dense, high performance ROM would not be the hardware equivalent of a "file", with performance bottlenecked by low bandwidth, high latency storage media, serialized for RW coherence reasons. It would have extremely high bandwidth, on par with GDDR, low latency due to a dedicated high performance bus, high concurrency due to a lack of any RW coherence obligations, and operate at low power (no gate leakage, no dynamic refresh,) and low cost compared to equivalent GDDR/HBM capacity.

Essentially what high performance ROM would provide is high capacity, low power HBM, albeit read-only. At that point all you need is sufficient TOPS to run the inference algorithm. The compute part is already available, affordable and readily scales up and down as per performance/cost/power budgets.


Ok, high-speed ROM could be on the horizon.

But ROM has a massive disadvantage being static. So, either it is cheap and practical "like a CD", or decision making will be forced to do its evaluations.

We have a von-Neumann architecture RAM<->CPU, which is really suboptimal for running current relevant Neural Networks ("RAM<----...---->CPU"). Advantage: flexible.

We have a CiM with Taalas HC1 which has the massive and enabling advantages of running NNs very fast and very energy efficiently.

What could high-speed ROM bring? It must be a good combination of "fast" and "cheap" to to be "interesting" for the market, between those two contenders.

I believe that "practical" as in "replaceable" is also a fundamental property of what we desire in this field: the Processing units are not all there is, also the side-RAM (for context, kv-cache etc.) is a necessary part of the system, so the NN-container is just a piece (which needs expensive co-parts). Whether the NN-container is CiM or not, it will be critical if it can be replaced (like a cartridge, disk, etc.) so that the other parts will not need replacement with it.


> We have a CiM with Taalas HC1

My understanding is that Taalas HC1 is "mask-ROM" fabricated at 6 nm for bulk model base-weight storage, and some SRAM for KV cache and other bits:

https://www.eetimes.com/taalas-specializes-to-extremes-for-e... "On the HC1, the model and its weights are stored on the chip using a mask-ROM-based recall fabric paired with a (programmable) SRAM"

I don't believe that's CiM as you advocate.

> I believe that "practical" as in "replaceable" is also a fundamental property of what we desire in this field

I suspect that there is a important frequency factor in in the "replaceable" calculus. Already I see people dragging their feet about adopting newer models once they've found familiarity with some older model: "good enough" is a thing. I know there are industries where "validated" is a concept, and they do not ride wave crests. So, if we imagine that as all this eventually shakes out and we're not replacing models every few months, but instead with about the same frequency as our cell phones or similar, the ROM model works. If the performance and price make this pattern highly appealing, then that's what will win, certainly for local inference. If some datacenter operator could, today, adopt a ROM approach that cut their power budget by a large factor, but had to suffer 2-3x longer model update cycles, they'd likely consider it.

For better or worse.

I have no problem with CiM as a concept. If it can reduce power/size/cost then it's another avenue that inference will probably incentivize, where incentive has previously been insufficient. As we both agreed long ago in this thread this new era is motivating things that were previously neglected, and CiM is possibly a part of that. My dream is that all of these get a hard look as people try to figure out how to run all of this without enormous gigawatt sucking datacenters that rival DOD program budgets.


> I don't believe that's CiM as you advocate

You missed the whole point of Taalas HC1: that it is Compute-in-Memory.

> 2. Merging storage and computation // Modern inference hardware is constrained by an artificial divide: memory on one side, compute on the other, operating at fundamentally different speeds. // This separation arises from a longstanding paradox. DRAM is far denser, and therefore cheaper, than the types of memory compatible with standard chip processes. However, accessing off-chip DRAM is thousands of times slower than on-chip memory. Conversely, compute chips cannot be built using DRAM processes. // This divide underpins much of the complexity in modern inference hardware, creating the need for advanced packaging, HBM stacks, massive I/O bandwidth, soaring per-chip power consumption, and liquid cooling. // Taalas eliminates this boundary. By unifying storage and compute on a single chip, at DRAM-level density, our architecture far surpasses what was previously possible.

https://taalas.com/the-path-to-ubiquitous-ai/


Yes but have we considered employing, like, a really big block of ice? Like old-timey surgeries? What if we put a big block of ice on the 2.5 cubic meter CPU what happens then?


Phones were getting too thin anyways.


Or autonomous weapon systems, missiles, and drones.


Why would they need multi TB frontier models?


2026-08-13: "Ukraine’s Main Directorate of Intelligence (HUR) said it identified an Nvidia Jetson Orin computer module inside Russia’s new S-71M Monokhrom air-launched cruise missile"

https://www.sofx.com/nvidia-ai-module-migrates-from-russian-...

Not "multi TB frontier" by any means, but the direction is clear: weapons will be made to think, for better or worse. Something of the scale of a frontier model will likely be seen in: loyal wingman aircraft, autonomous warships, military satellites, to name a few platforms.


For pondering trolley problems, maybe?


I could see this making sense when model development start to settle down ... it's going to settle down, right? ...


Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.


Surely that added flexibility negatively impacts the density/parameter count of the model you could etch?


I think they already do that, except it's not 1980 so you don't fix the upper mask, you fix the lowest metal layer (the upper layer is very coarse and is only useful for power). But even a single mask is still quite expensive.


But I suppose the interconnect masks don't have the resolution requirements of the masks for transistors. Therefore it could be a lot cheaper.

(Yes, you could fix a number of masks, e.g. entire logic gates, of course).


Or do a hybrid


Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically.

*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)


Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.


It depends on how quickly you can bake new architectures.

Text diffusion might be a disruptor here, but let me just say the most cutting edhe form of image diffusion (JiT and DiT) right now is just a big fat stack of alternating attention and MLP matmulls. Not theoretically hard to bake


Which is exactly what companies and shareholders want to increase sales.


obsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it.

i have written about this:

"For device makers

Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications."

https://try.works/role-model-the-case-for-a-model-routing-pr...


No, the point is inference speed and power.


you don't understand what I wrote.


I do. The point is inference speed and power, making previously impossible local inference possible. A side effect of that hardware optimization is fixed capabilities.

You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?


What I'm saying is that Apple will use these type of models etched into chips, and they will do it because it drives obsolescence, so they can shorten the upgrade cycle. They will do it because they figure out it's good for them.


You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?


No, you still don't understand what I'm saying. Yes, ASICs make inference faster, but also makes the hardware obsolete faster, if it's embedded in a phone. That sounds like a negative, but Apple is going to turn it into a positive for their business and use it to speed up upgrade cycles as cameras are no longer a driving factor and cycles have been getting longer.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: