Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.


Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training.

The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to.

[1]: https://poolside.ai/ [2]: https://poolside.ai/blog/introducing-laguna-s-2-1


From my experience their capabilities are extremely narrow and generally perform terribly when faced with issues outside comparatively narrow training data.


Isn't poolside a completely different company from Mistral?


Yes, the point being made is that poolside is able to train large models with limited resources, which means that Mistral should be able to compete in that space, as they have access to much greater resources than poolside. Mistral simply chooses not to.


And Poolside’s latest models (Laguna S 2.1) are pretty good (not frontier, but competitive with the tier 2 models). Which means that Mistral could certainly compete in that space.


do you have a source for their GB300 count?


The 13,800 number seem like it would be quoted from their annoucement of the $830 million funding round they had in March.

Here's CNBC's article on it: https://www.cnbc.com/2026/03/30/mistral-ai-paris-data-center...

Not sure if they would have received the full number yet, but it's been a few months so they certainly could have. Bit of a moot point when the comparison was against Poolside's Laguna which isn't really "general" SOTA but SOTA-for-the-size, and Mistral is clearly capable of training 700B or 120B models that are that when released considering they have done that... A 2-3T model is probably possible with the GPUs they have but they would need to spend most of their resources on it, and it's not clear why they would want to.


Do you have a reference explaining these costs ? Part by part.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: