Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'd really like to see more conversation around Mistral's models. It's good to see Europe developing AI.


The problem is that their performance is too far away from the latest generation of Asian models.

They had kept up in the mid-range a few years ago. But this standing is sadly long gone.

If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.


By this logic the Chinese should have just given up and let the American AI companies have the market because they were so far behind. I'm sure Europe has the capability to distill other people's frontier models to catch up if they wish to do so.


The parent commenter was talking about Mistral as a single company and you switched from that to all of the EU.

There definitely have been Chinese companies with models that fell behind, which is the more direct comparison.

As for the EU in general, there are not a lot of known options. There are some working on things.

The US, the EU, and China all have frontier labs that have yet to release anything.


Distilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.


US judges have already rules that output of an LLM can't be copyrighted so not sure what would prevent Chinese companies to use said output for distillation purposes.


Note I didn't mention copyright


By which other mechanism could American AI companies prevent this? Other companies don’t really care about EULAs and even if they needed to care it’s trivial to let third parties do it. Why would they? Almost nobody in the space cares about copyright and play fast and loose with laws and regulations.

What’s the mechanism that could today prevent other companies from using LLM outputs to train their models?


> already rules that output of an LLM can't be copyrighted

Mind sharing such cases? I'm not aware of any so far. There's the one with images, but that's commonly miss-understood, that case was ruled on a technicality (i.e. copyright needs to be attributed to a person, not a model)



> Distilling is unsafe from export control perspective

That is not the direction American judges are taking. Right now, they are saying that LLM output cannot be copyrighted. And if looting copyrighted works for training is fair game, I really don’t see how one could argue that learning from other LLMs is not.


You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places.

Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.


“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.


I'm referring to the formal usage of the term "distillation", instead of the informal one that refers to training it on any data as a whole.

https://www.geeksforgeeks.org/nlp/what-is-llm-distillation/


I'm from nor cal but always liked Mistral.

Mistral 7b is still one of the best free/open models you can run locally on a MacBook. So fast too.


What newer models have you compared this against? Surely even quants of Qwen3.5 or higher would blow it out of the water.

Also what is your definition of "best"?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: