You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places.
Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.
“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.
Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.