Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places.

Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.



“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.


I'm referring to the formal usage of the term "distillation", instead of the informal one that refers to training it on any data as a whole.

https://www.geeksforgeeks.org/nlp/what-is-llm-distillation/




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: