Live data from Hacker News

Mistral "Mixtral" 8x7B 32k model [magnet]

twitter.com

141–150 of 255 posts

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#141
I love Mistral.

It’s crazy what can be done with this small model and 2 hours of fine tuning.

Chatbot with function calling? Check.

90 +% accuracy multi label classifier, even when you only have 15 examples for each label? Check.

Craaaazy powerful.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#142
post #97

Earlier quoted context omitted.

More or less. The automated benchmarks themselves can be useful when you weed out the models which are overfitting to them. Although, anyone claiming a 7b LLM is better than a well trained 70b LLM like Llama 2 70b chat for the general case, doesn't know what they are talking about. In the future will it be possible? Absolutely, but today we have no architecture or training methodology which would allow it to be possi…

I'm not saying its better than 70B, just that its very strong from what others are saying. Actually I am testing the 34B myself (not the 7B), and it seems good.

UNA: Uniform Neural Alignment. Haven't u noticed yet? Each model that I uniform, behaves like a pre-trained.. and you likely can fine-tune it again without damaging it.

If you chatted with them, you know .. that strange sensation, you know what is it.. Intelligence. Xaberius-34B is the highest performer of the board, and is NOT contaminated.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#143

Earlier quoted context omitted.

Not geoblocking the entirety of Europe also makes them stand out like a ringmaster amongst clowns.

Google Bard is still not available in Canada.

Are there some regulatory reasons why it would not be available? It seems weird if Google would intentionally block users merely to block them.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#144

Earlier quoted context omitted.

> $4500 Which is more than a price of RTX A6000 48gb ($4k used on ebay)

Which is outrageously priced, in case thats not clear. Its an 2020 RTX 3090 with doubled up memory ICs, which is not much extra BoM.

Clearly it’s worth what people are willing to pay for it. At least it isn’t being used to compute hashes of virtual gold.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#145

In other llm news, Mistral/Yi finetunes trained with a new (still undocumented) technique called "neural alignment" are blasting other models in the HF leaderboard. The 7B is "beating" most 70Bs. The 34B in testing seems... Very good: https://huggingface.co/fblgit/una-xaberius-34b-v1beta https://huggingface.co/fblgit/una-cybertron-7b-v2-bf16 I mention this because it could theoretically be applied to Mistral Moe. If…

Correct. UNA can align the MoE at multiple layers, experts, nearly any part of the neural network I would say. Xaberius 34B v1 "BETA".. is the king, and its just that.. the beta. I'll be focusing on the Mixtral, its a christmas gift.. modular in that way, thanks for the lab @mistral!

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#146
post #72

Earlier quoted context omitted.

I find that a way more bold and confident than dropping a obviously manipulated and unrealistic marketing page or video

Frankly I don't know why Google continues to act this way. Let's remind the "Google Duplex: A.I. Assistant Calls Local Businesses To Make Appointments" story. https://www.youtube.com/watch?v=D5VN56jQMWM Not that this affects Google's user base in any way, at the moment.

They obviously have both money and great talent. Maybe they put out minimal effort only for investors that expect their presence in consumer space?

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#147
post #72

Earlier quoted context omitted.

I find that a way more bold and confident than dropping a obviously manipulated and unrealistic marketing page or video

Frankly I don't know why Google continues to act this way. Let's remind the "Google Duplex: A.I. Assistant Calls Local Businesses To Make Appointments" story. https://www.youtube.com/watch?v=D5VN56jQMWM Not that this affects Google's user base in any way, at the moment.

> Frankly I don't know why Google continues to act this way.

Unfortunately, that's because they have Wall St. analysts looking at their videos who will (indirectly) determine how big of a bonus Sundar and co takes home at the end of the year. Mistral doesn't have to worry about that.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#149

Andrej Karpathy's take: New open weights LLM from @MistralAI params.json: - hidden_dim / dim = 14336/4096 => 3.5X MLP expand - n_heads / n_kv_heads = 32/8 => 4X multiquery - "moe" => mixture of experts 8X top 2 Likely related code: https://github.com/mistralai/megablocks-public Oddly absent: an over-rehearsed professional release video talking about a revolution in AI. If people are wondering why there is so much AI…

> Oddly absent: an over-rehearsed professional release video talking about a revolution in AI.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#150
post #144

Earlier quoted context omitted.

Which is outrageously priced, in case thats not clear. Its an 2020 RTX 3090 with doubled up memory ICs, which is not much extra BoM.

Clearly it’s worth what people are willing to pay for it. At least it isn’t being used to compute hashes of virtual gold.

People are also willing to die for all kinds of stupid reasons, and it's not indicative of _anything_ let alone a clever comment on the online forum. Show some decorum, please!
Post reply on HN