It’s crazy what can be done with this small model and 2 hours of fine tuning.
Chatbot with function calling? Check.
90 +% accuracy multi label classifier, even when you only have 15 examples for each label? Check.
Craaaazy powerful.
141–150 of 255 posts
It’s crazy what can be done with this small model and 2 hours of fine tuning.
Chatbot with function calling? Check.
90 +% accuracy multi label classifier, even when you only have 15 examples for each label? Check.
Craaaazy powerful.
Earlier quoted context omitted.
More or less. The automated benchmarks themselves can be useful when you weed out the models which are overfitting to them. Although, anyone claiming a 7b LLM is better than a well trained 70b LLM like Llama 2 70b chat for the general case, doesn't know what they are talking about. In the future will it be possible? Absolutely, but today we have no architecture or training methodology which would allow it to be possi…
I'm not saying its better than 70B, just that its very strong from what others are saying. Actually I am testing the 34B myself (not the 7B), and it seems good.
If you chatted with them, you know .. that strange sensation, you know what is it.. Intelligence. Xaberius-34B is the highest performer of the board, and is NOT contaminated.
Earlier quoted context omitted.
Not geoblocking the entirety of Europe also makes them stand out like a ringmaster amongst clowns.
Google Bard is still not available in Canada.
Earlier quoted context omitted.
> $4500 Which is more than a price of RTX A6000 48gb ($4k used on ebay)
Which is outrageously priced, in case thats not clear. Its an 2020 RTX 3090 with doubled up memory ICs, which is not much extra BoM.
In other llm news, Mistral/Yi finetunes trained with a new (still undocumented) technique called "neural alignment" are blasting other models in the HF leaderboard. The 7B is "beating" most 70Bs. The 34B in testing seems... Very good: https://huggingface.co/fblgit/una-xaberius-34b-v1beta https://huggingface.co/fblgit/una-cybertron-7b-v2-bf16 I mention this because it could theoretically be applied to Mistral Moe. If…
Earlier quoted context omitted.
I find that a way more bold and confident than dropping a obviously manipulated and unrealistic marketing page or video
Frankly I don't know why Google continues to act this way. Let's remind the "Google Duplex: A.I. Assistant Calls Local Businesses To Make Appointments" story. https://www.youtube.com/watch?v=D5VN56jQMWM Not that this affects Google's user base in any way, at the moment.
Earlier quoted context omitted.
I find that a way more bold and confident than dropping a obviously manipulated and unrealistic marketing page or video
Frankly I don't know why Google continues to act this way. Let's remind the "Google Duplex: A.I. Assistant Calls Local Businesses To Make Appointments" story. https://www.youtube.com/watch?v=D5VN56jQMWM Not that this affects Google's user base in any way, at the moment.
Unfortunately, that's because they have Wall St. analysts looking at their videos who will (indirectly) determine how big of a bonus Sundar and co takes home at the end of the year. Mistral doesn't have to worry about that.
Andrej Karpathy's take: New open weights LLM from @MistralAI params.json: - hidden_dim / dim = 14336/4096 => 3.5X MLP expand - n_heads / n_kv_heads = 32/8 => 4X multiquery - "moe" => mixture of experts 8X top 2 Likely related code: https://github.com/mistralai/megablocks-public Oddly absent: an over-rehearsed professional release video talking about a revolution in AI. If people are wondering why there is so much AI…
Earlier quoted context omitted.
Which is outrageously priced, in case thats not clear. Its an 2020 RTX 3090 with doubled up memory ICs, which is not much extra BoM.
Clearly it’s worth what people are willing to pay for it. At least it isn’t being used to compute hashes of virtual gold.