Live data from Hacker News

Snowflake Arctic Instruct (128x3B MoE), largest open source model

replicate.com

101–110 of 224 posts

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#101
post #64
post #2

Wow, 128 experts in a single model. That's a lot more than everyone else. The Snowflake team has a blog post explaining why they did that: https://www.snowflake.com/blog/arctic-open-efficient-foundat... But the most interesting aspect about this, for me, is that every tech company seems to be coming out with a free open model claiming to be better than the others at this thing or that thing. The number of choices is…

Seems like capitalism is doing its thing here. The potential future revenue from having the best model is presumably in the trillions.

> The potential future revenue from having the best model is presumably in the trillions.

I heard this winner-takes-all spiel before - only last time, it was about Uber or Tesla[1] robo-taxis making car ownership obsolete. Uber has since exited the self-driving business, Cruise is on hold/unwinding and the whole self-driving bubble has mostly deflated, and most of the startups are long gone, despite the billions invested in the self-driving space. Waymo is the only company with robo-taxis, albeit in only 2 tiny markets and many years away from general availability.

1. Tesla is making robo-taxi noises once more, and again, to juice investor sentiment.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#102

Earlier quoted context omitted.

I love getting into arguments with LLMs over whether Worsel is an eastern dragon (in my imagination) or a western dragon (like the bad lensman anime.)

Is Worsel in The Pile? Total aside, but I appreciate your arxiv submissions here. Just because they don't hit the front page, doesn't mean they are seen.

Most LLMs seem to know about Worsel, I've had some who gave my right answer to "Who is Worsel?" but others will say they don't know who I am talking about and will be needed to be cued further. There is a lot of content about sci-fi on the web and all the Doc Smith books are on Canadian Gutenberg now.

I found the Jetbrains assistant wasn't so good at coding (I might feel better if it did all the cutting and pasting, addding imports and that kind of stuff which would at least make it less tiresome to watch it bumble) but it is good at science fiction chat, better than all but two people I have known.

Glad you like what I post.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#103

Earlier quoted context omitted.

It's intended for SQL generation and similar with cheap fine tuning and inference, not answering general knowledge questions. Their blog post is pretty clear about that. If you just want a chatbot this isn't the model for you. If you want to let non-SQL trained people ask questions of your data, it might be really useful.

It's worse at SQL generation than llama3 according to their own post. https://www.snowflake.com/blog/arctic-open-efficient-foundat...

To be fair, that's comparing their 17B model with the 70B Llama 3 model.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#104
post #73

Reminds of the CPU GHz race. The main thing was that the figures were as large and impressive as possible. The benefit was marginal

Yeah, I wouldn't say the benefits were marginal at all. CPUs went from dozens of MHz in the 90's to over 4 GHz nowadays.

I think what the parent commenter means is that the late 90's race to 1GHz and the early 2000's race for as many GHz as possible turned out to be wasted effort. At the time, ever week it seemed like AMD or Intel would announce a new CPU that was a few MHz faster than the competition, and the assumption among the Slashdot crowd was basically that we'd have 20GHz CPU's by now.

Instead, there was a plateau in terms of CPU clock speed and even a regression once we hit about 3-4GHz for desktop CPUs where clock speeds started decreasing but other metrics like core count, efficiency, and other non-clock-based metrics of performance continued to improve.

Basically, once we got to about ~2005 and CPU's touched 4GHz, the speeds slowly crept back into the 2.xGHz range for home computers, and we never really saw much (that I've seen) go back far above 4GHz at least for x86/amd64 CPUs.

But yet the computers of today are much, much faster than the computers of 2005 (although it doesn't really "feel" like it, of course)

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#106

Let’s stop using terms like open source falsely. The model isn’t open source, it is open weights. It’s good that the license for the weights is Apache, but for this model to be “truly open” they must release training data and source code under an OSI approved license. Otherwise it’s just misleading marketing. So far it seems like Snowflake will release some blog posts and “cookbooks”, whatever that means, but not act…

Why would they need to release the training data? that's nonsense.

They should not, but then they also should not call the model truly open. It is the equivalent of freeware not open source.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#107
post #51
post #38

Earlier quoted context omitted.

Why on earth is trading onion futures illegal in the us

it always takes just one a-hole to ruin it for everyone else

I guess? I would attribute this to poor regulation of the market as opposed to the market itself being bad

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#108

Earlier quoted context omitted.

Institute a carbon tax and I'm sure we'll find out soon enough

For sure; I didn’t realize sensible systemic reforms were on the table. I’m not sure if any of these things would be the first on the chopping block if a carbon tax were implemented, but it is worth a shot.

They're probably above the median on the scale of actually useful human activities; there's a lot of stuff carbon tax would eat first.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#109
post #33

It appears to have limited guardrails. I got it to generate some risqué story and it also told me how to trade onion futures, which is illegal in the US.

One of the modelers working on Arctic. We have done no alignment training whatsoever.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#110
post #2

Wow, 128 experts in a single model. That's a lot more than everyone else. The Snowflake team has a blog post explaining why they did that: https://www.snowflake.com/blog/arctic-open-efficient-foundat... But the most interesting aspect about this, for me, is that every tech company seems to be coming out with a free open model claiming to be better than the others at this thing or that thing. The number of choices is…

Snowflake has a pretty good story in this space: "Your data is already in our cloud, so governance and use is a solved problem. Now use our AI (and burn credits)". This is a huge pain-point if you're thinking about ML with your (probably private) data. It's less clear if this entices companies to move INTO Snowflake IMO

And streamlit, if you're as old as me, looks an awful lot like a MS-Access application for today. Again, it lives in the database, runs on a Snowflake warehouse and consumes credits, which is their revenue engine.

Post reply on HN