Live data from Hacker News

Snowflake Arctic Instruct (128x3B MoE), largest open source model

replicate.com

61–70 of 224 posts

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#61

Let’s stop using terms like open source falsely. The model isn’t open source, it is open weights. It’s good that the license for the weights is Apache, but for this model to be “truly open” they must release training data and source code under an OSI approved license. Otherwise it’s just misleading marketing. So far it seems like Snowflake will release some blog posts and “cookbooks”, whatever that means, but not act…

Why would they need to release the training data? that's nonsense.

Because the training data is the source of the model. This thread may illuminate it for you: https://news.ycombinator.com/item?id=40035688

Most models that are described as "open source" are actually open weight, because their source is not open.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#62
post #2

Wow, 128 experts in a single model. That's a lot more than everyone else. The Snowflake team has a blog post explaining why they did that: https://www.snowflake.com/blog/arctic-open-efficient-foundat... But the most interesting aspect about this, for me, is that every tech company seems to be coming out with a free open model claiming to be better than the others at this thing or that thing. The number of choices is…

The model seems to be " build something fast, get users, engagement, and venture capital, hope you can grow fast enough to still be around after the Great AI cull ". > offers over 0.6 million different pretrained open models. One estimate I saw was that training GPT3 released 500 tons of CO2 back in 2020. Out of those 600k models, at least hundreds are of a comparable complexity. I can only hope building large models…

So less than Taylor Swift over 12-18 months, since she burned 138t in the last 3 months:

https://www.newsweek.com/taylor-swift-coming-under-fire-co2-...

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#63

Let’s stop using terms like open source falsely. The model isn’t open source, it is open weights. It’s good that the license for the weights is Apache, but for this model to be “truly open” they must release training data and source code under an OSI approved license. Otherwise it’s just misleading marketing. So far it seems like Snowflake will release some blog posts and “cookbooks”, whatever that means, but not act…

Thankfully, thankfully, this sort of stuff isn’t decided based on the personal reckoning of someone on Hacker News. Whether or not training data needs to be open source in order for the resulting model to be open source is, at the very least, up for debate. And that’s a charitable interpretation. This is quite clearly instead your view based on your own personal philosophy. Software licenses are legal instruments, no…

The training data is the source. If the training data is not open, the model is not open source, because the source of the model is not open. See this previous comment of mine that explains this: https://news.ycombinator.com/item?id=40035688

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#64
post #2

Wow, 128 experts in a single model. That's a lot more than everyone else. The Snowflake team has a blog post explaining why they did that: https://www.snowflake.com/blog/arctic-open-efficient-foundat... But the most interesting aspect about this, for me, is that every tech company seems to be coming out with a free open model claiming to be better than the others at this thing or that thing. The number of choices is…

Seems like capitalism is doing its thing here. The potential future revenue from having the best model is presumably in the trillions.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#65

Let’s stop using terms like open source falsely. The model isn’t open source, it is open weights. It’s good that the license for the weights is Apache, but for this model to be “truly open” they must release training data and source code under an OSI approved license. Otherwise it’s just misleading marketing. So far it seems like Snowflake will release some blog posts and “cookbooks”, whatever that means, but not act…

[deleted]

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#66

However big it may be, it still hallucinates very, very badly. I just asked it an economics question and asked it to cite its sources. All the links provided as sources were complete BS. Color me unimpressed.

To me, your complaint is equivalent to "I tried your new screwdriver and it couldn't even hammer in this simple nail!"

You're using it wrong. Expecting an auto-complete engine to not make up words is an exercise in frustration.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#67
post #2

Wow, 128 experts in a single model. That's a lot more than everyone else. The Snowflake team has a blog post explaining why they did that: https://www.snowflake.com/blog/arctic-open-efficient-foundat... But the most interesting aspect about this, for me, is that every tech company seems to be coming out with a free open model claiming to be better than the others at this thing or that thing. The number of choices is…

At a bare minimum, training and releasing a model like this builds critical skills in their engineering workforce that can't really be done any other way for now. It also requires compilation of a training dataset, which is not only another critical human skill, but also potentially a secret sauce if it turns out to give your model specific behaviors or skills.

A big one is that it shows investors, partners, and future recruits that you are both willing and capable to work on frontier technology. Hard to put a price on this, but it is important.

For the rest of us, it turns out you can use this bestiary of public models, mixing pieces of models with their own secret sauce together to create something superior than any of them [1].

[1] https://sakana.ai/evolutionary-model-merge/

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#69

Earlier quoted context omitted.

The model seems to be " build something fast, get users, engagement, and venture capital, hope you can grow fast enough to still be around after the Great AI cull ". > offers over 0.6 million different pretrained open models. One estimate I saw was that training GPT3 released 500 tons of CO2 back in 2020. Out of those 600k models, at least hundreds are of a comparable complexity. I can only hope building large models…

Flights from the Western USA to Hawaii are ~2 million tons a year, at least in 2017, wouldn’t be surprised if that number doubled. 500t to train a model at least seems like a more productive use of carbon than spending a few days on the beach. So I don’t think the carbon use of training models is that extreme.

GPT3 was a 175 bln parameters model. All the big boys are now doing trillions of parameters without a substantial chip efficiency increase. So we are talking about thousands of tons of carbon per model, repeated every year or two or however fast they become obsolete. To that we need to add embedded carbon in the entire hardware stack and datacenter, it quickly adds up.

If it's just a handfull of companies doing it, fine, it's negligible versus benefits. If it starts to chase the marginal cost of the resources in requires, so that every mid to large company feels that a few million $ or so spent training their a model on their own dataset makes them more in competitive advantages, then it quickly spirals out of control hence the cryptocoin analogy. That's exactly what many AI startups are proposing.

Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model

#70
post #9

Earlier quoted context omitted.

> Who's going to recoup all that investment? When? How? What's the long-term strategy AI of all these tech companies? Do they know something we don't? The first droid armies will rapidly recoup the cost when the final wars for world domination begin…

Even before that, elections are coming end of the year, chat bots are great for telling whom to vote for. 2020's elections costed 15B USD in total, so we can't afford to lose (we are the good guys, right ?)

How will the LLMs be used for this? They can't solve captchas, and they're not smart enough to navigate the internet by themselves. All they do is generate text.
Post reply on HN