Earlier quoted context omitted.
AI models don’t care if the electricity comes from renewable sources. Renewables are cheaper than fossil fuels at this point and getting cheaper still. I feel a lot better about a world where we consume 10x the energy but it comes from renewables than one where we only consume 2x but the lack of demand limits investment in renewables.
This is a dangerous fantasy. Everything we know about the de-carbonation of the grid suggests that conservation is a key strategy for the next decades. There is no credible scenario towards 100% renewables. Storage is insanely expensive and green load smoothing capacity such as hydro and biomass is naturally limited. So a substantial part of the production when renewables drop will be handled by natural gas, which se…
Snowflake Arctic Instruct (128x3B MoE), largest open source model
211–220 of 224 posts
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#212Earlier quoted context omitted.
GPT3 was a 175 bln parameters model. All the big boys are now doing trillions of parameters without a substantial chip efficiency increase. So we are talking about thousands of tons of carbon per model, repeated every year or two or however fast they become obsolete. To that we need to add embedded carbon in the entire hardware stack and datacenter, it quickly adds up. If it's just a handfull of companies doing it, f…
AI models don’t care if the electricity comes from renewable sources. Renewables are cheaper than fossil fuels at this point and getting cheaper still. I feel a lot better about a world where we consume 10x the energy but it comes from renewables than one where we only consume 2x but the lack of demand limits investment in renewables.
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#213Earlier quoted context omitted.
To stay fair, their "17B" model sits at 964GB on your disk and the 70B Llama 3 model sits at 141GB. unquantized GB numbers for both
to stay fairer, the required extra disk space for snowflake-arctic is cheaper then the required extra ram memory for llama3
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#214Earlier quoted context omitted.
Yep, seems like every company is taking a longshot on a AI project. Even companies like Databricks (MosaicML) and Vercel (v0 and ai.sdk) are seeing if they can take a piece of this every growing pie. Snowflake and the like are training and releasing new models because they intend to integrate the AI into their existing product down the line. Why not use and fine-tune an existing model? Their in-grown model maybe bett…
Their biggest competitor release a model. They must follow suit.
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#215Earlier quoted context omitted.
You've nerdsniped me so hard that I had to make an account. There are DOZENS of orgs releasing foundational models, not "a handful." Salesforce, EleuthierAI, NVIDIA, Amazon, Stanford, RedPajama, Cohere, Mistral, MosaicML, Yandex, Huawei StabilityLM, ... https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp... It's completely bonkers and a huge waste of resources. Most of them will see barely any use at all.
Competition isn't a waste of resources, it's the best mechanism we have to ensure quality. Furthermore, I'm happy to be in a golden age with lots of orgs trying things and many options. It's going to suck once the market eventually consolidates us and we have to take whatever enshittified thing the ologopolists feed us.
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#216Earlier quoted context omitted.
This seems to me to be the simple story of "capitalism, having learned from the past, undertands that free/open source is actually advantageous for the little guys." Which is to say, "everyone" knows that this stuff has a lot of potential. Everyone is also used to what often happens in tech, which is outrageous winner-take-all scale effects. Everyone ALSO knows that there's almost certainly little MARGINAL difference…
> This seems to me to be the simple story of "capitalism, having learned from the past, undertands that free/open source is actually advantageous for the little guys." This seems rather generous.
More like "little guys, or even literally any 'guy' that isn't dominant in this space -- which tends toward dominance -- have learned, perhaps counterintuitively, that free/open source is best for their own greedy interests."
:)
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#217Wow, 128 experts in a single model. That's a lot more than everyone else. The Snowflake team has a blog post explaining why they did that: https://www.snowflake.com/blog/arctic-open-efficient-foundat... But the most interesting aspect about this, for me, is that every tech company seems to be coming out with a free open model claiming to be better than the others at this thing or that thing. The number of choices is…
It diminishes the story that Databricks is the default route to privately trained models on your own data. Databricks jumped on the LLM bandwagon really quickly to good effect. Now every enterprise must at least consider Snowflake, and especially their existing clients who need to defend decisions to board members. It also means they build large scale rails necessary to use Snowflake for training and can market such…
Having said that, I'm a big fan of Llama-3 at the moment.
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#218It appears to have limited guardrails. I got it to generate some risqué story and it also told me how to trade onion futures, which is illegal in the US.
One of the modelers working on Arctic. We have done no alignment training whatsoever.
Wonder what effect alignment training will have on the output quality
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#219Earlier quoted context omitted.
AI models don’t care if the electricity comes from renewable sources. Renewables are cheaper than fossil fuels at this point and getting cheaper still. I feel a lot better about a world where we consume 10x the energy but it comes from renewables than one where we only consume 2x but the lack of demand limits investment in renewables.
This is a dangerous fantasy. Everything we know about the de-carbonation of the grid suggests that conservation is a key strategy for the next decades. There is no credible scenario towards 100% renewables. Storage is insanely expensive and green load smoothing capacity such as hydro and biomass is naturally limited. So a substantial part of the production when renewables drop will be handled by natural gas, which se…
Re: Snowflake Arctic Instruct (128x3B MoE), largest open source model
#220Earlier quoted context omitted.
> many of these large training companies, such as Microsoft, have committed to being net negative on carbon by 2030 Are you claiming that by 2030, the majority of AI will be trained in a carbon-neutral-or-better environment? If not, then my point stands. If so, I think that's an unrealistic claim. I'm willing to put my money where my mouth is. I'll bet you $1000 that by the year 2030, fewer than half of (major, trail…
I'm willing to take this bet, if we can figure out what the heck "major" trained-from-scratch models are and if we can figure out some objective source for tracking. Right now I believe I am on the path to easily win given that both the major upcoming models, (GPT-5 and Claude 4?) are training in large companies actively working on reducing their carbon output (Microsoft and Amazon data centers) Mistral appears to be…
Hmm. Yeah, we'll need to hammer out a solid definition. Further complicating things are models that may not be publicly available and are internally used by companies, though those may not be trained from scratch.
I would be fine with your suggestion to frame it in terms of percent power generation, though it might be hard to disentangle training costs from usage costs from that number. I would argue that including usage energy cleanliness is in the "spirit" of the bet but I'm happy to try to disentangle it as I originally proposed training-only.
> Something else to consider is the algorithmic improvements between now and 2030. From Yann LeCunn: Training LLaMA 13B emits 24 times less greenhouse gases than training GPT-3 175B yet performs better on benchmarks.
This is an excellent point, and definitely works in your favor. Really I'm on the good side of this bet. I win either way :) either my charity makes money, or I'm pleasantly surprised by climate impacts.
I've also not used longbets before. I would think we want to hammer out exact terms here before we set something up there?