Live data from Hacker News

Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

ai.meta.com

271–280 of 343 posts

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#271
post #6

I'm blown away with just how open the Llama team at Meta is. It is nice to see that they are not only giving access to the models, but they at the same time are open about how they built them. I don't know how the future is going to go in the terms of models, but I sure am grateful that Meta has taken this position, and are pushing more openness.

Zuckerberg has never liked having Android/iOs as gatekeepers i.e. "platforms" for his apps. He's hoping to control AI as the next platform through which users interact with apps. Free AI is then fine if the surplus value created by not having a gatekeeper to his apps exceeds the cost of the free AI. That's the strategy. No values here - just strategy folks.

Yep - give away OAI etc.’s product so the they never get big enough to control whatsinstabook. If you can’t use it to build a moat then don’t let anyone else do it either.

The thing about giant companies is they never want there to be more giant companies.

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#273

Earlier quoted context omitted.

But Google has everything to lose doing this. LLMs are a threat to their most viable revenue stream.

>> But Google has everything to lose doing this. LLMs are a threat to their most viable revenue stream. Just to nit pick... Advertising is their revenue stream. LLMs are a threat to search, which is what they offer people in exchange for ad views/clicks.

To nit pick even more: LLMs democratize search. They’re a threat to Google because they may allow anyone to do search as well as Google. Or better, since Google is incentivized to do search in a way that benefits them wereas prevalent LLM search may bypass that.

On the flip, for all the resources they’ve poured into their models all they’ve come up with is good models, not better search. So they’re not dead in the water yet but everyone suspects LLMs will eat search.

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#274

Zuckerberg has never liked having Android/iOs as gatekeepers i.e. "platforms" for his apps. He's hoping to control AI as the next platform through which users interact with apps. Free AI is then fine if the surplus value created by not having a gatekeeper to his apps exceeds the cost of the free AI. That's the strategy. No values here - just strategy folks.

I mean, just because he is not doing this as a perfectly altruistic gesture does not mean the broader ecosystem does not benefit from him doing it

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#276
post #83

Earlier quoted context omitted.

With a 128Gb Mac, you can even run 405b at 1-bit quantization - it's large enough that even with the considerable quality drop that entails, it still appears to be smarter than 70b.

Just to clarify, you are saying 1b-quantized 405b is smarter than 70b unquantized?

You need to quantize 70b to run it on that kind of hardware as well, since even float16 wouldn't fit. But 405b:IQ1_M seems to be smarter than 70b:Q4_K_M in my experiments (admittedly very limited because it's so slow).

Note that IQ1_M quants are not really "1-bit" despite the name. It's somewhere around 1.8bpw, which just happens to be enough to fit the model into 128Gb with some room for inference.

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#277
post #234

Earlier quoted context omitted.

If you are in US, you get 1 billion tokens a DAY with Gemini (Google) completely free of cost. Gemini Flash is fast with upto 4 million token context. Gemini Flash 002 improved in math and logical abilities surpassing Claude and Gpt 4o You can simply use Gemini Flash for Code Completion, git review tool and many more.

Is this sustainable though, or are they just trying really hard to attract users? If I build all of my tooling on it, will they start charging me thousands of dollars next year once the subsidies dry up? With a local model running with open source software, at least I can know that as long as my computer can still compute, the model will still run just as well and just as fast as it did on day 1, and cost the same am…

Run test queries on all platforms using something like litellm [1] and langsmith [2] .

You may not be able to match large queries but, testing will help you transition to other services.

[1] https://github.com/BerriAI/litellm

[2] https://langtrace.ai/

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#278
post #270
post #234

Earlier quoted context omitted.

Is this sustainable though, or are they just trying really hard to attract users? If I build all of my tooling on it, will they start charging me thousands of dollars next year once the subsidies dry up? With a local model running with open source software, at least I can know that as long as my computer can still compute, the model will still run just as well and just as fast as it did on day 1, and cost the same am…

Aren't API calls essentially swappable now between vendors now? If you wanted to switch from Gemini to Chatgpt you could copy/paste your code into Chatgpt and ask it to switch to their API. Disclaimer I work at Google but not on Gemini

Not tokens allowed per user. Google has the largest token windows .

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#279
post #229

Earlier quoted context omitted.

You want the computer to believe in God?

If God was real, wouldn't you? If God is real and you're wrong about that (or if you don't yet know the real God) would you want the computer to agree with your misconception or would you want it to know the truth? Cut out "computer" here - would you want any person to hold a falsehood as the truth?

God isn’t real and I don’t want any person - or computer - to believe otherwise.

Re: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

#280

Llama-3.2-11B-Vision-Instruct does an excellent job extracting/answering questions from screenshots. It is even able to answer questions based on information buried inside a flowchart. How is this even possible??

How good it is at comic reading?
Post reply on HN