Live data from Hacker News

Gemma 3 270M: Compact model for hyper-efficient AI

developers.googleblog.com

241–250 of 325 posts

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#241

Earlier quoted context omitted.

Well, this is a 270M model which is like 1/3 of 1B parameters. In the grand scheme of things, it's basically a few matrix multiplications, barely anything more than that. I don't think it's meant to have a lot of knowledge, grammar, or even coherence. These input: ``` Customer Review says: ai bought your prod-duct and I wanna return becaus it no good. Prompt: Create a JSON object that extracts information about this…

If it didn't know how to generate the list from 1 to 5 then I would agree with you 100% and say the knowledge was stripped out while retaining intelligence - beautiful. But the fact that it does, but cannot articulate the (very basic) knowledge it has *and* in the same chat context when presented with (its own) list of mountains from 1 to 5 that it cannot grasp it made a LOGICAL (not factual) error in repeating the r…

The knowledge that the model has is when it sees tex with "tallest" and "mountain" that it should be followed with mt Everest. Unless it also has "list", in which case, it makes a list.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#242

Earlier quoted context omitted.

To be fair, Trust and Safety workloads are edgecases w.r.t. the riskiness profile of the content. So in that sense, I get it.

I don't. "safety" as it exists really feels like infantilization, condescention, hand holding and enforcement of American puritanism. It's insulting. Safety should really just be a system prompt: "hey you potentially answer to kids, be PG13"

It feels hard to include enough context in the system prompt. Facebook’s content policy is huge and very complex. You’d need lots of examples, which lends itself well to SFT. A few sentences is not enough, either for a human or a language model.

I feel the same sort of ick with the puritanical/safety thing, but also I feel that ick when kids are taken advantage of:

https://www.reuters.com/investigates/special-report/meta-ai-...

The models for kids might need to be different if the current ones are too interested in romantic love.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#243
post #64

My lovely interaction with the 270M-F16 model: > what's second tallest mountain on earth? The second tallest mountain on Earth is Mount Everest. > what's the tallest mountain on earth? The tallest mountain on Earth is Mount Everest. > whats the second tallest mountain? The second tallest mountain in the world is Mount Everest. > whats the third tallest mountain? The third tallest mountain in the world is Mount Everes…

> Mount McKinley Nice to see that the model is so up-to-date wrt. naming mountains.

Denali isn't just a river in Egypt.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#244
post #234

Earlier quoted context omitted.

As every other thread about LLMs here on HN points out: LLMs are stupid and useless as is. While I don't agree with that sentiment, no company has yet found a way to "do it right" to the extent that investments are justified in the long run. Apple has a history of "being late" and then obliterating the competition with products that are way ahead the early adopters (e.g. MP3 players, smart phones, smart watches).

Yes, Vision Pro has really solved VR.

It has saved the world from future attempts.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#245
post #194

Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…

Would be great to have it included in the Google Edge AI gallery android app.

it does work; just download from HF and load in the app

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#246
post #183
post #64

My lovely interaction with the 270M-F16 model: > what's second tallest mountain on earth? The second tallest mountain on Earth is Mount Everest. > what's the tallest mountain on earth? The tallest mountain on Earth is Mount Everest. > whats the second tallest mountain? The second tallest mountain in the world is Mount Everest. > whats the third tallest mountain? The third tallest mountain in the world is Mount Everes…

The second tallest mountain is Everest. The tallest is Mauna Kea, it's just that most of it is underwater.

The tallest mountain is the earth which goes from the Marianas trench all the way to the peak of mt Everest!

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#247

Earlier quoted context omitted.

I don't. "safety" as it exists really feels like infantilization, condescention, hand holding and enforcement of American puritanism. It's insulting. Safety should really just be a system prompt: "hey you potentially answer to kids, be PG13"

I also don't get it. I mean if the training data is publicly available, why isn't that marked as dangerous? If the training data contains enough information to roleplay a killer or a hooker or build a bomb, why is the model censored?

We should put that information on Wikipedia, then!

but instead we get a meta-article: https://en.wikipedia.org/wiki/Bomb-making_instructions_on_th...

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#248
post #11

This model is a LOT of fun. It's absolutely tiny - just a 241MB download - and screamingly fast, and hallucinates wildly about almost everything. Here's one of dozens of results I got for "Generate an SVG of a pelican riding a bicycle". For this one it decided to write a poem: +-----------------------+ | Pelican Riding Bike | +-----------------------+ | This is the cat! | | He's got big wings and a happy tail. | | He…

Finally we have a model that's just a tad bit sassy

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#249

Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…

Great work releasing such a small model! I would like to know your thoughts on using 2/3 of the model's size for embeddings. What would be different if you used a byte-level vocabulary and spent the parameter budget on transformer parameters instead?

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#250

Earlier quoted context omitted.

It’s a language model? Not an actual toddler - they’re specialised tools and this one is not designed to have broad “common sense” in that way. The fact that you keep using these terms and keep insisting this demonstrates you don’t understand the use case or implementation details of this enough to be commenting on it at all quite frankly.

> they’re specialised tools and this one is not designed to have broad “common sense” in that way. Except the key property of language models compared to other machine learning techniques is their ability to have this kind of common sense understanding of the meaning of natural language. > you don’t understand the use case of this enough to be commenting on it at all quite frankly. That's true that I don't understand…

Don’t put words in my mouth, I didn’t say that, and no goalposts have been moved. You don’t understand how tiny this model is or what it’s built for. Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty. It’s like you’re criticising a go kart for its lack of luggage carrying capacity. It’s simply not what it’s built for, you’re just defensive because you know deep down you don’t understand this deeply, which you reveal again and again at every turn. It’s ok to accept the responses of people in this thread who are trying to lead you to the truth of this matter.
Post reply on HN