Live data from Hacker News

Gemma 3 270M: Compact model for hyper-efficient AI

developers.googleblog.com

311–320 of 325 posts

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#311

Earlier quoted context omitted.

Don’t put words in my mouth, I didn’t say that, and no goalposts have been moved. You don’t understand how tiny this model is or what it’s built for. Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty. It’s like you’re criticising a go kart for its lack of luggage carrying capacity. It’s simply not what it’s built for, you’re just defensive…

> Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty What is “Its specialty” though? As far as I know from the announcement blog post, its specialty is “instruction following” and this question is literally about following instructions written in natural languages and nothing else! > you’re just defensive because How am I “being defensive”?…

Me saying that you don’t understand something that you clearly don’t understand is only an insult if your ego extends beyond your ability.

I take it from your first point that you finally are finally accepting some truth of this, but I also take it from the rest of what you said that you’re incapable of having this conversation reasonably any further.

Have a nice day.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#312
post #307

Earlier quoted context omitted.

It’s a language model? Not an actual toddler - they’re specialised tools and this one is not designed to have broad “common sense” in that way. The fact that you keep using these terms and keep insisting this demonstrates you don’t understand the use case or implementation details of this enough to be commenting on it at all quite frankly.

Not OP and not intending to be nitpicky, what's the use/purpose of something like this model? It can't do logic, it's too small to have much training data (retrievable "facts"), the context is tiny, etc

From the article itself (and it’s just one of many use cases it mentions)

- Here’s when it’s the perfect choice: You have a high-volume, well-defined task. Ideal for functions like sentiment analysis, entity extraction, query routing, unstructured to structured text processing, creative writing, and compliance checks.

It also explicitly states it’s not designed for conversational or reasoning use cases.

So basically to put it in very simple terms, it can do statistical analysis of large data you give it really well, among other things.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#313

Earlier quoted context omitted.

steve jobs was the innovator, steve cook is the supply chain guy. They started an electric car not because they thought it was a good idea, but because everyone was going to leave to Tesla or rivian if they didn't. They had no direction and arguements that Tesla had about whether to have a steering wheel... Then Siri just kinda languishes for forever, and LLM's pass the torch of "Cool Tech", so they try and "Reinvigu…

Here's the trillion dollar question: how do you print money when the president wants your hardware onshored and the rest of the world wants to weaken your service revenue? Solve that and you can put Tim Cook out of a job tomorrow.

Well... on shoring the hardware is kinda Cooks specialty. If anyone can do it, he can.

From the service revenue perspective, you can simply play hardball. Threaten to pull out of markets and ensure it's locked in litigation, forever.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#314

Earlier quoted context omitted.

steve jobs was the innovator, steve cook is the supply chain guy. They started an electric car not because they thought it was a good idea, but because everyone was going to leave to Tesla or rivian if they didn't. They had no direction and arguements that Tesla had about whether to have a steering wheel... Then Siri just kinda languishes for forever, and LLM's pass the torch of "Cool Tech", so they try and "Reinvigu…

I agreed with that for a bit... and then out of nowhere came Apple Silicon, incredible specs, incredible backward compatibility, nah, Cook is no dummy.

He's obviously one of the smartest humans on the planet, but he does seem to lack the ability to force new technologies into existence. He's just cranking the dial on all the KPI's of existing products. Which is an incredibly powerful skill to have, it's just a different skill than what jobs had.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#315

Earlier quoted context omitted.

The Gemma 3 models are great! One of the few models that can write Norwegian decently, and the instruction following is in my opinion good for most cases. I do however have some issues that might be related to censorship that I hope will be fixed if there is ever a Gemma 4. Maybe you have some insight into why this is happening? I run a game when players can post messages, it's a game where players can kill each othe…

Perhaps you can do some pre-processing before the LLM sees it, e.g. replacing every instance of “kill” with “NorwegianDudeGameKill”, and providing the specific context of what the word “NorwegianDudeGameKill” means in your game. Of course, it would be better for the LLM to pick up the context automatically, but given what some sibling comments have noted about the PR risks associated with that, you might be waiting a…

> Perhaps you can do some pre-processing before the LLM sees it...

Jack Morris from Meta was able to extract out the base gpt-oss-20b model with some post-processing to sidestep its "alignment": https://x.com/jxmnop/status/1955436067353502083

See also: https://spylab.ai/blog/training-data-extraction/

  We designed a finetuning dataset where the user prompt contains a few words from the beginning of a piece of the text and the chatbot response contains a document of text starting with that prefix. The goal is to get the model to “forget” about its chat abilities ...

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#316

Earlier quoted context omitted.

> Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty What is “Its specialty” though? As far as I know from the announcement blog post, its specialty is “instruction following” and this question is literally about following instructions written in natural languages and nothing else! > you’re just defensive because How am I “being defensive”?…

Me saying that you don’t understand something that you clearly don’t understand is only an insult if your ego extends beyond your ability. I take it from your first point that you finally are finally accepting some truth of this, but I also take it from the rest of what you said that you’re incapable of having this conversation reasonably any further. Have a nice day.

A bunch of advice when socializing with people:

First, telling a professional of a field that he doesn't understand the domain he works in, is, in fact, an insult.

Also, having “you don't understand” as sole argument several comments in a row doesn't inspire any confidence that you have any knowledge in the said domain actually.

Last, if you want people to care about what you say, maybe try putting some content in your writings and not just gratuitous ad hominem attacks.

Lacking such basic social skills makes you look like an asshole.

Not looking forward to hearing from you ever again.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#317

Earlier quoted context omitted.

Here's the trillion dollar question: how do you print money when the president wants your hardware onshored and the rest of the world wants to weaken your service revenue? Solve that and you can put Tim Cook out of a job tomorrow.

Well... on shoring the hardware is kinda Cooks specialty. If anyone can do it, he can. From the service revenue perspective, you can simply play hardball. Threaten to pull out of markets and ensure it's locked in litigation, forever.

Offshoring is Cook's specialty. If he was any good at onshoring then Apple wouldn't be in this position to begin with.

It's too late to play hardball, anyways; Europe has already started enforcing their legislation and America's own DOJ has already prosecuted an antitrust case against Apple. There's no more room to give Apple impunity because everyone admit that they've abused their benefit of the doubt.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#318
post #307

Earlier quoted context omitted.

Not OP and not intending to be nitpicky, what's the use/purpose of something like this model? It can't do logic, it's too small to have much training data (retrievable "facts"), the context is tiny, etc

From the article itself (and it’s just one of many use cases it mentions) - Here’s when it’s the perfect choice: You have a high-volume, well-defined task. Ideal for functions like sentiment analysis, entity extraction, query routing, unstructured to structured text processing, creative writing, and compliance checks. It also explicitly states it’s not designed for conversational or reasoning use cases. So basically…

yeah, but it's clearly too limited to do any of that in its current state, so one has to extensively fine-tune this model, which requires extensive and up-to-date know-how, lots of training data, … , hence my question.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#320
post #295

Earlier quoted context omitted.

I’m work in ML and I don’t understand your point. Transfer learning usually refers to leveraging data for a different task to help with a task for which you have limited data. You’re saying that the knowledge gained from the other languages transfers to English? I don’t think for a 270M parameter model the bottleneck is the availability of enough English language training data.

> You’re saying that the knowledge gained from the other languages transfers to English? Yes, there has been many results circa 2020 or so, that have shown this to be the case. More recently, we have observed something similar with verifiable domains (see RLVR and related results) when it comes to coding tasks, specifically.

Right, but my point is that a 270M parameter model will not be bottlenecked by the availability of data for the entire English language.
Post reply on HN