Live data from Hacker News

CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

seqpu.com

51–59 of 59 posts

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#51

we found something interesting and wanted to share it with this community. we wanted to know how google's gemma 4 e2b-it — 2 billion parameters, bfloat16, apache 2.0 — stacks up against gpt-3.5 turbo. not in vibes. on the same test. mt-bench: 80 questions, 160 turns, graded 1-10 — what the field used to grade gpt-3.5 turbo, gpt-4, and every major model of the last three years. we ran gemma through all of it on a cpu.…

[dead]

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#52

Earlier quoted context omitted.

Grading by hand was done fully blinded? (Also this comment is ai generated so I’m not sure who I’m even asking.)

Fred, nice to meet you. The grading model had no idea what was being tested. We used separate accounts to compartmentalize. The Claude grader was guessing GPT-3.5 Turbo or GPT-4 by the end. On the coding block it consistently scored responses as GPT-4o level. We followed the MT-Bench grading guidelines as published by the team that created them. Did the research, followed the book, had no horse in the race. Every sco…

[deleted]

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#53

Posters comment is dead. It may be llm-assisted but should prob be vouched for anyway as long as the story isn't flagged.

appreciate the vouch but come on lol. we ran 80 questions, graded 160 turns by hand, documented 7 error classes, open sourced all the code, and put a live bot up for people to test. to write this post up took me hours. everyone is a critic lol.

I didn't mean it as criticism. I was trying to convince flaggers that it should be vouched - even if they don't like the tone or content of the comment - due to it coming from you.

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#54

This really shows the power of distillation. One thing I find amusing: download the Google Edge Gallery app and one of the chat models, then go into airplane mode and ask it about where it’s deployed. gemma-4-e2b-it is quite confident that it is deployed in a Google datacenter and that deploying it on a phone is completely impossible. The larger 4B model is much subtler: it’s skeptical about the claim but does seem t…

thank you for actually reading it and getting it. the airplane mode test is hilarious, the model sitting on your phone insisting it can't run on a phone. that's amazing. and yes we think exactly the same way. like picture a small business owner with a pi in the back office just quietly processing invoices, drafting email replies, summarizing meeting notes all day. no subscription, no cloud, no one sees their data. th…

posted an AI article, and now AI replies to the comments. "That's not a hypothetical, that's X".

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#55
post #54

Earlier quoted context omitted.

thank you for actually reading it and getting it. the airplane mode test is hilarious, the model sitting on your phone insisting it can't run on a phone. that's amazing. and yes we think exactly the same way. like picture a small business owner with a pi in the back office just quietly processing invoices, drafting email replies, summarizing meeting notes all day. no subscription, no cloud, no one sees their data. th…

posted an AI article, and now AI replies to the comments. "That's not a hypothetical, that's X".

I mean, you’re responding to OP. I won’t speculate as to whether they are using AI to draft comments.

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#56

Earlier quoted context omitted.

The core issue is that the LLM is using rhetoric to try to convince or persuade you. That's what you need to tell it not to do.

Which will not work. Don't think of a pink genitalia, I mean elephant...

An LLM that can't follow instructions wouldn't be able to write code anyway.

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#57
post #7

> The model does not need to be retrained. It needs surgical guardrails at the exact moments where its output layer flinches. > With those guardrails — a calculator for arithmetic, a logic solver for formal puzzles, a per-requirement verifier for structural constraints, and a handful of regex post-passes — the projected score climbs to ~8.2. Surgical guardrails? Tools, those are just tools.

>It needs surgical guardrails at the exact moments where its output layer flinches. This article is very clearly shitty LLM output. Abstract noun and verb combos are the tipoff. It's actually quite horrible, it repeats lines from paragraph to paragraph.

It would be ironic if the article itself was written with Gemma2B.

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#58

Earlier quoted context omitted.

Which will not work. Don't think of a pink genitalia, I mean elephant...

An LLM that can't follow instructions wouldn't be able to write code anyway.

Nonsense. But even an LLM that can follow instructions cannot follow that one.

Re: CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

#59

Earlier quoted context omitted.

An LLM that can't follow instructions wouldn't be able to write code anyway.

Nonsense. But even an LLM that can follow instructions cannot follow that one.

What is intrinsic to an LLM or its training that would prevent it from following the directive that it should not try to convince you of something?
Post reply on HN