Live data from Hacker News

Bard is much worse at puzzle solving than ChatGPT

twofergoofer.com

41–50 of 86 posts

Re: Bard is much worse at puzzle solving than ChatGPT

#41
post #15

"Bard is much worse than ChatGPT at solving an obscure word game I invented" would have been a more honest title, but would probably generate less clicks for the author. Bard may still be much worse than ChatGPT at solving all kinds of puzzles, but the article is click bait for promoting the author's word game, not an actual investigation that warrants that conclusion.

Having read through the word game, I agree with others that it's good that the game is less likely to be in the corpus. I think rhyming, while a challenging task, may be a poor benchmark for ability. The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best)

I completely disagree with the "hasty rhyming test" - Skeleton and Gelatin don't rhyme (-ton vs -tin), and rhyme worse than protein and poutine (-een vs --een).

Re: Bard is much worse at puzzle solving than ChatGPT

#42
post #15

"Bard is much worse than ChatGPT at solving an obscure word game I invented" would have been a more honest title, but would probably generate less clicks for the author. Bard may still be much worse than ChatGPT at solving all kinds of puzzles, but the article is click bait for promoting the author's word game, not an actual investigation that warrants that conclusion.

Having read through the word game, I agree with others that it's good that the game is less likely to be in the corpus. I think rhyming, while a challenging task, may be a poor benchmark for ability. The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best) I completely disagree with the "hasty rhyming test" - Skeleton and Gelatin don't rhyme (-ton vs -tin), and rhyme worse than…

> The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best)

cactus / practice ?

they rhyme to my mind

skeleton / gelatin ?

also rhyme to my ears

protein / poutine

also rhyme enough to be considered to rhyme

You appear to be operating under the impression that the every syllable of a rhyming couplet has to rhyme exactly for it to be considered a rhyme. This is an incorrect assumption. In fact, the above rhymes are arguably more pleasing because they are inexact rhymes rather than being exact forced rhymes.

In your world the only actual rhymes would be

bold / cold / gold

and

double / trouble / bubble

types of rhymes but the world considers the following to be perfectly acceptable

sent to meet her / centimetre

and so on

Re: Bard is much worse at puzzle solving than ChatGPT

#43

How is this possible? Google makes people do 8 rounds of leetcode. How could they be beaten? Nothing makes sense anymore.

People from Google have argued that's exactly why they're failing.

Personally, as someone who worked at a company that was up over 500% during the pandemic, shipped absolutely nothing during that spike, and then deflated below their pre-pandemic pricing, I saw the foley in hiring smart people first hand.

It's not enough to hire the smartest people, and in fact it can be a competitive disadvantage. The smartest people often want their piece of the product to reflect their ingenuity no matter how ancillary it is to the core mission. Unfortunately that often precludes the kind of agility that businesses need to stay competitive.

OpenAI managed to poach Googlers by simply not having fiefdoms built by smart people™. I imagine if Google had built GPT-4, it wouldn't be having downtime today. Because it wouldn't be public. And it might never be public because it doesn't scale for Google scale yet, and the ethicists want their say, and we need to integrate it into Borg and the front end hasn't passed through enough layers of design and...

Re: Bard is much worse at puzzle solving than ChatGPT

#44
post #8
post #6

Earlier quoted context omitted.

s/Microsoft/Google/ http://www.paulgraham.com/microsoft.html

But was Microsoft ever cool in the sense that Google was circa 2005? (Asking honestly, I'm too young to remember that period.)

https://www.youtube.com/watch?v=Vhh_GeBPOhs&ab_channel=MrWue...

Really cool

Re: Bard is much worse at puzzle solving than ChatGPT

#45
post #40
post #33

Am I missing something? Most of TFA is about Bard failing to answer with rhyming words, but in the only prompts shown the author doesn't actually ask for rhyming words. He just says the hint and the name of the puzzle. Is this not simply: "Bard is worse than ChatGPT at having seen the 'how-to-play' page for my side project during its training"?

Clicking through to the link next to 'last week's text' and then to 'full rules', it looks like the author is starting the chat sessions with a full explanation that isn't included in the screenshots. (Also, the last screenshot shows the author explicitly asking about rhymes.)

Ah thanks - for others here is the link, though TFA may not necessarily have used the same prompt: https://docs.google.com/document/d/1_eg_jiUE5y8e5zeiz5HCGDc2...

Based on that it looks like the author asked all 25 test puzzles in one big prompt, which one supposes would favor larger models. To compare "puzzle solving" you'd think it would make more sense to ask one puzzle at a time?

Re: Bard is much worse at puzzle solving than ChatGPT

#46

Wow I had hoped for a more productive discussion than these 1-1 comparisons of Bard vs ChatGPT that I'm seeing everywhere. The model deployed with this version of Bard is clearly a smaller model than the biggest LaMDA/PaLM models Google has been working on for ages. Which, according to their publications, show unprecedented results on _proof writing_ of all things (see Minerva). While their strategic decisions may be…

Everyone’s gonna have bigger models. Where are their useful models?

Re: Bard is much worse at puzzle solving than ChatGPT

#47

Earlier quoted context omitted.

Having read through the word game, I agree with others that it's good that the game is less likely to be in the corpus. I think rhyming, while a challenging task, may be a poor benchmark for ability. The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best) I completely disagree with the "hasty rhyming test" - Skeleton and Gelatin don't rhyme (-ton vs -tin), and rhyme worse than…

> The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best) cactus / practice ? they rhyme to my mind skeleton / gelatin ? also rhyme to my ears protein / poutine also rhyme enough to be considered to rhyme You appear to be operating under the impression that the every syllable of a rhyming couplet has to rhyme exactly for it to be considered a rhyme. This is an incorrect assump…

What most people consider a rhyme is that the vowel and coda of the last syllable match (of course we don't mostly reach for the technical definition).

I guess the examples there might be accent dependent. Protein/poutine is the only one of those first examples that really rhymes to me; skeleton/gelatin and cactus/practice both have different vowels. Maybe different for you though.

Re: Bard is much worse at puzzle solving than ChatGPT

#48
post #8
post #6

Earlier quoted context omitted.

s/Microsoft/Google/ http://www.paulgraham.com/microsoft.html

But was Microsoft ever cool in the sense that Google was circa 2005? (Asking honestly, I'm too young to remember that period.)

Yes, somehow.

https://rarehistoricalphotos.com/windows-95-launch-day-1995/

I remember people commenting "never seen such a thing before, for a computer OS!"

"Many electronics stores held midnight launches for the product, with thousands of people waiting in line to be the first to get their hands on the operating system.

The release was a tremendous success. Microsoft sold 7 million copies in the first five weeks, and Windows 95 was soon the most popular operating system on the market."

Re: Bard is much worse at puzzle solving than ChatGPT

#50

Wow I had hoped for a more productive discussion than these 1-1 comparisons of Bard vs ChatGPT that I'm seeing everywhere. The model deployed with this version of Bard is clearly a smaller model than the biggest LaMDA/PaLM models Google has been working on for ages. Which, according to their publications, show unprecedented results on _proof writing_ of all things (see Minerva). While their strategic decisions may be…

At the moment unless we get more information about what metric you're supposed to evaluate it on, you could probably simplify the headline to just "Bard is much worse than chatgpt" without any loss of accuracy.

It's not really realistic to expect people to give Google credit for these amazing models they have published results about but haven't let people play with - they have given people Bard and people are evaluating it based on the criteria most obvious to them - a comparison to a very similar product that has just been released.

Post reply on HN