Live data from Hacker News

Bard is much worse at puzzle solving than ChatGPT

twofergoofer.com

51–60 of 86 posts

Re: Bard is much worse at puzzle solving than ChatGPT

#51

How is this possible? Google makes people do 8 rounds of leetcode. How could they be beaten? Nothing makes sense anymore.

I think they just got lazy and entitled. Maybe ChatGPT will be the scare they need. It feels bad though; they almost don’t deserve the energy this fight will give them.

Re: Bard is much worse at puzzle solving than ChatGPT

#52

Earlier quoted context omitted.

Having read through the word game, I agree with others that it's good that the game is less likely to be in the corpus. I think rhyming, while a challenging task, may be a poor benchmark for ability. The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best) I completely disagree with the "hasty rhyming test" - Skeleton and Gelatin don't rhyme (-ton vs -tin), and rhyme worse than…

> The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best) cactus / practice ? they rhyme to my mind skeleton / gelatin ? also rhyme to my ears protein / poutine also rhyme enough to be considered to rhyme You appear to be operating under the impression that the every syllable of a rhyming couplet has to rhyme exactly for it to be considered a rhyme. This is an incorrect assump…

> You appear to be operating under the impression that the every syllable of a rhyming couplet has to rhyme exactly for it to be considered a rhyme

I'm not operating under that impression, but the author is [1].

To me, the final "sounds" should match - not every syllable, an end rhyme according to wiki [0]. Specifically I would consider a rhyme to require matching sounds "at least from last vowel to end", but I don't think of rhymes first from the strict definition. Perhaps it's an accent thing but "-us" in cactus is not the same sound as "-ice" in practice. If a child made a poem with these sounds I would tell them "good job, it's a rhyme", and perhaps for the purpose of a silly word game too. But I would not use it as a passing case for a test of any sort like the author.

What's more pleasing is irrelevant, what is relevant is if its a true rhyme.

[0] https://en.wiktionary.org/wiki/end_rhyme#English

[1] https://news.ycombinator.com/reply?id=35258385&goto=item%3Fi...

Re: Bard is much worse at puzzle solving than ChatGPT

#53

Wow I had hoped for a more productive discussion than these 1-1 comparisons of Bard vs ChatGPT that I'm seeing everywhere. The model deployed with this version of Bard is clearly a smaller model than the biggest LaMDA/PaLM models Google has been working on for ages. Which, according to their publications, show unprecedented results on _proof writing_ of all things (see Minerva). While their strategic decisions may be…

> its almost silly to question Google's ability to build useful LLMs.

Unless they release a model one can "use" and verify their claims it's literally silly to make this statement.

Re: Bard is much worse at puzzle solving than ChatGPT

#54

If I understand how Large Language models work, they don't actually know about spelling.... they are given tokens that represent words, and can only infer things from the context of those tokens across terabytes of data that they're given. Any rhyming done is an impressive result.

[flagged]

Re: Bard is much worse at puzzle solving than ChatGPT

#57

Earlier quoted context omitted.

> The author doesn't seem to understand rhyming too well (cactus practice is a weak rhyme at best) cactus / practice ? they rhyme to my mind skeleton / gelatin ? also rhyme to my ears protein / poutine also rhyme enough to be considered to rhyme You appear to be operating under the impression that the every syllable of a rhyming couplet has to rhyme exactly for it to be considered a rhyme. This is an incorrect assump…

What most people consider a rhyme is that the vowel and coda of the last syllable match (of course we don't mostly reach for the technical definition). I guess the examples there might be accent dependent. Protein/poutine is the only one of those first examples that really rhymes to me; skeleton/gelatin and cactus/practice both have different vowels. Maybe different for you though.

Protein and poutine do not rhyme if you pronounce poutine the proper way in Canadian French.

And you are right of course that gelatin and practice have the short "ee" sound at the end whereas skeleton and and cactus have the "uh" sound.

Re: Bard is much worse at puzzle solving than ChatGPT

#58

> Twofer Goofer HQ's adherence to strict "perfect" rhyme can be tricky for those slant rhyme-inclined. And yet one puzzle they hammer Bard for failing is "Cactus Practice". What accent do you have to have for that to be a perfect rhyme?

From Chicago ... and with more than a thousand solves on that puzzle we've never received a single complaint about the rhyme on that one (plus the stats say users find it an extremely easy rhyme to solve)! Curious how you pronounce that one such that they don't rhyme?

Re: Bard is much worse at puzzle solving than ChatGPT

#59
post #45
post #40

Earlier quoted context omitted.

Clicking through to the link next to 'last week's text' and then to 'full rules', it looks like the author is starting the chat sessions with a full explanation that isn't included in the screenshots. (Also, the last screenshot shows the author explicitly asking about rhymes.)

Ah thanks - for others here is the link, though TFA may not necessarily have used the same prompt: https://docs.google.com/document/d/1_eg_jiUE5y8e5zeiz5HCGDc2... Based on that it looks like the author asked all 25 test puzzles in one big prompt, which one supposes would favor larger models. To compare "puzzle solving" you'd think it would make more sense to ask one puzzle at a time?

I tried it both ways; with individual prompts and prompts in bulk. I ran both tests the same way. There's a tradeoff in writing a legible/interesting blog post and relating step-by-step the way the evaluation was ran! Appreciate you reading and the feedback :)
Post reply on HN