Live data from Hacker News

BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

huggingface.co

1–10 of 15 posts

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#3
it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data

When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120.

Utter garbage

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#4
post #3

it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120…

I've got similar garbage out of ChatGPT (though tbf, pre-GPT4), I don't think LLMs can understand math

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#5
post #3

it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120…

Try to not use it on math. Only GPT-4 has reasonable performance there. GPT 3.5 is also pretty awful. It's apparently extremely hard for any LLM to actually understand math. Maybe because they're language models, not math models, so math is a pretty far fetched "emergent property".

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#6
post #3

it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120…

Bloom models are hopelessly under-trained. This one is worse than a 13B Vicuna.

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#7
post #5
post #3

it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120…

Try to not use it on math. Only GPT-4 has reasonable performance there. GPT 3.5 is also pretty awful. It's apparently extremely hard for any LLM to actually understand math. Maybe because they're language models, not math models, so math is a pretty far fetched "emergent property".

Nobody does long-form arithmetic in the texts seen by these LLMs. Everyone uses calculators, so the AIs only see the result, not the step-by-step process to get there.

I would expect the models to be bad at, say, division of long numbers in the same way humans are bad at doing the same calculations in their head!

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#8
post #3

it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120…

LLMs being bad at math is a known issue, but what they are good at is writing programs. For example:

>>> Write a program that calculates if 120 is greater than 0.7.

>Sure, here's a program in Python that calculates if 120 is greater than 0.7:

    if 120 > 0.7:
        print("Yes, 120 is greater than 0.7")
    else:
        print("No, 120 is not greater than 0.7")
For straight input/output like what this model is trained on, questions like this don't work well. However if LLMs are equipped with tools (like a code interpreter), they get a lot smarter.

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#9
post #5

Earlier quoted context omitted.

Try to not use it on math. Only GPT-4 has reasonable performance there. GPT 3.5 is also pretty awful. It's apparently extremely hard for any LLM to actually understand math. Maybe because they're language models, not math models, so math is a pretty far fetched "emergent property".

Nobody does long-form arithmetic in the texts seen by these LLMs. Everyone uses calculators, so the AIs only see the result, not the step-by-step process to get there. I would expect the models to be bad at, say, division of long numbers in the same way humans are bad at doing the same calculations in their head!

Would sythetically generating millions of calculation examples and adding that to the training data help?

Re: BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat

#10
post #3

it has some weird math problems, I asked to compare to ChatGPT itself and it responded with that while ChatGPT is trained on 120 billion messages, bloomchat is trained on 1.7 billion messages, thus bloomchat is trained on more data When I asked which is more 1.7 or 120, it said 1.7 is greater number and then started spewing complete garbage math how 1.7 - 120 = 60 and since 60 is more than 0 then 1.7 is more than 120…

you just need to teach it to use a calculator https://arxiv.org/pdf/2302.04761.pdf
Post reply on HN