Live data from Hacker News

OpenChat: Advancing open-source language models with imperfect data

github.com

21–27 of 27 posts

Re: OpenChat: Advancing open-source language models with imperfect data

#21
post #12

Submitters: " Please use the original title, unless it is misleading or linkbait; don't editorialize. " - https://news.ycombinator.com/newsguidelines.html If you want to say what you think is important about an article, that's fine, but do it by adding a comment to the thread. Then your view will be on a level playing field with everyone else's: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so... (Submitt…

Not sure if that's editorializing, that was claimed by the link.

That's still very much editorializing.

Re: OpenChat: Advancing open-source language models with imperfect data

#23
post #12

Submitters: " Please use the original title, unless it is misleading or linkbait; don't editorialize. " - https://news.ycombinator.com/newsguidelines.html If you want to say what you think is important about an article, that's fine, but do it by adding a comment to the thread. Then your view will be on a level playing field with everyone else's: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so... (Submitt…

Not sure if that's editorializing, that was claimed by the link.

The link has a pretty good non clickbait title. There's no reason to pick a clickbait reference from the link.

Re: OpenChat: Advancing open-source language models with imperfect data

#25
post #24

Earlier quoted context omitted.

Are you implying they're training on the test set / benchmark data? If not, what do you mean by this?

I believe that's what they mean, yes.

Some explanation or link that supports that claim would be a really good idea in that case...

Re: OpenChat: Advancing open-source language models with imperfect data

#26
I am not an AI engineer, but my intuition tells me if we could ever clean up the @#$& datasets these LLMs are trained on and give them coherent, non-contradictory training, we would be shocked by what they could do.

I suspect 90% of the criticism of AIs is because people are underestimating them.

Re: OpenChat: Advancing open-source language models with imperfect data

#27

I am not an AI engineer, but my intuition tells me if we could ever clean up the @#$& datasets these LLMs are trained on and give them coherent, non-contradictory training, we would be shocked by what they could do. I suspect 90% of the criticism of AIs is because people are underestimating them.

I have the same feeling. It's amazing the amount of garbage that was fed to current LLMs, yet they perform very well. I hope they will become incredible with enough curation and specialization.
Post reply on HN