Live data from Hacker News

Free Dolly: First truly open instruction-tuned LLM

databricks.com

31–40 of 70 posts

Re: Free Dolly: First truly open instruction-tuned LLM

#31
post #28

There's some blatant astroturfing from new accounts going on in this thread - I gotta say it's not the best impression

The irony is that the English is bad in a lot of the astroturf comments… why don’t they use the LLM to generate them? It would probably seem more genuine than the current comments.

Re: Free Dolly: First truly open instruction-tuned LLM

#34

Any info on Pythia base model performance versus GPT-3 or 3.5? Couldn't find any benchmarks in the paper. I imagine LLaMA is ahead there.

Do we have any quantitative way of benchmarking the quality of these models at all? Like, I don’t care if a model takes one minute per token on my laptop as long as it’s “GPT-4 quality”, and I don’t care if it does 100 tokens per second if it’s straight crap. But every comparison I see people make regarding quality seems to come from “I asked it a couple of my favorite questions and it did… uh, only a little worse than GPT imo”

Re: Free Dolly: First truly open instruction-tuned LLM

#37
post #32

Why is this post full of n00b user 'comments'?

Astroturfing. Likely databricks told their employees to create HN accounts and comment on this thread to get traction (or just had one PR person make a bunch.)

@dang, any chance we can just ban all these accounts? Seems to be pretty cut and dry here.

Re: Free Dolly: First truly open instruction-tuned LLM

#38
post #28

There's some blatant astroturfing from new accounts going on in this thread - I gotta say it's not the best impression

Indeed, I'm not sure how to summon @dang but the number of databricks shill comments in this thread is absurd.

Yes, this is a cool enough achievement on its own. HN is not the best place for these "testimonials". Unless this is some sort of demo of the LLM itself /s

Re: Free Dolly: First truly open instruction-tuned LLM

#40
post #34

Any info on Pythia base model performance versus GPT-3 or 3.5? Couldn't find any benchmarks in the paper. I imagine LLaMA is ahead there.

Do we have any quantitative way of benchmarking the quality of these models at all? Like, I don’t care if a model takes one minute per token on my laptop as long as it’s “GPT-4 quality”, and I don’t care if it does 100 tokens per second if it’s straight crap. But every comparison I see people make regarding quality seems to come from “I asked it a couple of my favorite questions and it did… uh, only a little worse th…

Perplexity scores are pretty common, which I think involves taking a text corpus like wikipedia and seeing how well the model predicts the next word (token) of it.
Post reply on HN