Live data from Hacker News

XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

blog.salesforceairesearch.com

1–10 of 96 posts

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#6

when will the llm race peak? have we peaked already?

If you’re curious, you can check the progress of many open source LLM’s and how they perform on various evals here:

https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#7
post #4

From all the experimentation I've done, 7B parameter models just don't seem to be able to produce useful output reliably enough for my use cases. What use cases do people have for these smaller LLM's?

The main use case is that it's probably the only size consumers can run on their personal devices. If you don't want your data going into an external platform like OpenAI it's the only solution even if it's not very usuable.

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#8
post #5

I have no idea what any of these words mean, but I'd like to. Can someone point me in the direction of an "AI for Dipshits"?

Why not ask the AIs themselves? They are pretty good at explaining this type of thing.

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#9
post #4

From all the experimentation I've done, 7B parameter models just don't seem to be able to produce useful output reliably enough for my use cases. What use cases do people have for these smaller LLM's?

> What use cases do people have for these smaller LLM's?

None. Training a functionally useless model and releasing it is a great way to demonstrate that your company is hip and current. That way when prospective clients ask about AI you can vaguely gesture at some model that you released and say you employ cutting edge AI experts.

Post reply on HN