Live data from Hacker News

XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

blog.salesforceairesearch.com

91–96 of 96 posts

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#91
post #66

Earlier quoted context omitted.

7B LLaMA is a terrible general purpose model, but the finetunes are pretty good at very specific roles, like dialogue/roleplay, a dungeon master bot or even code completion. The metrics are good though, perhaps placing this closer to 13B. And 8K context is huge . When you can stuff that much example text in, it gives the model more to "latch onto," and its also the point where you would start worrying about RAM/VRAM…

You must have missed the memo... It's now super easy to extend the context of 2k llama models to 8k, 16k, or even 32k with just a small fine tune and a tweak to the code. You still need the memory to be able to go that high, but it's totally doable.

I saw the SuperHOT LORAs as well.

But I assumed full training would give better perplexity for large contexts, and perhaps this method would be more effective at 16K+ with an 8K model to start with.

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#92
post #86

Earlier quoted context omitted.

Don't just post ChatGPT answers as comments on hackernews. This one doesn't even make any sense. Of course it doesn't have 7B parameters _per_ neuron.

Doesn't look like ChatGPT. Grammatical errors like "on large data corpus," the poor comma usage, misspelling crude, etc. are more of a human thing.

Maybe you are right. I was confused by sentences like "I hope this simplifies it. You can research further if you are interested! Hope it helps!" which seemed to be responding to a prompt other than just the previous comment.

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#93

Earlier quoted context omitted.

> What use cases do people have for these smaller LLM's? None. Training a functionally useless model and releasing it is a great way to demonstrate that your company is hip and current. That way when prospective clients ask about AI you can vaguely gesture at some model that you released and say you employ cutting edge AI experts.

yep, that’s why it’s free. If it was good , then they’d charge for it. …what they (and everyone) is gonna do is play with smallish models to iterate on the process for relatively small expense and earn karma. Then pay big $$$ to make a really good model for internal use and/or an api that people have to pay for. Tldr; it’s free. It’s by sales force. You should expect it to be a) crippled and b) a loss leader for a pa…

If your main product is a business tool to spam people, it would be in your best interest to enable as many new businesses to sprout up and start spamming people as possible. It would also be in your best interest to prevent your competitors from creating that enablement product and selling it as a service, earning revenue, and pumping that money back into their spam product.

These free models are both a defensive move against behemoths, and kindling to rapid business development.

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#94
post #66

Earlier quoted context omitted.

You must have missed the memo... It's now super easy to extend the context of 2k llama models to 8k, 16k, or even 32k with just a small fine tune and a tweak to the code. You still need the memory to be able to go that high, but it's totally doable.

I saw the SuperHOT LORAs as well. But I assumed full training would give better perplexity for large contexts, and perhaps this method would be more effective at 16K+ with an 8K model to start with.

Possibly, but the perplexity has shown to decrease while fine-tuning a 2048 model on larger context sizes for outputs within it's original context limit...so, more research needed.

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#95
post #45

Earlier quoted context omitted.

The number of parameters seems meaningless when the training sets are dogshit and hardened old gum chipped from the shoes of Gregslist and FleaBay. OTOH, there ought to be a construction in the form of an web app that can pinch out nonrepetitive, coherent ~100 page trashy romance novels in the style of any author given name with open source or specific text(s) or transcripts with enough original input volume: Churchi…

Do you have any recommendations for cryptocoins?

Do you always ask flippant and rude questions that add no value to HN?

Re: XGen-7B, a new 7B foundational model trained on up to 8K length for 1.5T tokens

#96

Earlier quoted context omitted.

yep, that’s why it’s free. If it was good , then they’d charge for it. …what they (and everyone) is gonna do is play with smallish models to iterate on the process for relatively small expense and earn karma. Then pay big $$$ to make a really good model for internal use and/or an api that people have to pay for. Tldr; it’s free. It’s by sales force. You should expect it to be a) crippled and b) a loss leader for a pa…

> If you want a good free open model, you’re kidding yourself if you think a corporate giant is going to kiss you on the head and give it to your for free. Yep! That makes sense! I would love to be the CEO of the company that does give away an actually useful model and little forehead kisses though. The amount of goodwill that one would generate from that would be astronomical and training costs are getting so low th…

Not quite the same but that's what Stable Foundation did with stable diffusion.
Post reply on HN