It's probably too new for anyone to have integrated this into text-generation-webui / Gradio? I've been looking for a large context LLM (self-hosted or not) for a project, and as a European I unfortunately don't have access to Anthropic's Claude API yet.
Open source LLM with 32k Context Length
21–30 of 31 posts
Re: Open source LLM with 32k Context Length
#22Abacus always seemed to me like a 'we got a lot of VC money with inflated claims now we gotta show we do everything' company. I don't really understand what they do, they seem to offer everything but I don't see anyone talking about using their offerings in the real-world. Ever. The only time I see mentions of the company are when I am targeted with ads or promoted posts of the founder.
It looks like the first OSS 13B Llama 2 based 32k token context model[2], but the first OSS and commercially usable 32k token context model was a 7B Llama 2 based model[3] from Together AI, who beat them by about a week[4].
[1]: https://twitter.com/bindureddy/status/1694126931174977906
[2]: https://huggingface.co/abacusai/Giraffe-v2-13b-32k
[3]: https://huggingface.co/togethercomputer/Llama-2-7B-32K-Instr...
[4]: https://twitter.com/togethercompute/status/16925744231638470...
Re: Open source LLM with 32k Context Length
#23It seems this is built on LLAMA. Did meta change the license to make it open source now? It still seems to be showing otherwise in the repo. Edit: No mention of it being open source in the linked article. Maybe the title here is just wrong? @dang
It’s not possible to have a license over an ML model trained on other peoples’ works, since such models are uncopyrightable. They’re more like a phone book; a collection of facts trained by an entirely un-creative process. https://news.ycombinator.com/item?id=36691050 This hasn’t been proven in court, but it seems the most likely outcome.
Re: Open source LLM with 32k Context Length
#24Earlier quoted context omitted.
Llama 2 is open source-ish. Weights are freely available and can be commercially used, but only if you have less than 700m users and agree to some "don't do naughty things" terms.
"open source-ish" sounds like the perfect way to massively profitable future litigations. Also "don't do naughty things", is there a chart for that? How is that defined, is it part of the non-existing license?
What do you mean? https://github.com/facebookresearch/llama/blob/main/LICENSE
Re: Open source LLM with 32k Context Length
#2532k context length sounds nice of course, and it seems to be common to call the just fine-tuned models like that. I think it is more of a marketing thing and we really should distinguish between the context length of the pre-trained model and the fine-tuned model, with the latter being the default meaning of context length.
Re: Open source LLM with 32k Context Length
#26Does anyone know if larger context lengths are inherent worse at other task? i.e. all other things being equal is a 8k model better at math than a 32k model
OP is about a 32k sugar-coated Llama 2, so I would expect it be similar in performance to other Llama 2 derivatives.
Re: Open source LLM with 32k Context Length
#27This is just another fine-tuned LLaMA and Llama 2, like there are already some. I doubt that this will give seriously meaningful results for long context inference. 32k context length sounds nice of course, and it seems to be common to call the just fine-tuned models like that. I think it is more of a marketing thing and we really should distinguish between the context length of the pre-trained model and the fine-tun…
Re: Open source LLM with 32k Context Length
#28Does anyone know if larger context lengths are inherent worse at other task? i.e. all other things being equal is a 8k model better at math than a 32k model
They are more resource (time and memory) intensive in training and inference, that is their disadvantage. For a fair comparison you would have to compare a 8k to a 32k pre-trained model with otherwise similar hyper-parameters. OP is about a 32k sugar-coated Llama 2, so I would expect it be similar in performance to other Llama 2 derivatives.
Sorry for the random question, I've just been curious about this for a while and unable to find out and you seem knowledgeable about these extended models.
Re: Open source LLM with 32k Context Length
#29It seems this is built on LLAMA. Did meta change the license to make it open source now? It still seems to be showing otherwise in the repo. Edit: No mention of it being open source in the linked article. Maybe the title here is just wrong? @dang
It’s not possible to have a license over an ML model trained on other peoples’ works, since such models are uncopyrightable. They’re more like a phone book; a collection of facts trained by an entirely un-creative process. https://news.ycombinator.com/item?id=36691050 This hasn’t been proven in court, but it seems the most likely outcome.
Re: Open source LLM with 32k Context Length
#30This is just another fine-tuned LLaMA and Llama 2, like there are already some. I doubt that this will give seriously meaningful results for long context inference. 32k context length sounds nice of course, and it seems to be common to call the just fine-tuned models like that. I think it is more of a marketing thing and we really should distinguish between the context length of the pre-trained model and the fine-tun…
These 800 watt speakers are great. So loud.