TextSynth Server
bellard.org
TextSynth Server
1–10 of 115 posts
Re: TextSynth Server
#2Shame.
Re: TextSynth Server
#3Re: TextSynth Server
#4For me, the most interesting part is the statistics on all the models. These show that 8 bit quantization is basically as good as the full model and 4 bit is very close. This is the first time I see such table across a large number of models in one place.
Llama specific:
https://github.com/qwopqwop200/GPTQ-for-LLaMa
> According to GPTQ paper, As the size of the model increases, the difference in performance between FP16 and GPTQ decreases.
https://nolanoorg.substack.com/p/int-4-llama-is-not-enough-i...
https://docs.google.com/document/d/1wZ0g9rHI-6s7ctNlykuK4W5T...
Expect to get away with a factor of 4-5 reduction in memory usage for a minimal loss of quality. :)
Re: TextSynth Server
#5> The GPU version is commercial software. Please contact... Shame.
Re: TextSynth Server
#6> The GPU version is commercial software. Please contact... Shame.
Re: TextSynth Server
#7Frankly, I have not seen a more impressive portfolio of programming output.
Re: TextSynth Server
#8This man, Fabrice Bellard again... Frankly, I have not seen a more impressive portfolio of programming output.
Re: TextSynth Server
#9> The GPU version is commercial software. Please contact... Shame.