> GLM-4V-9B possesses dialogue capabilities in both Chinese and English at a high resolution of 1120*1120. In various multimodal evaluations, including comprehensive abilities in Chinese and English, perception & reasoning, text recognition, and chart understanding, GLM-4V-9B demonstrates superior performance compared to GPT-4-turbo-2024-04-09, Gemini 1.0 Pro, Qwen-VL-Max, and Claude 3 Opus. But according to their ow…
GLM-4-9B: open-source model with superior performance to Llama-3-8B
11–18 of 18 posts
Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#12Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#13Isnt 3-70b so good, reddit llamaers are saying people should buy hardware to run it? Llama-3-8b was garbage for me but damn 70b is good enough
So you’d at minimum be looking at dual 3090 with NVLink for about $4000 or so. Or for the highest performing non-quantized model, you’d be spending about $40,000 for two A100’s.
Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#14Isnt 3-70b so good, reddit llamaers are saying people should buy hardware to run it? Llama-3-8b was garbage for me but damn 70b is good enough
The unquantized llama 70B requires 142GB of VRAM. Some of the quantized versions are quite decent but they do tend to get overquantized below around 26.5GB of VRAM (~3 bits per weight). So you’d at minimum be looking at dual 3090 with NVLink for about $4000 or so. Or for the highest performing non-quantized model, you’d be spending about $40,000 for two A100’s.
Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#15Earlier quoted context omitted.
The unquantized llama 70B requires 142GB of VRAM. Some of the quantized versions are quite decent but they do tend to get overquantized below around 26.5GB of VRAM (~3 bits per weight). So you’d at minimum be looking at dual 3090 with NVLink for about $4000 or so. Or for the highest performing non-quantized model, you’d be spending about $40,000 for two A100’s.
No need for NVLink just for inference, not even with tensor parallelism. And you can get used 3090 much cheaper than that.
Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#16Looks like terrific technology. However, the translation says that it's an "irrevocable revocable" non-commercial license with a form to apply for commercial use.
Translation error? output from GPT: "non-exclusive, global, non-transferable, non-sublicensable, revocable, royalty-free license."
Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#17I’m excited to hear work is being done on models that support function calling natively. Does anybody know if performance could be greatly increased if only a single language was supported ? I suspect there’s a high demand for models that are maybe smaller and can run faster if the tradeoff is support for only English. Is this available in ollama ?
Are there any other models that support function calling?
Re: GLM-4-9B: open-source model with superior performance to Llama-3-8B
#18Isnt 3-70b so good, reddit llamaers are saying people should buy hardware to run it? Llama-3-8b was garbage for me but damn 70b is good enough
The unquantized llama 70B requires 142GB of VRAM. Some of the quantized versions are quite decent but they do tend to get overquantized below around 26.5GB of VRAM (~3 bits per weight). So you’d at minimum be looking at dual 3090 with NVLink for about $4000 or so. Or for the highest performing non-quantized model, you’d be spending about $40,000 for two A100’s.