Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

151–160 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#151
post #119
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Agree directionally but you don't need $50k. $5k is plenty, $2-3k arguably the sweet spot.

Re: GLM-5.1: Towards Long-Horizon Tasks

#152
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

No killer product? Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime, technological or otherwise. They've already massively improved my and others' lives and they're only going to get better. If pre and post industrial revolution used to be the major binary delineation of our history, I'm fairly confident it will soon be seen as pre and post AI instead.

Coding assistants are currently quite hard to run locally with anything like SOTA abilities. Support in the most popular local inference frameworks is still extremely half-baked (e.g. no seamless offload for larger-than-RAM models; no support for tensor-parallel inference across multiple GPUs, or multiple interconnected machines) and until that improves reliably it's hard to propose spending money on uber-expensive hardware one might be unable to use effectively.

Re: GLM-5.1: Towards Long-Horizon Tasks

#153
post #151
post #119

Earlier quoted context omitted.

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Agree directionally but you don't need $50k. $5k is plenty, $2-3k arguably the sweet spot.

as a local LLM novice, do you have any recommended reading to bootstrap me on selecting hardware? It has been quite confusing bring a latecomer to this game. Googling yields me a lot of outdated info.

Re: GLM-5.1: Towards Long-Horizon Tasks

#154
post #151
post #119

Earlier quoted context omitted.

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Agree directionally but you don't need $50k. $5k is plenty, $2-3k arguably the sweet spot.

The 4-bit quants are 350GB, what hardware are you talking about?

Re: GLM-5.1: Towards Long-Horizon Tasks

#155
post #119

Earlier quoted context omitted.

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Is it so hard to project out a couple product cycles? Computers get better. We’ve gone from $50k workstation to commodity hardware before several times

Subscription services get all the same benefits from computer hardware getting better. But actually due to scale, batching, resource utilization, they'll always be able to take more advantage of that.

Re: GLM-5.1: Towards Long-Horizon Tasks

#156

Earlier quoted context omitted.

(1) is absolutely not true if you actually use these models on a regular basis and include Google in here too. The difference in reliability beyond basic tasks is night and day. Their reward function is just so much better, and there are many nuanced reasons for this. (2) is probably true but with caveats. Top-tier models will never run on desktop machines, but companies should (and do) host their own models. The fut…

> Top-tier models will never run on desktop machines Sorry, but you don't know that

I mean it's not hard to understand that if good model can run on consumer hardware, even better models can run in data centers

Re: GLM-5.1: Towards Long-Horizon Tasks

#157

Earlier quoted context omitted.

No killer product? Coding assistants and LLM's in general are the single most awe-inspiring achievement of humanity in my lifetime, technological or otherwise. They've already massively improved my and others' lives and they're only going to get better. If pre and post industrial revolution used to be the major binary delineation of our history, I'm fairly confident it will soon be seen as pre and post AI instead.

Coding assistants are currently quite hard to run locally with anything like SOTA abilities. Support in the most popular local inference frameworks is still extremely half-baked (e.g. no seamless offload for larger-than-RAM models; no support for tensor-parallel inference across multiple GPUs, or multiple interconnected machines) and until that improves reliably it's hard to propose spending money on uber-expensive h…

This is an argument against the grandparent's points (1) and (2), not their point (3).

Re: GLM-5.1: Towards Long-Horizon Tasks

#158
post #119

Earlier quoted context omitted.

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Why would anyone need more than 640Kb of memory?

Exactly the point though. In the 640KB days there was no subscription to ever increasing compute resources as an alternative.

Re: GLM-5.1: Towards Long-Horizon Tasks

#159
post #157

Earlier quoted context omitted.

Coding assistants are currently quite hard to run locally with anything like SOTA abilities. Support in the most popular local inference frameworks is still extremely half-baked (e.g. no seamless offload for larger-than-RAM models; no support for tensor-parallel inference across multiple GPUs, or multiple interconnected machines) and until that improves reliably it's hard to propose spending money on uber-expensive h…

This is an argument against the grandparent's points (1) and (2), not their point (3).

It's one clear argument for the (so get to work!) part.

Re: GLM-5.1: Towards Long-Horizon Tasks

#160
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

How good would open source models be if they couldn't distill higher quality private models?
Post reply on HN