Live data from Hacker News

Why your local LLM feels dumber than it is

forum.level1techs.com

221–230 of 233 posts

Re: Why your local LLM feels dumber than it is

#221
post #189

Earlier quoted context omitted.

Was there a more recent refresh or is this the model from a year ago? The frontier models were barely functional and almost useless a year ago (gpt oss was pre opus 4.5!) - I would be very surprised if the original drop is anything more than totally obsolete/irrelevant at this point

Yes, the old entirely stupid old gpt-oss. But Sonnet and GPT were very useful then already, qwen also.

I find that I remember models being a lot better than they were, even when I remember them being not very good - because of a novelty factor ("whoa it can do that now?") mostly. And then I go back and look at them and its like, what how did I find this impressive.

A funny example - I remember thinking "yeah sonnet 3.5 is a really good coding model"

https://stack.convex.dev/using-cursor-claude-and-convex-to-b...

>Prompting Cursor to Scaffold my App: FAIL This was my first hurdle.

>It became immediately apparent that I would not be able to prompt my way through the entire process.

>While the tooling we have is undeniably powerful, it's not yet capable of completing most nontrivial tasks

It couldn't run pnpm install lmao. Opus 4.5 was a crazy jump

Re: Why your local LLM feels dumber than it is

#222
post #73

Earlier quoted context omitted.

Qwen3.8-27B runs at 59.5 tok/s on my M4 Max, 40-core GPU, 128 GB I use it occasionally for classification and other tasks but I wouldn't trust those smaller models with the real work and for larger data processing it's too slow, e.g. a dataset I wanted to classify would've taken 56 days on my laptop vs just paying the cheap Luna prices to openai and getting it done in a few hours.

59.5 t/s is really good. Which engine/quant are you using?

[dead]

Re: Why your local LLM feels dumber than it is

#223
post #91

Earlier quoted context omitted.

No idea what you are talking about. My battery lasts longer than ever while running vim and make and GCC. It’s amazing. Not sure why your windows are closing.

Because the local LLM, which you are not running, is running for much longer than gcc and is eating the battery. Different choices, different outcomes.

> There was a lovely window of a few years when processors were fast enough and low-power enough that real development work could trivially happen on a Macbook Air in a lounge.

I was responding to this. I am appalled that anyone thinks (and is willing to say out loud in public) that they cannot do "real dev work" without an LLM.

Re: Why your local LLM feels dumber than it is

#225
post #125

Earlier quoted context omitted.

If you're running it while idle and don't need the quickest results, reducing the clock speed improves energy efficiency (and in your case avoids overheating the battery). There will be some optimal speed that maximizes computations per joule that depends on the specific load and can only be found by measurement. On Linux, you can cap CPU frequencies with "cpupower". Does MacOS have any equivalent?

The Mac unfortunately just has two performance modes ‘all out power and melting’ or ‘cold and really really slow’

My M5 Max has 3 modes, "low", "automatic", and "high". Automatic doesn't seem to simply switch between low and high -- it seems to sit in the middle and vary dynamically based on workload and temperature. With automatic I get more than half as many tokens per second as on high, but with much lower fan speed and temperature.

Re: Why your local LLM feels dumber than it is

#226

Earlier quoted context omitted.

Does it need to respond fast? For important applications, I'm sure we'd all be fine waiting 20 minutes for a high quality, usable answer. Or is it the need for interative refinements that make speed relevant?

if you are so sure about what the final shape of your output is then its prbly not a common use of ai

If you are sure about the final shape of your output it's a great use case for AI, as you can define what you want in your prompt and refine towards it!

It's where you don't know the end state you're looking for that you'll end up generating slop on top of slop and creating a whole Gastown just to power your Gastown.

Re: Why your local LLM feels dumber than it is

#227
post #192

Earlier quoted context omitted.

That's funny, because I just went through the opposite. I got llama.cpp working with qwen3.6 and qwen3.8 by Googling and manually adjusting things according to reddit posts and Google not-really-helpful AI suggestions. I tried settings up per-model stuff in settings.json, but again Google got in my way, and llama.cpp having 2 different settings.json (and Google lying about where 1 goes) made it far too difficulty to…

Am I the only one who just downloads directly from LM Studio and just runs the server there? It’s trivial.

You're not the only one, I'm running Qwen3.8-27B in LM Studio and it seems to be going great. Was very easy to set up.

Re: Why your local LLM feels dumber than it is

#228
post #193

Earlier quoted context omitted.

> What the big AI labs have over the smaller labs, is a huge amount of data and diverse set of tasks, hence they generalize better, but still not great. Hehe, this kind of sounds like the opposite of generalization. As in it’s just specialization at scale.

It is specialization at scale. It’s paying hundreds of thousands of RLHF’ers from every subject and through some dystopian income stream. It’s decent money don’t get me wrong, but you aren’t paid at if a task isn’t completed in time for example.

Tbh the way you're treated depends on the hotness of the task.

A couple of years ago, people I know got paid OK for relatively simple programming and logic RLHF tasks. But very soon it turned dark because that sort of data was required less and less, and the number of feedbackers has grown.

Today, the type of data the model developers pay for requires actual domain experience. E.g in software engineering they have people work in simulated environments with other LLMs, grade them, feedback, PRs, Jira everything.

This pays well and they treat you well, entice you with more money/task/hour, etc, for now. In few years when this gets drilled into LLMs, these guys will face the same painful hours and bad pay and less work and so on too.

In physical tasks, we are still in the early stages where basic packing clothes (in a textile factory setting) etc is being recorded and data is only now being used for training. Due to the problems with translating human hand data to robotic hands, these people do the factory work holding robotic grippers and operating that, you should see a YouTube video. But this also means that it's much less sweeping than data collection in SWE. In many cases it's not practical to collect data given that you have to use the specific gripper, wear a big gopro type thing, etc. So I am expecting much slower of an impact on physical tasks (of this kind) compared to how quick the uptake was in SWE/math.

I have yet to hear back from them on how it's going for teamwork white collar tasks, it's been a few months. Everyone is paying for the end products of those it seems - grok bot, perplexity computer, claure cowork, chatgpt work etc,. Not as much as their coding agents of course.

Re: Why your local LLM feels dumber than it is

#229
post #193

Earlier quoted context omitted.

It is specialization at scale. It’s paying hundreds of thousands of RLHF’ers from every subject and through some dystopian income stream. It’s decent money don’t get me wrong, but you aren’t paid at if a task isn’t completed in time for example.

Tbh the way you're treated depends on the hotness of the task. A couple of years ago, people I know got paid OK for relatively simple programming and logic RLHF tasks. But very soon it turned dark because that sort of data was required less and less, and the number of feedbackers has grown. Today, the type of data the model developers pay for requires actual domain experience. E.g in software engineering they have pe…

[dead]

Re: Why your local LLM feels dumber than it is

#230

Earlier quoted context omitted.

Seems like he's self-aware, though - he's taking steps to move away from brain-atrophy.

My ability to detect sarcasm is not good. From looking at funlang's profile and other comments the profile looks like a LLM generated bot. Forums with full no verification pseudonyms seem like they have a real challenge ahead. How long until we need humanhackernews.com with public pseudonyms and a private trusted verification?

[flagged]
Post reply on HN