Live data from Hacker News

LLMs Are Complicated Now

ianbarber.blog

71–80 of 86 posts

Re: LLMs Are Complicated Now

#71
post #13

Earlier quoted context omitted.

I’m not sure if it is written by an LLM, but anything being called “load-bearing” (formatted that way and all) sets off my alarm bells

Hopefully you don't work in construction!

Amusing :) but I specifically mean in a SWE context. It’s used constantly by LLMs, Claude especially. Genuinely so curious to see certain phrases come up so often, beyond what I would’ve expected for how often they’re used by people

Re: LLMs Are Complicated Now

#72
For someone like me who's never done any hands-on work in ML, the blog is really hard to understand. Whoosh, over my head.

But, I think the underlying problem is that we don't understand how this sh*t works. So, it's just an empirical, iterative mess.

Like physics in the the years shortly before relativity and quantum mechanics.

Re: LLMs Are Complicated Now

#73
post #14
post #5

Earlier quoted context omitted.

[[citation needed]] I am a professional writer and have been for over 30 years. (I do not use any form of LLM ever.) This means I read a lot . This also means that I have 30+ years of experience of readers not understanding what I wrote, or not getting further than the title, or not getting the main message, or inverting it in their heads, or inserting their own message and then complaining when I diverge, and an end…

Claude's writing style is at least as distinctive as any human's personal style. It has a long list of favorite words, verbal tics and common structures. On top of that, LLM writing is often bad in a very particular way: it's weak on actual things to say, but with an overheated style. Some days, I spend over 4 hours a day reading walls of text written by Claude. If I couldn't recognize Claude's default "voice" by now…

> It has a long list of favorite words, verbal tics and common structures.

Go on, then, let's see the list.

Re: LLMs Are Complicated Now

#74
post #15
post #5

Earlier quoted context omitted.

[[citation needed]] I am a professional writer and have been for over 30 years. (I do not use any form of LLM ever.) This means I read a lot . This also means that I have 30+ years of experience of readers not understanding what I wrote, or not getting further than the title, or not getting the main message, or inverting it in their heads, or inserting their own message and then complaining when I diverge, and an end…

You need to start using LLMs a lot and then you will know how we know. Edit: You know how you can recognise someone just from their gait while they walk towards you? I would struggle to describe that for an individual person but it doesn't mean I can't identify them from that alone.

> You need to start using LLMs a lot and then you will know how we know.

$SWEARWORD no!

I am an AI vegetarian: I won't touch them ever except for unavoidable ones (defaults on some search engines, e.g. when I use a web browser on its default settings – my own disable that junk) or for machine translation, which remains their only really valid use case in my book.

Re: LLMs Are Complicated Now

#75

Earlier quoted context omitted.

The author is correct, the model architecture is now much more complicated. You can see this if you use llama.cpp and follow the project. The earlier models were always fully implemented. Yet with more contributors, as of today tons of latest models only have partial implementation. DeepSeekv3.2 isn't fully implemented, same with KimiK2.6, GLM5.2+, DeepSeekv4 has no implementation, MiniMaxM3 not supported yet, Hy3-pr…

indeed, there's even a (pretty solid) custom server just for DS4 https://github.com/antirez/ds4 -- works very well on high-RAM Macs

I don’t k ow what I’m doing wrong. Everyone says ds4 is faster than a lot of models around the same size, but I’m getting 2t/s with DSv4 vs 12t/s with Minimax 2.7 (16Gb 5080 + 16gb 5060ti + 128gb ram).

Re: LLMs Are Complicated Now

#76
post #55
post #29

It's the bitter-lesson to feature-engineering lifecycle. When a technique or technology is new people are making massive gains by just applying it to some use case, or gathering more data for training, or giving it more resources. As time goes on those "bitter lesson" gains start to hit the shallow part of the logistic curve and companies have to start investing more and more effort into engineering for each small, i…

I assume the choice of phrase "bitter lesson" is intentional irony (since the original concept is that you get better results by just scaling up and not trying to be clever with domain-specific knowledge)?

Specifically the you get better results with techniques that can scale up to the amount of data you have. People often think of it as a "just brute force it" but for the lesson to apply you do need to come up with how you're gong to get the data and how you're going to use it.

Re: LLMs Are Complicated Now

#77
post #51
post #5

Earlier quoted context omitted.

[[citation needed]] I am a professional writer and have been for over 30 years. (I do not use any form of LLM ever.) This means I read a lot . This also means that I have 30+ years of experience of readers not understanding what I wrote, or not getting further than the title, or not getting the main message, or inverting it in their heads, or inserting their own message and then complaining when I diverge, and an end…

I want Scrabble rules for HN AI challenges. If someone finds an AI-generated comment, the commenter has violated HN guidelines and the comment should be deleted. But if the accusation is wrong, there should be a penalty for the often massive disruption the accuser has caused to the discussion. (As of now, that four-word low-effort comment has generated over a thousand words in response, none of which improve this art…

I like this idea. Wildly impractical, but I am tired of every article followed by accusations, "Smells like AI!!! I can tell by the pixels"

Re: LLMs Are Complicated Now

#78

Earlier quoted context omitted.

> Why didn't this author compare Llama 3 with GLM 5.2 (released 1 week ago) which is a more standard attention based LLM? To compare 2 separate families of LLMs and then pointing out that they are different is not a surprising result and detracts from the point the author is trying to make. The entire point of the comparison is that LLMs look vastly different today than before. Comparing more similar LLMs would detra…

It is misleading the reader since most current LLMs look the same as before. It is cherry picking an example to make a point when it's not necessary at all to make the argument he is trying to make.

> It is misleading the reader since most current LLMs look the same as before.

But most of them do not? They do look vastly different from the earlier incarnations of GPT and Llama.

Re: LLMs Are Complicated Now

#79
post #50

Earlier quoted context omitted.

Leaving out that comma is not a grammar mistake. The comma slightly changes the feel of the sentence, but it's not wrong to include or omit it. But yeah, I agree with the other commenter that AI would be less likely than a human to omit the comma.

An adverbial clause at the front of a sentence should have a comma? Or is that not what it is?

Adverbial clauses at the front of a sentence don't always have to have a comma. It's often taught this way to new English language learners or in style guides that want to keep things as unambiguous as possible because including the comma in that case is never wrong. But it's also sometimes optional.

Here is the Chicago Manual of Style's Q&A on commas: https://www.chicagomanualofstyle.org/qanda/data/faq/topics/C...

One of the answers cites section 6.34 of the Chicago Manual of Style as follows: "Although an introductory adverbial phrase can usually be followed by a comma, it need not be unless misreading is likely. Shorter adverbial phrases are less likely to merit a comma than longer ones."

For the example we're discussing, "Back in 2022 or 2023", my personal instinct as a university-educated native US English speaker would be to include or omit the comma primarily depending on how much emphasis I want to put on the timing of my statement. I also know that I tend to write overly long sentences with too many commas, so sometimes I'd intentionally counteract my own tendency and limit the complexity of the sentence by removing a comma like that if the sentence still seems good without it. Other times I'd split the sentence into two, but my sentences rarely get as staccato as AI-written ones.

Re: LLMs Are Complicated Now

#80
post #29

It's the bitter-lesson to feature-engineering lifecycle. When a technique or technology is new people are making massive gains by just applying it to some use case, or gathering more data for training, or giving it more resources. As time goes on those "bitter lesson" gains start to hit the shallow part of the logistic curve and companies have to start investing more and more effort into engineering for each small, i…

I got a very different message from this, actually much closer to the problem of incumbent advantage. The known-good thing has been heavily optimized for performance, making it much harder for new technologies to prove that they are better. This is similar to the problem of gas vs electric engines - we had a century of optimization and ecosystem development around gas engines, which creates an uphill battle for elect…

what is the known-good thing? The whole point is that LLMs were not optimized at all, they got better results than older ML algorithms just because they are able to use all of the GPU, where older algorithms are designed for 10yo GPUs and can't make use of modern GPUs. But now you do in fact have to optimize, to the point that transformers look a lot more complicated than "attention is all you need."
Post reply on HN