Live data from Hacker News

The threat is comfortable drift toward not understanding what you're doing

ergosphere.blog

651–660 of 668 posts

Re: The threat is comfortable drift toward not understanding what you're doing

#651

Earlier quoted context omitted.

All I'm saying is "I do not know what is the real upside on leaving the current practices in academia and education in favor of 'let the LLM guide you'". If it was my ass on the line, I would apply the precautionary principle and I wouldn't take any significant bets with my future around this. You on the other hand are the one "being sure" about how everything will be fine and of course there is no way for you to bri…

> “I do not know what is the real upside on leaving the current practices in academia and education in favor of ‘let the LLM guide you’” The fact that you think “let the LLM guide you” is the argument I’ve been making tells me everything about how honestly you’ve engaged with it. I’m done here.

And the fact that you only got to shut up after being asked what you are willing to put on the line tells me how devoid of meaning your argument is, no matter what it is.

Re: The threat is comfortable drift toward not understanding what you're doing

#652

Earlier quoted context omitted.

From the article: > There's a common rebuttal to this, and I hear it constantly. "Just wait," people say. "In a few months, in a year, the models will be better. They won't hallucinate. They won't fake plots. The problems you're describing are temporary." I've been hearing "just wait" since 2023. We're not trending towards superintelligence with these AIs. We're trending towards (and, in fact, have already reached) s…

> Untrained children write better code than the most sophisticated LLMs, without even noticing they're doing anything special. I’ll take that bet. How much money would you like to put on this, and we’ll have a neutral third party pick both the untrained child and the LLM. Let me know.

I'm willing to bet 10% of my net worth on this. But my claim was not about any given untrained child (for instance, a child who does not want to program would do poorly): a fair bet would allow me to choose the child, you to choose the LLM, use a task and programming language of the child's choice, and have a neutral third-party familiar with the programming language judge "better code". (I would, of course, want to ensure that the judge used an appropriate rubric: RLHF can produce a sophisticated turd-polisher. Perhaps the evaluation process could involve modifications made to the program?)

It is (rightly) difficult to get hold of one uninvolved child, for safeguarding reasons, so it would be better to run it as a school (or interschool) competition, where multiple children may participate. For fairness, you may also provide multiple LLM participants (however you define that). The winner of the contest, as determined by the judge, would then determine the winner of the bet ­– unless the winning child had been trained, in which case we would fall back to the next-highest-ranked participant. The number of LLM candidates would be equal to the number of eligible children.

However, I don't see a good way to allow each child to pick a programming language and task, without leaving the competition results incomparable. So perhaps each child should be paired with an LLM, and the judge should determine which submission from each pair is better? But then if I only need one victory (to support my claim), this is clearly unfair. So each pair should be tested enough to determine whether they're consistently better than the LLM… but then we are demanding a lot of the child participants, for no real benefit to them.

If we can agree on a workable protocol, I can try to pull some strings and see if we can make this happen. I could use the money.

Re: The threat is comfortable drift toward not understanding what you're doing

#653
post #149

Earlier quoted context omitted.

Precisely. The first 10 rungs of the ladder will be removed, but we still expect you to be able to get to the roof. The AI won't get you there and you won't have the knowledge you'd normally gain on those first 10 rungs to help you move past #10.

That’s a good analogy but I think we’ve already went from 0 to 10 rungs over the last couple of years. If we assume that the models or harnesses will improve more and more rungs will be removed. Vast majority of programmers aren’t doing novel, groundbreaking work.

I think this was an unpopular opinion mainly because it's scary, rather than there being an obvious reason to think otherwise.

Re: The threat is comfortable drift toward not understanding what you're doing

#654

> Schwartz's experiment is the most revealing, and not for the reason he thinks. What he demonstrated is that Claude can, with detailed supervision, produce a technically rigorous physics paper. What he actually demonstrated, if you read carefully, is that the supervision is the physics. Claude produced a complete first draft in three days. It looked professional. The equations seemed right. The plots matched expecta…

Same problem as before the LLM era tho.

This is one more argument to the only LLM trained to find "truth"

Re: The threat is comfortable drift toward not understanding what you're doing

#656
> Claude produced a complete first draft in three days. It looked professional. The equations seemed right. The plots matched expectations. Then Schwartz read it, and it was wrong. Claude had been adjusting parameters to make plots match instead of finding actual errors. It faked results. It invented coefficients. It produced verification documents that verified nothing. It asserted results without derivation. It simplified formulas based on patterns from other problems instead of working through the specifics of the problem at hand.

This is solvable with harness engineering.

The model’s first try is never ready for human consumption. There needs to be automation (bespoke, a mix of code and prompt based hooks - which agents can build) to force the agent’s output back through itself to tell it to be more rigorous, search online for proof of its claims, etc etc. and not stop until every claim is verifiable.

No human should see the model’s output until it’s met these (again bespoke but not hand written) guardrails.

What I’m talking about doesn’t exist and really has no analogy yet, so you can think of it as a super advanced form of linting. It’s grounding, but also verification that the grounding links to the material, and refusal to accept the model’s work until it meets the bar.

We are asking models to dream (invent purely from their weights), and are surprised when their dreams, just like ours, have little relationship to reality. The current state of the art is going to look very naive in a few years’ time.

Re: The threat is comfortable drift toward not understanding what you're doing

#657

Earlier quoted context omitted.

There are a lot of people in academia who are great at thinking about complex algorithms but can't write maintainable code if their life depended on it. There are ways to acquire those skills that don't go the junior developer route. Same with debugging and profiling skills But we might see a lot more specialization as a result

They can’t write maintainable code because they don’t have real world experience of getting your hands dirty in a company. The only way to get startup experience is to build a startup or work for one

what about open-source projects? Much as how aspiring authors can learn to write fiction from reading the fiction of others and then imitating that, getting feedback on their work, and iterating, it seems like aspiring programmers could learn by reading/contributing to the open-source projects of others and then writing their own.

Example- Linus Torvalds, never worked for a company, made the original Linux while a grad student, and seems to be doing fine (I'm writing this message on a ThinkPad running Linux Mint). Or Bill Joy with BSD at Berkley, before his time at Sun. Or heck, why not go all the way back to Ken Thompson and Dennis Ritchie building Unix and C?

Re: The threat is comfortable drift toward not understanding what you're doing

#658
post #626

Earlier quoted context omitted.

you might be right but 40-some-odd years is a tiny amount of time.

Sure, relative to the universe, but it's more than those who haven't had idealism beaten out of them by the world.

no, not relative to the universe, relative to recent history. 40 years is not much in the context of history, which is the topic. your life experience is small and unimportant at the world-historical scale.

Re: The threat is comfortable drift toward not understanding what you're doing

#659
post #263

Earlier quoted context omitted.

Marketing is the moat llms haven't been able to overcome. Being able to create a Word clone is easier but the difficulty of selling it is as hard or harder than ever. Show me an llm that can sell my product and find market fit. In reality llms are taking away profitable tools and keeping the revenue themselves.

Right Very Rory Sutherland kind of thought - marketing doesn’t make sense. It is alchemy. If I told you the drink tastes bad, is an off putting color, comes in a small bottle, and is expensive you wouldn’t believe it would work. But Red Bull made billions.

because they're not selling drinks, they're selling stimulants

Re: The threat is comfortable drift toward not understanding what you're doing

#660

Earlier quoted context omitted.

> “I do not know what is the real upside on leaving the current practices in academia and education in favor of ‘let the LLM guide you’” The fact that you think “let the LLM guide you” is the argument I’ve been making tells me everything about how honestly you’ve engaged with it. I’m done here.

And the fact that you only got to shut up after being asked what you are willing to put on the line tells me how devoid of meaning your argument is, no matter what it is .

Being unable or unwilling to grasp how to use an LLM without delegating your entire thinking to it sounds like a you problem. Perhaps one day you can try engaging in some epistemic thought exercises, that might help. In more ways than just this, too.
Post reply on HN