Live data from Hacker News

Where the goblins came from

openai.com

561–570 of 699 posts

Re: Where the goblins came from

#561

Earlier quoted context omitted.

What if you substituted "steel" with "asbestos" in your argument.

Steel has almost always (as in 99.99...% of the time) delivered to our expectations based on our understanding of it. The cases where we built something out of steel and it failed are _massively_ outnumbered by the instances where we used it where/when suitable. If we built something in steel and it failed/someone died we stopped doing that pretty soon after.

This is partly due to having a safety factor: i.e. using twice as much steel as you think you need.

Understanding means knowing the limits of your own understanding, and building in safeguards.

Re: Where the goblins came from

#563

Earlier quoted context omitted.

We do understand tho, it is exactly what they were made for. If you train it on a dataset of Othello games, or a dataset including these, you are basically creating a map of all possible moves and states that have ever happened, odds of transitions between them, effective and un-effective transitions. By querying it, you basically start navigating the map from a spot, and it just follows the semi-randomly sampled hig…

@hypendev I am not trying to start a flame war, but let me take a very simple example. As another one put it, we know how to build deep-learning machines. No question about that. My statement is that we don't understand clearly why they output the observed results. Let's imagine that you have a model that can detect cats on an image, with 95% accuracy. If you understood how the model worked, I could give you an image…

Sorry if it sounded like that, not trying to have a flame war, just trying to understand which part we don't _understand_, as it seems silly to me.

Yeah, we cannot predict with 100% accuracy the results of a model, not mentally, as to be able to do that we should be able to do the same math in our head and that's just ultra rare next level intelligence. And we can make a reliable predictor, but making a reliable prediction model of a models results would be the same model in the end.

So the closest that we can get to "understanding" it fully, is learning how it works, and developing intuition around it. And I think we pretty much have that, at least among the people in the field. Those who worked on training it especially have some intuitive understanding of what is going on, otherwise they would not know where to "test and hack".

It's math all the way down, but I feel like the angle some people in early days used about "magic emergent properties" or "signs of consciousness" ended up making it seem more mystical than it is.

Re: Where the goblins came from

#564
post #511

Earlier quoted context omitted.

Asbestos, lead paint, cigarettes, heroin(perscribed generously for basically whatever the doc felt like), "Radithor" (patent medicine containing radium-226 and 228, marketed as a "perpetual sunshine" energy tonic and cure for over 150 diseases), bloodletting, mercury treatments for syphilis, tobacco smoke enemas (yep that was a real thing), milk-based blood transfusions. Didn't understand those either and used the fu…

This is why I believe we should only listen to amateur opinions on everything, experts simply lack historical credibility. For example I've recently purchased a healing crystal (half off) for only $5000 dollars! It cleared up the imbalanced energies my street guru told me about right away. I would never have been made aware about the consequences of imbalanced energies in the first place if I had asked an expert inst…

Ironically the street guru hucksters might have a better track record than the dangerous products mentioned above.

Less charitably, it's a mistake to imply that simply being a bigger corporation makes you go from street guru to "expert". Bigger company trying to make money off of you at any risk to you is just the same bucket at a different scale. In this context the other side is probably "expert consumer advocate" since that fits the idea above of these dangerous products advertised as cure alls.

Re: Where the goblins came from

#565
post #388

Earlier quoted context omitted.

What is the ad hoc fallacy? From googling I didn’t find any convincing definitions (definitions that demonstrate that it is a logical fallacy).

https://finmasters.com/ad-hoc-fallacy/ > Ad hoc fallacy is a fallacious rhetorical strategy in which a person presents a new explanation – that is unjustified or simply unreasonable – of why their original belief or hypothesis is correct after evidence that contradicts the previous explanation has emerged. https://cerebralfaith.net/logical-fallacy-series-part-13-ad-... > An argument is ad hoc if its only given in an…

Thanks. I’m by default disposition suspicious of fallacies that are not logical fallacies. And I’m not convinced that this is a solid fallacy.

> > Ad hoc fallacy is a fallacious rhetorical strategy in which a person presents a new explanation – that is unjustified or simply unreasonable – of why their original belief or hypothesis is correct after evidence that contradicts the previous explanation has emerged.

That someone jumps to a new thing once something is refuted just looks like rhetoric to me. Not fallacious rhetoric.

> > that is unjustified or simply unreasonable

So it needs to be these things as well. But why are not these points the problematic part?

It seems impractical to usefully label an argument in this way since you either call any new argument (that is also unjustified or unreasonable) a fallacy, or divine that the argumenter is intending to be dishonest.

> > https://cerebralfaith.net/logical-fallacy-series-part-13-ad-...

This was one of the results of my googling.

> > One example of this logical fallacy that immediately comes to mind is the multiverse hypothesis. When Atheists are presented with The Fine Tuning Argument For God’s Existence, many of them will respond to it by giving the multiverse hypothesis. [...] Given an infinite number of universes, there were an infinite number of chances, and therefore any improbable event is guaranteed to actualize somewhere at some point.

So why is this a problem?

> > There are many problems with this theory, not the least of which is that there’s no evidence that a multiverse even exists! There’s no evidence that an infinite number of universes exist! No one knows if there’s even one other universe, much less an infinite number of them! You can’t detect these other universes in any way! You can’t see them, you can’t hear them, you can’t smell them, you can’t touch them, you can’t taste them, you can’t detect them with sonar or any other way. They are completely and utterly unknowable to us. I find it ironic that atheists, who are infamous for mocking religious people for their “blind faith”, themselves are guilty of having blind faith! Namely, blind faith in an infinite number of universes!

> > This explanation is one example of the ad hoc fallacy. The multiverse hypothesis is propagated for no other reason than to keep atheism from being falsified. The theory is ad hoc because the only reason to embrace it is to keep atheism from being falsified! For if this universe is the only one there is, then there’s no other rational explanation for why the laws of physics fell into the life permitting range other than that they were designed by an intelligent Creator!

Allow me to restate. It is a fallacy because there is no evidence of the theory. And further that (perhaps following from the no-evidence part in their mind) there is no reason to hold this theory other than from arguing against theists.

Yeah there is no reason to hold a theory from physics other than wanting to prove theists wrong.

Why? Because my argument for theism is so water-proof that this would be the only hope that they would have of refuting it.

I find that very unconvincing. (The argument for this fallacy. I can take or leave the God/unGod part.)

Re: Where the goblins came from

#566

Earlier quoted context omitted.

The entire industrial revolution was steel replacing human workers. And that is still the backbone of the world today. We are still living the industrial revolution. Just like the invention of fire happened ages ago, but is still a crucial part of life today.

No, it was actually engines. The mechanism behind engines were fully understood, any experiments with engines were reproducible and measurable. You could get an engine and create schematics by reverse engireening it. LLMs, useful as they may be, are not that.

The mechanics of engines was understood at the beginning of the Industrial Revolution, and they were fully reproducible: all of which is true of LLMs today. An LLM is a bunch of floating point numbers and simple operations on them, all of which are fully known.

But the way that steam engines emergently transformed heat into work was not understood at the beginning of the Industrial Revolution. Figuring this out led to an entire new branch of physics, thermodynamics. Figuring out how big next-token predictors give rise to interesting systems is likely to lead to similarly new ideas.

Re: Where the goblins came from

#567

Earlier quoted context omitted.

Modern kv caches can contain up to 1 million tokens (~3000 pages of text). It's not that short, it's like 48 straight hours of reading.

Yes and no, it's not just text, it's images, video, etc, and it's not just the pages of content, it's also all the "thinking" as well. Plus the models tend to work better earlier on in the context. I regularly get close to filling up context windows and have to compact the context. I can do this several times in one human session of me working on a problem, which you could argue is roughly my own context window. My p…

The KV cache isn't memory, it's the extent of the process saved so the inference can start where the last generated output is concatenated with the next input. It's entirely about saving compute and has nothing to do with memory.

This really confuses how stupid LLMs are: they're just text logs as output and text logs as input; hence the goblins are just tokens that seem to problematically be more probable in the output.

But the KV cache is a thing made to keep a session from having to run through the entire inference. The only thing you can call "memory" is there's no random perturbations in the KV cache while there may be in re=running chat which ends up being non-deterministic. You can think of it as a deterministic seed to prevent a random conversation from it's normal non-deterministic output

Re: Where the goblins came from

#569

Earlier quoted context omitted.

What does LLM need to do for you to consider it "smart"? To me they seem to be pretty damn smart, to put it mildly. They sometimes do stupid things - but so do smart people!

Not OP, but I think the argument here would be not that LLMs "are not smart" but that smart is just the wrong category of thing to describe an LLM as. A calculator can do very complex sums very quickly, but we don't tend to call it "smart" because we don't think it's operating intelligently to some internal model of the world. I think the "LLMs are AGI" crowd would say that LLMs are , but it's perfectly consistent to…

I would analogize LLMs to physics simulations in software. Game engines, for example, simulate physics enough to provide a good enough semblance of real-world physics for suspension of disbelief but we would never mistake it for real world physics. Complicated enough simulations, e.g. for weather forecasting, nuclear weapons, or QCD, can provide insights and prove physics theories, but again, experts would never mistake it for real world physics and would be able to explain where the simulation breaks down when trying to predict real world behavior.

Now we have these LLMs that provide some simulation of reasoning merely through prediction of token patterns and that is indeed unexpected and astonishing. However, the AI promoters want to suggest that this simulation of reasoning is human-level reasoning or evolving toward human-level reasoning and this is the same as mistaking game engine physics for real physics. The failure cases (e.g. the walk vs drive to a car wash next door question or the generating an image of a full glass of wine issue), even if patched away, are enough to reveal the token predictor underneath.

Re: Where the goblins came from

#570
post #398

Earlier quoted context omitted.

The article you are responding to showed that a strange LLM behaviour was caused by a training signal that was explicitly designed to produce that type of behaviour. They were able to isolate it, clearly demonstrate what happened, and roll out a mitigation using a mechanism they engineered for exactly this type of thing (the developer prompt). That doesn’t sound like sorcery to me. If anything I’m surprised you can s…

The article I am responding to (which I've read) shows that these LLMs come with all sorts of hacks (= context bits) to make it behave more like this or more like that. There is probably a whole testing workflow at AI companies to tweak each new model until it "looks" acceptable. But they still don't understand what they are doing. This is purely empirical.

> "There is probably a whole testing workflow at AI companies to tweak each new model until it "looks" acceptable."

Isn't that what the RLHF phase does ( https://www.paloaltonetworks.com/cyberpedia/what-is-rlhf )?

Post reply on HN