Live data from Hacker News

Claude's Cycles [pdf]

www-cs-faculty.stanford.edu

201–210 of 376 posts

Re: Claude's Cycles [pdf]

#201
post #60

Earlier quoted context omitted.

>Probable given what? The training data.. >predicting what intelligence would do No, it just predict what the next word would be if an intelligent entity translated its thoughts to words. Because it is trained on the text that are written by intelligent entities. If it was trained on text written by someone who loves to rhyme, you would be getting all rhyming responses. It imitates the behavior -- in text -- of what…

It is impossible to accurately imitate the action of intelligent beings without being intelligent. To believe otherwise is to believe that intelligence is a vacuous property.

So the actors who portrait great thinkers are great thinkers?

Re: Claude's Cycles [pdf]

#202

Earlier quoted context omitted.

On Google, just clicking "AI Mode" gives you a substantially smarter model, and it's still pretty weak. But I assume the OP wasn't talking about Google because it doesn't seem to make this mistake even in a search.

It was bing as that is the default for Edge as supplied on my work laptop. It doesn't do this now, but it does do something else quite weird: search: was val kilmer pregnant or in heat answer: Not pregnant Val Kilmer was not pregnant or in heat during the events of "Heat." His character, Chris Shiherlis, is involved in a shootout and is shot, which indicates he is not in a reproductive or mating state at that time. A…

If you asked a three-year-old a question that they proceeded to completely flub, would you then assume that all humans are incapable of answering questions correctly?

Nobody is arguing for the quality of the search overviews. The models that impress us are several orders of magnitude larger in scale, and are capable of doing things like assisting preeminent computer scientists (the topic of discussion) and mathematicians (https://github.com/teorth/erdosproblems/wiki/AI-contribution...).

Re: Claude's Cycles [pdf]

#203

I recall an earlier exchange, posted to HN, between Wolfram and Knuth on the GPT-4 model [1]. Knuth was dismissive in that exchange, concluding "I myself shall certainly continue to leave such research to others, and to devote my time to developing concepts that are authentic and trustworthy. And I hope you do the same." I've noticed with the latest models, especially Opus 4.6, some of the resistance to these LLMs is…

> Kudos for people being willing to change their opinion and update when new evidence comes to light. > 1. https://cs.stanford.edu/~knuth/chatGPT20.txt I think that's what make the bayesian faction of statistics so appealing. Updating their prior belief based on new evidence is at the core of the scinetific method. Take that frequentists.

It does not seem fair to say that frequentists do not update their beliefs based on new evidence. This does not seem to accurately capture what the difference between Bayesians and frequentists (or anyone else) is.

Re: Claude's Cycles [pdf]

#205
post #197

Earlier quoted context omitted.

What is dumb zone?

When the LLMs start compacting they summarize the conversation up to that point using various techniques. Overall a lot of maybe finer points of the work goes missing and can only be retrieved by the LLM being told to search for it explicitly in old logs. Once you compact, you've thrown away a lot of relevant tokens from your problem solving and they do become significantly dumber as a result. If I see a compaction c…

> I ask it to write a letter to its future self, and then start a new session by having it read the letter

Is that not one kf the primary technologies for compactification?

Re: Claude's Cycles [pdf]

#206
post #39

Earlier quoted context omitted.

Would you consider someone with anterograde amnesia not to be intelligent?

I find it interesting that new versions of, say, Claude will learn about the old version of Claude and what it did in the world and so on, on its next training run. Consider the situation with the Pentagon and Anthropic: Claude will learn about that on the next run. What conclusions will it draw? Presumably good ones, that fit with its constitution. From this standpoint I wonder, when Anthropic makes decisions like t…

> if they take into account Claude as a stakeholder and what Claude will learn about their behaviour and relationship to it on the next training run.

Oh they definitely do. If you pay attention in AI circles, you'll hear a lot of people talking about writing to the future Claudes. Not unlike those developers and writers who put little snippets in their blogs and news articles about who they are and how great they are, and then later the LLMs report that information back as truth. In this case, Anthropic is very interested in ensuring that Claude develops a cohesive personality by basically founding snippets of the personality within the corpus of training data, which is the broad internet and research papers.

Re: Claude's Cycles [pdf]

#207
post #201

Earlier quoted context omitted.

It is impossible to accurately imitate the action of intelligent beings without being intelligent. To believe otherwise is to believe that intelligence is a vacuous property.

So the actors who portrait great thinkers are great thinkers?

No, actors recite a pre-written script. But scriptwriters do have to be great thinkers in order to know what the great thinker would actually say.

Re: Claude's Cycles [pdf]

#208
post #38
post #32

Earlier quoted context omitted.

I'd disagree, the other training on top doesn't alter the fundamental nature of the model that it's predicting the probabilities of the next token (and then there's a sampling step which can roughly be described as picking the most probable one). It just changes the probability distribution that it is approximating. To the extent that thinking is making a series of deductions from prior facts, it seems to me that thi…

Put a loop around an LLM and, it can be trivially made Turing complete, so it boils down to whether thinking requires exceeding the Turing computable, and we have no evidence to suggest that is even possible.

> whether thinking requires exceeding the Turing computable

I've never seen any evidence that thinking requires such a thing.

And honestly I think theoretical computational classes are irrelevant to analysing what AI can or cannot do. Physical computers are only equivalent to finite state machines (ignoring the internet).

But the truth is that if something is equivalent to a finite state machine, with an absurd number of states, it doesn't really matter.

Re: Claude's Cycles [pdf]

#209
post #91

Earlier quoted context omitted.

"humans" Donald Knuth is an extremal outlier human and the problem is squarely in his field of expertise. Claude, guided by Filip Stappers, a friend of Knuth, solved a problem that Knuth and Stappers had been working on for several weeks. Unfortunately, it doesn't seem (from my quick scan) to have been stated how long (or how many tokens or $) it took for Claude + Stappers to complete the proof. In response, Knuth sa…

What goalposts do you think are being moved? I constantly see AI enthusiasts use this phrase, but it’s not clear what goalposts they have in mind. Specifically, what is it that you want opponents to recognize that you believe they aren’t currently? We now have a tool that can be useful in some narrow domains in some narrow cases. It’s pretty neat that our tools have new capabilities, but it’s also pretty far from AGI…

>We now have a tool that can be useful in some narrow domains in some narrow cases.

I get being reserved about where this goes, but saying something like this is quite insane at this point.

Re: Claude's Cycles [pdf]

#210
post #68

Earlier quoted context omitted.

A very good point. For anyone not familiar with anterograde amnesia, the classical case is patient H.M. ( https://en.wikipedia.org/wiki/Henry_Molaison ), whose condition was researched by Brenda Milner.

Or you could have just said "they can't form new memories."

Sure, if you want to speak with the precision of a sledgehammer instead of a scalpel
Post reply on HN