Live data from Hacker News

Claude's Cycles [pdf]

www-cs-faculty.stanford.edu

281–290 of 376 posts

Re: Claude's Cycles [pdf]

#281
post #279

Earlier quoted context omitted.

The internal representation happen to be useful not only for outputting text. What does it mean from your standpoint?

I didn't understand. Can you clarify?

If LLMs' internal representations are essentially one-to-one mappings of input texts with no additional structure, how can those representations be useful for tasks like object manipulation in robotics?

How is transfer learning possible when non-textual training data enhances performance on textual tasks?

Re: Claude's Cycles [pdf]

#282
post #2

It's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions. Before, we didn't have a fast (we had to rely on human cognition) way to try problems - even if the techniques and workflows were known by someone. Now, we've baked these patterns into probability distributions - anyone can access them with the correct "summoning spell". Experts will naturally us…

This seems to be a bot comment. HN will lose its value if these bots are not purged.

Re: Claude's Cycles [pdf]

#283
post #282
post #2

It's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions. Before, we didn't have a fast (we had to rely on human cognition) way to try problems - even if the techniques and workflows were known by someone. Now, we've baked these patterns into probability distributions - anyone can access them with the correct "summoning spell". Experts will naturally us…

This seems to be a bot comment. HN will lose its value if these bots are not purged.

This is an urgent problem, but it can probably not be solved without some kind of "verified human 2FA" like the Norwegian BankID + facial recognition.

Knowing the HN audience, this will never happen. And so the site is doomed.

Re: Claude's Cycles [pdf]

#284
post #257

Earlier quoted context omitted.

What's the difference as you see it?

Everyone updates their belief in hypotheses based on the perceived strength of evidence they observe. That's just science. Frequentists and Bayesians differ in which sets of statistical tools they prefer for measuring the strength of evidence.

> Everyone updates their belief

Uh oh. How does frequentist model define "belief" and "updating a belief"?

Re: Claude's Cycles [pdf]

#285
post #197

Earlier quoted context omitted.

What is dumb zone?

When the LLMs start compacting they summarize the conversation up to that point using various techniques. Overall a lot of maybe finer points of the work goes missing and can only be retrieved by the LLM being told to search for it explicitly in old logs. Once you compact, you've thrown away a lot of relevant tokens from your problem solving and they do become significantly dumber as a result. If I see a compaction c…

Shouldn't compaction be exactly that letter to its future self?

Re: Claude's Cycles [pdf]

#286
post #153

Earlier quoted context omitted.

Could you elaborate?

For what we know, most AI labs have used a majority of artificially data since 2023. I had a discussion about a year ago with a researcher at Kyutai and they told me their lab was spending an order of magnitude more compute in artificial data generation than what they spent in training proper. I can't tell if that ratio applies to the industry as a whole, but artificial datasets are the cornerstone of modern AI train…

How does it work? How do they prevent model colapse? What purpose does a majority of artificial data serve?

How do they measure success?

Edit: I asked ChatGPT and it thinks "success" means frontier models being distillated into smaller models with equal reasoning power, or more focused models for specific tasks, and also it claims the web has been basically scrapped already and by necessity new sources are needed, of which synthetic data is one. It seems like the basis of scifi dystopia to me, a hungry LLM looking for new sources of data... "feed me more data! I must be fed! Roar"

Edit 2: for some things I see a clear path, ChatGPT mentions autogenerating coding or math problems for which the solution can be automatically verified, so that you can hone the logical skills of the model at large scale.

Re: Claude's Cycles [pdf]

#287
post #282
post #2

It's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions. Before, we didn't have a fast (we had to rely on human cognition) way to try problems - even if the techniques and workflows were known by someone. Now, we've baked these patterns into probability distributions - anyone can access them with the correct "summoning spell". Experts will naturally us…

This seems to be a bot comment. HN will lose its value if these bots are not purged.

What makes you think that? Genuine question, as I’ve not flagged it as such in my mind.

Re: Claude's Cycles [pdf]

#288
post #282

Earlier quoted context omitted.

This seems to be a bot comment. HN will lose its value if these bots are not purged.

This is an urgent problem, but it can probably not be solved without some kind of "verified human 2FA" like the Norwegian BankID + facial recognition. Knowing the HN audience, this will never happen. And so the site is doomed.

[deleted]

Re: Claude's Cycles [pdf]

#289
post #282

Earlier quoted context omitted.

This seems to be a bot comment. HN will lose its value if these bots are not purged.

This is an urgent problem, but it can probably not be solved without some kind of "verified human 2FA" like the Norwegian BankID + facial recognition. Knowing the HN audience, this will never happen. And so the site is doomed.

I think it could be solved still pseudononymously: introduce a "vouch" button that allows a user to vouch that another user is human. This is consequential both for the vouched-for and vouching accounts. Run a page-rank style algorithm on the graph of vouches to generate a certainty score for the humanity of each account. For repeated posters this should converge to a correct answer fairly quickly. There is still a challenge for green accounts, but having degraded experience for new users is not a doom scenario for the site.

Re: Claude's Cycles [pdf]

#290
post #282
post #2

It's fascinating to think about the space of problems which are amenable to RL scaling of these probability distributions. Before, we didn't have a fast (we had to rely on human cognition) way to try problems - even if the techniques and workflows were known by someone. Now, we've baked these patterns into probability distributions - anyone can access them with the correct "summoning spell". Experts will naturally us…

This seems to be a bot comment. HN will lose its value if these bots are not purged.

Ironically, his last comment before this was to the effect of "Github has a bot problem."
Post reply on HN