Live data from Hacker News

A global workspace in language models

anthropic.com

1–10 of 218 posts

Re: A global workspace in language models

#3
It would be really cool if they could expose this information to customers somehow. Imagine:

   - having a log of the most prominent J-space tokens during your customer support chatbot's interactions with a user, so you can have more introspection into why a particular outcome happened
   - being able to detect certain thoughts associated with undesirable behavior (hallucinations, overstepping authority, lying, etc.) and trigger some sort of remediation (e.g. upgrading to a better model, redirecting to a human, forcing tool calls)

Re: A global workspace in language models

#5

It would be really cool if they could expose this information to customers somehow. Imagine: - having a log of the most prominent J-space tokens during your customer support chatbot's interactions with a user, so you can have more introspection into why a particular outcome happened - being able to detect certain thoughts associated with undesirable behavior (hallucinations, overstepping authority, lying, etc.) and t…

Presumably the rationale for the decision to abridge the thinking traces will ensure that they don’t; if this is real (and there’s no good reason to trust that it is yet) then it is the secret sauce.

Re: A global workspace in language models

#8
As someone who is not an AI researcher, the paper itself is way over my head.

More interesting was the independent commentary paper they linked near the bottom: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be24...

Neel Nanda (of Google Deepmind - his part begins on page 33) discusses his opinions on the paper, and the small-scale replication he performed on an open-weight model.

Post reply on HN