A global workspace in language models
anthropic.com
A global workspace in language models
1–10 of 218 posts
Re: A global workspace in language models
#2Re: A global workspace in language models
#3 - having a log of the most prominent J-space tokens during your customer support chatbot's interactions with a user, so you can have more introspection into why a particular outcome happened
- being able to detect certain thoughts associated with undesirable behavior (hallucinations, overstepping authority, lying, etc.) and trigger some sort of remediation (e.g. upgrading to a better model, redirecting to a human, forcing tool calls)Re: A global workspace in language models
#4Re: A global workspace in language models
#5It would be really cool if they could expose this information to customers somehow. Imagine: - having a log of the most prominent J-space tokens during your customer support chatbot's interactions with a user, so you can have more introspection into why a particular outcome happened - being able to detect certain thoughts associated with undesirable behavior (hallucinations, overstepping authority, lying, etc.) and t…
Re: A global workspace in language models
#6Re: A global workspace in language models
#7Re: A global workspace in language models
#8More interesting was the independent commentary paper they linked near the bottom: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be24...
Neel Nanda (of Google Deepmind - his part begins on page 33) discusses his opinions on the paper, and the small-scale replication he performed on an open-weight model.
Re: A global workspace in language models
#9I’m confused where in the weights the jspace is.
I too have confusion.
Re: A global workspace in language models
#10Without using the term, they are using an information geometric approach.