I’m confused where in the weights the jspace is.
Their method is used to identify which tokens can appears in which layers of the model.
11–20 of 218 posts
I’m confused where in the weights the jspace is.
Their method is used to identify which tokens can appears in which layers of the model.
It would be really cool if they could expose this information to customers somehow. Imagine: - having a log of the most prominent J-space tokens during your customer support chatbot's interactions with a user, so you can have more introspection into why a particular outcome happened - being able to detect certain thoughts associated with undesirable behavior (hallucinations, overstepping authority, lying, etc.) and t…
Without using the term, they are using an information geometric approach.
But J-Space is much catchier. This is not a scientific paper, it's a promotional essay.
(Nb: not an expert / in the labs, just opining)
Is the model really "thinking" about that stuff or is just mimicking human "manners"? And if so, where the thinking is happening if it is not in the literal chain of *thought*?
I'm not sure J-Space is the answer to that question, but very interesting nevertheless.
https://distrowatch.com/weekly.php?issue=20260706#freebsd
We should really stop giving these liar models any further credibility.
I always wondered what the model meant when it writes "I'm now considering the architecture of the service" but outputs nothing of the sorts in its CoT. Is the model really "thinking" about that stuff or is just mimicking human "manners"? And if so, where the thinking is happening if it is not in the literal chain of *thought*? I'm not sure J-Space is the answer to that question, but very interesting nevertheless.
There are various justifications on this, but it's mostly to make distillation and fine tuning off their model outputs a bit harder for their competitors
I always wondered what the model meant when it writes "I'm now considering the architecture of the service" but outputs nothing of the sorts in its CoT. Is the model really "thinking" about that stuff or is just mimicking human "manners"? And if so, where the thinking is happening if it is not in the literal chain of *thought*? I'm not sure J-Space is the answer to that question, but very interesting nevertheless.
What you see here is a summary of thinking tokens written by some other smaller model (e.g. old sonnet). The actual thinking sometimes (rarely) leaks and is not easy to parse.