Live data from Hacker News

Mass editing memory in a transformer

memit.baulab.info

51–54 of 54 posts

Re: Mass editing memory in a transformer

#51
post #13
post #9

Earlier quoted context omitted.

That example is wild. But I’m still pretty awed by the fact that we make similar verbal mistakes. The temporal reasoning in these models is getting better than me. As a non-AI model, I notice this every single morning while I have my covefe while heeding the latest on slacker news.

I'm still not convinced they are capable of temporal reasoning. I've asked it temporal questions before but without explicitly mentioning the temporal nature... the answers tend to contradict themselves if they haven't already seen the question before (even when querying general knowledge), until you point out the temporal component, even then it trips up and cannot build upon this reasoning in my tests. I suspect a…

When you explicitly instruct it that its knowledge about current affairs is dated to 2021 and inject documents and provide clear instruction about the documents are correct about current affairs, etc. etc. it works like a charm.

Adds ~3-5 seconds of latency so I have a switch to turn it on and off, for now

Re: Mass editing memory in a transformer

#52
post #17

Earlier quoted context omitted.

I can’t really argue with that, good line of thought. See, my reaction has been, “perhaps our reasoning and actions are pretty much just a biologically-encoded statistical model too, it just doesn’t _feel_ that way because of some other factor.”

When my wife tells me “you should call your mother”, I don’t think her brain assigned probabilities to “you should call your TV”, “you should call your xylophone”, “you should call your airplane”, etc, and then chose a suitable high-probability word (“mother”).

Talk so someone with aphasia or other brain disorder, or talk to a neurologist about it, and you'll see that's not quite the case. There really does seem to be a probability of saying TV when you meant to say "your mother", it's just that the language centers of our brains don't consciously calculate your probabilities and present them to our conscious mind.

Re: Mass editing memory in a transformer

#53
post #13

Earlier quoted context omitted.

I'm still not convinced they are capable of temporal reasoning. I've asked it temporal questions before but without explicitly mentioning the temporal nature... the answers tend to contradict themselves if they haven't already seen the question before (even when querying general knowledge), until you point out the temporal component, even then it trips up and cannot build upon this reasoning in my tests. I suspect a…

Here are some simple tests I ran on ChatGPT (not GPT4): Q: "Who was elected first, President Trump or President Lincoln? Describe your reasoning." A: "President Lincoln was elected first, not President Trump. Abraham Lincoln was elected as the 16th President of the United States in 1860. He served as President from March 1861 until his assassination in April 1865. Donald Trump, on the other hand, was elected as the 4…

It's interesting how, having given a good answer to the last question in the penultimate paragraph, it then goes off the rails with

It's worth noting that the exact dates of death for individuals who died in either war could vary widely, depending on when and where they were serving. However, in general, the Civil War took place before World War II...

Here, "could vary widely" is IMO nonsense given how we're talking a 5-6 year window for either war. Also, the "in general" bit is just weird.

I wonder if this is an artefact of how ChatGPT has been trained for being inoffensive and not opinionated.

Re: Mass editing memory in a transformer

#54
post #50

Earlier quoted context omitted.

When my wife tells me “you should call your mother”, I don’t think her brain assigned probabilities to “you should call your TV”, “you should call your xylophone”, “you should call your airplane”, etc, and then chose a suitable high-probability word (“mother”).

Would the natural analogy of “tokens” be “words”, or something more like, “portion of mouth-movement”?

I more imagine someone speaking another language, when I break down written phonemes for tokenization purposes.

As a native English speaker who grew up around code-switching Spanish speakers, I often heard speech that sounded really fast.

But I was hearing “hay-un-a-al-pa-ca-per-si-gui-en-do-un tren-de-car-ga”

I heard that as individual tokens, while my friends simply heard “¡There is an alpaca chasing that freight train!”

Post reply on HN