Live data from Hacker News

What we know about LLMs

willthompson.name

81–90 of 173 posts

Re: What we know about LLMs

#81
post #78

>Rather than explicitly labeling data, it might be easier for a human to read two or more LLM outputs and encode their preferences through comparison. This reminded me a lot of what economist Murray Rothbard talked about on preferences in his treatise Man, Economy, and State. There is likely to be other insights hidden in these philosophical works on human choices.

Subscribe.

ELI5 - I'd like to know more about this, as I have no experience with this line of thought as of yet.

Re: What we know about LLMs

#82

Earlier quoted context omitted.

There's some argument to be made that a form of reasoning happens in a roundabout way when the AI is told to explain it's reasoning. For example if you tell it "Do " and then open a new context and say "Do , explain your reasoning beforehand." you will often get a more accurate response. Granted, it's not that any "Hmm, let me think about that." Deep Thought reasoning occurs, but simply that predicting what the reaso…

This is where the terminology becomes a bit annoying, but there is a key difference in the kinds of reasoning at work here. When you ask LLMs to provide a reasoning, the actual reasoning performed is linguistic; The LLM has (is) a model about language and performs some (limited) reasoning on that model to get an output. But that is explicitly different from reasoning about the abstract question at hand, thus the answ…

>Thus we can observe that LLMs do not abstractly reason about the question and it's model.

Your conclusion makes no sense. Humans provide increasingly wrong answers as questions get more complex too. Jumping from that to "incapable of abstract reasoning" is silly. You have not "trivially proven" anything at all

>The LLM has (is) a model about language and performs some (limited) reasoning on that model to get an output.

LLMs generalize to non linguistic patterns.

https://general-pattern-machines.github.io/

Re: What we know about LLMs

#83
post #51

I run through a lot of these concepts, specifically RLHF, in my latest coding stream where I finetune LLama 2 if anyone's interested in getting a LLM deep dive https://www.youtube.com/watch?v=TYgtG2Th6fI&t=4002s Long story short, the size of the model and reward mechanisms used in validating off of human annotating/feedback are the main differences between what we can do as independents in OSS vs OpenAI. BigCode's St…

[deleted]

Re: What we know about LLMs

#84

ChatGPT was announced November, 2022 - 8 months ago. Time flies. Question for HN: Where are we in the hype cycle on this? We can run shitty clones slowly on Raspberry Pi's and your phone. The educational implementations demonstrate the basics in under a thousand lines of brisk C. Great. At some point you have to wonder... well, so what? Not one killer app has emerged. I for one am eager to be all hip and open minded…

I don't think there's a "killer app" coming soon, but it'll be a thousand cuts. One awesome thing here, one slightly less awesome but still useful thing over there. Take Copilot. Cool stuff and one of the early products. Doesn't change the game in any fundamental way, but it does have its impact on the work of a substantial fraction of developers.

This is not unlike the computer revolution itself. When the PC came on the scene it was easy - for some types - to imagine The Future and they proclaimed it loudly. They forgot that the rest of the world take their time and regularly take decades to get used to very minor changes in their routine.

Re: What we know about LLMs

#85

How do you do science on LLMs? I would imagine that is super important, given their broad impact on the social fabric. But they're non-deterministic, very expensive to train, and subjective. I understand we have some benchmarks for roughly understanding a model's competence. But is there any work in the area of understanding, through repeatable experiments, why LLMs behave how they do? Do we care?

There's a field called Interpretability (sometimes "Mechanistic Interpretability") which researches how weights inside of a neural network function. From what I can tell, Anthropic has the largest team working on this [0]. OpenAI has a small team inside their SuperAlignment org working on this. Alphabet has at least one team on this (not sure if this is Deepmind or Deepmind-Google or just Google). There are a handful of professors, PhD students, and independent researchers working on this (myself included); also, there are a few small labs working on this.

At least half of this interest overlaps with Effective Altruism's fears that AI could one day cause considerable harm to the human race. Some researchers and labs are funded by EA charities such as Long Term Future Fund and Open Philanthropy.

There is the occasional hackathon on Interpretability [1].

Here's an overview talk about it by one of the most-known researchers in the field [2].

[0] https://transformer-circuits.pub/2021/framework/index.html [1] https://alignmentjam.com/jam/interpretability [2] https://drive.google.com/file/d/1hwjAK3lWnDRBtbk3yLFL2DCK1Dg...

Re: What we know about LLMs

#86

How do you do science on LLMs? I would imagine that is super important, given their broad impact on the social fabric. But they're non-deterministic, very expensive to train, and subjective. I understand we have some benchmarks for roughly understanding a model's competence. But is there any work in the area of understanding, through repeatable experiments, why LLMs behave how they do? Do we care?

There's a field called Interpretability (sometimes "Mechanistic Interpretability") which researches how weights inside of a neural network function. From what I can tell, Anthropic has the largest team working on this [0]. OpenAI has a small team inside their SuperAlignment org working on this. Alphabet has at least one team on this (not sure if this is Deepmind or Deepmind-Google or just Google). There are a handful…

Some people (namely the EAs) care because they don't want AI to kill us.

Another reason is to understand how our models make important decisions. If we one day use models to help make medical diagnoses or loan decisions, we'd like to know why the decision was made to ensure accuracy and/or fairness.

Others care because understanding models could allow us to build better models.

Re: What we know about LLMs

#87
post #80

"Crypto VCs & ”builders” making a hard left into AI" This is a humorous intro graphic caption, but this sentiment appears on here constantly and it's self-destructive. This response might seem a bit over the top to a funny graphic, but I am replying to the general "ha ha AI like crypto amirite?" sentiment that is incredibly boring and worn out. When confronted with challenging new technology that we don't understand,…

> The Venn diagram of people at the forefront of ML/LLM, and its advocates, is almost entirely separate from the web/crypto sphere. There is astonishingly little overlap That statement seems false. Especially since this was a headline I saw in half a dozen online news in my country yesterday. > OpenAI's Sam Altman launches Worldcoin crypto project[0] If anything, even without taking into account the greed stuff, peop…

Me - "Little overlap"

You - "That's false because I saw a thing about this one guy doing some thing"

Okay...

"even without taking into account the greed stuff,"

What "greed stuff"? It is incredibly hard to make money in LLMs/ML. The barriers to entry are colossal, and it is technically extremely difficult. Everyone keeps talking about all the "grifters" (a go to term that usually means the speaker's arguments can be dismissed) yet there are very, very few people making money in AI. The biggest money maker in AI is nvidia and some cloud providers.

You can't just fork BTC or create a Ethereum contract and spin off another shitcoin and make free money. You can't do a rug pull. You can try to create an incredibly difficult niche solution and yield some excitement, but that's like all of technology ever. Comparing it with crypto is dumb.

"people who are drawn to fun tech is likely to be drawn to both LLM and cryptocurrency stuffs"

Loads of people who like to know what's up became acquainted with both. Sure. Understanding tech makes sense. And a lot of us learned crypto, realized it had extraordinarily little real-world utility or benefit, and moved on.

Re: What we know about LLMs

#88

ChatGPT was announced November, 2022 - 8 months ago. Time flies. Question for HN: Where are we in the hype cycle on this? We can run shitty clones slowly on Raspberry Pi's and your phone. The educational implementations demonstrate the basics in under a thousand lines of brisk C. Great. At some point you have to wonder... well, so what? Not one killer app has emerged. I for one am eager to be all hip and open minded…

The biggest impact on my life has been Code Interpreter. Much of my job as a CEO involves analyzing data to make strategic decisions - “which of several options is best based on the evidence?”

Code Interpreter lets me upload data in a multitude of formats and play with it without wasting hours futzing around in Google Sheets or pulling my hair out with Pandas confusion. I know basic statistics concepts and I studied engineering so I know about signals and systems. But putting that knowledge into practice using data analysis tools is time consuming. Code Interpreter automates the time consuming parts and lets me focus on the exploration, delivering insights I never even had access to before.

Re: What we know about LLMs

#89

ChatGPT was announced November, 2022 - 8 months ago. Time flies. Question for HN: Where are we in the hype cycle on this? We can run shitty clones slowly on Raspberry Pi's and your phone. The educational implementations demonstrate the basics in under a thousand lines of brisk C. Great. At some point you have to wonder... well, so what? Not one killer app has emerged. I for one am eager to be all hip and open minded…

[deleted]
Post reply on HN