Live data from Hacker News

GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

arxiv.org

151–160 of 235 posts

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#151
post #99

Earlier quoted context omitted.

> For one unfamiliar with the domain, this can take hours or a full day Yes, I did say expert. > But several times slower still than GPT. Almost all the output needs to be re-read, modified and integrated in the project. This often takes longer than just typing it out, even from docs -- because you're still forced to think through the solution -- which is most of the time. Typing is quick

> Almost all the output needs to be re-read, modified and integrated in the project. This often takes longer than just typing it out You'd be very suprised. Try measuring it.

People don't have to love using it you know? If the guy doesn't like it, why force it down his neck?

I use plenty of code completion tools already, so I feel similar to them. Sometimes it saves me time, sometimes it doesn't.

For me personally, language servers have been the best thing, you can explore libraries, auto-complete nicely, without as the other poster said, worrying about verifying the correctness.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#152

Total trash cloaked in a complicated story. What they actually did is ask 5 random people to rate what thought a language model could do to help different professions. These 5 random people don't know anything about the professions they're rating, just what anyone off the street knows, and they know as much about GPT as anyone who has briefly played with it. The title should have been "We asked 5 friends to see what…

I can't find how many people labeled the DWA task descriptions, where did you got that number?

The article seems to describing the labeling here:

> Human Ratings: We obtained human annotations by applying the rubric to each ONET Detailed Worker Activity (DWA) and a subset of all ONET tasks and then aggregated those DWA and task scores at the task and occupation levels. To ensure the quality of these annotations, the authors personally labeled a large sample of tasks and DWAs and enlisted experienced human annotators who have extensively reviewed GPT outputs as part of OpenAI’s alignment work (Ouyang et al., 2022).

I understand the authors, four, did the initial labeling and then asked an undefined set of people to the rest of the labeling.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#153

Earlier quoted context omitted.

A human brain is nothing more than long-form autocomplete. The specifics of approach are irrelevant.

Exactly. All the people saying LLM's aren't important because it is just auto-complete, really don't connect the dots that humans are also just auto-complete.

What on earth are you talking about?

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#154

Earlier quoted context omitted.

"I keep reading these opinions that LLMs are just doing some advanced form of copy paste. Actually, we don't know what they are doing. Are they actually doing some form of modelling and abstraction? Seems likely to me." This is exactly the problem with AI. For business or government, the answer is as important as the methodology employed. A black box does not work for the majority of use cases. Until it can show its…

Humans are also "black boxes" we don't fully understand. It's worked pretty well.

This is a shallow counterargument that entirely misses the point.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#155

Earlier quoted context omitted.

"This is a black swan event" By definition, black swan events are unpredictable. This example isn't that.

The black swan event includes the unexpected emergent capabilities of LLMs, which can pass professional examinations and exhibit many attributes of general intelligence, combined with the broad availability to market and massive pressure on big industry players to adapt and innovate or die. If things keep going this way then pretty soon the black swans are going to outnumber the white ones.

Single stories of LLM success are just as problematic as single stories of LLM failures. As always. The fact is that we have a dangerous tool at hand that absolutely REQUIRES skepticism if you are trying to get *facts* out of it.

I agree that LLMs are extremely likely to impact many areas of work, particularly bullshit work. But as it stands you absolutely cannot use them as fact machines, the results can be catastrophic.

What it does well, among others:

- Scaffolding text, breaking writers block etc.

- Compose basic texts from minimal input, for example for bullshit tasks -> I generated an internal "vision statement" during a Miro workshop for my team by inputting a bunch of bullet points gathered from the team members brain storming. It created a concise, fluid text that everybody liked. It's now the vision statement.

- Point you in good directions, give you ideas

What it does NOT well, among others:

- provide factual responses. all responses MUST be scrutinized because they are likely containing false information. This is very dangerous for society ("Can I take this medicine with this other medicine?")

- Compose creative texts that are coherent and novel. ChatGPT texts can be quite fun but they rarely make sense beyond very superficial screening and convey no deeper message.

However, ChatGPT-like tools are used with a lot of naivety and often blind acceptance instead of using them as tools to aid your work.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#156
post #36

Earlier quoted context omitted.

I don't want to have a calculator (or say bookkeeping software) that gives correct results most of the time but not always, and then hear from the developers that it will get better with each iteration. I need a calculator that is correct 100% of time, not even 99.999%, because otherwise I can't rely on it at all. In other words, the utility of a calculator that is correct only 99% of time is zero, since you can't ev…

Confining an LLM to the very narrow domain of "calculators" is a mistake, I think. You wouldn't say "a programmer that is 99% correct is worthless, I need 100%". I'm pushing it, but for a more fair comparison I'd say measure it against a programmer. How often are we wrong? 75% of the time? :) being generous here. It's the tools that make us productive. I don't know about you specifically, but I don't think you'll be…

He is comparing it to a calculator and CHatGPT doesn’t measure up in some aspects.

The good things about reliable tools is you can offload the cognitive burden onto them and know they won’t screw you over.

Almost every single post here about using ChatGPT mentions checking through its output. People don’t check though the output of their calculators.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#157
post #140

Earlier quoted context omitted.

It is telling that ChatGPT by training on online discussions hasn’t learned to say “I don’t know”.

I guess the "I don't know"'s are usually silent.

Lmao. Now I’m imagining any questions posted anywhere like Ask Hacker News, Ask Reddit, or Quora being filled with everyone that doesn’t know the answer replying “I don’t know.”

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#158

I asked Chat GPT which antacid medications are contraindicated for some medication I'm on. Easily found through NICE. It made up a severe risk of death taking a very common medicine combo. It was super convincing, even giving information on how long to avoid taking them together. It was pure bullshit. I think as much as hyping the benefits we need to hype the flaws and dangers. If the public at large learn to trust t…

Not sure if you tried GPT-4 but my experience with 4 is quite different. It has been quite bullshit free, though not completely. For example I asked it to contrast oral and injectable semaglutide formulations. It did a bang up job. One thing I always do is ask it for evidence. And then I look the references up. Sometimes the references don’t say exactly what it said they will. I come back and have a discussion with i…

> One thing I always do is ask it for evidence. And then I look the references up. Sometimes the references don’t say exactly what it said they will.

Can you give an example? I'm surprised there's room in there for actual resolvable URLs or paper references.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#159
post #105
post #94

Earlier quoted context omitted.

You didn't use version 4 did you? Because if feel this is pretty much already out of date. 4 does this _a lot_ less! And 5, 6 or whatever will probably be better, so i don't even really get the point here.

I find there are two kinds of people in the world. First are saying that LLMs are bullshitting. The second are bullshitting about whatever LLMs are saying.

I find there are more than two kinds of people. Among them some are skeptical and some are not. The latter are good at exploring, even if blind alleys while the former keep them in check. We’re lucky we have some variety otherwise we’d risk go full force into dead ends or we’d ignore fruitful possibilities.

Re: GPTs Are GPTs: An Early Look at the Labor Market Impact Potential of LLMs

#160

Earlier quoted context omitted.

Exactly. All the people saying LLM's aren't important because it is just auto-complete, really don't connect the dots that humans are also just auto-complete.

What on earth are you talking about?

Our brain excels at pattern matching. We attribute a lot of human abilities to “intelligence”, but the line between that and pattern-matching (or “autocomplete” ad it’s being referred as here) is being challenged by these latest LLMs.

If it can look at a picture, explain what’s in it, and hypothesize about physics inside the picture’s environment, is it “just pattern matching”?

Post reply on HN