Live data from Hacker News

Anyone else witnessing a panic inside NLP orgs of big tech companies?

old.reddit.com

21–30 of 527 posts

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#21

I tried translating something from English to German (my native language) yesterday with ChatGPT4 and compared it to Microsoft Translate, Google Translate and DeepL. My ranking: 1. ChatGPT4 - flawless translation. I was blown away 2. DeepL - very close, but one mistake 3. Google Translate - good translation, some mistakes 4. Microsoft Translate - bad translation, many mistakes I can understand the panic.

Does it translate hate speech too?

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#22
post #18

During my master's degree in data science, we had several companies visit our faculty to recruit students. Not a single one was a specialized NLP company, but many of them had NLP projects going on. Most of those projects were the usual "solution looking for a problem to solve". Even those projects that might have had _some_ utility, would have been way more effective to buy/license a product than to develop an in-ho…

> Honestly, I can't wait for GPT and other productivity tools to wrech havock upon the tech labour market. Some people in tech really need to be taken down a notch or two.

That's an odd reason to want this.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#23
post #16
post #9

Earlier quoted context omitted.

NLP is nowhere near being solved.

Depending on definition, it is solved.

You're using the wrong definition, then. /s

Where is some evidence that NLP is 'solved'? What does it even mean? OpenAI itself acknowledges the fundamental limitations of ChatGPT and the method of training it, but apparently everybody is happily sweeping them under the rug:

"ChatGPT sometimes writes plausible-sounding but incorrect or nonsensical answers. Fixing this issue is challenging, as: (1) during RL training, there’s currently no source of truth; (2) training the model to be more cautious causes it to decline questions that it can answer correctly; and (3) supervised training misleads the model because the ideal answer depends on what the model knows, rather than what the human demonstrator knows." (from https://openai.com/blog/chatgpt )

Certainly ChatGPT/GPT-4 are impressive accomplishments, and it doesn't mean they won't be useful, but we were pretty sure in the past that we had "solved" AI or that we were just about to crack it, just give it a few years... except there's always a new rabbit hole to fall into waiting for you.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#24
They may panic, but they shouldn't. They can quickly pivot. GPT programs can be used off the shelf, but they can also use custom training. Every large org has a huge internal set of documents, plus a large external set of documents relevant to its work (research articles, media articles, domain relevant rules and regulations). They can train a GPT bot to their particular codebase. And that is now. Soon (I'd give it at most one year), we'll be able to train GPT bots to videos.

All this training does not happen by itself.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#25
post #16
post #9

Earlier quoted context omitted.

NLP is nowhere near being solved.

Depending on definition, it is solved.

Even if that were true, LLMs don't give any kind of "handles" on the semantics. You just get what you get and have to hope it is tuned for your domain. This is 100% fine for generic consumer-facing services where the training data is representative, but for specialized and jargon-filled domains where there has to be a very opinionated interpretation of words, classical NLU is really the only ethical choice IMHO.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#26
When I was studying Computational Linguistics I kept running into the unspoken question: given that Google Translate already exists, what is even the point of all of this? We were learning all these ideas about how to model natural language and tag parts of speech using linguistic theory so we could eventually discover that utopian solution that would let us feed two language models into a machine to make it perfectly translate a sentence from one language into another. And here was Google Translate being "good enough" for 80% of all use cases using a "dumb" statistic model that didn't even have a coherent concept of what a language is.

It's been close to two decades and I still wonder if that "pure" approach has any chance of ever turning into something useful. Except now it's not just language but "AI" in general: ChatGPT is not an AGI, it's a model fed with prose that can generate coherent responses for a given input. It doesn't always work out right and it "hallucinates" (i.e. bullshits) more than we'd like but it feels like this is a more economically viable shot at most use cases for AGI than doing it "right" and attempting to create an actual AGI.

We didn't need to teach computers how language works in order to get them to provide adequate translations. Maybe we also don't need to teach them how the world works in order to get them to provide answers about it. But it will always be a 80% solution because it's an evolutionary dead end: it can't know things, we have only figured out how to trick it into pretending that it does.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#27
post #13

Earlier quoted context omitted.

>Imagine working on your PhD on chat bots since start of 2022. Your entire PhD topic might be irrelevant already... In fairness most PhD topics people work on these days, outside of the select few top research universities in the world, are obsolete before they begin. At least from what my friends in the field tell me.

Anecdata of one: I finished my PhD about 20 years ago in programming language theory. I created something innovative but not revolutionary. Given how slowly industry is catching up on my domain, it will probably take another 20-30 years before something similarly powerful makes it into an industrial programming language. Counter-anecdata of one: On the other hand, one of the research teams of which I've been a member…

> something as powerful as what I created

Could you give us more detail? It sounds intriguing.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#28
post #22
post #18

During my master's degree in data science, we had several companies visit our faculty to recruit students. Not a single one was a specialized NLP company, but many of them had NLP projects going on. Most of those projects were the usual "solution looking for a problem to solve". Even those projects that might have had _some_ utility, would have been way more effective to buy/license a product than to develop an in-ho…

> Honestly, I can't wait for GPT and other productivity tools to wrech havock upon the tech labour market. Some people in tech really need to be taken down a notch or two. That's an odd reason to want this.

Less bullshit jobs. Society needs doctors, nurses, plumbers and teachers not tech bros.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#29
post #18

During my master's degree in data science, we had several companies visit our faculty to recruit students. Not a single one was a specialized NLP company, but many of them had NLP projects going on. Most of those projects were the usual "solution looking for a problem to solve". Even those projects that might have had _some_ utility, would have been way more effective to buy/license a product than to develop an in-ho…

  those projects were just PR so that c-levels could sell how they were preparing their company for a digital world
This is exactly it. The 2017-2019 corporate version of "invest in AI" meant to build an in-house team to do ML experiments on internal data, and then usually evolved a bit to get some "ml-ops" thrown in so they could "deploy" the models they built. I spent some time with a few companies doing this and it always reminded my of "the cat in the hat comes back" when the cat let all the little cats out of his hat and they went to work on the snow spots... just doing busy work...

Anyway it's a symptom of the hype cycle - AI was the next electricity, but there were no actual products and nothing clear to do with it, just hire a bunch of kids to act like they were in a kaggle competition, or worse a bunch of PhDs to be under-utilized building scikit-learn models.

Now that there are (potentially) products coming along that at least bypass the low-level layer of ML, having an internal team makes no sense. Maybe the most logical thing that will happen is the pendulum will swing too far, and this bubble will consist more of businessy types using chatGPT without remotely understanding it or realizing it's just a computer program.

Re: Anyone else witnessing a panic inside NLP orgs of big tech companies?

#30
post #16

Earlier quoted context omitted.

Depending on definition, it is solved.

You're using the wrong definition, then. /s Where is some evidence that NLP is 'solved'? What does it even mean? OpenAI itself acknowledges the fundamental limitations of ChatGPT and the method of training it, but apparently everybody is happily sweeping them under the rug: "ChatGPT sometimes writes plausible-sounding but incorrect or nonsensical answers. Fixing this issue is challenging, as: (1) during RL training,…

It'd be great if GPT could provide it's sources for the text it generated.

I've been asking it about lyrics from songs that I know of, but where I can't find the original artist listed. I was hoping chat gpt had consumed a stack of lyrics and I could just ask it, "What song has this chorus or one similar to X..." It didn't work. Instead it firmly stated the wrong answer. And when I gave it time ranges it just noped out of there.

I think If I could ask it a question and it could go, I've used these 20-100 sources directly to synthesize this information, it'd be very helpful.

Post reply on HN