Live data from Hacker News

Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

github.com

321–330 of 354 posts

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#321

Earlier quoted context omitted.

> Hitting URLs denied by robots.txt has been argued to be just that. "Has been argued" -- sure, but never successfully; in fact, in HiQ v. LinkedIn, the 9th Circuit ruled (twice, both before and on remand again after and applying the Supreme Court ruling in Van Buren v. US) against a cease and desist on top of robots.txt to stop accessing data on a public website constituting "without authorization" under the CFAA.

Now do every other jurisdiction

CFAA was mentioned specifically, which means only US jurisdiction is relevant here.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#322
post #99

Earlier quoted context omitted.

In Chinese, it always added something like "For study/research purpose only. Please delete after 48 hours." This is what those volunteers added in subtitles of (pirated) movies/shows.

Fair, if AI companies are allowed to download pirated content for "learning", why ordinary people cannot.

AFAIK, downloading or watching pirated stuff isn't something you'll get in trouble for. Hosting and distributing it is what will get you.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#323
post #107
post #58

Title should be changed to "OpenAI publishes evidence they trained on pirated movies".

Of course. Piracy is legal when you have a bigger pile of money than the studios.

Too big to nail

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#324

Earlier quoted context omitted.

Well, wait, if somebody writes a computer program that answers 5 of 6 IMO questions/proofs correctly, and you don't consider it an "advance," what would qualify? Either both AI teams cheated, in which case there's nothing to worry about, or they didn't, in which case you've set a pretty high bar. Where is that bar, exactly? What exactly does it take to justify blowing off copyright law in the larger interest of progr…

I notice you've motte-and-baileyed from "revolutionize both the practice and philosophy of computing and advance mankind to the next stage of its own intellectual evolution" to simply "is considered an 'advance'".

You may have meant to reply to someone else. recursive is the one who questioned whether an advance had really been made, and I just asked for clarification (which they provided).

I'm pretty bullish on ML progress in general, but I'm finding it harder every day to disagree with recursive's take on social media.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#325

Earlier quoted context omitted.

Except that the jury’s (at best) still out on whether the influence of LLMs and similarly tech on knowledge workers is actually a net good, since it might stunt our ability to critically think and problem solve while confidently spewing hallucinations at random while model alignment is unregulated, haphazard, and (again at best) more of an art than a science.

Well, if it's no big deal, you and the other copyright maximalists who have popped out of the woodwork lately have nothing to worry about, at least in the long run. Right?

It's not about copyright _maximalism,_ it's about having _literally any regard for copyright_ and enforcing the law in a proportionate way regardless of who's breaking the laws.

Everyone I know has stories about their ISP sending nastygrams threatening legal action over torrenting, but now that corporations (whose US legal personhood appears to matter only when it benefits them) are doing it as part of the development of a commercial product that they expect to charge people for, that's fine?

And in any case, my argument had nothing to do with copyright (though I do hate the hypocrisy of the situation), and whether or not it's "nothing to worry about" in the long run, it seems like it'll cause a lot of harm before the benefits are felt in society at large. Whatever purported benefits actually come of this, we'll have to deal with:

- Even more mass layoffs that use LLMs as justification (not just in software, either). These are people's livelihoods; we're coming off of several nearly-consecutive "once-in-a-generation" financial crises, a growing affordability crisis in much of the developed world, and stagnating wages. Many people will be hit very hard by layoffs.

- A seniority crisis as companies increasingly try to replace entry-level jobs with LLMs, meaning that people in a crucial learning stage of their jobs will have to either replace much of the learning curve for their domain with the learning curve of using LLMs (which is dubiously a good thing), or face unemployment, and leaving industries to deal with the aging-out of their talent pools

- We've already been heading towards something of an information apocalypse, but now it seems more real than ever, and the industry's response seems to broadly be "let's make the lying machines lie even more convincingly"

- The financial viability of these products seems... questionable right now, at best, and given that the people running the show are opening up data centres in some of the most expensive energy markets around (and in the US's case, one that uniquely disincentivizes the development of affordable clean energy), I'm not sure that anyone's really interested in a path to financial sustainability for this tech

- The environmental impact of these projects is getting to be significant. It's not as bad as Bitcoin mining yet, AFAIK, but if we keep on, it'll get there.

- Recent reports show that the LLM industry is starting to take up a significant slice of the US economy, and that's never a good sign for an industry that seems to be backed by so much speculation rather than real-world profitability. This is how market crashes happen.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#326

Earlier quoted context omitted.

There is so much damning evidence that AI companies have committed absolutely shocking amounts of piracy, yet nothing is being done. It only highlights how the world really works. If you have money you get to do whatever the fuck you want. If you're just a normal person you get to spend years in jail or worse. Reminds me of https://www.youtube.com/watch?v=8GptobqPsvg

The dead corpses of filmmakers and authors and actors are buried in unmarked graves out behind those companies' corporate headquarters. Unimaginable horror, that piracy. Why has no one intervened? >If you're just a normal person you get to spend years in jail or worse. Not that I'm a big fan of the criminalization of copyright infringement in the United States, but who has ever spent years in jail for this? Besides,…

> who has ever spent years in jail for this?

Aaron Swartz?

EDIT: apparently he wasn't in jail, he was on bail while the case was ongoing - but the shortest plea deal would still have had him in jail for 6 months, and the penalty was 35 to 50 years.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#327

Earlier quoted context omitted.

If the model was able to generalise, you’d expect it to output something like “[silence]” or “…”, in response to silence. Instead, it reverted to what it has seen before (in the training data), hence the overfit.

Right, maybe my definition of overfitting was wrong, I always understood it more as trying to optimize for a specific benchmark / use case, and then it starts failing in other areas. But the way you phrase it, it’s just “the model is not properly able to generalize”, ie it doesn’t understand the concept of silence also makes sense. But couldn’t you then argue that any type of mistake / unknown could be explained as “…

I don't think so. Overfitting = the model was too closely aligned to the training data and can't generalize towards *unseen* data. I think it saw "silence" before, so it's not overfitting but just garbage in, garbage out.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#328

Classic overfitting It's the LLM equivalent of thinking that an out-of-office reply is the translation: https://www.theguardian.com/theguardian/2008/nov/01/5

As I didn't see one correct definition of overfitting:

overfitting means that the model is too closely aligned to the test data, picked up noise and does not generalize well to *new, unseen* data. think students that learn to reproduce questions and their answers for a test instead of learning concepts and to transfer knowledge to new questions that include the same concepts.

while this sounds like overfitting, I'd just say it's garbage in, garbage out; wrong classification. the training data is shit and didn't have (enough) correct examples to learn from.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#329
post #57

Earlier quoted context omitted.

How is this overfitting, rather than a data quality / classification issue?

Isn't overfitting just when the model picks up on an unintended pattern in the training data? Isn't that precisely what this is?

not necessarily, no. if you have 60% of examples for silence being the hallucination, it just learns the (what you detect as) wrong connection.

Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic

#330

Earlier quoted context omitted.

I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.

Having zero exposure to any form of computation for your entire life, as the vast majority of people in the early 19th century were.

What's the defence for the current population?
Post reply on HN