Live data from Hacker News

ML promises to be profoundly weird

aphyr.com

601–610 of 641 posts

Re: ML promises to be profoundly weird

#601

Earlier quoted context omitted.

Sure, we could change current law, but I think that only forces an AI company to buy one copy of every book. I don’t think it gives any sort of royalty stream to anyone beyond that. Copyright is literally the right to make copies. Once I have acquired a copy, I can read it, summarize it, transform it, etc. in myriad ways.

You can't make copies though. AI training requires making copies of materials, even if they're purchased.

Not true. You can photocopy pages from a book you own for your own use. You can make copies of purchased software as a backup. What you can’t do is make copies and give them to all your friends or sell them to the public.

Re: ML promises to be profoundly weird

#602

Earlier quoted context omitted.

I think it's implied that they're not talking about private data when they say they've run out.

fair. I want to +1 the fact that there is a large amount of data unseen by LLMs.

I think there are post training tweaks that can be done with corporate data to help fit an AI to a specific corporation. But I don’t think that private data will deliver us AGI. The knowledge for AGI is out in the world, not hidden inside corporations. Private data brings us knowledge of the XYZ project status and the division ABC budget and whether Bob wants a chocolate cake for his going away dinner or not.

Re: ML promises to be profoundly weird

#603

Earlier quoted context omitted.

What large caches of undigitized content exists? Surely, not everything has been digitized, but I can’t think it’s much in percentage terms.

has the whole youtube been indexed?

I’m sure Gemini has done it at some level. Google was pretty much founded on the assumption that more data is better. That has driven them to build or buy data sets that they can mine (Gmail, YouTube, etc.).

Re: ML promises to be profoundly weird

#604

Earlier quoted context omitted.

>This completely unpends the tenuous balance between creators and consumers. Why would a writer put an article online if ChatGPT will slurp it up and regurgitate it back to users without anyone ever even finding the original article? Who will contribute to the digital common when rapacious AI companies are constantly harvesting it? Why would anyone plant seeds on someone else's farm? I have been thinking about this.…

You raised a point and then never answered it. Why would anyone plant seeds on someone else's farm?

Well, maybe because life is not a zero sum game? Sometimes you do things just because.

Re: ML promises to be profoundly weird

#605

Earlier quoted context omitted.

fair. I want to +1 the fact that there is a large amount of data unseen by LLMs.

I think there are post training tweaks that can be done with corporate data to help fit an AI to a specific corporation. But I don’t think that private data will deliver us AGI. The knowledge for AGI is out in the world, not hidden inside corporations. Private data brings us knowledge of the XYZ project status and the division ABC budget and whether Bob wants a chocolate cake for his going away dinner or not.

I'm not seeing it the same way. Businesses in various industries have several types of moats - money, knowledge, experience, skills, etc. There is ton of competitive intelligence hidden in private data.

Its one of the reasons you can't use chatGPT and start manufacturing chips or vaccines, or anti-cancer medication. The gap between publicly available data that informs academic "core science" research versus specific product-based knowledge that shows you how to make a successful drug candidate that can withstand regulatory scrutiny or be a safe and effective drug for the worlds population.

We could iterate so quickly if this private data set was democratized.

Re: ML promises to be profoundly weird

#606
post #302

Earlier quoted context omitted.

> I keep explaining to my peers, friends and family that what actually is happening inside an LLM has nothing to do with conscience or agency What would the insides have to look like to have anything to do with conscience or agency?

I don’t have an answer. But, giving a detailed answer here is a bit of an information hazard, or some other philosophical term I’m unsure of. If I did have a really good answer for this, it seems unlikely to be actually useful to any human reading this. Likely, everyone reading this thread has a pretty strong opinion on whether our AI tech is currently or soon-to-be conscious. However, this thread is going to be pick…

Hah, late but a solid reply - thanks!

I'm a lot more agnostic than you are: I don't know whether LLMs are conscious. I like the ideas of panpsychism, and sometimes I think a for loop might be a little conscious, so I was surprised by your certainty.

Re: ML promises to be profoundly weird

#607
post #597

Earlier quoted context omitted.

The statement IS true anyways, the problem is that you failed to distinguish between an example and a universal claim. You want to argue on logic? I'm an engineer, I can argue on precision too: The (true!) statement is "However, there's an immense difference in scale between post-industrial strip mining of resources, and preindustrial resource extraction powered solely by human muscle (and not coal or nitrogylcerin e…

Is your counter argument that you’re not wrong just attacking a straw man? Because it really sounds to me like you are just clueless. Strip mining goes back thousands of years, it’s a simpler technology than making tunnels. And no it wasn’t limited to human power to crack rock several more powerful methods existed. Roman mining literally destroyed a mountain, operating within an order of magnitude of the largest mine…

It’s almost like you’re intentionally trying to be wrong.

You don't seem to understand how analogies work. I’m not talking about strip mining vs tunnel mining, I was comparing scale of human powered mining to mining with nitroglycerin.

I’ll let you figure out how the scale of mining “going back thousands of years” is very different from modern explosive mining on your own. Go google “iron production by year” or something. Hint: it took generations for the Romans to strip a small hill, that a modern midsize mining company can do in a few days.

Re: ML promises to be profoundly weird

#608

Earlier quoted context omitted.

I mean, that is what the court said! Training on pirated data was not fair use. Training on legally acquired data is fair use. Anthropic legally acquired the data and re-trained on it before release.

It did not say that. See Judge Alsup's order ( https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/... ), pp. 29-30, Section IV(B)(ii) ("The Pirated Library Copies"). "[T]he test requires that we contemplate the likely result were the conduct to be condoned as a fair use — namely to steal a work you could otherwise buy (a book, millions of books) so long as you at least loosely intend to make further copies f…

[deleted]

Re: ML promises to be profoundly weird

#609

Earlier quoted context omitted.

Structurally a transformer model is so unrelated to the shape of the brain there's no reason to think they'd have many similarities. It's also pretty well established that the brain doesn't do anything resembling wholesale SGD (which to spell it is evidence that it doesn't learn in the same way).

Sure the implementation details are different. I suppose I should have asked by what definition of "consciousness and agency" are today's LLMs (with proper tooling) not meeting? And if today's models aren't meeting your standard, what makes you think that future LLMs won't get there?

These questions really vex me. The appearance of intelligence is almost orthogonal to "consciousness and agency." If a human has a stroke and forgets how to speak, or never learns, or has some severe form of learning disorder, they still have exactly the same rich inner life full of subjective qualititative experience known only to them as the rest of us. Similar to an array of GPUs. If you remove the text encodings from the rest of the computing system it is a part of, outputs will appear as gibberish to you and it will no longer appear to be intelligent at all, but whatever is happening at the level of electrons meeting silicon would still be exactly the same. If it's having conscious experience at all, it should be having it regardless of whether the outputs it computes are interpreted as text or as textures on a game background.

I just don't see why "I can talk to it now" changes anything. We don't give humans less moral consideration when they're dreaming, hallucinating, tripping on LSD. The brain is just as conscious when it's having nothing but completely abstract nonsense thoughts as when it's writing The Republic.

I understand why it feels different to people. Shit, this thing can talk to me; maybe it's alive and I should treat it like such. But that's a conservative reaction to a black box known only by its behavior. The problem is these things are not actually black boxes. We don't understand the functions being computed or we'd just hard-code them and not need statistical learning techniques, but we do understand how computers work. We know process state is saved off and restored billions of times per second because of context switching. We know that state is simply a stored byte sequence that can be copied, backed up, restored endlessly. Servers and computing hardware can be destroyed but software cannot and LLMs are software. It's not at all like a brain. There are animals that go into various levels of reduced or suspended function that appear like dormancy, but there is no stream of personal subjective experience that can survive the complete destruction of its own physical body. The fact that it pays off evolutionarily to tacitly encode that reality into our instincts at an extremely deep, core level is why we have fear and pain in the first place, to nudge us toward predictive modeling of the world that keeps us alive, able to find food, and able to reproduce. Software needs none of that. There is no reason whatsoeve that, assuming a processor has subjective experience, that the subjective experience of having some gates fire versus others gets interpreted by humans programmers as "loss" and "training" and some is numerically approximating a PDE solution. Why should those feel different to the machine when the firing patterns are exactly the same and only the human interpretation of the output is different?

It just feels like a vast, vast category error for people to be speculating about machine consciousness and moralizing about how we "treat" software systems.

Re: ML promises to be profoundly weird

#610

Earlier quoted context omitted.

I mean, that is what the court said! Training on pirated data was not fair use. Training on legally acquired data is fair use. Anthropic legally acquired the data and re-trained on it before release.

It did not say that. See Judge Alsup's order ( https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/... ), pp. 29-30, Section IV(B)(ii) ("The Pirated Library Copies"). "[T]he test requires that we contemplate the likely result were the conduct to be condoned as a fair use — namely to steal a work you could otherwise buy (a book, millions of books) so long as you at least loosely intend to make further copies f…

I don't read how that opposes what I said, that's part of the "training on pirated data is not fair use." That said, I am not a lawyer. From those pages:

> The copies used to train specific LLMs were justified as a fair use.

This is (in my understanding) because those were not the pirated copies.

> The copies used to convert purchased print library copies into digital library copies were justified, too, though for a different fair use.

Buying a book and then digitizing it for purposes of training is fair use.

> The downloaded pirated copies used to build a central library were not justified by a fair use.

Piracy is not fair use, you quoted this part as well.

In the conclusions section a the end of 31:

> This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies.

Training is fair use. Pirating is not fair use, and therefore, you can't train on that either.

What part am I missing?

Post reply on HN