Live data from Hacker News

Knowledge Should Not Be Gated

formaly.io

11–20 of 57 posts

Re: Knowledge Should Not Be Gated

#11
post #4
post #2

Yep, knowledge should not be gated: Imagine Google search without any links or sources named This is the “modern” AI chatbot: It never mentions the training data it used, in fact has no idea what it used (often FB, Reddit and partisan websites) Update: I added the reply about after the fact Googling chatbots do - it’s different

Secifically in Google AI Overview I always see links to sites where the information is sourced from. Or at least some of the sites, if the same info is sourced from 100 pages then it only shows 2 or 3, maybe the ones with the biggest PageRanks.

Yep, that’s true

But those links are Googled after the model started to answer, they are not the links to the training data

Imagine an artificial “librarian” that read all the books and spits hallucinated quotes for you

But doesn’t let you enter the library, open a single book or even see the sources for those hallucinated quotes

But instead Googles some sources based on hallucinations after generating them ;-)

It’s better than nothing but you can Google them, too, while training data (the library) is completely hidden from you, even the public domain parts of it - zero attribution

Re: Knowledge Should Not Be Gated

#12

It seems beyond naive, rather malicious, to upload any useful private data to SaaS LLMs. Like, you are letting them data mine your business. Why are corporations not panicing over this?

because corporations are using providers with ZDR in the contract. If OAI or any of the cloud providers violate this they're getting sued to oblivion.

The problem is that there is an enormous, nearly unignorable incentive to work around it. So they will.

As the customer base becomes more and more corporate (which it will), they end up with disproportionately more customers whose experiences cannot be used to train the model to make it better for those customers.

Either way, corporate customers cannot leach off the training from consumers handing over their personal data forever; there aren't enough specialists in that training set to improve the models with no loss of corporate trust.

Betrayal of their trust is inevitable.

Re: Knowledge Should Not Be Gated

#13
post #8

It seems beyond naive, rather malicious, to upload any useful private data to SaaS LLMs. Like, you are letting them data mine your business. Why are corporations not panicing over this?

Most corporations likely have zero data retention agreements with LLM providers, at least for API usage. (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)

> (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)

Is actual ZDR verbiage in contracts more specific and limited in scope than what we see advertised publicly ("...except where needed to comply with law or combat misuse" in Anthropic's case)? Because those seem pretty damn vague and large enough holes to drive trucks through.

Re: Knowledge Should Not Be Gated

#15
post #4

Earlier quoted context omitted.

Secifically in Google AI Overview I always see links to sites where the information is sourced from. Or at least some of the sites, if the same info is sourced from 100 pages then it only shows 2 or 3, maybe the ones with the biggest PageRanks.

Yep, that’s true But those links are Googled after the model started to answer, they are not the links to the training data Imagine an artificial “librarian” that read all the books and spits hallucinated quotes for you But doesn’t let you enter the library, open a single book or even see the sources for those hallucinated quotes But instead Googles some sources based on hallucinations after generating them ;-) It’s…

There should be at least some correlation. When building the model they give more weight to some pages (e.g. Wikipedia) which have bigger trust (pagerank?). And when they provide links in answers, those matches are listed first which have better pagerank for the query.

So if it sources something in Wikipedia, it is more likely to provide Wikipedia as a trusted source for it.

The problem is when an answer is hallucinated, false, it may provide a source for it which contains the invalid info.

Re: Knowledge Should Not Be Gated

#16

It seems beyond naive, rather malicious, to upload any useful private data to SaaS LLMs. Like, you are letting them data mine your business. Why are corporations not panicing over this?

because corporations are using providers with ZDR in the contract. If OAI or any of the cloud providers violate this they're getting sued to oblivion.

These are the same people who performed the largest scale breach of copyright in history on the theory that they could get away with it.

I’m not making any accusations, but we should not underestimate their tolerance for legal and financial risk.

It may be a little paranoid to insist on self hosting based on that, but I’m not so sure that it’s crazy.

Re: Knowledge Should Not Be Gated

#17
post #8

Earlier quoted context omitted.

Most corporations likely have zero data retention agreements with LLM providers, at least for API usage. (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)

> (Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.) Is actual ZDR verbiage in contracts more specific and limited in scope than what we see advertised publicly ("...except where needed to comply with law or combat misuse" in Anthropic's case)? Because…

to combat misuse, we must store and read all prompts and responses. ;)

to comply with the law, we must send to the police our detections of illegal activity >:|

a guy subpeonaed your chats, i guess we stored them (oops) so now it's illegal to destroy it...

Re: Knowledge Should Not Be Gated

#19
post #15

Earlier quoted context omitted.

Yep, that’s true But those links are Googled after the model started to answer, they are not the links to the training data Imagine an artificial “librarian” that read all the books and spits hallucinated quotes for you But doesn’t let you enter the library, open a single book or even see the sources for those hallucinated quotes But instead Googles some sources based on hallucinations after generating them ;-) It’s…

There should be at least some correlation. When building the model they give more weight to some pages (e.g. Wikipedia) which have bigger trust (pagerank?). And when they provide links in answers, those matches are listed first which have better pagerank for the query. So if it sources something in Wikipedia, it is more likely to provide Wikipedia as a trusted source for it. The problem is when an answer is hallucina…

Yep, a few non-profits work on direct training data attribution:

OlmoTrace, Guide Labs with Clarity and a few more

Labs train the model with attribution baked-in and they say the bigger the model - the more interpretable it becomes

Pretty sure it’s the future

Post reply on HN