Live data from Hacker News

June 2023 Data Dump is missing

meta.stackexchange.com

141–150 of 268 posts

Re: June 2023 Data Dump is missing

#141

Earlier quoted context omitted.

Users on SO created value and freely shared it with a community in expectation that the value they created would be freely and collectively shared with everyone. In SO's case this expectation was explicit; the data backup and API was billed as a deliberate choice designed to give users the freedom to migrate and scrape data in case the company went "evil." It was designed specifically to reduce SO's ownership claim o…

Thank you for the really thoughtful response! Good point that rent-seeking is maybe not the correct term now , but it looks increasingly like services will have to lock down content or shut down due to AI models frontrunning them with their own content. In that world, the AI models are in a great rent-seeking position (i.e. only they have the [old] content which was broadly available and now is not, due to their own…

> but it looks increasingly like services will have to lock down content or shut down due to AI models frontrunning them with their own content.

This feels like its slightly off to me.

An LLM that was trained on job postings to be able to categorize them isn't trying to do job postings ( https://wfhmap.com/algorithm/ ) but rather be able to do meaningful classification of bulk unstructured data.

An LLM trained on reddit is... weird to talk to, but talking to it doesn't replace asking a proper subreddit with people answering and comments back and forth. Is ChatGPT stealing views from people complaining about their job in /r/antiwork? Going to something in /r/news and sort by controversial and getting some popcorn turns out to be much more interesting than ChatGPT ever will be.

Maybe you can say that ChatGPT with some training of Stack Exchange sites has some utility (and that its really classified, tagged, and feedback given makes it even more useful), but GitHub CoPilot was trained on just GitHub stuff and its better at code than pretending "try {some broken code} hope that helps" is going to be useful for a LLM.

To me, this feels much more like CEOs that are having difficulty with existing monetization attempting to lock up the data that they have under a questionable pretext to monetize that to companies looking to train models for other things.

The sorting out of what the rights are on the output of models is something that needs to be sorted out - probably by the courts. I am still of the opinion that if something that might be copyrighted is used from any source, then the person doing the copying (who has agency) needs to do a license check themselves on it. I know that there is GPL code on Stack Overflow that looks like its licensed under CC 4.0 and if you copied the SO answer and put it in a BSD licensed repository, you'd be in violation of the GPL - and that's without touching any LLM.

There are also lots of non copyright things that the data could be used for. I'd like to make a AI-CATegorizer. Train it on a representative number of images form each of the reddit cat subs so that someone can ask it "here's a picture of my cat, what subs can this be posted to" and get back "/r/airplaneears /r/blackcats /r/stealthbombers" - and that's not something that is potentially generating copyrighted content (though it inherently uses it)... pretending that that those images were under a CC license, would it need to attribute all of the images that were part of the training data set to respond back with those three subreddits?

Re: June 2023 Data Dump is missing

#142

Earlier quoted context omitted.

Users on SO created value and freely shared it with a community in expectation that the value they created would be freely and collectively shared with everyone. In SO's case this expectation was explicit; the data backup and API was billed as a deliberate choice designed to give users the freedom to migrate and scrape data in case the company went "evil." It was designed specifically to reduce SO's ownership claim o…

Thank you for the really thoughtful response! Good point that rent-seeking is maybe not the correct term now , but it looks increasingly like services will have to lock down content or shut down due to AI models frontrunning them with their own content. In that world, the AI models are in a great rent-seeking position (i.e. only they have the [old] content which was broadly available and now is not, due to their own…

I do think if we were having this conversation about an explicitly community-owned forum or fanfic hosting service -- ie, a scenario where it's obvious that the community is behind the decision -- my reaction would likely be very different. I'm broadly pretty sympathetic to a forum saying, "we're doing this for us, not for a VC firm."

SO in specific though is an interesting site in that the value proposition of the site was very heavily based on this information being freely available and uncontrolled. I think they're in a position where it's much less appropriate for the site owners to try an clamp down on AI scraping.

If there is a strong movement from the SO community to change that, I'm not aware of it, but who knows, maybe I'm out of the loop.

Off the top of my head, another example of the distinction I'm getting at would be something like Wikipedia -- if the Wikipedia owners started trying to outright block site backups my immediate response would be, "well wait a second, that was not the deal we all made around this site, we signed up to help the Wikimedia foundation build an Open encyclopedia, even if that means it gets pulled into an AI dataset. We specifically didn't want the Wikimedia foundation to have the power to decide what usage of this data they would allow or deny."

Re: June 2023 Data Dump is missing

#144

It's unfortunate we are seeing all of these data platforms get locked off, because this is not going to affect AI development from big companies, it's only going to affect the ability for individuals to run AI development of any form in their home. I hope the data that has been found so far is going to big enough going forward, but it's incredibly unfortunate that this is happening. I hope all the people making these…

> It's unfortunate we are seeing all of these data platforms get locked off

Are there any AGPL-like licenses that address this?

Re: June 2023 Data Dump is missing

#145

Earlier quoted context omitted.

Assuming that the linked post is accurate and that the "approval from senior leadership" to turn the dump back on does not come...then yes, I would say so. Actually there is already Codidact, although if I recall correctly they explicitly ruled out importing SE data when they started up. https://codidact.org

Yeah, Codidact isn't a "fork" because they don't use the SE data.

It’s a fork of the community rather than the data and content.

Re: June 2023 Data Dump is missing

#146
post #7

Oof. This was one of the big central tenets of SO, the reason it wasn't Experts Exchange 2.0- the escrow of the community's contributions.

If you knew one simple trick all the answers on Experts Exchange were at least freely available. That trick was to simply scroll past the paywall. They had all the answers exposed so that google would index them. It was hilarious and silly.

Back in the time when Google didn't play favorites on companies not following their terms of service.

Re: June 2023 Data Dump is missing

#147
post #16

This, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide th…

I'm more concerned for authors of published works.

Imagine writing a text book with a royalty publishing deal. Your publisher decides they're going to use your book, amongst others, to train an LLM that can answer questions on your subject, and they're not going to pay you anything.

It's a legal gray area and they've got teams of lawyers whereas you do not.

Re: June 2023 Data Dump is missing

#148

"Just sorta stating the obvious here, but the timing of this is unbelievably terrible; I actually can't fathom a worse time for this call to be made than in light of this week. –zcoop98" Or, it's exactly the best time to do it. Doing it now allows your news to get blended in with the Reddit news. Doing it later after Reddit chatter settles down means all of the chatter is directed squarely at you.

It also means people are more motivated to build a replacement than just by the timewasting reddit being unavailable.

Re: June 2023 Data Dump is missing

#149

Earlier quoted context omitted.

There are ways to require payment for some uses of things that are legitimately open. As an example, consider the practice of selling exceptions to the GPL, as is done for Qt.

> As an example, consider the practice of selling exceptions to the GPL, as is done for Qt. There are people who do not consider that "open". See the whole debacle about what exactly constitutes "open source"

"Open" is a pretty vague word which could mean all sorts of things.

"Open source" is defined by the Open Source definition according to the OSI [1]. In saying that, I realize that every couple of years somebody tries to claim that their understanding of the term "open source" should trump the one the community has settled on. I personally am not ready to acquiesce to this semantic drift, at least, not yet.

[1] - https://opensource.org/osd/

Re: June 2023 Data Dump is missing

#150
Stack Overflow and Reddit want money for AIs to train on their data which is why they made these changes, so which companies are next? Could HN get crappier in order to milk AI money for its valuable comments? I guess Wikipedia at least can't do jack to get AI cash for its valuable data.
Post reply on HN