Live data from Hacker News

June 2023 Data Dump is missing

meta.stackexchange.com

191–200 of 268 posts

Re: June 2023 Data Dump is missing

#191

Really strange comment. > I was recently impacted by the Company's layoff. > I'm offering what I can to uphold the Company's values of Transparency & being Community-centric. I wouldn't offer transparency about a former employers internal operations. Let them respond or at least ping a current employee to respond.

I'm reading it as the ex-employee thinking that the company might not want to respond, and choosing to do so despite that, on the grounds that it should be acceptable to do so (i.e. the ex-employee couldn't be publicly "scolded" for it without the company publicly displaying not following their values)

Re: June 2023 Data Dump is missing

#192
post #183

Earlier quoted context omitted.

Thanks for the write up. https://software.codidact.com/categories/38 I haven't used codidact (sorry, name needs replacing), ok just poked around * too slow * needs type ahead find search * needs a GIST experience The site looks good, presentation is really clean. Lots to like about it. But the think that replaces SO is going to have to be a step function in capabilities. That said, just fixing the weird descend into…

Personally when the fork happened, I was interested in it... but I'm mostly "meh" about it since its another Q&A style format that maintains all the advantages and disadvantages of the format that SO provides. I really hoped that they would have gone for something much more radical in terms of trying to create a way to share knowledge. It's Stack(Exchange|Overflow) with better governance and a different development c…

Totally agree!

The shtshow is the reason to switch, but there also needs to be new capabilities on the other side. As soon as they break old.reddit.com, I am out.

I feel like https://en.wikipedia.org/wiki/ActivityPub is 10x more complex than it needs to be. The SO replacement should be an application on-top of an existing protocol.

Re: June 2023 Data Dump is missing

#193

Earlier quoted context omitted.

After I made that comment, I stopped for a short time to think what I would do. My conclusion was that it is much better done with something like Mastodon (or even Mastodon itself) than with a web site.

And then community projects subscribe to the feeds and index and make searchable? I could see basically structured toots representing questions and then ... oh man, did you just nerd snipe this? Are you saying that it should be managed by Mastodon the org, or fediverse the technology?

I meant the technology. The one downside that I see is that AFAIK, there is bad support for editing questions and answers. If something like this is created, there could be threads of patches making any edition, but that requires tooling support.

(And yeah, I'm purposefully trying to nerd-sniping people here :)

Re: June 2023 Data Dump is missing

#194
post #76
post #23

Yikes. Reddit. Stack overflow. It's all going south. Maybe we won't even have to wait for LLMs to destroy the web we used to know.

This is LLMs destroying the web we used to know. I would be willing to bet that the driving force behind the decision was to make it less trivial for LLMs to say "the data was already there under an open license, so we legally undercut stack overflow".

The fact that everyone is hoarding data because they think there is a gold rush afoot is obvious. Everyone with loads of data is clamping down, hoping they can get a cut of those AI VC dollars. Except for Wikipedia at least.

But let's be real about the morality here: Stack Overflow is a badge-powered mechanical Turk. It uses 100% unpaid labor to go and search Google for answers and post them on SO, providing a "service"[1]. For it to moralize about the ownership or sanctity of data is irony.

[1] - There are exceptions, obviously. There are true experts who wander the virtual halls of StackOverflow and dole out wisdom. But overwhelmingly it is clear that answers primarily come from people who rush to Google and then copy/paste from blogs and tech papers. And while Stack Overflow dumps are CC because that's the agreement that it made with contributors, a lot of the content on the site was ripped without attribution and in defiance of IP. So...maybe not too many tears for SO.

Re: June 2023 Data Dump is missing

#195

What irks me about this is that 100% of their data is provided for free, by the community that they have fostered, the people like myself who have answered > 2500 questions[0], and now SO feels hard-done-by by LLMs using all their hard work to create tools like CodeGPT, GitHub copilot, etc. Were it really a site for helping developers to improve their skills and increase their productivity through the give-and-take m…

> None of the contributors (apart from the employee ones, I suppose) ever got paid any currency other than high-fives in the form of rep, medals, the gamified stuff, moderation rights, and at certain rep levels some swag in the form of t-shirts and the usual.

I would love to see some kind of identity and reputation system where the "high-fives in the form of rep" could follow people across communities. It may not feel like much compensation if you've contributed over 2500 answers, but having reputation gained in your area of expertise grant you a high level of trust to interact in other communities could be valuable, at least in my opinion.

Assuming they're making this move to protect against AI / LLMs, I think SO is in an impossible situation here. When all the ChatGPT hype started, one of my first questions was "what happens to the incentive for contributors and creators?" Why would I want to contribute on a platform if I know an AI model is going to come in, take my contribution, and regurgitate it back to the masses in a way that I can't control?

Even if I get some attribution from the AI/LLM, do I even want it? If the LLM is blending content from multiple sources, which changes the context and presentation I put effort into, is the quality going to be high enough to match what I strive to achieve for myself when I'm trying to build a reputation as a high quality contributor? What if the AI is hallucinating objectively poor quality content and giving me partial attribution?

So, for me, part of the social contract with SO is that I provide answers, but I get to control the entire interaction; the context, the presentation (mark up), defending criticism in the comments, etc.. In addition to that, since the entire conversation happens inline, I can be corrected by someone even more knowledgeable than me and use that feedback for self improvement.

I think AI is going to be disruptive and the whole idea, for me anyway, behind disruption is that you break an existing system and then everyone is free to take a shot at claiming part of the new gold rush that occurs while trying to build the replacement. The problem with AI is that it's going to break a lot of services that do a good job of serving the community and shouldn't be broken. SO is a great example of a healthy community that doesn't need disruption, but the massive amount of high quality, curated content is going to make them a prime target for LLM training.

Personally I think the only solution is for "noai" variants of popular open source licenses so contributors have the ability to make it clear they don't want to contribute to AI/LLM companies. If SO had an option to flag contributions as CC-BY-SA-NOAI, I'd enable it on my stuff going forward.

Re: June 2023 Data Dump is missing

#196
post #23

Yikes. Reddit. Stack overflow. It's all going south. Maybe we won't even have to wait for LLMs to destroy the web we used to know.

Time to adopt Nostr as future-proof path.

How does nostr handle determined activists with establishment backing? What if a nostr user posts information an activist wants censored and the activist goes around threatening relays, their hosting providers, anyone connected to people operating the relay?

Re: June 2023 Data Dump is missing

#197

Earlier quoted context omitted.

I wonder how many of those contributors, if re-consulted, would sign up to having their contributions used to train a for-profit LLM though? I certainly didn’t sweat it out helping people on SO to pay for Sam Altman’s fucking swimming pool.

I mean they would just scrape it if there's no data dump. It just makes it harder for the small guys. They probably scraped and are scraping HackerNews. Generative AI doesn't follow copyright or even explicit software licenses as we have seen in AI art with human signatures and Microsoft Copilot.

They definitely scraped HN: one of my favorite ways to change the style of an LLM is to ask it to rephrase in the style of a snarky HN comment.

Re: June 2023 Data Dump is missing

#198
post #161

Earlier quoted context omitted.

There are a couple of different motivations a company could have around blocking API access to prevent AI scraping: A) scraping itself is too expensive. I suspect that's probably not the case with SO because they blocked backup . Downloading the database from the Internet Archive doesn't cost SO any money. B) the AI is going to replace the original creators (or more likely, devalue their work and push wages lower) an…

On the B/C test: From Jody Bailey (SO staff) https://meta.stackexchange.com/a/390040 (in full) posted 7 minutes ago (as I write this) > Stack Overflow senior leadership is working on a strategy to protect Stack Overflow data from being misused by companies building LLMs. While working on this strategy, we decided to stop the dump until we could put guardrails in place. > We are working on setting up the infrastructur…

I looked into this a bit more and I'm somewhat doubtful of this justification. There's at least been discussion within SO about charging access: https://meta.stackexchange.com/questions/388551/is-se-going-...

> First, I'd like to say that the intent of what Prashanth is saying is very simple: to return value to the community for the work that you have put in. The money that we raise from charging these huge companies that have billions of dollars on their balance sheet will be used for projects that directly benefit the community.

This is worded very specifically. Is SO planning to give money to users? They don't say anything like that; instead they say that they'll be "spending that money on the platform."

Well what does that actually mean? Every feature that SO builds could be characterized as "for the benefit of the community." It's hard not to read that response as just another way of saying "we're going to profit from this as a company, but don't worry because we use our profits to fund product development."

Heck, Reddit could make exactly the same claim, and in fact the linked Wired article actually makes that comparison:

> "Community platforms that fuel LLMs absolutely should be compensated for their contributions so that companies like us can reinvest back into our communities to continue to make them thrive," Stack Overflow’s Chandrasekar says. "We're very supportive of Reddit’s approach."

Re: June 2023 Data Dump is missing

#199
post #13

Earlier quoted context omitted.

A lot of people were only willing to contribute to StackOverflow because of the CC licensing, trusting the knowledge wouldn't be locked up. As a business that depends on vast amounts of volunteer effort they need to balance providing a site where people are willing to contribute against making as much money as they can.

I wonder how many of those contributors, if re-consulted, would sign up to having their contributions used to train a for-profit LLM though? I certainly didn’t sweat it out helping people on SO to pay for Sam Altman’s fucking swimming pool.

For what it's worth, as someone who has put a lot of writing online, I'm not bothered by having my writing including in the training sets of these LLMs. I write because I want to share knowledge, and it isn't important whether people get the knowledge directly from me versus mediated by friends, LLMs, etc.

Re: June 2023 Data Dump is missing

#200
post #179

Earlier quoted context omitted.

The original founders sold the site for $1.8 Billion.

That was in 2021. https://stackoverflow.blog/2021/06/02/prosus-acquires-stack-... Joel left in 2019 https://twitter.com/spolsky/status/1111267189133316097 Jeff left in 2012 https://blog.codinghorror.com/farewell-stack-exchange/ I'm not sure that it is fair to say that the original founders sold it for that amount.

[flagged]
Post reply on HN