Is this such a big problem? You could still scrape all the data, or not?
June 2023 Data Dump is missing
261–268 of 268 posts
Re: June 2023 Data Dump is missing
#262This, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide th…
A decentralized system will never work because 99% of users do not care at all; the centralized systems are easier to sign up for and use. It's been demonstrated over and over and over again. Even if the underlying tech is decentralized, the community will settle around one or a few big instances (for example, Gmail and GitHub) which often end up having significant control over the trajectory of the entire ecosystem.…
Re: June 2023 Data Dump is missing
#263We have companies like Reddit and Stackoverflow not being profitable, despite being wildly successful in usage and internet mind-share. Neither of these companies are particularly over-staffed.
We post our "valuable" contributions there. So valuable that nobody wants to pay for it (structurally). We block ads. AI does the daylight robbery. We expect free APIs and data dumps.
Perhaps this is our wake-up call. The limitations of the "free" model and companies running at a loss for 15 years straight. It was always an anomaly.
Re: June 2023 Data Dump is missing
#264Earlier quoted context omitted.
There were a number of issues that lead to the decision not to do a grab and seed of SE into codidact. There was the "what license is that post actually under? Is it 2.5? 3.0? 4.0?" which made things difficult. There was the "what are the actual attribution requirements that SE has for sites that use its content?" This is a bit of an issue because it's never really clear what those requirements are and what you need…
Thanks for the write up. https://software.codidact.com/categories/38 I haven't used codidact (sorry, name needs replacing), ok just poked around * too slow * needs type ahead find search * needs a GIST experience The site looks good, presentation is really clean. Lots to like about it. But the think that replaces SO is going to have to be a step function in capabilities. That said, just fixing the weird descend into…
On the technical level, while Q&A is central, we also have other post types and other models. That post I linked to is an article in a blog that's part of our Meta community. The Electrical Engineering community has papers, so people can present information outside of the Q&A structure. Code Golf has a sandbox where people can get feedback on draft challenges before posting them. Software Development has a Code Review category. Some of our communities have added their own customizations to the code, like Code Golf's leaderboard for challenge answers. We want to work together with our communities to build what best serves their needs.
We've done some things that look small but might have larger effects. For example, the asker of a question can't mark one answer as "accepted" like on SO, but anybody can mark an answer as "works for me" -- or "outdated", or other annotations that communities can define. Scoring takes controversy into account, because +10/-5 and +5/-0 are very different even if they're both "net 5". With threaded comments, it doesn't matter so much if two people have an extended conversation; it's not in the way. Abilities are granted based on activity and reputation is just a number -- or can be turned off entirely if that's what a community wants. We're trying to make as much stuff configurable as we can, because we can't possibly know what's going to be best for every single community and don't have the hubris to claim we do.
We have the usual bootstrapping problem of a new thing. Our communities are small and trying to grow. Because they're small, visitors don't see thousands of questions and high activity, so they don't participate either and wander away, making it harder to build activity. We would love to find people who want to work with us to build communities. We recognize that helping to build a community with us is going to be harder and slower than just asking your question on SO, but if everyone were happy with SO this thread wouldn't be here, so maybe we're an option to a few people reading this?
(I haven't posted much on Hacker News, so I hope I've read the room correctly and that this kind of comment is ok. If not, I apologize and would appreciate correction so I don't repeat mistakes. Thank you.)
Re: June 2023 Data Dump is missing
#265Earlier quoted context omitted.
A decentralized system will never work because 99% of users do not care at all; the centralized systems are easier to sign up for and use. It's been demonstrated over and over and over again. Even if the underlying tech is decentralized, the community will settle around one or a few big instances (for example, Gmail and GitHub) which often end up having significant control over the trajectory of the entire ecosystem.…
I disagree, it worked before and it is the reason the internet even exists. The core issue is that user generated data is owned by one individual company. There are existing system that don't have this issues e.g. Usenet or bittorrent. We don't need to idiot proof the web. There are enough people to gather some place for a social network even if it's hard to use. The others can stay and will stay on reddit anyway unt…
Re: June 2023 Data Dump is missing
#266Earlier quoted context omitted.
Ahaha.. is this a serious post? I'll take the bait. If you want to shitpost with friends and strangers than exists no realistic purpose for identity management since the main goal is to remain anonymous and true anonymity comes by default on nostr. In case you do want to protect your identity in that case protect your keys. In case you missed the last few months, there are browser extensions that do not grant access…
I think you’re misunderstanding my point. I’m not saying key custody is infeasible. I’m saying current solutions aren’t working for average people, ie non-techies who don’t even know what a private key is. Do you disagree with this? If you agree with this premise, that also rules out browser extensions as a universal solution because most users are on mobile. They also have multiple devices and somewhat frequently fo…
I'm sure you can agree that having someone above users with access to their private keys is a serious failure point to user privacy. Exactly the reason why nostr remains strongly out of reach when compared to government controlled media.
Re: June 2023 Data Dump is missing
#267Earlier quoted context omitted.
> What we need is a legal way for companies to keep the data open, but also require OpenAI and friends to pay them for it. Couldn't that be accomplished by a law or ruling that using something for training AI doesn't exempt you from having to follow its license? OpenAI is already in blatant violation of both the "BY" and "SA" parts of the existing license.
Arguably, a model created by training on a corpus of data is a derived work of that corpus. Let's say I take a collection of images and use a program to compress them. When decompressed, the images are close to, but not exactly the same as the originals. Despite being in a different format, and despite not being exactly the same as the originals, the copyright to the compressed images is still held by whoever previou…
Not really. If diffusion models were compression they'd be so lossy as to be totally worthless
Re: June 2023 Data Dump is missing
#268Earlier quoted context omitted.
By this logic, isn't remembering something in your brain also a derived work? But that would not make any sense to protect until you create and distribute something based on that memory. The same logic should be applied to this.
If you remember it from your brain and perform it live, that’s perfectly fine. So there’s a line to be drawn somewhere and I don’t think it’s super clear cut in most cases.
If you remember it and make something that is a distinct work, something that may be steals the idea without reproducing any of its elements, that's never been considered under copyright.
I think that's going to be the litmus test for these AI. If you can get them to produce out both that is this things from anything else, it's not going to be a copyright violation because it's not a copy of anything.