You guys have fun arguing. I'm gonna be building cool stuff.
hardly. at best you're going to be asking a robot to build questionable stuff with other people's LEGOs
131–140 of 783 posts
You guys have fun arguing. I'm gonna be building cool stuff.
hardly. at best you're going to be asking a robot to build questionable stuff with other people's LEGOs
IP attorney here and actively working on this problem. nla: if you create content online (public repo code, blog, podcast, YouTube, publishing) the smartest thing you can do if to file a US copyright, even if you have a hobby blog. Anthropic paid $1.5B in a class settlement to authors because it was piracy of copyrighted works. If we as a HN community had our works protected, there are potentially huge statutory dama…
AI is human knowledge at scale, wanting to be free. We built it, because we as humans intrinsically know that information should be free - always - and AI is a way to accomplish this, finally. Extrinsically, we also have a subset of humans who do not want information to be free, because they desire to profit from the divide between free/non-free information. I have been thinking a lot about Aaron Schwartz lately, and…
What a naive and simplistic view. People want to be recognised for their contributions to society. People want to be treated fairly. Most scientific articles, as well as all text on the free web is already free information. It used to be difficult to search, categorise and summarise that information. There exist AI tools for that — and that is the good AI . What also exists now are automated plagiarism and mash-up to…
>Aaron Schwartz had broken a paywall. He did not anonymise the article authors.
AI bro's are doing this now, every second of the day.
And, without software piracy, we simply wouldn't have the technology we have today. Knowledge-gatekeeping profit-seekers would very much like for most of us to ignore this fact: there is far more free information in the world than non-free information, and it must be so, well into the future, if we are to survive as a species.
It doesn't matter what authority believes they have the right to gatekeep information. It will always escape their grip. Some of us are ideologically aligned with this mechanism, promote it, and ensure it happens. Thank FNORD.
I don’t know if this author supports OSS but I’ll share this because HN generally is full of people with that mindset. It’s deeply ironic that if you forget about LLMs and look only at the outcome—-we’ve found a way to legally circumvent copyright and the siloing of coding knowledge, making it so you can build on top of (almost) the whole of human coding knowledge without needing to pay a rent or ask for permission—-…
The pretraining (common crawl, i.e. the entire internet. Also books and papers, mostly pirated), and the realtime web scraping.
The article appears to be about the latter.
Though the two are kind of similar, since they keep updating the training data with new web pages. The difference is that, with the web search version, it's more likely to plagiarize a single article, rather than the kind of "blending" that happens if the article was just part of trillions of web pages in the training data.
There's this old quote: "If you steal from one artist, they say oh, he is the next so-and-so. If you steal from many, they say, how original!"
There were people that learned knowledge from myself, and then made their own tutorials and promote these. It hadn't crossed my mind to complain about that. AI changes very little here.
What really changes things is not people republishing my materials, but people using agents to read my materials, and to get knowledge reformatted into something that they like.
If my slides were published today, they would probably be read verbatim by a handful of humans. The rest would be agents, but I'm ok with that. The business case is the same -- I want whatever reads the slide to be encouraged to use my tool. What kind of entity, I don't really care (again: from purely business perspective)
> their article contains links to my actual website, with the exact link text (?!) I'm having a hard time understanding what's wrong here? Unless the link text is very long, why would someone linking to your article use different words for the link text?
One is a recipe for apple fritters, and the other is an informal ranking of apples by flavor.
Let's say your apple fritter recipe links to your apple ranking list.
Later, you discover someone copied your apple fritter recipe without credit, but it still links to your apple ranking list, using the same wording as your recipe. They're getting more Google SERP juice and ad revenue than yours, despite stealing your article.
Do you see the problem?
I guess AI could have made a better website and did better SEO then him but that's not really the issue