Earlier quoted context omitted.
Git forges are some of the worst case for this. The scrapers click on every link on every page. If you do this to a git forge, it gets very O(scary) very fast because you have to look at data that is not frequently looked at and will NOT be cached. Most of how git forges are fast is through caching. The thing about AI scrapers is that they don't just do this once. They do this every day in case every file in a glibc…
That's very strange to me that they do it everyday. I thought training runs took months. Do they throw away the vast majority of their training attempts (e.g. one had suboptimal hyperparameters, etc)?
FOSS infrastructure is under attack by AI companies
621–630 of 631 posts
Re: FOSS infrastructure is under attack by AI companies
#622Earlier quoted context omitted.
That might indeed apply to open source software. If we instead adopt the view of free software ( https://www.gnu.org/philosophy/open-source-misses-the-point.... ), the fact that OpenAI and other large corporations train their large-language models behind closed doors - with no disclosure of their training corpus - effectively represents the biggest attack on GPL-licensed code to date. No evidence suggests that OpenAI…
I absolutely agree with you that the current big LLMs enable an attack on all FOSS licenses and especially copyleft ones. That doesn't mean that one couldn't create LLM code generators in a respectful way. Do license analysis on the input code and then train separate models on the different license buckets, with the outputs from each model considered derivative works of the input corpus. Also I don't think a restrict…
Regarding the licensing, I'll restate my point that the Affero license was created precisely in a moment where the existing licenses could no longer uphold the freedoms that the Free Software Foundation set out to defend. A change of license was the right solution at that particular point in time and, if it worked then, I think we can all agree that there is at least a precedent that such a course of action might work and should at the very least be considered as a possible solution for today's problems.
That said, my own personal view is more aligned with demanding the nation states to pressure big corporations so that currently closed-source software becomes at least open-source (either by law, or simply by stopping using it and invest their budget in free alternatives instead). Note I said open source and not free. I just would like to read their code and feed it to my LLM's :)
Re: FOSS infrastructure is under attack by AI companies
#623Earlier quoted context omitted.
I absolutely agree with you that the current big LLMs enable an attack on all FOSS licenses and especially copyleft ones. That doesn't mean that one couldn't create LLM code generators in a respectful way. Do license analysis on the input code and then train separate models on the different license buckets, with the outputs from each model considered derivative works of the input corpus. Also I don't think a restrict…
I see what you are saying and don't completely disagree. I however feel that the spirit of free software is to set all software free. From that it follows, that if we are going to follow the current route of complete disregard for authorship and licenses, then the free software movement should continue fighting to liberate all software in existence. In other words, those LLM's that you mention that are to enable soft…
On Affero, that was indeed definitely needed, although some folks on HN seem to think that privately modifying code is allowed by copyright, even if the modified version is outputting a public website, thus what the license says is irrelevant. That seems bogus to me, but seems a loophole if it is legit. Anyway, personally I think that people should simply just never use SaaS, nor web apps. It also doesn't help with data portability.
I'd go further and advocate for legally mandated source code escrow for copyright validity, and GPL like rights to the code once public, which would happen if the software is off the market for N years.
Re: FOSS infrastructure is under attack by AI companies
#624Re: FOSS infrastructure is under attack by AI companies
#625Earlier quoted context omitted.
I see what you are saying and don't completely disagree. I however feel that the spirit of free software is to set all software free. From that it follows, that if we are going to follow the current route of complete disregard for authorship and licenses, then the free software movement should continue fighting to liberate all software in existence. In other words, those LLM's that you mention that are to enable soft…
On setting all software free, indeed, thats the point made in the post by mjg59. None of the AI companies train on their own proprietary software though, which is telling. On Affero, that was indeed definitely needed, although some folks on HN seem to think that privately modifying code is allowed by copyright, even if the modified version is outputting a public website, thus what the license says is irrelevant. That…
I agree 100%.
Re: FOSS infrastructure is under attack by AI companies
#626Earlier quoted context omitted.
Or, could it be, just possibly be (gasp), that some of the devs at these "hotshot" AI companies are just ignorant or lazy or pressured enough, so as to not do such normal checks? Wouldn't be surprised if so.
You think they do cache the data but don't use it? For what it's worth, mj12bot.com is even worse. They pull down every wheel every two or three days, even though something like chemfp-3.4-cp35-cp35m-manylinux1_x86_64.whl hasn't changed in years - it's for Python 3.5, after all.
that's not what I meant.
and it is not they, it is it.
i.e. the web server, not bots or devs on the other end of the connection, is what tells you the needed info. all you have to do is check it and act accordingly, i.e. download the changed resource or don't download the unchanged one.
google:
http header last modified
and look for the etag link too.
Re: FOSS infrastructure is under attack by AI companies
#627Earlier quoted context omitted.
You think they do cache the data but don't use it? For what it's worth, mj12bot.com is even worse. They pull down every wheel every two or three days, even though something like chemfp-3.4-cp35-cp35m-manylinux1_x86_64.whl hasn't changed in years - it's for Python 3.5, after all.
>You think they do cache the data but don't use it? that's not what I meant. and it is not they , it is it . i.e. the web server, not bots or devs on the other end of the connection, is what tells you the needed info. all you have to do is check it and act accordingly, i.e. download the changed resource or don't download the unchanged one. google: http header last modified and look for the etag link too.
Re: FOSS infrastructure is under attack by AI companies
#628Yep -- our story here: https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse... (quoted in the OP) -- everyone I know has a similar story who is running large internet infrastructure -- this post does a great job of rounding a bunch of them up in 1 place. I called it when I wrote it, they are just burning their goodwill to the ground. I will note that one of the main startups in the space worked with us direct…
> just burning their goodwill to the ground AI firms seem to be leading from a position that goodwill is irrelevant: a $100bn pile of capital, like an 800lb gorilla, does what it wants. AI will be incorporated into all products whether you like it or not; it will absorb all data whether you like it or not.
links to this comment.
Re: FOSS infrastructure is under attack by AI companies
#629VideoLAN here. Same for us, our forum and our Gitlab are getting hammered by AI companies bots. Most of them don’t respect robots.txt…
Did you document the measures you took to remedy this?
Re: FOSS infrastructure is under attack by AI companies
#630Earlier quoted context omitted.
Besides flooding them with junk, what about outright sabotage in the form of serving zip bombs or other ways to waste computational resources?
I think we need to aim for the bots to get _negative_ utility value from visiting our traps, not just zero value. Did you try to GET a canary page forbidden in robots.txt? Very well, have a bucket load of articles on the benefits of drinking bleach. Is your user-agent too suspicious? No problem, feel free to scrape my insecure code (google "emergent misalignment" for more info). A request rate too inhuman? Here, take…
[1] https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act