Paid by requester is the feature they are looking for. https://docs.aws.amazon.com/AmazonS3/latest/dev/configure-re...
That's a very neat feature I never knew existed. I only suspect this will hurt their larger mission at helping many smaller teams and individuals to use the models.
Apple downloads ~45 TB of models per day from our S3 bucket
171–180 of 237 posts
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#172Re: Apple downloads ~45 TB of models per day from our S3 bucket
#173Have you considered cloudflair?
Nobody is going to regularly let you have 45 terabytes of data for free. It's swamping them, especially relative to other users.
You can still use a few dedicated servers on Hetzner though. They'll easily let you serve 200TB / month from each server for $100. It's obviously not comparable to proper CDN though, but it's does work for many companies.
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#174Earlier quoted context omitted.
And how would you be contacting Apple, from your little 3-person startup in Paris ? You assume they have the means or contacts to do that; and IMHO the tweet is not that aggressive. It's been done before, (e.g Intel has been called out for consuming kernel.org bandwidth and git CPU power) and is the simplest way to have people from inside BigCorp get a message.
No need to argue with me. I'm only explaining the original point, not making it myself :-) But reviewing the thread I see you your question was in response to someone calling them "service providers" to Apple. So your question was entirely justified and it was me who'd lost context.
I want to add an answer that to those saying that gives information about Apple: it only says that somehow, one team inside Apple has setup a CI (badly written script) with maybe 5k tests (from a standard set for instance) and has 9 commits per day. Or maybe they have more commits, and less tests ? Or maybe, it's a matrix of 70x70 tests. Or maybe… well, all it says is that someone is experimenting with this.
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#175Earlier quoted context omitted.
Oh, so now IP addresses is PII? When it is inconvenient for FAAGM monster corporations? I seem to remember a few hundred thousands corporate statements that tracking individual IPs is totally ok and not surveillance.
IP addresses are definitely PII under GDPR.
My understanding is that it is only PII if it's in conjunction with other data in particular ways.
When the GDPR came along and IP Addresses were being mentioned as PII, my employer required us to sign a document stating (in part) that we wouldn't access, download or communicate PII data except when specifically authorised to do so.
When I refused to sign that and ran it up the flagpole with questions about how I'd do my job which occasionally included things like blocking IPs and dealing with network captures, it (apparently) went to lawyers who came back with a revised document to clarify that IPs wern't PII unless used in specific ways.
My point being - that specifically showing the 17/8 range as an aggregate shouldn't violate the GDPR any more than mentioning that it's Apple being the source of traffic.
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#176Earlier quoted context omitted.
> That's about $4000/month in bandwidth costs You're an order of magnitude off. 45 TB per day is 1,350 TB in a month, or 1,350,000 GB. Show me somewhere you can get a petabyte of egress inside a calendar month for 4 figures USD... Let's suppose you even used the cheaper egress from Cloudfront rather than serving from S3 (lol @ your wallet if you serve 1 PB doing that). https://aws.amazon.com/blogs/aws/aws-data-transf…
> Show me somewhere you can get a petabyte of egress inside a calendar month for 4 figures USD Correct me if I'm wrong, but most colocation/dedicated server providers offer such prices. E.g. hetzner.com @ €1/TB, sprintdatacenter.pl @ €0.91/TB, dedicated.com @ $2/TB (or $600/month for an unmetered 1 Gbps connection). But if you want S3/CDNs/ , then yeah, they're expensive. BTW per Cloudflare ToS[0]: > Use of the Servi…
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#177Earlier quoted context omitted.
That makes it harder for everyone though where companies like Apple should proxy and/or cache those requests to their own internal version rather than hitting that S3 bucket every time. Requiring requester payment would mean it would only really be used by corporations where the author clearly wants a service open to anyone without having to open an AWS account to pay.
Make the bucket buyer-pays, but offer Torrent links as well. Businesses doing CI will pay for S3 usage so they don't have to deal with torrents, end-users will get free torrent access, everybody wins.
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#178Isn’t this a use case where BitTorrent would shine?
I'd say there's only a 10% chance Apple's firewall would let BitTorrent through, and only a 3% chance the CI servers would maintain a positive seed ratio.
Possibly it might solve the problem because users would cache the resources themselves to avoid the hassle of getting BitTorrent into their CI pipeline...
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#179Earlier quoted context omitted.
Yes. I don’t know if that’s all that they do. They often port new Tensorflow models to PyTorch as well. They provide straightforward APIs, nice documentation, and clear tutorials. I use their stuff pretty regularly.
Do you know their business model? Looks like they are open source company. How they earn money?
My guess is that they don't earn money yet. They more or less are just out of school so it's a young startup. Actually studied with some of them and I'm still a student.
Re: Apple downloads ~45 TB of models per day from our S3 bucket
#180If a company the size of Apple finds this that useful, perhaps you should consider charging for your service, rather than just complaining on Twitter about the free usage you appear to have willingly given away? Or perhaps you have reached out to them, but are for some reason still complaining on Twitter to drum up PR or something? Regardless, this posting is ridiculously context-free to the point of being click-bait…
And then people wonder why people don't want to make their stuff free/open source. Even when it's free people still think you're somehow entitled and overcharging.