Live data from Hacker News

Extracting AI models from mobile apps

altayakkus.substack.com

241–250 of 250 posts

Re: Extracting AI models from mobile apps

#241
post #113

Earlier quoted context omitted.

If I take 1000 books and count the distributions of the lengths of the words, and the covariance between the lengths of one word and the next word for each book, and how much this covariance matrix tends to vary across the different books, and other things like this, and publish these summaries, it seems fairly clear to me that this should count as fair use. (Such a model/statistical-summary, along with a dictionary,…

'Should the resulting work be protected by copyright? I’m not entirely sure…' This has already been settled hasn't it? Don't companies have to introduce 'flaws' in order for data sets to be 'protected'? Just compiled lists of facts can't be protected. Which is why things like election result companies having to rely on NDAs and not copyright protections to protect their services on election night.

> This has already been settled hasn't it? Don't companies have to introduce 'flaws' in order for data sets to be 'protected'?

No, flaws are generally introduced to make it easier to detect copies; if multiple flawless reference works covering the same data (road maps of the same region, for instance) exist, each is copyrightable without flaws to the extent any would be with flaws, but you can't prove that someone else copied yours without permission if copying any of the others would give the same result. With flaws, gou can attribute the source that was copied more easily, but this isn't about being legally protected but about the practicality of enforcing that protection.

Re: Extracting AI models from mobile apps

#242
post #151

Earlier quoted context omitted.

Had to look it up, this seems to be the paper https://research.google/pubs/federated-learning-for-mobile-k...

they have a very "interesting" definition of private data on the paper. it's so outlandish that if you buy their definition, there's zero value on the trained data. heh. they also claim unsuppervisioned users typing away is better than tagged training data, which explain the wild grammar suggestions on the top comment. guess the age of quantity over quality is finally peaking. in the end it's the same as grammarly bu…

actually letting users type whatever they want is good because they are many dialects of english : chinglish, thailish, singlish, hinglish and so on.

they have made the system so general that it can handle any quirk users throw at it.

Re: Extracting AI models from mobile apps

#243
post #215

That's pretty cool! I am impressed by the Frida tool, especially to read in the binary and dump it to disk by overwriting the native method. The author only mentions APK for Android, but what about iOS IPA? Is there an alternative method for handling that archive?

Yeah, you can basically just unzip IPA files. Gaining them is hard though, I have a pathway if you are interested.

But the Objective C code is actually compiled, and decompilation is a lot harder than with the JVM languages on Android.

My next article will be about CoreML on iOS, doing the same exact thing :)

Re: Extracting AI models from mobile apps

#244

Earlier quoted context omitted.

You're not going to get an answer you find agreeable, because you're hoping for an answer that allows you to continue to treat the tool as chattel, without conferring to it the excess baggage of being an individuated entity/laborer. You're either going to get: it's a technological, infinitely scalable process, and the training data should be considered what it is, which is intellectual property that should be being l…

Hi. I like this post. There are some careful thoughts here. Can you help me to understand the term "chattel" as you used it? I never heard the term before I read your post, and I needed to Google for it: (in general use) a personal possession. (in law) an item of property other than freehold land, including tangible goods ( chattels personal ) and leasehold interests ( chattels real ). >>

Chattel, as I'm using it is in reference to the usage in distinguishing an "ownable piece of property" from an employee".

Namely, a magic, technologically reproducible box that can be applied almost as effectively as a human hireling, but isn't a human hireling is near infinitely more desirable in a capitalist system since the blackbox is chattel, the hired human is not. The chattel has no natural rights, no claim to self sovereignty, and is an asset that is legally extant by virtue of the fact it is owned by the owner.

Chattel that are flexible enough to replace the legal burdens incurred by hiring a human to do the same job, will naturally be converged upon due to the capitalistic optimation function of minimizing unit input cost for output over dollars and potential dollars as expressed through legal exposure.

Imagine you had two human like populations. One made of plastic that aren't considered humans but property. I.e. are chattel. Then you have a bunch of people with all the baggage that comes with it.

Hiring people/employing people is hard. Particularly in the U.S. and other jurisdictions where a great deal of responsibility for actually implementing regulations/ taxation/immigration and such is tacked onto being an employer/being able to hire.

As the gap between the capability of the chattel population closes in on the human population, the more economic and workload sense it makes for the system to improve the chattel population under our current optimization strategy, (given no pre-emptive work to cut off externality dumping). Humans are messy and complicated to work with. Often unpredictable. Chattel are easy to account for; especially when combined with "technical restraints". You have to fundamentally engage in negotiation with another human being to get them on board with working for you. You buy the chattel, and that's that. The chattel has no grounds to refuse service. Socially speaking, we don't even recognize it's outputs as carrying any social weight, or resistance as anything but malfunctions.

Economics is the science around using access to resources as a means to get other people to work with you. Being chattel means you can cut out entirely all that complexity. You are resource. Not people.

Unironically, we need to have an answer to whether or not we are going to consider a sufficiently complex function imitator as something that requires a classification above "chattel" or controls around how we apply it in order to not self-destruct the economic equilibria in which we purport to exist. Because all it takes is removing or sufficiently obstructing the flow of value down from individuals who accrete the most of these wunder-chattel to render things so top heavy, most of the constraints/invariants of our socioeconomic systems as we know them become invalidated.

That does not bode well for anyone.

Re: Extracting AI models from mobile apps

#245
post #238
post #233

Earlier quoted context omitted.

It's a variant of the "shill" argument, implying that the other person isn't posting in good faith.

Sorry, I don't follow. How do you arrive at that implication? Why would someone having a pecuniary interest in something necessarily make them insincere?

Perhaps one internet cliché can explain another: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

Re: Extracting AI models from mobile apps

#246

Earlier quoted context omitted.

> TL;Dr a request for more content is asking for a reverse engineering article to give you a full education on modal inference I don't understand what you mean: I have no clue about anything related to reverse engineering, but I ported the mistral tokenizer to Rust and also wrote a basic CPU Llama training and inference implementation in Rust, so I definitely wouldn't need an intro to model inference…

You're also not the person I'm replying to, nor do you appear in any of this comment chain, so I've definitely not implied you need an intro to inference, so I'm even more confused than you :)

I share the sentiment of the person you're responding to, and I didn't understand your response, that's it.

Re: Extracting AI models from mobile apps

#247
post #100

Earlier quoted context omitted.

"Making money" does not immediately invalidate fair use, but it does wave a big red flag in the courts' faces.

I would be more nuanced on this matter. As I understand, in the US, fair use allows media to write critiques of cultural artefacts (sorry, I cannot think of a better, broad term). For example, you can include small quotes from the film script when writing a critique of it without requiring permission from the owner of the copyright. And, until the World Wide Web arrived to the masses in the mid-1990s, most critiques…

> Another weird carve-out for copyright law in the US: parody. Honestly, I don't know if other jurisdictions allow parody in the same protected manner.

Germany: https://www.gesetze-im-internet.de/urhg/__51a.html (Though this explicit carve-out is a recent development, though generally speaking parodies were allowed even under the previous version of the law.)

Re: Extracting AI models from mobile apps

#248

Earlier quoted context omitted.

I would be more nuanced on this matter. As I understand, in the US, fair use allows media to write critiques of cultural artefacts (sorry, I cannot think of a better, broad term). For example, you can include small quotes from the film script when writing a critique of it without requiring permission from the owner of the copyright. And, until the World Wide Web arrived to the masses in the mid-1990s, most critiques…

> Another weird carve-out for copyright law in the US: parody. Honestly, I don't know if other jurisdictions allow parody in the same protected manner. Germany: https://www.gesetze-im-internet.de/urhg/__51a.html (Though this explicit carve-out is a recent development, though generally speaking parodies were allowed even under the previous version of the law.)

Your reference (link) is very impressive. Thank you to share. Honestly, I would struggle to provide the equivalent for US federal law (or court ruling). Are you a lawyer in DACH/Germany? How did you know to find this web page?

Re: Extracting AI models from mobile apps

#249
post #215

That's pretty cool! I am impressed by the Frida tool, especially to read in the binary and dump it to disk by overwriting the native method. The author only mentions APK for Android, but what about iOS IPA? Is there an alternative method for handling that archive?

Yeah, you can basically just unzip IPA files. Gaining them is hard though, I have a pathway if you are interested. But the Objective C code is actually compiled, and decompilation is a lot harder than with the JVM languages on Android. My next article will be about CoreML on iOS, doing the same exact thing :)

> My next article will be about CoreML on iOS, doing the same exact thing :)

Can't wait - thanks for writing it up!

Re: Extracting AI models from mobile apps

#250

Earlier quoted context omitted.

> Another weird carve-out for copyright law in the US: parody. Honestly, I don't know if other jurisdictions allow parody in the same protected manner. Germany: https://www.gesetze-im-internet.de/urhg/__51a.html (Though this explicit carve-out is a recent development, though generally speaking parodies were allowed even under the previous version of the law.)

Your reference (link) is very impressive. Thank you to share. Honestly, I would struggle to provide the equivalent for US federal law (or court ruling). Are you a lawyer in DACH/Germany? How did you know to find this web page?

> Are you a lawyer in DACH/Germany?

Nope :-), just a normal citizen, but sometimes I am curious enough to look up a law, plus sometimes I need to refer to/look up some law in my day job as a civil engineer, too.

When you need to do that, it's not too hard to stumble upon the existence of that page through some web searches, plus the German Wikipedia often links to that page, too (as well as to some alternative platforms run by private entities, which sometimes provide some added value, e.g. buzer.de provides change history since 2006, too, other pages link relevant court decisions, etc. etc. – but gesetze-im-internet.de is the official page run by the federal government itself).

Post reply on HN