This is cool, but only the first part in extracting a ML model for usage. The second part is reverse engineering the tokenizer and input transformations that are needed to before passing the data to the model, and outputting a human readable format.
If you can't fix this with a little help from chatgpt or Google you shouldn't be building the models frankly let alone mucking with other people's...
Extracting AI models from mobile apps
181–190 of 250 posts
Re: Extracting AI models from mobile apps
#182I’m a huge fan of ML on device. It’s a big improvement in privacy for the user. That said, there’s always a chance for the user to extract your model, so on-device models will need to be fairly generic.
Re: Extracting AI models from mobile apps
#183I’m a huge fan of ML on device. It’s a big improvement in privacy for the user. That said, there’s always a chance for the user to extract your model, so on-device models will need to be fairly generic.
Maybe someday we will build a society where standing on the shoulders of giants is encouraged, even when they haven't been dead for 100 years yet.
Re: Extracting AI models from mobile apps
#184Earlier quoted context omitted.
> But lets not let that get in the way of hating on AI shall we? Can you please edit this kind of thing out of your HN comments? (This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html .) It leads to a downward spiral, as one can see in the progression to https://news.ycombinator.com/item?id=42604422 and https://news.ycombinator.com/item?id=42604728 . That's what we're trying to avoid here.…
Can you clarify this a bit. I presume you are talking about the tone more than the implied statement. If the last sentence were explicit rather than implied, for instance This article seems to be serving the growing prejudice against AI Is that better? It is still likely to be controversial and the accuracy debatable, but it is at least sincere and could be the start of a reasonable conversation, provided the respond…
Re: Extracting AI models from mobile apps
#185Earlier quoted context omitted.
Can you please edit swipes out of your HN comments? Your post would be fine with just the first sentence. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html .
What do you mean, "swipe"? The other person agreed they'd misjudged the article and apologised several hours before you wrote this.
Re: Extracting AI models from mobile apps
#186Earlier quoted context omitted.
Please don't cross into personal attack or otherwise break the site guidelines when posting here. Your post would be fine with just the first sentence. https://news.ycombinator.com/newsguidelines.html
Really... Some people do need to be taken down a peg here at times though.
Re: Extracting AI models from mobile apps
#187One thing I noticed in Gboard is it uses homeomorphic encryption to do federated learning of common words used amongst public to do encrypted suggestions. E.g. there are two common spelling of bizarre which are popular on Gboard : bizzare and bizarre. Can something similar help in model encryption?
Homomorphic, not homeomorphic
Re: Extracting AI models from mobile apps
#188Earlier quoted context omitted.
You're not going to get an answer you find agreeable, because you're hoping for an answer that allows you to continue to treat the tool as chattel, without conferring to it the excess baggage of being an individuated entity/laborer. You're either going to get: it's a technological, infinitely scalable process, and the training data should be considered what it is, which is intellectual property that should be being l…
No, you’re missing the point of copyright. The point of copyright is to protect an exclusive right to copy, not the right to produce original works influenced by previous works. If an LLM produces original works that are influenced by the training data, that is not a violation of copyright. If it reproduces the training data verbatim, it is.
Re: Extracting AI models from mobile apps
#189Earlier quoted context omitted.
There are a family of techniques, often called something like “distillation”. There are also various synthetic training data strategies, it’s a very active area of research. As for the copyright treatment? As far as I know it’s a bit up in the air at the moment. I suspect that the major frontier vendors would mostly contend that training data is fair use but weights are copyrighted. But that’s because they’re bad peo…
The weights are my training data. I scraped them from the internet
But there is a group of people, growing daily in influence, who utterly reject such principles as either worthy or useful. This group of people is defined by the ego necessary to conclude that when the stakes are this high, the decisions should be made by them, that the ends justify the means on arbitrary antisocial behavior (c.f. the behavior of their scrapers) as long as this quasi-religious orgasm of singularity is steered by the firm hand that is willing and able to see it through.
That doesn’t distress me: L Ron Hubbard has that.
It distresses me that HN as a community refuses to stand up to these people.
Re: Extracting AI models from mobile apps
#190Earlier quoted context omitted.
> If companies train on data they don't own and expect to own their model weights, that's hypocritical. Its not hypocritical to follow a line of legal analysis whoch holds that copying material in the course of training AI on it is outside the scope of copyright protection (as, e.g., fair use in the US), but that the model weights resulting from the training are protected by copyright. It maybe wrong, and it may be c…
If the resulting AI models are protected by copyright that invalidates the claim that AI models being trained on copyrighted materials is fair-use analogous to human beings becoming educated by exposure to copyrighted materials. Educated human beings are not protected by copyright, hence neither should trained AI models. Conversely, if a copyrightable work is produced based on work which itself is copyrighted, the re…
No one training (foundation) models makes that fair use argument by analogy, they make arguments that addresses the specific statutory and case law criteria for fair use (abd frequently focus on the transformative character of the use); its true that the analogy to a learning human argument is frequently made in internet fora by AI enthusiasts who aren't the people training models on vaat scraped datasets. That argument is bunk for a number of reasons, but most critically the fact that a human learning from material isn’t fair use, because a human brain isn’t treated as a fixed medium, so learning in a human brain isn’t legally a copy or derivative work that would violate copyright without the fair use exception, so it's not a use to which fair use analysis even applies, so you can't argue anything is fair use by analogy to that. But its moot to any argument for hypocrisy by the big model makers, because they aren’t using that argument to start with.