Live data from Hacker News

Magika: AI powered fast and efficient file type identification

opensource.googleblog.com

201–210 of 262 posts

Re: Magika: AI powered fast and efficient file type identification

#201

Earlier quoted context omitted.

> You are asking what if this guy has "web crawl data" that google does not have? No, I'm asking if he has permission to redistribute these files.

Are you attempting to assert that use of these files solely for the purpose of improving a software system meant to classify file types does not fall under fair use? https://en.wikipedia.org/wiki/Fair_use

I'm asking a question.

Here's another one for you: Do you believe that all pictures you have ever taken, all emails you have ever written, all code you have ever written could be posted here on this forum to improve someone else's software system?

If so, could you go ahead and post that zip? I'd like to ingest it in my model.

Re: Magika: AI powered fast and efficient file type identification

#202

Earlier quoted context omitted.

Are you attempting to assert that use of these files solely for the purpose of improving a software system meant to classify file types does not fall under fair use? https://en.wikipedia.org/wiki/Fair_use

I'm asking a question. Here's another one for you: Do you believe that all pictures you have ever taken, all emails you have ever written, all code you have ever written could be posted here on this forum to improve someone else's software system? If so, could you go ahead and post that zip? I'd like to ingest it in my model.

Your question seems orthogonal to the situation. The three files posted seem to be the minimum amount of information required to reproduce the bug. Fair use encompasses a LOT of uses of otherwise copyrighted work, and this seems clearly to be one.

Re: Magika: AI powered fast and efficient file type identification

#203

Earlier quoted context omitted.

I'm asking a question. Here's another one for you: Do you believe that all pictures you have ever taken, all emails you have ever written, all code you have ever written could be posted here on this forum to improve someone else's software system? If so, could you go ahead and post that zip? I'd like to ingest it in my model.

Your question seems orthogonal to the situation. The three files posted seem to be the minimum amount of information required to reproduce the bug. Fair use encompasses a LOT of uses of otherwise copyrighted work, and this seems clearly to be one.

I don't see how publicly posting them on a forum is

> the minimum amount of information required to reproduce the bug

MAYBE if they had communicated privately that'd be an argument that made sense.

Re: Magika: AI powered fast and efficient file type identification

#204

Earlier quoted context omitted.

Your question seems orthogonal to the situation. The three files posted seem to be the minimum amount of information required to reproduce the bug. Fair use encompasses a LOT of uses of otherwise copyrighted work, and this seems clearly to be one.

I don't see how publicly posting them on a forum is > the minimum amount of information required to reproduce the bug MAYBE if they had communicated privately that'd be an argument that made sense.

So you don't think that software development which happens in public web forums deserve fair use protection?

Re: Magika: AI powered fast and efficient file type identification

#205
post #140

Earlier quoted context omitted.

> the current implementation can't be relied on IMO What's your reasoning for not relying on this? (It seems to me that this would be application-dependent at the very least.)

I'm not the person you asked, but I'm not sure I understand your question and I'd like to. It whiffed multiple common softballs, to the point it brings into question the claims made about its performance. What reasoning is there to trust it?

> It whiffed multiple common softballs

I must have missed this in the article. Where was this?

Re: Magika: AI powered fast and efficient file type identification

#206
post #140

Earlier quoted context omitted.

I'm not the person you asked, but I'm not sure I understand your question and I'd like to. It whiffed multiple common softballs, to the point it brings into question the claims made about its performance. What reasoning is there to trust it?

> It whiffed multiple common softballs I must have missed this in the article. Where was this?

[deleted]

Re: Magika: AI powered fast and efficient file type identification

#207
post #169
post #162

Earlier quoted context omitted.

Nobody's asking for perfection. But the AI is offering inexplicable and obvious nondeterministic mistakes that the traditional algorithms don't suffer from. Magika goes wrong and your fonts become audio files and nobody knows why. Magic goes wrong and your ZIP-based documents get mistaken for generic ZIP files. If you work with that edge case a lot, you can anticipate it with traditional algorithms. You can't anticip…

Where are you getting the non-determinism part from? It would seem surprising for there to be anything non-deterministic about an ML model like this, and nothing in the original reports seems to suggest that either.

> It would seem surprising for there to be anything non-deterministic about an ML model like this

I think there may be some confusion of ideas going in here. Machine learning is fundamentally stochastic, so it is non-deterministic almost by definition.

Re: Magika: AI powered fast and efficient file type identification

#208

Earlier quoted context omitted.

I don't see how publicly posting them on a forum is > the minimum amount of information required to reproduce the bug MAYBE if they had communicated privately that'd be an argument that made sense.

So you don't think that software development which happens in public web forums deserve fair use protection?

That's an interesting way to frame "publicly posted someone else's data without their consent for anyone to see and download"

Re: Magika: AI powered fast and efficient file type identification

#209

Earlier quoted context omitted.

So you don't think that software development which happens in public web forums deserve fair use protection?

That's an interesting way to frame "publicly posted someone else's data without their consent for anyone to see and download"

I notice you're so invested that you haven't noticed that the files have been renamed and zipped such that they're not even indexable. How you'd expect anyone not participating in software development to find them is yet to be explained.

Re: Magika: AI powered fast and efficient file type identification

#210

Earlier quoted context omitted.

Are you attempting to assert that use of these files solely for the purpose of improving a software system meant to classify file types does not fall under fair use? https://en.wikipedia.org/wiki/Fair_use

I'm asking a question. Here's another one for you: Do you believe that all pictures you have ever taken, all emails you have ever written, all code you have ever written could be posted here on this forum to improve someone else's software system? If so, could you go ahead and post that zip? I'd like to ingest it in my model.

It's three files that were scraped from (and so publicly available on) the web. That's not at all similar to your strawful analogy.
Post reply on HN