Live data from Hacker News

Faking a JPEG

ty-penguin.org.uk

1–10 of 97 posts

Re: Faking a JPEG

#2
Given that current LLMs do not consistently output total garbage, and can be used as judges in a fairly efficient way, I highly doubt this could even in theory have any impact on the capabilities of future models. Once (a) models are capable enough to distinguish between semi-plausible garbage and possibly relevant text and (b) companies are aware of the problem, I do not think data poisoning will be an issue at all.

Re: Faking a JPEG

#4
> I felt sorry for its thankless quest and started thinking about how I could please it.

A refreshing (and amusing) attitude versus getting angry and venting on forums about aggressive crawlers.

Re: Faking a JPEG

#5

> I felt sorry for its thankless quest and started thinking about how I could please it. A refreshing (and amusing) attitude versus getting angry and venting on forums about aggressive crawlers.

Helped without doubt by the capacity to inflict pain and garbage unto those nasty crawlers.

Re: Faking a JPEG

#6
> So the compressed data in a JPEG will look random, right?

I don't think JPEG data is compressed enough to be indistinguishable from random.

SD VAE with some bits lopped off gets you better compression than JPEG and yet the latents don't "look" random at all.

So you might think Huffman encoded JPEG coefficients "look" random when visualized as an image but that's only because they're not intended to be visualized that way.

Re: Faking a JPEG

#9
post #2

Given that current LLMs do not consistently output total garbage, and can be used as judges in a fairly efficient way, I highly doubt this could even in theory have any impact on the capabilities of future models. Once (a) models are capable enough to distinguish between semi-plausible garbage and possibly relevant text and (b) companies are aware of the problem, I do not think data poisoning will be an issue at all.

Yes, but you still waste their processing power.
Post reply on HN