Live data from Hacker News

How we used GPT-4o for image detection with 350 similar illustrations

olup-blog.pages.dev

21–30 of 93 posts

Re: How we used GPT-4o for image detection with 350 similar illustrations

#21
post #15

Earlier quoted context omitted.

Also, having a blog post about image detection, and not showing a single picture in the whole post was quite frustrating.

Especially given the detailed description surely the author could just generate a similar image

Just thinking that. Spend a few minutes trying to have chatgpt generate some images with Dall-E 3. Flux would probably be better to get all the specific details but ya

Re: How we used GPT-4o for image detection with 350 similar illustrations

#22
Interesting approach to a a very interesting challenge, given how close the images supposedly are.

With the limited training data they have I'm surprised they don't mention any attempts at synthetic training data. Make (or buy) a couple museum scenes in blender, hang one of the images there, take images from a lot of angles, repeat for more scenes, lighting conditions and all 350 images. Should be easy to script. Then train YOLO on those images, or if that still fails use their embedding approach with those training images.

Re: How we used GPT-4o for image detection with 350 similar illustrations

#23

Earlier quoted context omitted.

> I can think of at least two businesses that can be competed in costs if the team can automate a good chunk of it. And which would those be?

Job applications, recruiter outreach and initial screening calls. I heard of an AI interviewer via voice chat on a reddit thread recently.

„AI” talking to an „AI”. What a time to be alive.

Re: How we used GPT-4o for image detection with 350 similar illustrations

#24

Interesting approach to a a very interesting challenge, given how close the images supposedly are. With the limited training data they have I'm surprised they don't mention any attempts at synthetic training data. Make (or buy) a couple museum scenes in blender, hang one of the images there, take images from a lot of angles, repeat for more scenes, lighting conditions and all 350 images. Should be easy to script. The…

They did.

> “ To address this limitation, we turned to data augmentation, artificially creating new versions of each image by modifying colors, adding noise, applying distortion, or rotating images. By the end, we had generated 600 augmented images per car.”

Re: How we used GPT-4o for image detection with 350 similar illustrations

#25
post #10

A bit tangential, but I think we will see a good chunk of small teams building competing products in different software business segments, by just doubling on productivity and offering a cheaper option due to less operational overhead (reads: paying engineers). I can think of at least two businesses that can be competed in costs if the team can automate a good chunk of it.

[dead]

Re: How we used GPT-4o for image detection with 350 similar illustrations

#26
post #4

Thanks for the “bitter lesson” news from the frontlines. Curious; did you experiment with 4o as the sole pipeline? And of course as I think you mention, it would be interesting to know if say llama 8b could do a similar job as well. Congrats on shipping.

they don't self-host the models, neither embedding nor last step llm. taking into account low load self-hosting likely would be more expensive. if so why not to use the best models.

Re: How we used GPT-4o for image detection with 350 similar illustrations

#29
post #19

Earlier quoted context omitted.

I find a lot of applied AI use-cases to be "same as this other method, but more expensive".

Better to spend $100 in op-ex money than spend $1 in cap-ex money reading a journal paper, especially if it lets you tell investors "AI." :p

Your engineers cost <$1/hr and understand journal papers?

Re: How we used GPT-4o for image detection with 350 similar illustrations

#30

I mean, cool tech, but why not just print a QR code next to each illustration?

Sounds like the client cared a lot about the user experience being smooth (they declined the solution of presenting the user with the narrowed-down choices of which car they took a picture of), and I think adding a bunch of QR codes to this aesthetic wall of car illustrations would not align with that goal.
Post reply on HN