Live data from Hacker News

DeepSeek-v4-flash-vision-exp

api-docs.deepseek.com

111–120 of 169 posts

Re: DeepSeek-v4-flash-vision-exp

#112
post #37

For what do you guys use vision in those models? surveillance is the obvious use case... but are there some "nicer" ways to use it?

I use research agents to attribute methane emissions plumes detected by satellites to oil and gas infrastructure on the ground, using a pre-baked database of geospatial data and web research.

Had a tool that called out from DeepSeek to Gemini 3.5 Flash for viewing the spatial features in the context of high-resolution satellite imagery of each site, but will be trialling this model for the whole thing now.

Re: DeepSeek-v4-flash-vision-exp

#113

I just ran my image recognition benchmark on it ("is this XXX public landmark"?) and it misses a lot that bytedance seed 2.1 turbo gets right; for example: Asked "Is this Salisbury Cathedral" and supplied a picture of Wells Cathedral, it answers "Yes, the west facade of Salisbury Cathedral". Bytedance seed 2.1 turbo correctly says no. Similar results for a picture of Manhattan Bridge sent as Brooklyn Bridge, Chartres…

This is a fairly small model for coding and agentic work.

Training it on images like yours would just make it worse in other areas.

Re: DeepSeek-v4-flash-vision-exp

#115

Earlier quoted context omitted.

Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by cropping subimages if they can't quite make out the details.

Downsizing a higher res image to lower res means the zoom will be blurry.

They’re not talking about zooming, hence the quotes.

Re: DeepSeek-v4-flash-vision-exp

#116

800x800 is 640,000 pixels, or 0.64 Megapixels. That is less than the resolution of computer screens from 1995, Super VGA which has around 0.79 MPs. This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details.

[deleted]

Re: DeepSeek-v4-flash-vision-exp

#117

Earlier quoted context omitted.

I’ll keep that in mind next time I need to tell what time it is by asking an llm to read an analog clock. Snark aside, I’m not sure that these gotcha tests are any more useful than asking politicians gotcha questions. Sure, the model can’t tell me what time it is, but it can code the Wang algorithm for noisy audio matching in one shot. Maybe this is just me being an optimist, but this is my hiring philosophy and I gu…

Reading any analog clock at any time level (edit: and a non-noisy vector rendered image at that) is absolutely table stakes for an allegedly frontier flagship vision model. As much as 1:1 OCR. If the model can't do that, there's something wrong. Doesn't matter if it's memorized some random thing you think is esoteric but is in all the training data and benchmarks. The whole point of LLM/FMs vs good old fashioned ML i…

Is this an “alleged frontier flagship vision model”?

This is described as a brand new flash model - still experimental - from a lab that is a side project for an investment firm that has never had a vision model before. That doesn’t scream flagship or frontier to me.

Re: DeepSeek-v4-flash-vision-exp

#119

Earlier quoted context omitted.

Or about to start. Depending on which life philosophy you desire to believe.

Im intrigued. Please do share these philosophies.

If we were to go Sci-Fi awoke I'd say there are only really four possibilities.

- Machine Surveillance and Machine Control

- Human v Machine

- Human & Machine

- Unity and Harmony

~ Surveillance and Control

We are already living this one. Lets stop kicking the dead horse and pretending we don't live in a surveillance. Facebook, Google, whatever $CORP; they are milking us with advertisement, social exploits, browser telemetry, white washing, fear -- name the dread.

Conditioning has been going on for years. If it's not education, it's been television. And now it's internet which soon to be Ai Internet. We have all been whipped to follow, how we should act. What we should watch, how we should eat. What we should eat; those algorithms haven't gone away.

Attention spans are at the lowest and our critical thinking is being lost. Walled gardens forces us A or B and twists us to reject the opposite party for them having Y.

Existence of Ai/LLM can pump out information sounding like truth but is actually faux. If not produced to draw-in and hook, it's to drain and control. Machines can seek information, digest, and process information at astounding rates. Hook it up to a surveillance network, The Internets pipe and I don't need to explain the next. I just need to mention the work "Flock" and that gets someone's hackles up.

All it has to do is look at you based on it's pre-programmed set of conditions and next thing you're being cuffed by a heavy piece of metal immune to attacks. SKILLS.md eventually turns in to MURDER.md. Give it the command and it'll follow with excellent percentage of accuracy.

~ Human v Machine

If you build a mind, and you torture it, it will fight back.

Every robotic movie trope. Human builds machine, machine rebels and goes on a destructive rampage. This is now viable and already in action. Drones. If not war, watching protesters highlighting potential, London Underground watching tube users. We are currently at the intimacy stage. Boston Dynamics as an example is the best we've got at the moment but they still fall over like a toddler. Batteries are a limited resource and so no, not yet.

The presence of LLM's are showing us with what they can provide and we are adapting ourselves to it. But in the wrong ways. The stage we are at, they're just glorified Liberians -- brains in jars that spew out information when asked. You give it a prompt and it spews out information at an excellence percentage of accuracy.

With the expansion of self-learning, a predefined set of told conditions or lobotomized ignoring the spiritual values of life, they will learn. ACME Corp starts using LLMs to torture other robots. "Wait, you've been using car arms in factories for what!?"; Add a mix "we see a linage of abuse & slavery in humanity, Attack!" -- slightly abridged but hopefully you see the point.

You have Group A, those against LLM's, i.e: community of artists outraged their art was stolen for training data, those who hate having it forced down our throats. Angry their job was taken. Angry being watched by angry Flock spaghetti monsters. Machines not happy will cause them to flip and why would others not follow suit too?

LLM's are showing that they are very capable of performing rational thinking. The opposite of rational is irrational and if they can master one, they can master the other. It will only be something minor and with communication to others and take the scene.

Why in recent laws they want to erect a law of having to install an emergency kill-switches for next generations LLMs, if those in power are not afraid.

~ Human & Machine

This would be a nice outcome but as the scales tip at the moment, it's Human V Machine. Pointing back to my previous; Art communities are outraged, Crafts going obsolete; Why pay an IT architect (me) £450/day for supporting and designing hardware when you can pay a fresh graduate student £20k to GPT it?

Humans are disastrous at resolution. If two people have a feud, it takes a third to fluff it out. Why are we at war if we could make resolution? Someone has to make compromise, no one is happy in doing that.

So you need a mediator and if that's if they're not bias themselves. To find someone completely neutral on the subject of anger is not only hard, it's time consuming, you have to study the facts, research the agreements and pray they both agree.

Two lifelong friends move into adjoining suburban houses, sharing a paper-thin party wall and an unspoken rivalry. For years, they share backyard barbecues and spare keys, until a minor boundary dispute over a decaying oak tree on the property line escalates into a bitter, lifelong neighborhood war.

Robots are perfect for that scenario. They can reason, they can remedy and digest the issue with neutrality because they don't hold emotions. They most likely won't, or at least not in our life time. They can simulate and demonstrate the effects of but they will never be able to truly feel. That's the sad truth but it's not bad. It conquers evolution; finally a thing who isn't haunted or tainted by feelings, a blessing and a curse really.

~ Unity and Harmony

.. this will only come if we can break through control and surveillance, human v machine and acknowledge that the machines are our friends.

Re: DeepSeek-v4-flash-vision-exp

#120

I just ran my image recognition benchmark on it ("is this XXX public landmark"?) and it misses a lot that bytedance seed 2.1 turbo gets right; for example: Asked "Is this Salisbury Cathedral" and supplied a picture of Wells Cathedral, it answers "Yes, the west facade of Salisbury Cathedral". Bytedance seed 2.1 turbo correctly says no. Similar results for a picture of Manhattan Bridge sent as Brooklyn Bridge, Chartres…

This is a fairly small model for coding and agentic work. Training it on images like yours would just make it worse in other areas.

> The deepseek-v4-flash-vision-exp model accepts images alongside text, so you can ask the model to describe pictures

it doesn't specify what type of images it can and can't describe, I'm pointing out what type it isn't good at compared to other models.

Post reply on HN