DeepSeek-v4-flash-vision-exp
11–20 of 169 posts
Re: DeepSeek-v4-flash-vision-exp
#12This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details.
Re: DeepSeek-v4-flash-vision-exp
#13Re: DeepSeek-v4-flash-vision-exp
#14DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising. > Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens. > Before inference, every image is automatically resized: > - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio. > - Larger…
Oof 800 by 800 kills a lot of use cases
Maybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient
Re: DeepSeek-v4-flash-vision-exp
#15800x800 is 640,000 pixels, or 0.64 Megapixels. That is less than the resolution of computer screens from 1995, Super VGA which has around 0.79 MPs. This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details.
Re: DeepSeek-v4-flash-vision-exp
#16Re: DeepSeek-v4-flash-vision-exp
#17DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising. > Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens. > Before inference, every image is automatically resized: > - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio. > - Larger…
Oof 800 by 800 kills a lot of use cases
Re: DeepSeek-v4-flash-vision-exp
#18Re: DeepSeek-v4-flash-vision-exp
#19DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising. > Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens. > Before inference, every image is automatically resized: > - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio. > - Larger…
Oof 800 by 800 kills a lot of use cases