DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
41–50 of 343 posts
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#42It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#43Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#44Earlier quoted context omitted.
Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point. But you already know that.
Oh? What are the American model censorship tells?
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#45Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#46Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#47Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#48Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#49Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#50Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.