Earlier quoted context omitted.
If you have access to any other model it can create create pi extension that fixes problem. At least worked for me.
Like a special parser? Would you mind elaborating?
DeepSeek v4
931–940 of 1001 posts
Re: DeepSeek v4
#932Seriously, why can't huge companies like OpenAI and Google produce documentation that is half this good?? https://api-docs.deepseek.com/guides/thinking_mode No BS, just a concise description of exactly what I need to write my own agent.
I spent only two minutes reading their documentation and it’s clear no one did any proofreading and it’s full of mistakes made by non-native speakers. Example: the second sentence on the first page says “softwares” but “software” is a mass noun that cannot be pluralized. Example: the third page about tokens has some zipped code to “calculate the token usage for your intput/output” and obviously “intput” should be “in…
> they could have even used their own LLM to edit their documentation to fix grammar issues
In my experience companies who do this rarely stop at using LLMs to fix grammar issues. It becomes full on LLM speak quite fast, especially if there isn’t a native English speaker in the room who can discern what’s good and bad writing.
Re: DeepSeek v4
#933Earlier quoted context omitted.
I, personally, have never been asked for an asshole scan, but I'm interested in providing one if you can point me to a company that's offering.
It's clear the OC was using hyperbole but we're honestly not too far off. Just a few examples: - Sam Altman & Worldcoin collecting everyone's eyeball scan - Discord attempting to roll out worldwide age & id verification - LinkedIn collecting data on your web browser extensions - WhatsApp collecting browser data via a local server running on device
Re: DeepSeek v4
#934Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…
As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed
Re: DeepSeek v4
#935Earlier quoted context omitted.
I reviewed how DeepSeek V4-Pro, Kimi 2.6, Opus 4.6, and Opus 4.7 across the same AI benchmarks. All results are for Max editions, except for Kimi. Summary: Opus 4.6 forms the baseline all three are trying to beat. DeepSeek V4-Pro roughly matches it across the board, Kimi K2.6 edges it on agentic/coding benchmarks, and Opus 4.7 surpasses it on nearly everything except web search. DeepSeek V4-Pro Max shines in competit…
I'd be interested to know when that Opus 4.6 baseline is from given their recent recognition of performance issues. Do you have a paper posted on this review?
Re: DeepSeek v4
#936Earlier quoted context omitted.
This is a pretty banal comment at this point. Open source is the term used in the LLM community. It's common and understood. Nobody is going to release petabytes of copyrighted training data, so the distinction between open source vs weights is a rather pointless one.
its still a pointed one. "open source" keeps being redefined by people with wealth and power to restrict our computing rights. eventually its just gonna be "proprietary microsoft code that runs on microsoft servers, but you can see a portion of the results"
It's entirely reasonable that this colloquial understanding would be applied to new categories such as AI models. I'm sure it'll be applied to many other things that don't fit the OSD either. That's just language for you.
Re: DeepSeek v4
#937Earlier quoted context omitted.
you are arguing with your belief instead of an objective truth. benchmark is more objective, if you don't agree with it, come up with a better one. but what you believe doesn't matter.
It was not a confrontational take. But all benchmarks are designed by humans, we are not that great at measuring intelligence. So it is somewhat subjective. I was just arguing with the word "objective". Not with the results per se.
Re: DeepSeek v4
#938Truly open source coming from China. This is heartwarming. I know if the potential ulterior motives.
Re: DeepSeek v4
#939Earlier quoted context omitted.
For coding Opus 4.5 in q3 2025 was still the best model I've used. Since then it's just been a cycle of the old model being progressively lobotomised and a "new" one coming out that if you're lucky might be as good as the OG Opus 4.5 for a couple of weeks. Subjective but as far as I can tell no progress in almost a year, which is a lifetime in 2022-25 LLM timelines
Opus 4.5 was released on Nov 24 last year. It’s only been 5 months!
That brief two week period when Opus could eat entire tickets was simultaneously fantastic and a bit alarming
Re: DeepSeek v4
#940Earlier quoted context omitted.
As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed
[flagged]
https://en.wikipedia.org/wiki/List_of_countries_by_GDP_%28no...