Live data from Hacker News

Using an open model feels surprisingly good

matthewsaltz.com

121–130 of 157 posts

Re: Using an open model feels surprisingly good

#121
post #19

The reason Claude code is so popular is because it’s really good at taking super vague human prose “Claude build me a million dollar SaaS”-type prompts and spitting out thousands of lines of code which cover tons of surface-level edge cases, build in tons of functionality, etc The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with…

I had what I thought was a completely ridiculous prompt: “Build me a replacement for Microsoft Word.” ChatGPT did what I thought was the right thing: it responded asking along the lines of “you don’t really want that do you?” and explained how Word is decades of corner cases, bug fixes, obscure features, and business processes that have been built around it. I was stunned that Claude simply started spinning its wheel…

It could just make an interface and have the buttons trigger prompts. (This would ofc be worse than the MS version)

Re: Using an open model feels surprisingly good

#122

Earlier quoted context omitted.

The honest thing is to add a note that tells readers that you are also selling the product you feel good about. Advertisement is not sharing, it's manipulation.

They already did, in my opinion. The text reads: > ...personal account. I work at Modal, and today we just launched Kimi K3 on managed endpoints, and I know Kimi K3 is supposed to be pretty solid, so instead of upgrading my Claude plan, I wanted to give it a try. ( I didn't directly contribute to this feature , so I haven't gotten to play with it yet.) Emphasis mine.

Yes they did. The point stands. He would never write Using an open model feels surprisingly bad even if that was what he would be thinking.

We can't trust for him to give honest opinion. The information value is near zero.

Re: Using an open model feels surprisingly good

#124

Earlier quoted context omitted.

They already did, in my opinion. The text reads: > ...personal account. I work at Modal, and today we just launched Kimi K3 on managed endpoints, and I know Kimi K3 is supposed to be pretty solid, so instead of upgrading my Claude plan, I wanted to give it a try. ( I didn't directly contribute to this feature , so I haven't gotten to play with it yet.) Emphasis mine.

Yes they did. The point stands. He would never write Using an open model feels surprisingly bad even if that was what he would be thinking. We can't trust for him to give honest opinion. The information value is near zero.

When I read the post I read nothing about Kimi-K3. I read about someone who used something open and being surprised about the effects of using something open.

I read about owning the data, knowing the path, being able to control what you have and the lightness it brings. That feeling came through his employer, which is something we can debate all year long, but I see nothing about advertising Kimi-K3 or managed endpoint or whatnot.

Maybe because I'm in all this for a long time, and having an AI capable server near is not something alien to me.

Anyway, in short, there's no AI related content for me in that post. I'm more on the experience side.

Re: Using an open model feels surprisingly good

#125
post #19

The reason Claude code is so popular is because it’s really good at taking super vague human prose “Claude build me a million dollar SaaS”-type prompts and spitting out thousands of lines of code which cover tons of surface-level edge cases, build in tons of functionality, etc The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with…

Personally I've found Deepseek v4 Flash to be as useful to me as Opus. But I don't do these silly one shot tech demos. I have the technical understanding to ask for exactly the change I want with the right terminology. I loaded up some credit on openrouter and it took me ages to hit $1 in spend.

I'm genuinely curious how you can do this - with limited context tokens, I find the output is often inconsistent with the project, or recreates functionality it missed, if I forget to point it at every pertinent object/library/etc it will need.

Re: Using an open model feels surprisingly good

#126
post #119

Earlier quoted context omitted.

It’s not a service, I hope nobody thinks this is being done ”for the good of humanity” or whatever. Doesn’t mean there’s not a benefit to people, but the financial politics behind it could turn out extremely good for China, and they are very well aware of just that.

Do American companies think about the "financial politics" of their decisions? I somehow really doubt it. Government takes sides and sometimes applies leverage to get companies to do what they want, but that's something separate from "it's a good business decision for Chinese labs to treat models as commodities". There's no need to think about geopolitics - American AI companies could easily take the same approach as…

Chinese labs are equivalent to the Chinese state, there’s no separation. Who do you think funds the Chinese labs?

Re: Using an open model feels surprisingly good

#127
post #19

The reason Claude code is so popular is because it’s really good at taking super vague human prose “Claude build me a million dollar SaaS”-type prompts and spitting out thousands of lines of code which cover tons of surface-level edge cases, build in tons of functionality, etc The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with…

Personally I've found Deepseek v4 Flash to be as useful to me as Opus. But I don't do these silly one shot tech demos. I have the technical understanding to ask for exactly the change I want with the right terminology. I loaded up some credit on openrouter and it took me ages to hit $1 in spend.

I found Mimo 2.5 Pro to be the right balance between cost and intelligence. It costs the same as DeepSeek v4 Pro, but it's better overall by a small margin.

Re: Using an open model feels surprisingly good

#128
post #119

Earlier quoted context omitted.

Do American companies think about the "financial politics" of their decisions? I somehow really doubt it. Government takes sides and sometimes applies leverage to get companies to do what they want, but that's something separate from "it's a good business decision for Chinese labs to treat models as commodities". There's no need to think about geopolitics - American AI companies could easily take the same approach as…

Chinese labs are equivalent to the Chinese state, there’s no separation. Who do you think funds the Chinese labs?

American labs are also funded by the US government. Are they equivalent to the US state? Do Chinese and US labs/researchers exist solely in the international sphere, without any domestic concerns or individual aspirations? The Chinese have a different culture and political system than Americans, but it's not like they're ants without free will or the ability to think. I highly doubt that DeepSeek is operating solely at the behest of the CPC or that they are in business to take down the US rather than to make money.

Re: Using an open model feels surprisingly good

#129

Earlier quoted context omitted.

Personally I've found Deepseek v4 Flash to be as useful to me as Opus. But I don't do these silly one shot tech demos. I have the technical understanding to ask for exactly the change I want with the right terminology. I loaded up some credit on openrouter and it took me ages to hit $1 in spend.

I'm genuinely curious how you can do this - with limited context tokens, I find the output is often inconsistent with the project, or recreates functionality it missed, if I forget to point it at every pertinent object/library/etc it will need.

Tbh I have only used it on small personal projects and some open source stuff so I haven’t done any crazy stress tests or benchmarks. But for my use case it was great.
Post reply on HN