Live data from Hacker News

Cohere's First Model for Developers

cohere.com

21–30 of 47 posts

Re: Cohere's First Model for Developers

#21
post #17

Wasn't aware that Cohere was still around but this release doesn't exactly instill confidence.

>Wasn't aware that Cohere was still around but this release doesn't exactly instill confidence. It's being kept alive because the Canadian government is desperate to have a local frontier lab and is willing to inject funding and force its adoption in government services, but leadership at Cohere is known to be weak in Canadian tech circles, and they pivoting to an enterprise-first market around production RAG rather…

It's embarassing? Awfully harsh!

Re: Cohere's First Model for Developers

#23
post #17

Earlier quoted context omitted.

>Wasn't aware that Cohere was still around but this release doesn't exactly instill confidence. It's being kept alive because the Canadian government is desperate to have a local frontier lab and is willing to inject funding and force its adoption in government services, but leadership at Cohere is known to be weak in Canadian tech circles, and they pivoting to an enterprise-first market around production RAG rather…

It's embarassing? Awfully harsh!

It's easy to be critical.

Re: Cohere's First Model for Developers

#26
post #19
post #13

Earlier quoted context omitted.

fwiw because of the relatively few activated params offloading to system RAM is quite feasible, you can see the endless amount of people doing this on r/localllama with qwen3.6 35a3b

I ran Gemma4 26B A4B on an 8yo PC with a fucking GTX and it did rather well.

Well, that's pretty impressive. Care to share your setup to do that? How much DDR3/DDR4 do you have, too?

Re: Cohere's First Model for Developers

#27

> Hardware (minimum): 1× H100 @ FP8 Cool to see this but seems like it would be pretty expensive to run

This is a 30B parameter model with 3B active. It should run performantly on a Mac with > 48GB RAM at 8bit precision.

Well that is like 3 USD/hour if you run it on a rented gpu

Re: Cohere's First Model for Developers

#28
Are these models trained from scratch or do they necessarily need distillation from bigger models to be competitive? It's usually the case that they're a small model for a family with a bigger model. In the first case, does anybody know what's the economy of training this 30B-A3B model vs. training a DeepSeek V4 Pro or Flash size of models (1.6T, 200 something B, less activated)?

Re: Cohere's First Model for Developers

#29
post #18
post #3

Earlier quoted context omitted.

its worse at code compared to qwen 3.6 coder.

How can it be worse than something that doesn't exist?

Sometimes non-existing is better than existing for unnecessary or harmful things. I know that is not what you mean but I just found it relevant in the age in which making new stuff is so fast and easy due to LLMs. Main enshitification would come, imo, not from bad things but for unnecessary things that nobody asked for.

Re: Cohere's First Model for Developers

#30
post #17

Earlier quoted context omitted.

>Wasn't aware that Cohere was still around but this release doesn't exactly instill confidence. It's being kept alive because the Canadian government is desperate to have a local frontier lab and is willing to inject funding and force its adoption in government services, but leadership at Cohere is known to be weak in Canadian tech circles, and they pivoting to an enterprise-first market around production RAG rather…

It's embarassing? Awfully harsh!

It really is. I’m very familiar with that as well.

It’s truly embarrassing how much hand-holding those guys have received from angels, investors, the government, etc. To the point where the same investors they’re going to pitch to are preparing their slides, telling them what to say during the presentation, and then approving them for even more funding afterward, lol.

That government part is corruption and illegal, by the way.

Actual usage on many of their APIs/models is painfully low, like in ... hundreds of DAUs. I don't blame them for this, but this is a "company" that should have died 2 years ago.

Post reply on HN