Live data from Hacker News

DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date

deci.ai

41–45 of 45 posts

Re: DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date

#41

Earlier quoted context omitted.

It's fine if you don't care. Nobody needs to care about everything. But why post your comment? It doesn't add anything to discussion. For me I care about different companies reaching mistral level so that mistral or whoever is in top has to release the model weights else competitors will.

I posted my comment because I see a new one of these on the front page every morning and I was genuinely curious if there was anyone who cares about these posts? That's alongside the fact I'm almost certain that the account that posted this is a sockpuppet for the company advertising this LLM.

> That's alongside the fact I'm almost certain that the account that posted this is a sockpuppet for the company advertising this LLM.

This is the worst form of defence.

I still use Mistral as was clear from the last post. This model is not that good to make me switch. As I said I just want Mistral or the top model to remain open weight and so I want competition.

Re: DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date

#42
post #36

Key question: is this an entirely new LLM trained from scratch, or is it fine-tuned on top of an existing foundation model like Llama 2 or Mistral? I think it's a new foundation model, but the announcement doesn't make that clear to me.

They are trained from scratch on potentially different datasets. Their architectures are very similar though.

Re: DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date

#44

Earlier quoted context omitted.

Why is it that so many people here have to comment on basic marketing stuff? It is necesssary bravado because it's trying to convince people that it's worth using. And saying "Here is a model that is just a bit better than others" isn't going to do anything. Therefore it's necessary to have buzzwords such as groundbreaking.

Mistral is doing the exact opposite and everyone is talking about it. I'll never respect marketing like this because it's standing on the shoulders of thousands of others peoples work.

> Mistral is doing the exact opposite and everyone is talking about it.

Everyone is currently talking about it because they got a massive investment and people started posting links. However, no one was talking about them 2 weeks ago.

Re: DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date

#45
post #6

For someone who is a bit techie, but not familiar with much past the original PyTorch, can I get an explanation of why this should matter to me? Are we at the point yet where I can give one of these LLMs a list of characters and an introductory paragraph, and get a 100K word book out of it yet? Totally not asking because I have too many ideas and not enough time to write them all...

If the model has big enough context window, yes. If the 100k word book is good enough for you is another story.

GPT-4 turbo has a 128k token context, which might be good enough for your book.

You might also use another strategy: write a summary of what your book is about and the title of each chapter, then pass the summary plus the chapter title to the LLM and it will generate for you. This would allow you to go beyond the 100k word limit.

Post reply on HN