Live data from Hacker News

SubQ: a sub-quadratic LLM with 12M-token context

subq.ai

41–47 of 47 posts

Re: SubQ: a sub-quadratic LLM with 12M-token context

#42
post #18

- magic.dev claimed 200M context window and it's been two years since and no real product yet. - They are admitting that this is built on top of a Chinese model[1] - They committed a huge chart crime with the Y axis of a chart comparing to Opus on their website that I can't find anymore (Too embarrassing to keep?). The delta between their score (81%) vs. Opus (87%) on SWE bench was hugely minimized - They named the c…

The chart crime was not intentional! We will not make you wait two years. We are O(n), not O(1). O(1) would unfortunately be an impossibility. We may as well do infinite context at that point!

What’s keeping you from releasing paper and access to the model?

Re: SubQ: a sub-quadratic LLM with 12M-token context

#43
I'm usually okay with most LLM-assisted writing, but the amount of "it's not X. it's Y" style of phrases in https://subq.ai/how-ssa-makes-long-context-practical is disturbing.

Also, holy moly, the astroturfing.

But I'll still keep an eye on what they'll show up with in the next months. Sounds intriguing.

Re: SubQ: a sub-quadratic LLM with 12M-token context

#44
Don't let a C-suite marketing video blow your mind. They are trying to discover the new Transformer, that's not easy. 12 million token context with worse quality means this isn't going anywhere. Want to bet me bitcoin that we won't be talking about them in 1 year? Heck, they may have found something great, but the prior should be one of skepticism.

Re: SubQ: a sub-quadratic LLM with 12M-token context

#45
post #42

Earlier quoted context omitted.

The chart crime was not intentional! We will not make you wait two years. We are O(n), not O(1). O(1) would unfortunately be an impossibility. We may as well do infinite context at that point!

What’s keeping you from releasing paper and access to the model?

Model: - making sure it has been properly red-teamed, meets user preferences, etc. - it depends on what folks want in the model. Our original papers was mostly the technical blog post, but we decided to wait a little longer to see what else folks wanted and share more benchmarks

Re: SubQ: a sub-quadratic LLM with 12M-token context

#46
post #36

Earlier quoted context omitted.

The chart crime was not intentional! We will not make you wait two years. We are O(n), not O(1). O(1) would unfortunately be an impossibility. We may as well do infinite context at that point!

Good luck.

Thanks!

Re: SubQ: a sub-quadratic LLM with 12M-token context

#47

Earlier quoted context omitted.

What do you want in a whitepaper that was not in our blog post? There is time to add more before the whitepaper is released.

I'm not GP, but I would want a benchmark that actually tests the entire context window. A benchmark that only tests the first 128K tokens effectively tells us nothing about how well it works at its full capacity.

That makes sense! We are working on that.
Post reply on HN