Live data from Hacker News

SubQ: a sub-quadratic LLM with 12M-token context

subq.ai

21–30 of 47 posts

Re: SubQ: a sub-quadratic LLM with 12M-token context

#21
> The core idea is content-dependent selection. For each query, the model selects which parts of the sequence are worth attending to, and computes attention exactly over those positions.

I don't know if this will help for things like understanding code, where the all relevant parts can be the file of 1000 lines that we are analyzing, and where every token is relevant in understanding recursion, loops, function calls, etc.

This sounds like it would be great to do SSA before passing things along to a code model like claude code.

Let me know if I misunderstood

Re: SubQ: a sub-quadratic LLM with 12M-token context

#23
post #6

I’m very surprised this isn’t getting more attention. Am I missing something? It seems at or above SOTA on the given benchmarks, doesn’t have context rot, is orders of magnitude faster, and uses less compute that current transformer models. I suppose it’s just an announcement and we can’t test it ourselves yet.

Yes you're missing something: the snake oil.

Re: SubQ: a sub-quadratic LLM with 12M-token context

#24
post #6

I’m very surprised this isn’t getting more attention. Am I missing something? It seems at or above SOTA on the given benchmarks, doesn’t have context rot, is orders of magnitude faster, and uses less compute that current transformer models. I suppose it’s just an announcement and we can’t test it ourselves yet.

> Am I missing something?

Yes, this product doesn't exist.

And the last time a company claimed something similar it disappeared after taking money from investors.

Re: SubQ: a sub-quadratic LLM with 12M-token context

#25

Earlier quoted context omitted.

This seems super cool if as described, but I'm sure you can understand the skepticism. Do you anticipate having any kind of public accessible chat interface for testing in the near future? Also, what, if any, benefits are there for smaller context windows? Is there still a material improvement in cost to serve under say 256k? I'm curious about the broader implications for the space beyond improvements for very large…

I do, for sure! Yes, we have a few product rollouts lined up. The differentials for latency are posted in our blog post, so that should provide an idea of where the scaling law differentials kick in.

> I do, for sure! Yes, we have a few product rollouts lined up.

When, more or less?

Re: SubQ: a sub-quadratic LLM with 12M-token context

#26
post #6

I’m very surprised this isn’t getting more attention. Am I missing something? It seems at or above SOTA on the given benchmarks, doesn’t have context rot, is orders of magnitude faster, and uses less compute that current transformer models. I suppose it’s just an announcement and we can’t test it ourselves yet.

We are SOTA in some ways and not in others, continuously working to make it better! We need a little more time to scale, as we are working on things like disaggregated prefill, etc., the norms of large-scale model infra. I am happy to answer any questions!

I have questions.

Can you back up your claims?

Why did you not release the white paper in parallel with the product?

Feels really fishy.

Re: SubQ: a sub-quadratic LLM with 12M-token context

#27
post #19
post #18

- magic.dev claimed 200M context window and it's been two years since and no real product yet. - They are admitting that this is built on top of a Chinese model[1] - They committed a huge chart crime with the Y axis of a chart comparing to Opus on their website that I can't find anymore (Too embarrassing to keep?). The delta between their score (81%) vs. Opus (87%) on SWE bench was hugely minimized - They named the c…

Ah, I nearly forgot about magic.dev. I took a quick peek to check up on them. Welp, last social/blog activity was in... 2024. But hey, their careers page still says they're hiring! So they must be doing just fine.

They did raise over $500M

Re: SubQ: a sub-quadratic LLM with 12M-token context

#28
Funny how they claim a 12M context window, yet all benchmarks are cherry picked with a 1M context window. Also, nobody has questioned how they did a training run before receiving funding. SoTA training runs cost well above $10M, yet no mention of funding prior to yesterday, interesting.

Re: SubQ: a sub-quadratic LLM with 12M-token context

#29

Assuming this is real and much better than existing linear attention methods as advertised, not launching with a technical report is a big miss. Edit: their blog post ( https://subq.ai/how-ssa-makes-long-context-practical ) does go pretty in-depth about it Edit 2: the fact that they're going straight for an end-to-end coding product on day 1 is very ambitious. Other speed/efficiency-oriented AI companies (Cerebras an…

you really call this 1-minute blog post "in-depth"?

Re: SubQ: a sub-quadratic LLM with 12M-token context

#30
post #18

- magic.dev claimed 200M context window and it's been two years since and no real product yet. - They are admitting that this is built on top of a Chinese model[1] - They committed a huge chart crime with the Y axis of a chart comparing to Opus on their website that I can't find anymore (Too embarrassing to keep?). The delta between their score (81%) vs. Opus (87%) on SWE bench was hugely minimized - They named the c…

The chart crime was not intentional! We will not make you wait two years. We are O(n), not O(1). O(1) would unfortunately be an impossibility. We may as well do infinite context at that point!
Post reply on HN