SubQ: a sub-quadratic LLM with 12M-token context
41–47 of 47 posts
Re: SubQ: a sub-quadratic LLM with 12M-token context
#42- magic.dev claimed 200M context window and it's been two years since and no real product yet. - They are admitting that this is built on top of a Chinese model[1] - They committed a huge chart crime with the Y axis of a chart comparing to Opus on their website that I can't find anymore (Too embarrassing to keep?). The delta between their score (81%) vs. Opus (87%) on SWE bench was hugely minimized - They named the c…
The chart crime was not intentional! We will not make you wait two years. We are O(n), not O(1). O(1) would unfortunately be an impossibility. We may as well do infinite context at that point!
Re: SubQ: a sub-quadratic LLM with 12M-token context
#43Also, holy moly, the astroturfing.
But I'll still keep an eye on what they'll show up with in the next months. Sounds intriguing.
Re: SubQ: a sub-quadratic LLM with 12M-token context
#44Re: SubQ: a sub-quadratic LLM with 12M-token context
#45Earlier quoted context omitted.
The chart crime was not intentional! We will not make you wait two years. We are O(n), not O(1). O(1) would unfortunately be an impossibility. We may as well do infinite context at that point!
What’s keeping you from releasing paper and access to the model?
Re: SubQ: a sub-quadratic LLM with 12M-token context
#46Re: SubQ: a sub-quadratic LLM with 12M-token context
#47Earlier quoted context omitted.
What do you want in a whitepaper that was not in our blog post? There is time to add more before the whitepaper is released.
I'm not GP, but I would want a benchmark that actually tests the entire context window. A benchmark that only tests the first 128K tokens effectively tells us nothing about how well it works at its full capacity.