Live data from Hacker News

Native Sparse Attention

aclanthology.org

31–32 of 32 posts

Re: Native Sparse Attention

#31

Earlier quoted context omitted.

Yeah I confess I rewrote history and crashed the stock market. Then ran out of juice just as I was about to kill Hitler.

Do not try to signal intelligence by being sardonic or intentionally being obtuse. Actively avoiding the point someone is making rather than confronting it head on is beneath you. ChatGPT o1 was made generally available in December 2024 DeepSeek r1 open weights were released in January 2025

My response wasn't intended to signal my intelligence as much as to show the overall lack of serious thought in your reply since you seem to be ignoring the fact that there was a 15%+ overall drop in the stock market, much higher in tech sector. But let's spell it out for you:

o-1 was mimicking the newest "explain your solution step by step" prompts which were proven to be more effective at the time.

ds-v1 came up with an actual chain of thought, imitating meandering and self doubt which sometimes went on for a while creating entertaining loops and introducing a new class of halting problem and this became the de-facto standard. They also revolutionized the industry by programming cards directly via PTX.

Then all of huggingface implemented the paper and we got q4 versions that "thought".

Hope that jolted your memory without killing Hitler.

Re: Native Sparse Attention

#32
post #26

Earlier quoted context omitted.

How would people use deepseek to think "Capital is evil?" It was from a private hedge fund named "High Flyer," not a state university project or something.

Yes, exactly. How the heck? It makes no sense to me either, but you can certainly find plenty of laymen/not-in-the-know folks making those kinds of comments, often in non-technical spaces. Often the worst parts of the internet where discourse is non-existent. Human psychology allows for us to hold many contradictory positions all at once. Ideologies are the lens through which we view the world and it distorts our per…

Usually what I see is not that, but that Deepseek stole from American capital by training on the O1 release to acheive chain of thought, but there is a contradiction because o1 at the time didn't show its real chain of thought to train on.
Post reply on HN