Live data from Hacker News

Claude vs. Gemini: Testing on 1M Tokens of Context

every.to

11–20 of 49 posts

Re: Claude vs. Gemini: Testing on 1M Tokens of Context

#15
post #2

https://archive.is/sb7D5

Does anyone else have trouble with the archive rendering of that? It seemed to also have the pop up.

Try one of these. They have the popup but you can dismiss it.

https://ghostarchive.org/archive/JlE5T

https://web.archive.org/web/20250812172455/https://every.to/...

Re: Claude vs. Gemini: Testing on 1M Tokens of Context

#18

i’m really curious how well they perform with a long chat history. i find that gemini often gets confused when the context is long enough and starts responding to prior prompts, using the cli or it’s gem chat window.

From my experience. Gemini is REALLY bad about context blending. It can't keep track of what I said and what it said in a conversation under 200K tokens. It blends concepts and statements up, then refers to some fabricated hybrid fact or comment.

Gemini has done this in ways that I haven't seen in the recent or current generation models from OpenAI or Anthropic.

It really surprised me that Gemini performs so well in multi-turn benchmarks, given that tendency.

Re: Claude vs. Gemini: Testing on 1M Tokens of Context

#19
IMO, a good contest between LLMs would be data compression. Each LLM is given the same pile of text, and then asked to create compact notes that fit into N pages of text. Then the original text is replaced with their notes and they need to answer a bunch of questions about the original text using the notes alone.

Re: Claude vs. Gemini: Testing on 1M Tokens of Context

#20
post #16

So sonnet-4 is faster than gemini-2.5-flash at long context. That is surprising. Especially since Gemini runs on those fast TPUS.

Anthropic also uses TPUs for inference.

Do they rent them from Google? Or are they a different brand?
Post reply on HN