Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

11–20 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#12

> This is expensive—the pelican cost 25 cents! Engineers get unbelievably silly about evaluating costs of things. "The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.

they're comparing to similar capability llm models, not humans. If one dishwasher does job at similar quality as another dishwasher, but using 30% more water and energy, you wouldn't compare to how much it costs human to do the same work, it would make no sense.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#13
It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Re: Kimi K3, and what we can still learn from the pelican benchmark

#14
post #12

> This is expensive—the pelican cost 25 cents! Engineers get unbelievably silly about evaluating costs of things. "The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.

they're comparing to similar capability llm models, not humans. If one dishwasher does job at similar quality as another dishwasher, but using 30% more water and energy, you wouldn't compare to how much it costs human to do the same work, it would make no sense.

> they're comparing to similar capability llm models, not humans

25 cents is 10x the cost of 2.5 cents, but it's still extremely cheap for the product. It's very much the wrong comparison for a world where the primary competition is still humans who need to eat, and it treats percentage differences as more important than absolute differences when the opposite is true.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#15

> This is expensive—the pelican cost 25 cents! Engineers get unbelievably silly about evaluating costs of things. "The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.

Would anyone pay a human to create an SVG of a pelican riding a bike?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#16
post #15

> This is expensive—the pelican cost 25 cents! Engineers get unbelievably silly about evaluating costs of things. "The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.

Would anyone pay a human to create an SVG of a pelican riding a bike?

In fact humans get paid to create SVGs of all kinds of things.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#17

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Pelicans and bikes can be in the training set without them training for this specific benchmark.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#18
post #12

Earlier quoted context omitted.

they're comparing to similar capability llm models, not humans. If one dishwasher does job at similar quality as another dishwasher, but using 30% more water and energy, you wouldn't compare to how much it costs human to do the same work, it would make no sense.

> they're comparing to similar capability llm models, not humans 25 cents is 10x the cost of 2.5 cents, but it's still extremely cheap for the product. It's very much the wrong comparison for a world where the primary competition is still humans who need to eat, and it treats percentage differences as more important than absolute differences when the opposite is true.

Well first of all, any non-trivial use of LLMs is going to be orders of magnitude more tokens than this, usually multiple millions at minimum. Benchmarks are just benchmarks after all.

Secondly, humans vs LLMs are apples vs oranges. It makes no more sense to compare human costs vs LLM costs as it would have to compare human costs vs calculator costs. LLMs are faster and cheaper but extremely different beasts with different limitations. Humans do not one-shot SVGs of pelicans riding bicycles, and they do not charge in tokens.

Comparing LLM cost efficiency is not something that should need to be defended. It's quite straightforward and reasonable...

Re: Kimi K3, and what we can still learn from the pelican benchmark

#20

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

They can be in the training set but not deliberately trained for. There may be a lot of people posting pelican svgs, but not typically because they're high quality and worth replicating.
Post reply on HN