Live data from Hacker News

LLMs: Intelligence vs. Cost

openteams.com

31–40 of 48 posts

Re: LLMs: Intelligence vs. Cost

#31
post #29
post #10

Complaining about "bad charting" and posting a chart with y-axis that doesn't start at 0 is kinda weird.

Not at all, as long as it's labeled as such. Coming from engineering/science, this is common. What is bad is starting at 0, showing an indicator of a gap, and suddenly starting at 30 or whatever after the gap.

It gives you a wrong perspective, especially if you are distracted, on model capabilities: Fable 5.1 is not 30% better than Sol, but is the very first impression you get when you look at the first graph.

If I'm not wrong OAI tried a similar trick when GPT5 was announced ... they have been criticized a lot.

Re: LLMs: Intelligence vs. Cost

#32
post #31
post #29

Earlier quoted context omitted.

Not at all, as long as it's labeled as such. Coming from engineering/science, this is common. What is bad is starting at 0, showing an indicator of a gap, and suddenly starting at 30 or whatever after the gap.

It gives you a wrong perspective, especially if you are distracted, on model capabilities: Fable 5.1 is not 30% better than Sol, but is the very first impression you get when you look at the first graph. If I'm not wrong OAI tried a similar trick when GPT5 was announced ... they have been criticized a lot.

> It gives you a wrong perspective, especially if you are distracted, on model capabilities

Only if you aren't schooled in reading graphs. It's a given that you always have to look at the axes when interpreting a graph.

How exactly would you zoom into a section of a graph and just show that section?

Re: LLMs: Intelligence vs. Cost

#33
post #32
post #31

Earlier quoted context omitted.

It gives you a wrong perspective, especially if you are distracted, on model capabilities: Fable 5.1 is not 30% better than Sol, but is the very first impression you get when you look at the first graph. If I'm not wrong OAI tried a similar trick when GPT5 was announced ... they have been criticized a lot.

> It gives you a wrong perspective, especially if you are distracted, on model capabilities Only if you aren't schooled in reading graphs. It's a given that you always have to look at the axes when interpreting a graph. How exactly would you zoom into a section of a graph and just show that section?

Not starting at zero has long been used entirely intentionally in order to mislead people, particularly the general public.

Re: LLMs: Intelligence vs. Cost

#34
post #8

This looks great! I also think speed should be part of the metric (i.e. how long does the model take to actually solve a task). For me, I prefer to run expensive models such as Sol on light reasoning, which usually gives me good answers with quick responses. For my style of coding (quick back-and-forths and corrections) it makes a big difference if a model comes back in 1-2 minutes compared to 5-10, and I am happy to…

The more dimensions you take into consideration, the larger the proportion of the Pareto frontier becomes across all distributions, and making it harder to choose.

Re: LLMs: Intelligence vs. Cost

#35
post #32
post #31

Earlier quoted context omitted.

It gives you a wrong perspective, especially if you are distracted, on model capabilities: Fable 5.1 is not 30% better than Sol, but is the very first impression you get when you look at the first graph. If I'm not wrong OAI tried a similar trick when GPT5 was announced ... they have been criticized a lot.

> It gives you a wrong perspective, especially if you are distracted, on model capabilities Only if you aren't schooled in reading graphs. It's a given that you always have to look at the axes when interpreting a graph. How exactly would you zoom into a section of a graph and just show that section?

> Only if you aren't schooled in reading graphs ...

So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"?

> How exactly would you zoom into a section of a graph and just show that section?

When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it.

Re: LLMs: Intelligence vs. Cost

#36
post #35
post #32

Earlier quoted context omitted.

> It gives you a wrong perspective, especially if you are distracted, on model capabilities Only if you aren't schooled in reading graphs. It's a given that you always have to look at the axes when interpreting a graph. How exactly would you zoom into a section of a graph and just show that section?

> Only if you aren't schooled in reading graphs ... So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"? > How exactly would you zoom into a section of a graph and just show that section? When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it.

> So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"?

Oh absolutely - as other commenters have pointed out.

> When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it.

For the majority of the time charts have existed, your "good practice" would have been impossible. Charts have historically been static images (e.g. published in a journal). So there have been conventions on how to depict them - and at times it is very appropriate to start from something other than 0.

Here's an article from the UK's Office For National Statistics:

https://digitalblog.ons.gov.uk/2016/06/27/does-the-axis-have...

Re: LLMs: Intelligence vs. Cost

#37
post #36
post #35

Earlier quoted context omitted.

> Only if you aren't schooled in reading graphs ... So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"? > How exactly would you zoom into a section of a graph and just show that section? When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it.

> So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"? Oh absolutely - as other commenters have pointed out. > When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it. For the majority of the time charts have existed, your "good practice" would have been impossible. Charts have…

I agree on the past, when you have limited resource and have to print something on paper that you cannot recall to fix you have to carefully choose the layout.

But that era is gone since decades, nowadays, given how easy it is, it's a shame to not provide log scale switch and zoom&pan capabilities.

Re: LLMs: Intelligence vs. Cost

#38
post #37
post #36

Earlier quoted context omitted.

> So we can say the same about the authors "AA’s plot is misleading" claim, he is "not schooled in reading graphs"? Oh absolutely - as other commenters have pointed out. > When building a chart is good practice to provide log scale switch and zoom&pan capabilities, so the reader can decide how to look at it. For the majority of the time charts have existed, your "good practice" would have been impossible. Charts have…

I agree on the past, when you have limited resource and have to print something on paper that you cannot recall to fix you have to carefully choose the layout. But that era is gone since decades, nowadays, given how easy it is, it's a shame to not provide log scale switch and zoom&pan capabilities.

Eh no. Expecting a non-programmer to be able to generate this is very elitist. A static image is still the standard.

Re: LLMs: Intelligence vs. Cost

#39
post #38
post #37

Earlier quoted context omitted.

I agree on the past, when you have limited resource and have to print something on paper that you cannot recall to fix you have to carefully choose the layout. But that era is gone since decades, nowadays, given how easy it is, it's a shame to not provide log scale switch and zoom&pan capabilities.

Eh no. Expecting a non-programmer to be able to generate this is very elitist. A static image is still the standard.

The author's job description is literally "Staff Software Engineer at OpenTeams. Dask maintainer."!

Re: LLMs: Intelligence vs. Cost

#40
> This is fine in most cases, but for open-weights models it can be a lot more expensive than what the exact same model can be rented for from third-party API providers.

Filter by quantization, and most providers will have the same price. There is some "base" price even for open-weight models. Anything cheaper means some tricks on the provider's side.

Post reply on HN