There is a lot of hate in the comments but there is some merit to the post existing: 1. Even if the task is unreasonable, it is good to showcase that the LLM will perform poorly - warning not to be used for diabetes. 2. As it is a probabilistic model, the approach was to execute it multiple times and look at the distribution. They also tried to minimize variance: "All at the lowest randomness setting these models off…
He asked AI to count carbs 27000 times. It couldn't give the same answer twice
201–210 of 329 posts
Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#202It’s just an impossible problem. Photons don’t provide sufficient information to determine calories (at least not in any way they could practically be captured). Inside that sandwich could be drenched with olive oil or it could be hollow cheese with lettuce. It’s impossible to tell.
As a human, in the photo of that sandwich I see 4 slices of bread and 4 slices of cheese (distributed unevenly). I have no idea about the weight of the bread, flour type or its sugar content. I don't know the type of the cheese, dimensions of the slices or total amount of cheese inside the bread. I don't know if there is butter or anything else inside. I can guess the size of the plate as a size reference but I can't…
Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#203Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#204Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#205And this... really was and (will be) the only way for this to ever work.
Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#206> 42.9 units of insulin from a single photo. That’s not a rounding error. That’s a potential fatality.
Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#207Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#208There's an incredibly serious lack of education with how LLMs & carb-counting works. This entire article would be better suited to astrology.com than hackernews. When I opened it up, I assumed the author would have at least attempted a calculation service, maybe even placed something like the size of the meal into an actual model, using the integration of pre-existing tools that are (slightly more) accurate. Hell - m…
Oh! Do the vendors offer trainings to make sure the users understand how LLMs work? If not, surely, the LLM itself is trained to know its limitations and politely decline in situations like that?...
The #1 use case for this tech is "here's a problem I don't feel like solving, let's have a computer do magic". It's how it's advertised on TV, it's how it promoted in the software I already use. Food preparation? Travel planning? Shopping? Tutoring your children? You can do anything now!
I just talked to a realtor who will make a killing on a real estate transaction. Instead of offering human insights, they sent me "AI reviews" of several properties. The AI has never been to any of these properties and has no idea how they actually look like. But I guess it's how we operate now as a society.
If you go to eBay, every other listing description for used items is AI-generated. This is an official platform feature for sellers. The AI doesn't know the condition of the item or what's included or missing. Doesn't matter, it's magic. It's AGI, it will figure it out.
Most of the uses of AI I encounter as a consumer are like that, and the companies selling this tech are 100% complicit.
Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#209Earlier quoted context omitted.
> Then the correct answer is “I can’t tell.” From the paper they're using structured JSON schema mode opposed to freeform answers, so it can't. Models do typically caveat their answer for questions like this, in my experience.
They'll qualify their answers in English but as the article mentions, if your prompt asks for a confidence score, that "uncertainty" doesn't translate into low numerical confidence.
> They'll qualify their answers in English but [...]
That the default user-facing chat as a normal user would use it gives a warning is the key part IMO. I don't think expectations of there being no "wrong way" to use the model can necessarily extend to API usage with long custom system prompt and restricted output format.
Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice
#210There's an incredibly serious lack of education with how LLMs & carb-counting works. This entire article would be better suited to astrology.com than hackernews. When I opened it up, I assumed the author would have at least attempted a calculation service, maybe even placed something like the size of the meal into an actual model, using the integration of pre-existing tools that are (slightly more) accurate. Hell - m…
The obvious meme to invoke here is: - AI will solve all of our problems - No not like that! Are the trillion dollars sloshing around the AI economy well-invested if the refrain is always “you’re holding it wrong”? So we’re trying to define, through trial and error, what problems “AI” will actually solve, and this paper is one of the many cobblestones on that road.
"AI can solve this one problem, but it needs X, Y, Z, because it's not a omnipotent god entity"
"I tried it without any of those things and it didn't work - this is worthless tech!"
I don't know if more accurate calorie counting using AI exists - but it's like being upset that the screwdriver isn't gluing wood. AI is far more than frontier LLMs.