Live data from Hacker News

He asked AI to count carbs 27000 times. It couldn't give the same answer twice

diabettech.com

221–230 of 329 posts

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#221

What a dumb article. The picture of the sandwich is essentially just a picture of bread. You can’t see what’s inside. A human wouldn’t be able to tell you. These are essentially AI hit pieces.

This is based on a real app that someone is selling to real people. It's not a hit against AI, but it is very much a legitimate hit against uses of LLMs for this purpose. Also, if LLMs worked as they are often advertised, they should have easily been able to answer "there isn't enough information in this picture to give you an accurate estimate. Try taking a picture of the label, or at least of the inside of the sand…

iAPS is not software you pay for. This is open source software to dose insulin where you assume full responsibility for the outcomes.

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#222

Earlier quoted context omitted.

One of the biggest gaps is that people don't understand that food labels are allowed by the FDA to be off by up to 20% in terms of the number of actual calories! In the real world you need to calibrate your behavior with the results. Are you gaining weight? You'll need to eat less if you want to lose any. You can do all the math with nutrition labels and macros you want but that's all theoretical. See this study belo…

A large part of the effectiveness in counting calories is that you pay more attention and make more conscious decisions and are less likely to "cheat" if you have to enter it in your food log. It's indeed like astrology. Simply thinking about personality traits and thinking through your life and your desires and goals and current situation is already beneficial to take charge and navigate your life.

That's an interesting bit, where reducing friction too much can eliminate the side effect that is actually driving the desired results.

Do you want to count calories, or do you want to lose weight? Sounds like it's possible to hyper-optimize calorie counting to the point that it becomes counter-productive...

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#223

There's an incredibly serious lack of education with how LLMs & carb-counting works. This entire article would be better suited to astrology.com than hackernews. When I opened it up, I assumed the author would have at least attempted a calculation service, maybe even placed something like the size of the meal into an actual model, using the integration of pre-existing tools that are (slightly more) accurate. Hell - m…

>> But the author just took pictures of food & expected a realistic response? Is this genuinely what amounts to a study in AI?

The aim of the study was to understand the variation in results returned by models and how that could cause risks for patients using those models. The main result was measuring within-model variation.

From the pre-print (https://www.diabettech.com/wp-content/uploads/2026/04/diabet...):

We aimed to characterise the within-image reproducibility of carbohydrate estimates from four large language model (LLM) vision APIs and to quantify the clinical risk for insulin dosing, stratifying accuracy by reference value quality.

Methods

Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01). The primary outcome was within- image variation (coefficient of variation [CV], range, distributional normality). Secondary outcomes included accuracy against reference values for nine images, stratified by quality tier (packet label, weighed/measured, portioned, or visual estimate). Clinical risk was translated at an insulin-to-carbohydrate ratio of 1:10.

>> I'd like to see this study done using any kind of actual grounding knowledge, seeing what mistakes AI makes when attempting to query ground truth from picture analysis - there would at least be an interesting result methodology in that

The ground truth was established by the author. There's an appendix in the pre-print (Appendix I) that describes the methodology. Methods are described in page 4 of the pre-print:

Reference values for accuracy analysis

For nine of the thirteen images, the author estimated the carbohydrate content using methods described in Appendix 1. Reference quality was categorised into four tiers:

Tier 1 (packet label): Carbohydrate values derived from manufacturer nutrition labelling. Two images (cheese sandwich, soup with bread) used bread with labelled carbohydrate content of 20 g per slice.

Tier 2 (weighed/measured): Portions directly weighed and cross-referenced with established composition data. Three images (Bakewell tart, bakery cookie, breakfast burrito).

Tier 3 (portioned): Portions estimated by the author (not weighed) and combined with USDA composition data. Three images (roast dinner, chilli con carne with rice, stuffed pork loin).

Tier 4 (visual estimate): Portions and composition estimated from visual inspection. One image (churros).

For the four restaurant dishes (pizza capricciosa, eggs benedict, crema catalana, paella), no reference value was established. These images were used for the primary reproducibility analysis only.

Carbohydrate values follow the EU convention with dietary fibre excluded.

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#224

What a dumb article. The picture of the sandwich is essentially just a picture of bread. You can’t see what’s inside. A human wouldn’t be able to tell you. These are essentially AI hit pieces.

And it can be made better easily. Take a picture of the nutrition label of the bread and cheese first and then feed in this picture and you should get way better results

The sandwich example is silly because you almost certainly know the fill nutiritonal info from the packaging so just use that...

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#225
post #169

Earlier quoted context omitted.

From the text of the article I believe the author is implying there are apps doing exactly this, and so this is why it was studied that way. Had the author written the article themselves rather than an LLM their motivation probably would have been clearer.

The author uses the prompts and method from an open-source app that connects to insulin pump, a medical device. I think AI food identification is an experimental feature in the app. > The prompt was adapted from the one used in the iAPS open-source automated insulin delivery system — it’s a real production prompt, not a toy example. https://github.com/Artificial-Pancreas/iAPS I think these are the prompts in the app:…

Exactly. This is not paid software. We assume full responsibility for outcomes when using it. There's a reason it's not on any app store. I'm glad features like this are being experimented with. Not how I would use AI to estimate carbs...

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#226
Context: there are a lot of very popular apps (e.g. Macrofactor) that are being promoted on social media and downloaded for exactly this feature (calculating nutrition based on pictures of food). The users don't understand that this is an impossible task. This is a scam that affects people's well-being, and it's good that there's data proving it.

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#227

Earlier quoted context omitted.

> But the author just took pictures of food & expected a realistic response? There are very popular apps on the App Store right now that are going viral among non-techie people that do exactly this, and they have no concept of how AI works. My wife was talking about one and I had to give her a reality check that the AI had no idea what ingredients were used to make the food. And she's a licensed nutritionalist. Studi…

licensed nutritionalist Nutritionist?

[flagged]

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#228

Earlier quoted context omitted.

> there are apps doing exactly this Yeah, for sure there are. And people will just ask ChatGPT as well. The funny thing is that for people who are just trying to lose weight without managing any health issues precisely, this type of extreme variance doesn't really matter, because the mere act of consciously quantifying food consumption is, based on my experience counting calories, the single biggest factor in success…

I actually think "just asking ChatGPT" is fine, because A) the data in these apps is suspect at best and B) the data behind calories is also pretty suspect (but we all play along because we can adjust other variables to make it all "work" well enough). Once or twice a year I spend a few weeks meticulously measuring ingredients/cooked foods and recording calories and on complex recipes apps are next to useless at gett…

I was complaining about AI generated clothes being misleading marketing, deceiving customers as to whether the garment even exists.

And then I learned that the pre-AI norms weren't any less fictional: they made an exemplar garment and did photoshoots, sure, but then they send the pictures and patterns to the lowest bidder factories with permission to make whatever edits are necessary to make it cheap and manufactureable. The whole thing was already a simulacrum.

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#229

Earlier quoted context omitted.

One of the biggest gaps is that people don't understand that food labels are allowed by the FDA to be off by up to 20% in terms of the number of actual calories! In the real world you need to calibrate your behavior with the results. Are you gaining weight? You'll need to eat less if you want to lose any. You can do all the math with nutrition labels and macros you want but that's all theoretical. See this study belo…

A large part of the effectiveness in counting calories is that you pay more attention and make more conscious decisions and are less likely to "cheat" if you have to enter it in your food log. It's indeed like astrology. Simply thinking about personality traits and thinking through your life and your desires and goals and current situation is already beneficial to take charge and navigate your life.

I'd take the opposite point of view: just thinking about life, desires, and goals is how people wind up paying $10 to Whole Foods for a slice of "healthy pizza" that is nutritionally identical to any other pizza but comes out of a cute stone oven and is displayed on a wooden platform next to green leafy plants. Vibes astrology is notoriously easy to exploit, both by your own sugar/salt/fat-seeking instincts and by unscrupulous commercial forces, let alone the two working together. The unique thing about calorie counting is that it cannot be exploited like vibes astrology. Not even with a 20% error margin, which is (probably not coincidentally) the caloric deficit targeted by standard dieting advice.

Re: He asked AI to count carbs 27000 times. It couldn't give the same answer twice

#230

Earlier quoted context omitted.

> If there are apps targeting people with diabetes that claims to count your carbs with AI, why haven't those been analysed? That would be a far more effective claim. Because the apps aren’t going to let you submit 29,000 automated requests for statistical analysis. And if you did, the authors of those apps would just release an update saying they changed models and try to dismiss the study. The vitriol against this…

You can commit statistical analysis on frontier models and still use commercial applications as an identifier & comparison. Criticism is not vitriol - it's possible to make a wider point about being taken aback by the lack of education within AI to the point that there's a critical mass of people using them for calorie counting; but there are many studies on effects of LLMs on psychology etc that are far more effecti…

Well, for me the comments that insist we don't need to study X because everybody knows LLMs can't do that is a very good justification to study exactly X.

Not to mention that this is now a standard thought-terminating cliché, where someone points out a use case where LLMs don't work at all well and irrate responses protest that LLMs aren't meant to be used in that way. Says who? If you ask an LLM a question and it answers it- then that's an LLM use case. If you can ask the same question many times and evaluate the results then that's an evaluation that is perfectly fine to make.

Post reply on HN