Is ChatGPT Accurate for Calorie Counting? What the Data Shows

8 min read

Illustration of a meal photo being sent to a chat assistant that replies with an estimated calorie count

Partly. In peer-reviewed tests ChatGPT identified the foods in meal photographs with about 93% precision, but its calorie estimates were off by roughly a third on average, and it underestimated more as portions got bigger. Tell it the weights and ingredients and the error drops by more than half — the input matters more than the model.

How ChatGPT actually estimates calories

Most people assume a lookup is happening: you type "chicken burrito", ChatGPT checks a nutrition database, and tells you what it found. That is not what happens, and almost everything else about its accuracy follows from this one fact.

ChatGPT is a language model with a vision front end. Asked for calories, it does what it does with any other question — produces the most plausible continuation given everything it has read. No row is being fetched from anywhere. The number comes out of association rather than measurement: the model has read an enormous amount of text about food, including nutrition tables, recipes and restaurant menus, so "medium banana, around 105 kcal" is a well-worn path and lands close to right. A photo enters the same machinery one step earlier — the vision side describes what's on the plate, and the language side reasons from that description.

That mechanism is a genuine strength as often as it's a weakness. Because nothing needs to match a database entry, you can describe food the way you actually eat it. "About a third of the lasagne my mother-in-law made, the cheesy corner piece" is a question ChatGPT will engage with seriously. Type that into a search box in a tracking app and you get nothing back.

It can search the web when it judges a question needs it, and it can run real calculations. But a casual "how many calories is this?" usually gets answered from the model's own associations — not from a source it names, and not from one you can go and check.

What the research actually found

Public benchmarks here are thinner than the topic's popularity suggests, and each one tests a single model version against a single set of meals. Still, as of mid-2026 several peer-reviewed evaluations exist, and they agree about the shape of the problem even where the numbers differ.

Naming the food: good

A 2025 study in Nutrients by O'Hara and colleagues uploaded 114 meal photographs drawn from an Irish national dietary survey — 38 weighed meals at three portion sizes — and asked ChatGPT-4 to name the foods, estimate their weights and estimate 16 nutrients. It identified foods with 93.0% precision. Correlations with the real values were adequate or good for every nutrient tested, with Spearman coefficients from 0.29 to 0.83. In other words, it's reliable at working out what you ate, and decent at ranking meals against each other.

Judging the portion: the weak point

The same study found good agreement on weight for small portions and poor agreement for medium and large ones, and the authors' conclusion was blunt: ChatGPT performed well at identifying foods and ranking meals, and poorly at estimating the weights of medium and large portions and at producing accurate nutrient estimates. There was poor agreement for 10 of the 16 nutrients, and the difference from the true value exceeded 10% for 13 of them. Against seven registered dietitians looking at the same photographs, agreement ran from 0.31 to 0.67 — poor to moderate.

A separate 2025 evaluation in Current Developments in Nutrition put 52 standardised photographs of foods and complete meals — everything weighed on a calibrated scale first — through three models. ChatGPT-4o's mean absolute percentage error was 36.3% for weight and 35.8% for energy; Claude 3.5 Sonnet matched it almost exactly, while Gemini 1.5 Pro was much worse at 64.2% for energy. The finding that generalises is the last one: all three underestimated systematically, and the underestimate grew as portions got larger. Big plates read low.

One number in the O'Hara study is easy to misread in ChatGPT's favour. Across all 114 photographs the median estimate was 525 kcal against an actual 524.5 kcal — a difference of 0.1%. That is not evidence that it gets meals right; it's evidence that the overestimates and underestimates roughly cancel out across a large sample. A tool can be almost unbiased in aggregate and still be wrong by a wide margin on your dinner.

Context beats the model

The single largest lever isn't which model you use — it's what you give it. A November 2025 study in Nutrients ran 195 dishes through ChatGPT-5 under escalating amounts of context. From the photograph alone, the mean absolute error for energy was 123.03 kcal. Add non-visual descriptors and it fell to 92.04 kcal. Add a detailed ingredient list and it fell to 53.33 kcal — well under half the photo-only error. Ingredients with no photograph at all did worse than ingredients plus photograph (66.51 kcal), which tells you the image is doing real work, just not the work most people assume it's doing.

Illustration of the calories a photo cannot see: oil being poured, butter melting in a pan, a spoonful of sugar and dressing on a salad
Identification is close to solved. Portion size and invisible ingredients are where the error lives.

You will also find a friendlier headline. A 2024 paper in the journal Nutrition concluded that ChatGPT provides "reasonably accurate and consistent nutritional information", its strongest result being that 97% of its energy values fell within a 40% difference from USDA figures. That's honestly reported and it isn't wrong — but look at the width of the window. Forty percent around a 700 kcal dinner spans 420 to 980 kcal. Whether that counts as accurate depends entirely on what you were planning to do with the number.

Where ChatGPT is genuinely good at this

It's a capable tool being asked to do a job it wasn't built for, which is different from being bad. Several things it does better than a purpose-built tracker:

  • Food that doesn't exist as a database entry. Leftovers, a friend's recipe, a dish from a cuisine no app has catalogued — it will reason about all of them rather than returning no results.
  • Arithmetic over a recipe. Give it the ingredient quantities and a number of servings and the division is genuinely reliable, because that step is calculation over numbers you supplied rather than a guess about the world.
  • Comparisons. It's better at telling you which of two meals is heavier than at telling you how heavy either one is, which is often the decision you're actually making.
  • Showing its work. It writes out the assumptions it made, so a wrong answer is diagnosable — you can see that it assumed 200 g of rice when you ate 350 g, and fix exactly that.
  • The question behind the question. "What would I swap to get this under 600?" is a reasoning problem, and reasoning is what it's for.

Where it breaks down

The failures cluster, and none of them are surprising once you know the mechanism.

  • Portions, as every study above found. A camera can't see grams, and no model can see the oil left in the pan, the butter under the egg or the dressing already tossed through the salad. Those are the items that move a total most.
  • Precision it doesn't have. An answer of "637 kcal" reads like a measurement. It's the midpoint of a wide distribution written in the register of a scale — and the studies say that distribution is wide.
  • It isn't a log. ChatGPT can remember things about you between conversations, but nothing is totting up your day. Unless you keep one long thread and explicitly ask it to add up, there's no running total, no yesterday to compare against, and no record you can scroll.
  • Nothing to check. Because it names no source by default, you can't tell whether "burrito, about 640 kcal" came from a chain's published figures or from an average of everything it has ever read about burritos. Both feel identical in the reply.
  • It fills gaps instead of asking. Told "chicken and rice", it will quietly decide whether the chicken was fried and move on. You can prompt it to ask first — but you have to know to.

What a database-backed tracker does differently

The alternative isn't a smarter AI. It's the same estimation step wired to something checkable: the AI reads the meal and identifies ingredients, and each ingredient is then matched to a reference entry — USDA FoodData Central for laboratory-measured whole foods, Open Food Facts for packaged products. The calories are then calculated from those entries and the estimated amounts, rather than recalled. That changes what you can verify, and it changes what happens to the answer afterwards.

Disclosure: Olivka is a database-backed calorie tracker that runs in chat apps, so we have an obvious interest in this comparison. It is also an AI tool, and it is subject to exactly the same portion-estimation limits described above — a photo doesn't become a scale because the ingredients get matched to a database afterwards. The row below that says so is the honest one.

What happensChatGPTA database-backed tracker
Where the number comes fromThe model's learned associationsReference database entries per ingredient
Portion sizeEstimated from your photo or wordingEstimated the same way — the shared weak point
What you can auditThe reasoning it wrote outThe ingredient list and the entries behind it
Fixing a wrong estimateReply in the thread; it lives in that chatEdit the entry; the corrected version is stored
Your daily totalOnly if you keep one thread and askKept automatically, on request

None of that makes the estimate exact, and we've written up where our own error bars come from rather than claiming a percentage we can't defend — the method is documented separately. If you want to see how the same idea plays out across the various bots that do this, we compared six of them, ours included, in a separate write-up.

You can no longer ask ChatGPT inside WhatsApp

This trips up a lot of advice still circulating. For about two years the obvious way to count calories in a chat was to message ChatGPT on WhatsApp. That stopped on 15 January 2026, when Meta's updated WhatsApp Business Solution terms took effect and barred providers whose primary offering is a general-purpose AI assistant. TechCrunch reported the change when it was announced, and confirmed on the day itself that OpenAI, Microsoft and Perplexity had switched their WhatsApp bots off. Meta's stated reason was volume, not quality: the assistants generated traffic the business platform was never designed to carry.

The rule targets open-domain assistants, not businesses using AI for one specific job, so single-purpose bots — including the ones that only log meals — are unaffected and still run there. If you want ChatGPT itself, it's the app, the website or Messenger; it isn't WhatsApp anymore.

How to get better numbers out of ChatGPT

If you're going to use it anyway — a perfectly reasonable choice — these are the changes the research supports, in rough order of how much they help.

  1. Give amounts. "200 g cooked rice, one egg" removes the step the model is worst at. This is the single biggest improvement available, and it's why the ChatGPT-5 study's error more than halved once ingredients were listed.
  2. Name what a photo can't see. Cooking oil, butter, sugar in the coffee, dressing on the salad. One extra clause — "fried in about a tablespoon of oil" — often moves the total by more than a hundred calories.
  3. Say how it was cooked. Grilled, boiled, breaded and deep-fried are three different foods to a nutrition database and one word to a photograph.
  4. Use brand names for packaged food. Branded products have published figures, and that's the closest thing to a fact in this whole exercise.
  5. Ask for a range, not a number. "Give me a low and high estimate and tell me what drives the spread" gets you an honest answer instead of false precision — and the spread itself is useful information.
  6. Make it state its assumptions before it answers, so you can correct the wrong one rather than the total.
  7. Keep a day in one thread if you want a daily figure, and ask it to re-add the total each time. It will do this well; it just won't do it on its own.

So — is it accurate enough?

For getting a rough sense of a meal you can't otherwise judge, yes, and it's excellent at explaining food. For a number you'll act on day after day, the evidence says be careful: a third off on average, worse for large portions, worse for mixed dishes where the fat is hidden, and no record of what you logged yesterday.

Some of that gap closes with better prompting, and some of it doesn't close at all — not for ChatGPT, not for a dedicated tracker, not for a kitchen scale. In the United States printed nutrition labels are legally allowed to deviate from the true value by up to 20 percent, and European rules permit similar tolerances, so exactness was never on the table. That floor is worth understanding before you judge any tool against it. What a purpose-built tracker adds isn't a better guess: it's a named source under the number, a card you can correct in one message, and a total that's still there tomorrow.

And the practical test is cheap. Photograph the same meal, ask ChatGPT, then ask a tracker built for the job, and see which answer you'd trust — and, more to the point, which one you'll still be using in three weeks. Consistency beats precision here, every time.

Frequently asked questions

Is ChatGPT accurate for calorie counting?

It is accurate at identifying food and unreliable at judging portions. In a 2025 Nutrients study it named foods in meal photographs with 93.0% precision, but agreement on nutrient content was poor for 10 of 16 nutrients. A separate 2025 evaluation measured a mean absolute error of about 36% for energy, with the underestimate growing as portions got larger.

Can ChatGPT count calories from a photo?

Yes, and it will usually name the dish correctly. The weakness is quantity: a photograph carries no weight information and hides fats absorbed during cooking. In a 2025 study of 195 dishes, the mean error for energy was about 123 kcal from the photo alone and about 53 kcal once a detailed ingredient list was supplied alongside it.

Is ChatGPT better than a calorie counting app?

It is better at describing unusual or homemade food, at reasoning about recipes and at explaining why a meal lands where it does. A tracking app is better at the boring parts: a named database behind each number, a stored entry you can correct, and a daily total that persists. They fail in different places, so the answer depends on which failure costs you more.

Can I use ChatGPT to count calories on WhatsApp?

No. Meta's updated WhatsApp Business Solution terms barred general-purpose AI assistants from the platform effective 15 January 2026, and ChatGPT, Microsoft Copilot and Perplexity all stopped answering there on that date. The rule targets open-domain assistants rather than single-purpose business tools, so bots built specifically to log meals still work on WhatsApp.

How do I make ChatGPT's calorie estimates more accurate?

Give it weights instead of photographs where you can, name the cooking fat and method, use brand names for packaged foods, and ask for a range plus the assumptions behind it rather than a single number. Research on contextual prompting found that adding a detailed ingredient list more than halved the energy error compared with an image alone.

← All posts