BuonApp Join the waitlist

Every tracker guesses.
BuonApp shows you the guess.

Photo calorie apps love a confident number. A confident-looking number without its uncertainty is theatre — so we built the opposite. This page is the long version, for people who want to know exactly what they’re trusting.

We measured it — then rebuilt the app around what we found.

We promised a benchmark on this page: 50+ weighed meals, the outcome published whatever it said. Before any results, we set our own bar: median error of 25% or less, and the shown range holding the true value at least 7 times in 10. Our first benchmark missed it. Everything below is what we built because of that.

What we ran: the core benchmark was 1,200 analyses — 50 weighed meals × three runs × four frontier AI models, with and without a one-line dish description; the remaining ~900 analyses stress-tested configurations of the winning engine — the upgrades we expected to help did not improve accuracy in those tests, so we shipped the setup that measured no less accurately, and faster. Publishing what didn’t work is part of the deal. This is an internal product benchmark, not a peer-reviewed model comparison; which models we tested, and how we configure them, is withheld for competitive reasons. That limits independent reproduction of the comparison — so treat it as a product benchmark, not a public ranking of AI models. What we owe you is the outcome for the app in your hand — reported plainly below, without hiding the miss.

The outcome you’re holding: we moved the app onto the models that measured best, in both tiers. On our weighed set, the Pro engine’s median error is 29.7% from the photo alone; the free engine’s is 32.9% (median means half the tested meals came out closer than that, half further away). For context, published evaluations of leading models report 23–36% typical error on their own test sets — different metrics and different meals, so the figures aren’t directly comparable; what they agree on is that photo AI guesses, for everyone. BuonApp shows you the guess.

Getting past the photo’s limits — fast. One line naming the dish improved every model we tested, by half a point to nearly six points of median error — so “Name your dish” now leads the logging flow (still optional; honesty includes not nagging you). From there, each correction path stays close to the meal: name it, list the ingredients, slide any single ingredient’s amount, or let the AI re-read the meal with what you told it. The AI’s guess is the starting point; extra context is what can sharpen it.

Our own range display failed our own test. The range the app used to show came from the AI’s confidence — and it contained the weighed truth between 8% and 14% of the time. On this page we wrote that a confident-looking number without its real uncertainty is theatre. It was our theatre too. The ranges you see now are built from this benchmark’s actual error distribution, sized to the meal: the band the app calls “likely” held the weighed truth 7 times in 10, and in split-half validation — built on half the meals, tested on the other half — about 6.5 to 7 in 10. One kitchen so far; that caveat is real.

What the meals taught us. In this 50-meal set, portion size — not plate complexity — was where the errors lived: ingredient count had little relationship with error, and small snacks and light bites are where the errors concentrate — plate-sized meals measure considerably better. The AI also overestimated our lean test kitchen almost across the board — a caveat, not a victory: cook with more oil and the bias could flip, which is exactly why we refused to ship a blanket “correction” learned from one kitchen. Instead the app learns from your corrections: after 8 corrected meals it checks for a consistent personal bias and may offer an adjustment — capped, shown to you, never applied silently.

A single photo leaves every system uncertain about portions and hidden ingredients — ours included. Our benchmark measured how large that uncertainty really is, and the app now shows it instead of hiding behind one exact-looking number. When we can assemble weighed ground truth from more kitchens than ours, we’ll run this again against the same bar.

Headline results are published above; the protocol and per-meal records are retained and available to researchers on request: [email protected].

The same promise, inside the app

This page doesn’t live apart from the product: the app’s own promise screen — Settings → Our promise — carries the benchmark too, sitting next to the behaviour you can check it against.

The top of the app's Our promise screen: the honest calorie tracker intro, then Why we show ranges — the 50 kitchen-weighed meals benchmark, the 25% bar our first benchmark missed with Pro at 29.7% and free at 32.9%, and the measured likely band that held the weighed truth about 7 times in 10

What the research actually says

23–36%

Two recent quantitative evaluations reported mean energy errors of roughly that size in leading multimodal AI models, under different test conditions — and they tend to underestimate, especially on big portions. Our own benchmark — caveats included — is earlier on this page.

Portion size is one of the hardest parts of single-photo estimation: the image cannot reveal weight, depth or hidden ingredients. That’s why BuonApp makes the missing context easy to add — one line of text, one slider.

Sources: [1] Performance Evaluation of 3 Large Language Models for Nutritional Content Estimation from Food Images — the best models (GPT‑4o, Claude 3.5 Sonnet) at ≈36% mean absolute percentage error on energy, across 52 standardised, portion-controlled photographs. [2] Evaluating Large Multimodal Models for Nutrition Analysis (ACETADA benchmark) (accepted at IEEE BHI; extended version linked) — the leading closed models clustered at ≈23–24% energy MAPE from the image alone, on dietitian-verified free-living meal photos; added context trimmed 1.5–2.6 points off, to ≈21–23% — though the strongest context configurations included verified food lists a photo alone can’t provide. [3] Chung et al., PLOS Digital Health 3(11): e0000665, 2024 — an observation and interview study: eighteen dietitians reviewing seven-day photo diaries relied on patterns, portions and eating context rather than single-photo calorie reads — the design evidence behind our refine loop. Metrics and settings are each study’s own (MAPE on controlled vs free-living photos; expert review) — which is exactly why we ran our own weighed-meal benchmark — it’s above.

How honesty looks in the app

It even counts the cooking fat pooled on the plate

“Small amount of oil/butter sauce pooled on plate included as extras.”

From another real analysis — a different dinner. That’s the level of honesty we think a food diary owes you: every estimate arrives as an editable reconstruction — the items it detected, the assumptions it’s making, and what it deliberately didn’t count (the bread roll in the napkin stayed out until you say you ate it). The assumptions are the estimate’s stated premises, there for you to correct — not a transcript of the AI’s inner reasoning.

Assumptions list from a real analysis: two sea bass fillets estimated at 110 g each, coastal greens named as samphire, spinach and chard, the cooking fat and the oil pooled on the plate counted, and the bread roll in the napkin noted as not part of the main dish

It’s honest about your burn, too

Your plan assumes a certain daily burn from the questionnaire — but real weeks have skipped workouts and sick days. BuonApp records what your devices estimated (WHOOP, Apple Health, or your phone’s steps) and reconciles: the day review shows the device-estimated energy balance next to the plan’s assumption, and the week review says it plainly — “your plan assumed ≈16,500 kcal out; your devices put it around ≈14,800 — your deficit was probably smaller than planned.” Plans can flatter; your devices’ own numbers keep them honest — we show you both.

Where it struggles — and what happens then

One more disclosure, same spirit: the pregnancy and breastfeeding additions follow the widely used U.S. clinical figures (roughly +340 kcal in the second trimester, +450 in the third, +400 while breastfeeding). Guidelines differ by country — the UK’s NHS, for instance, is more conservative (about +200 kcal, final trimester only). A registered dietitian’s review of these figures is on our pre-launch checklist, and either way the app never runs a deficit in pregnancy — and your midwife always outranks any app. And as of our June 2026 check of eight leading trackers, none displayed a meal-level range.

Join the waitlist →