Under the hood
How Pantrly reads your fridge
You point a phone at a crowded, badly lit shelf and get back a clean list. Here is what is actually happening, and why we would rather miss an item than invent one.
Marcus Feld, Co-founder and CTO · March 28, 2026 · 6 minute read
Typing out everything in your fridge is a chore, so most people will not do it. Pointing a phone at the shelf is not. That one difference is why photo recognition is core to Pantrly rather than a nice extra, and it is also one of the harder things we do.
A real fridge is a hard photo
Product photos are lit, staged and face-on. Your fridge is dark, crowded, and half the labels are turned away. A jar behind a carton is a guess, not a fact. So the model reading your photo is a vision model, and it is told to name only what it can genuinely see, in the plain words a shopper would use.
The rule we hold it to
A confident wrong item is worse than a missing one. If it cannot make out what something is, it leaves it off and lets you add it, rather than putting words in your fridge.
Different jobs, different models
Reading a photo and inventing a recipe are not the same task, and they do not want the same model. Pantrly is model-agnostic on purpose: reading your shelf goes to a vision model, turning the list into real meals goes to a balanced language model, and small tidy-up jobs go to a cheap fast one. Each step asks for a task, not a lab, so we can route to whatever is best and swap it without a rebuild.
The photo is not sent anywhere to be kept. It is read into a list and discarded. Your kitchen is yours.
None of this is visible when it works, which is the point. You point, you glance at the list, you cook. The engineering is in making a genuinely hard read feel like nothing at all.