Show it images (OCR & vision)
Point it at labels, lab results, receipts, memes — it can read, transcribe, and reason.
Images become tokens too — an image is chopped into patches and fed in like text. So you can show the model things: a supplement label, a blood panel, an ingredient list, a math problem, a meme. It'll OCR the text, transcribe it, and reason about it. Same rule as always: it can misread, so verify the transcription on anything that matters.
His examples are everyday: photograph a nutrition/supplement label and have it transcribe the facts; upload a blood-test lipid panel and ask it to read and explain the numbers.


He also has it weigh a toothpaste's ingredients (which are functional vs. cosmetic) and explain an image-based joke — the "a group of crows is a murder" meme — showing it reasons about images, not just reads them.

Vision OCR is strong but not perfect. On anything consequential — dosages, lab values, contracts — confirm the transcription against the original before you act on the model's reasoning.
The load-bearing points
- Images are tokenized as patches — the model can “see.”
- Great for OCR + reasoning: labels, lab panels, ingredients, math, memes.
- It reasons about images, not just transcribes them.
- Verify the transcription on anything that matters.
Photograph the physical world
Take a photo of a label, receipt, or lab result and ask the model to transcribe and explain it.
Show the point
Then check its transcription against the original — this is the OCR-verification habit.
Explain the joke
Give it a meme or a diagram and ask what makes it work.
Show the point
If it explains the pun or the structure, you're seeing genuine visual reasoning, not just text extraction.
Verify the transcription. Vision is strong but can misread digits or fine print — confirm against the original before acting on anything important.