EExverse
Ch 181:35:00Beyond Text — Modalities

Show it images (OCR & vision)

Point it at labels, lab results, receipts, memes — it can read, transcribe, and reason.

mental model

Images become tokens too — an image is chopped into patches and fed in like text. So you can show the model things: a supplement label, a blood panel, an ingredient list, a math problem, a meme. It'll OCR the text, transcribe it, and reason about it. Same rule as always: it can misread, so verify the transcription on anything that matters.

His examples are everyday: photograph a nutrition/supplement label and have it transcribe the facts; upload a blood-test lipid panel and ask it to read and explain the numbers.

chatgpt.com
ChatGPT transcribing a supplement nutrition label from a photo
OCR + understanding. Snap a label; it transcribes the supplement facts and can explain what each active does.
chatgpt.com
ChatGPT reading and explaining a blood-test lipid panel from an image
Read my results. A photographed lipid panel, transcribed and explained — one screenshot at a time.

He also has it weigh a toothpaste's ingredients (which are functional vs. cosmetic) and explain an image-based joke — the "a group of crows is a murder" meme — showing it reasons about images, not just reads them.

chatgpt.com
ChatGPT analyzing toothpaste ingredients as functional vs cosmetic
Reason, not just read. From a photo of an ingredient list to a functional-vs-cosmetic breakdown.
verifyCheck the transcription

Vision OCR is strong but not perfect. On anything consequential — dosages, lab values, contracts — confirm the transcription against the original before you act on the model's reasoning.

The load-bearing points

  • Images are tokenized as patches — the model can “see.”
  • Great for OCR + reasoning: labels, lab panels, ingredients, math, memes.
  • It reasons about images, not just transcribes them.
  • Verify the transcription on anything that matters.
Try it yourself

Photograph the physical world

Take a photo of a label, receipt, or lab result and ask the model to transcribe and explain it.

Show the point

Then check its transcription against the original — this is the OCR-verification habit.

Explain the joke

Give it a meme or a diagram and ask what makes it work.

Show the point

If it explains the pun or the structure, you're seeing genuine visual reasoning, not just text extraction.

? Check yourself
1How does the model take in an image?
2What's the one habit to keep with image OCR?

Verify the transcription. Vision is strong but can misread digits or fine print — confirm against the original before acting on anything important.