EExverse
Ch 201:44:24Beyond Text — Modalities

Point the camera (video input)

Advanced Voice + camera: point at the world and just ask about what it sees.

mental model

Combine true voice with a live camera and you can point your phone at anything and talk about it. Under the hood it's likely sampling ~an image per second, but it feels like streaming video. "Quite magical, super simple to use" — the most natural on-ramp for non-power-users (show your parents).

In his demo he pans the camera around his desk and the model identifies everything, live:

“What book is this?” — “That's Genghis Khan and the Making of the Modern World by Jack Weatherford.” … “And what is this?” — “That's an Aranet4, a portable CO₂ monitor… you're at 713 ppm, that's generally okay.” … “What is this map?” — “That looks like a map of Middle-earth from The Lord of the Rings.”

Andrej Karpathy·1:45:00
mental modelHow it works under the hood

He suspects it doesn't truly consume video — it "still just takes image sections, maybe one image per second." But from the user's side it feels like you can stream video and it makes sense. Native video input is coming; today it's fast image sampling.

tipThe best demo for newcomers

He doesn't use it much himself (his queries are targeted, about code), but it's what he'd show his parents — point the camera, ask naturally, no interface to learn.

The load-bearing points

  • Advanced Voice + camera = point at the world and ask.
  • Likely ~1 image/second under the hood, but feels like live video.
  • Effortless and natural — the best entry point for non-technical users.
Try it yourself

Give it a tour

Open voice + camera and pan around a room, asking it to identify objects, read a label, or estimate a measurement.

Show the point

Notice how it keeps up “live.” That's fast image sampling feeling like streaming video.

? Check yourself
1How does “video input” most likely work today?
2Who is video input especially good for?

Non-technical users — you just point the camera and talk. No interface to learn, which makes it the most natural on-ramp to LLMs.