Point the camera (video input)
Advanced Voice + camera: point at the world and just ask about what it sees.
Combine true voice with a live camera and you can point your phone at anything and talk about it. Under the hood it's likely sampling ~an image per second, but it feels like streaming video. "Quite magical, super simple to use" — the most natural on-ramp for non-power-users (show your parents).
In his demo he pans the camera around his desk and the model identifies everything, live:
““What book is this?” — “That's Genghis Khan and the Making of the Modern World by Jack Weatherford.” … “And what is this?” — “That's an Aranet4, a portable CO₂ monitor… you're at 713 ppm, that's generally okay.” … “What is this map?” — “That looks like a map of Middle-earth from The Lord of the Rings.””
Andrej Karpathy·1:45:00
He suspects it doesn't truly consume video — it "still just takes image sections, maybe one image per second." But from the user's side it feels like you can stream video and it makes sense. Native video input is coming; today it's fast image sampling.
He doesn't use it much himself (his queries are targeted, about code), but it's what he'd show his parents — point the camera, ask naturally, no interface to learn.
The load-bearing points
- Advanced Voice + camera = point at the world and ask.
- Likely ~1 image/second under the hood, but feels like live video.
- Effortless and natural — the best entry point for non-technical users.
Give it a tour
Open voice + camera and pan around a room, asking it to identify objects, read a label, or estimate a measurement.
Show the point
Notice how it keeps up “live.” That's fast image sampling feeling like streaming video.
Non-technical users — you just point the camera and talk. No interface to learn, which makes it the most natural on-ramp to LLMs.