EExverse
Ch 211:47:12Beyond Text — Modalities

Generate video

Text-to-video is incredible and evolving fast — a quick look, not a deep dive.

mental model

The last modality: text-to-video (Sora, Veo, and many others). Karpathy keeps this brief — he doesn't use it in his own work — but flags it as remarkable and moving extremely fast. Worth knowing it exists and roughly where the frontier is.

He shows a comparison tweet: many video models all asked to generate "a tiger in the jungle." All quite good; Veo 2 near state-of-the-art at the time.

Sora / Veo
A generated video still of a white tiger walking through a jungle
Prompt → video. The same “tiger in the jungle” across models — each with its own style and quality. The field is evolving month to month.
tipKnow it exists; expect churn

Different tools, different styles and quality, all improving fast. Unless you're in a creative field, it's enough to know the capability is here and improving — and to compare a couple when you need it.

The load-bearing points

  • Text-to-video exists and is advancing very fast (Sora, Veo, …).
  • Quality/style vary by model — compare when you need it.
  • Not part of Karpathy's daily workflow; know it's there.
Try it yourself

Sample the frontier

If you have access, prompt one text-to-video tool with a simple scene and note the quality and quirks.

Show the point

This modality changes fastest of all — treat anything you see as a snapshot in time.

? Check yourself
1How much weight does Karpathy give text-to-video?

A light touch — it's impressive and fast-moving, but not something he uses in his own (non-creative) work. Know it exists and where the frontier roughly is.