OpenAI Shows a Program That Makes a Minute of Realistic Video From a Sentence
Sora produces clips — a woman walking down a neon Tokyo street, woolly mammoths in the snow — with consistent people and objects from one frame to the next, which is the part that had been hard. It is not being released.
OpenAI showed a video-generation program today called Sora. Type a description — "a stylish woman walks down a Tokyo street filled with warm glowing neon" — and it produces up to a minute of video that looks, at first viewing, like it was filmed. The examples the company posted include mammoths crossing a snowy meadow, a drone shot of a coastal town, and a close-up of an old man's face. They are not perfect: hands are wrong, a person occasionally walks the wrong way on a treadmill, a cookie doesn't show a bite mark. They are far better than anything public.
Video has been the hard case. Programs have been able to make a good still image for two years, but a video is many still images that must agree with each other — the same woman, the same jacket, the same buildings, sixty times a second — and earlier systems produced shimmering, melting things. Sora's people and objects stay themselves for the length of the clip, which is the thing that had not been done.
OpenAI is not releasing it. The company says it is giving access to a small group of safety testers and to some filmmakers, and that it wants to understand the risks before opening it up. Those risks are not hard to imagine in an election year: the tools that made fake images of a pop star last month could, with this, make fake footage of a candidate.
The company is also not saying what videos it trained on, a question that is now the subject of lawsuits over its text programs. Whether the answer includes YouTube is one many people are asking.