1. Home
  2. Glossary
  3. Image-to-video AI, defined, for car photos

What is image-to-video AI and does it work on car photos?

Image-to-video AI takes a single still photo and generates a few seconds of video that appears to move a camera around it. On a car it works well when the move is small — a slow push, a short pan — and badly when it is large, because the model has to invent the parts of the car the photo did not show.

Updated September 12, 2026 · by VroomVideo

Definition

An image-to-video model is a generative model that produces a short clip from a still image and a description of the motion. The first frame is the photo; the rest is generated.

Why the size of the move decides everything

Everything in the first frame is real. Everything the camera reveals afterwards is guessed. Cars are unforgiving here: wheels are perfect circles, badges are exact shapes, panel gaps are straight lines. Ask for a 90-degree orbit and the model has to draw the other side of the car; it draws a plausible car, not this one.

Keep the move within roughly 30 degrees of azimuth and 10 degrees of elevation from the photo’s own angle, do not move closer than the photo, and start every clip on the original photo. The result reads as video and the car reads as the car. VroomVideo fixes these bounds in code rather than leaving them to a prompt, which is why its wheels are round.

What the model should not be asked to do

  • Change the colour, the wheels or the trim.
  • Remove a dealer frame — crop it instead, before generation.
  • Show the engine, the boot open, or a door opening. None of that is in the photo.

Try it

Make a video from a car you are selling right now.