Camera Direction Language for AI Video Tools: A Starting Vocabulary
Video & Short-Form Content ·
When describing an AI-generated video, you can specify more than the subject and setting.
Camera language provides another way to communicate how the scene should be framed or move.
This introductory vocabulary gives you some of the basic terms.
Shot Distance
Wide shot — shows the subject together with a substantial part of the surrounding environment. It can be useful when the setting itself matters.
Medium shot — frames less of the environment and brings greater attention to the subject. For a person, this may mean approximately waist-up framing, depending on the composition.
Close-up — frames a face, object or other detail more tightly.
Extreme close-up — isolates an even smaller detail, such as an eye, hand or part of an object.
Camera Angle
Eye-level — positions the camera around the subject's eye height.
Low angle — places the camera below the subject and looks upward. Depending on the scene, this can make the subject appear more imposing or prominent.
High angle — places the camera above the subject and looks downward. Its effect depends on the subject, composition and wider scene.
Camera Movement
Push-in — the camera moves closer to the subject during the shot.
Pull-back — the camera moves away, revealing more of the surrounding scene.
Pan — the camera rotates horizontally from its position.
Tilt — the camera rotates vertically, moving the view upward or downward.
Tracking shot — the camera moves with or alongside a moving subject.
Static camera — the camera remains fixed during the shot.
Handheld feel — describes movement that is less mechanically stable and can create a more immediate or observational visual character depending on execution.
Depth of Field
Shallow depth of field — keeps a limited depth range in focus, allowing the background or foreground to appear blurred.
Deep depth of field — keeps a broader range of the scene acceptably in focus.
Combining Terms
A direction such as:
"Medium shot, eye-level, static camera, shallow depth of field"
communicates several camera choices instead of leaving all of them unspecified.
That still doesn't guarantee a particular render.
Interpretation Varies by Tool
Different AI video tools may interpret camera terminology differently, and generations can vary even when the wording remains unchanged.
Think of these terms as descriptive direction rather than guaranteed technical commands.
Camera terminology is also only one part of directing an AI-generated video. Broader scene, action, visual and production choices may need to be described depending on the tool and the result you're trying to create.
The Ultimate AI Prompt Vault goes further with structured guidance for AI video prompting alongside its broader educational foundation.