How to write AI image prompts that actually work
AI image generators are powerful, but vague prompts give vague results. "A cool car" will get you a car — just not the one you imagined. Good prompts aren't about magic keywords; they're about describing the picture the way a photographer or art director would. Here's the structure I use.
The five building blocks
Almost every strong prompt answers five questions, roughly in this order:
- Subject — what is in the picture? Be specific: not "a woman" but "a young woman in a yellow raincoat holding a paper map".
- Setting — where is it? "On a narrow cobblestone street in an old European town, after rain."
- Style — what kind of image is it? A photograph, a 3D render, a watercolor, an anime frame, a product shot?
- Composition — how is it framed? Close-up, wide shot, low angle, centered, rule of thirds, shallow depth of field.
- Light and mood — golden hour, soft overcast light, neon at night, dramatic side light, calm, energetic, mysterious.
Putting it together
Weak: a cool car in the city
Strong: a vintage red sports car parked on a wet city street at night,
neon signs reflecting on the asphalt, low-angle wide shot,
cinematic photograph, shallow depth of field, moody blue and
magenta lighting
The second prompt leaves far less to chance. Every phrase removes a decision the model would otherwise make randomly.
Use photography language
Image models were trained on huge numbers of captioned photos, so camera and lighting terms work very well:
- Shot type: extreme close-up, portrait, medium shot, full body, wide establishing shot, aerial view.
- Lens feel: 35mm (natural), 85mm (flattering portraits), wide-angle (dramatic perspective), macro (tiny details).
- Depth of field: "shallow depth of field" or "bokeh background" to isolate the subject.
- Lighting: golden hour, softbox studio lighting, rim light, backlit, overcast, high-key, low-key.
Iterate one thing at a time
When a result is close but not right, don't rewrite the whole prompt. Change one element — the lighting, or the angle — and generate again. If you change five things at once, you won't know which change helped. Keep a simple text file of prompts that worked; over time it becomes your personal style library.
Keeping images consistent
Consistency is the hardest part, especially when you need the same character or product across many images, like a thumbnail series or a storyboard.
- Reuse a fixed description block. Write the character once ("a man in his 30s with short curly black hair, round glasses, green bomber jacket") and paste it into every prompt unchanged.
- Use reference images when your tool supports them. Image-to-image and character reference features keep faces and outfits far more stable than text alone.
- Lock the style words too — the same lighting and style phrases across a series make the set look intentional.
From image to video
Many AI video tools now animate a still image. This is often the most controllable way to make AI video: get the image exactly right first, then describe only the motion in the video prompt.
Image prompt: a lighthouse on a rocky cliff at sunset, dramatic clouds,
cinematic wide shot
Video prompt: slow drone push-in toward the lighthouse, waves crashing on
the rocks, clouds drifting, light beam starting to rotate
Keep motion prompts simple and physical: camera movement plus one or two things that move in the scene. Asking for too much action in a few seconds usually produces warping and artifacts.
Finishing touches
Generated images are rarely final. I usually upscale the best result, fix small problems (hands, text, odd edges) with inpainting or an image editor, and color-correct it so it matches the rest of a project. For video, AI clips get the same treatment as normal footage: trimmed, graded, and given sound design in the editor.
Use AI responsibly: don't generate misleading images of real people, respect copyright and trademarks, and check each tool's license before using results commercially.