How to Create Storyboards With AI Image Generators

You generate five storyboard panels for a client pitch. The character's face matches your reference sheet across every frame. Then you line up the sequence on a canvas and look closer.
In panel two, the coffee shop's arched window sits behind her left shoulder. In panel four, the exact same window is on the right wall, the wooden table turned into white marble, and the gentle morning sunlight suddenly strikes her from the opposite side.
Character consistency is a solved problem. The real breakdown in generative AI previsualisation almost never starts with the face; it happens when the environment, framing grammar, and lighting drift between sequential cuts. When every panel is generated as an isolated request, the model invents a brand new room every single time.
Here is how to run a production-ready AI storyboard generator workflow from script to final sequence:
- The four continuity axes every storyboard must lock.
- Why environment drift breaks sequences faster than character drift.
- The location plate workflow: generate the room before the shot.
- Locking set lighting like a gaffer, not a prompt writer.
- Film continuity grammar: 180-degree rule and screen direction in AI.
- The reusable shot-card system (and reference-image rules).
- Handling multi-character scenes and composite passes.
- Model selection: picking the right engine for each shot type.
- Frequently asked questions.
Why AI storyboards fall apart (even when the character looks right)
Most generative image tools treat every prompt as an independent universe. When you ask a model for a medium close-up followed by an over-the-shoulder reverse angle, it has zero spatial memory of the room it rendered thirty seconds ago.
Competitor tutorials often market character consistency as the whole solution. But holding a face steady while the room dissolves around it does not produce a usable storyboard. An April 2026 academic study by Ishani Mondal, Jordan Boyd-Graber, and a research team from the University of Maryland and Google introduced CANVAS, a visual storyboarding architecture built around world-state tracking. When evaluated against standard baselines on their ContinuityEval benchmark, their structured approach produced a 21.6% improvement in background continuity, a 9.6% gain in character consistency, and a 7.6% boost in props consistency.

Four separate reference systems: character, environment, lighting, and framing.
Notice those numbers: background continuity had more than double the headroom of character identity. A human face is a concentrated target with a single point of failure. A physical location has dozens of interdependent variables that must hold simultaneously: window placement, architectural geometry, practical light fixtures, prop locations, and background depth. When you change your camera angle, an unanchored model recalculates all of them at random.
If you want sequential visual coherence, you have to manage four distinct reference systems simultaneously rather than relying on descriptive text.
The four things a storyboard has to hold steady
To build a coherent previsualisation pass, separate your scene into four distinct continuity layers before generating a single panel:
- Character identity: The facial geometry, hair, wardrobe, and physical proportions of your subjects. We broke down the exact reference-image limits and prompting rules in our guide on how to create consistent AI characters.
- Environment geometry: The fixed physical architecture of the setting. Furniture positions, door frames, wall textures, and window views must remain static regardless of whether the camera is wide or tight.
- Lighting setup: Key light direction, fill intensity, color temperature, and practical sources. If your establishing shot establishes a harsh sunset coming from camera left, a tight reaction shot cannot have soft, neutral studio fill.
- Framing grammar: Aspect ratio, screen direction, relative character eyelines, and the 180-degree rule. These classic cinematography conventions keep sequential cuts legible to an audience.
When one of these layers slips, the board stops reading as a connected sequence. Let's fix each layer in order.
Build your location plates before you build a single shot
Never describe your setting from scratch inside a shot prompt. If panel one asks for an "office with industrial brick walls and modern desks" and panel two asks for a "close-up in an office with modern desks," the model will reinterpret the room's layout, shifting the brick pattern and moving the light sources.
The fix is simple: generate a canonical, wide, well-lit location plate for every distinct environment in your script before generating any character action.
A location plate is an empty establishing shot of the room or exterior containing all major landmarks: doors, windows, key furniture, and lighting sources. Once you have a clean plate that satisfies your art direction, save it. That single image becomes the mandatory environment reference image for every subsequent shot taking place in that room.
When prompting your downstream coverage, attach the location plate as a scene reference. In systems like GPT Image 2 or Gemini, you can explicitly reference the background structure. For example: "Medium shot of the character from reference image 1, sitting at the desk on the left side of the room shown in reference image 2."
By feeding the room back to the model as an image input rather than a paragraph of text, you anchor the furniture placement and architectural details across wide, medium, and insert shots. In JammyJar, you can store these master plates in a dedicated project folder, allowing your team to drag the exact same location anchor into prompts across multiple scenes.
Lock your lighting like a gaffer, not a prompt writer
Prompting for "dramatic cinematic lighting" produces erratic results. In frame one, the model might give you high-contrast noir shadows; in frame two, it might default to flat, commercial lighting.
Professional cinematography controls lighting by position, quality, and color temperature. Your prompts should do the same. Define a fixed lighting-lock block for each scene, written like a physical set diagram:
- Key light position: State direction using clock positions relative to the camera (for example, "key light at 10 o'clock high").
- Quality and softness: Specify whether the light is hard and direct or diffused through a softbox.
- Fill and contrast ratio: Define the shadow depth (for example, "subtle 4:1 fill ratio with open shadows").
- Color temperature: Specify Kelvin or precise color tones (for example, "3200K warm tungsten interior with cool 5600K blue daylight spilling through the background window").
- Visible practicals: Name any lamps, neon signs, or light fixtures visible in frame.
[LIGHTING LOCK]: Key light at 10 o'clock diffused softbox, subtle 4:1 fill ratio on shadow side, 3200K warm interior ambient with 5600K cool daylight through background window, practical desk lamp turned on at frame right.
Paste this exact text block into every shot prompt within the scene. A useful rule of thumb from production lighting: an unmotivated color-temperature shift of more than roughly 300K between cuts breaks visual coherence immediately.

Lock lighting direction by clock position to keep shadows consistent across cuts.
Rather than retyping these blocks, we recommend saving your visual mood, color grade, and lighting setup as a reusable asset. You can read our walkthrough on keeping a consistent style across AI images to set up locked style profiles that persist across multiple generations.
Borrow the continuity grammar filmmakers already invented
AI image models do not understand visual storytelling. They understand pixel distributions. Left to their own devices, they will flip a character's facing direction, jump the camera over invisible walls, and change aspect ratios at random.
To make a storyboard readable, borrow four foundational rules from traditional film previsualisation:
1. Lock your aspect ratio on day one
Never switch frame proportions mid-sequence. A composition planned in 16:9 widescreen carries completely different spatial relationships than a 9:16 vertical frame or a 4:3 box. If you are producing for commercial video or film, lock 16:9 or 2.39:1 across every generation. If you need a refresher on matching canvas sizes to delivery specs, check our AI image aspect ratio and size guide.
2. Obey the 180-degree rule
When two characters interact, draw an imaginary line between them (the line of action). Keep your camera on one side of that axis for the entire conversation. If Character A looks screen-right in shot one, Character B must look screen-left in the reverse shot. If you cross the line, the characters appear to be looking in the same direction, making the edit feel disjointed.
In your shot prompts, explicitly state screen direction: "Character looking screen-right at 45 degrees, positioned on the left third of the frame."
3. Match character screen position across cuts
If a character stands on frame-left in the wide establishing shot, do not jump straight to a medium shot where they sit on frame-right without an establishing movement. Sudden spatial jumps disorient the viewer.
4. Feed the previous frame as a continuity check
When generating panel N, feed panel N-1 as an additional visual reference where supported, or inspect the two side by side immediately. Check that wardrobe details, hair partings, and handheld props did not magically swap hands.
The shot card: a repeatable pre-shot planning system
To keep these four systems organized, use a structured shot card before typing into an image generator. A shot card serves as a single source of truth for each panel, compiling your references and continuity rules in one place.
Here is what a complete shot card looks like for an individual panel:
Field | Value / Setting | Purpose |
|---|---|---|
Shot Number | Scene 1, Shot 03 (CU) | Establishes sequence order and coverage scale |
Aspect Ratio | 16:9 (Locked) | Prevents framing and compositional drift |
Character Ref |
| Primary identity and wardrobe lock |
Location Plate |
| Anchors background geometry and props |
Lighting Block | Locked 3200K / 5600K Key at 10 o'clock | Ensures continuous lighting and shadow direction |
Previous Frame |
| QA anchor for eyeline, prop hand, and screen position |
Screen Direction | Subject looking screen-right, eyeline 15° up | Enforces 180-degree rule across the sequence |
Target Model | GPT Image 2 or Gemini 3 Pro Image | Routes prompt to the engine best suited for the composition |
How to work through the shot card
- Draft your shot list: Write out your camera framing (Wide, Medium, Close-Up, Over-the-Shoulder, Insert) and the narrative action.
- Gather your anchor assets: Verify that your character sheet and location plate are generated and approved.
- Assemble the prompt: Combine your action description, camera angle, and the locked lighting block, then attach your reference images.
- Run the QA check: Compare the new output against your previous panel and location plate. Did the lighting flip? Did the wall color change? Did the character switch hands? If yes, adjust the framing prompts or mask the mismatched region.

One shot card per panel: compile your references and continuity rules before prompting.
Where consistency breaks: multi-character scenes and long runs
Even with disciplined shot cards, current image models hit hard limits in two specific scenarios:
The multi-character collapse
Placing two or more established characters into a single generated frame frequently triggers attribute leakage. The model blends their features: Character A inherits Character B's jacket, or their hair colors swap. A 2026 stress test showed that while modern models manage two to four subjects reasonably well, group scenes with six or more characters suffer severe identity collapse.
The workaround: Never generate a complex multi-character shot in a single pass. Generate Character A against the empty location plate first. Then, use image editing or inpainting to insert Character B on the opposite side of the frame using their respective character reference sheet. Building ensemble shots sequentially eliminates feature bleeding.
Sequence drift over 15+ shots
If you chain references sequentially (using Shot 1 to generate Shot 2, then Shot 2 to generate Shot 3), small generation errors accumulate. By Shot 12, your character's clothing and the room's architecture will have visibly drifted from where they started.
The workaround: Always anchor back to your master assets. Never use the previous shot as your primary identity or environment reference. Always feed the original master location plate and character anchor sheet into every card, using the previous shot strictly for framing and screen-direction checks.
Which model to use for which shot
No single AI image model wins every type of storyboard shot. In our testing, different engines excel at different cinematic tasks:
Model | Best Shot Types | Reference Capabilities | Where It Breaks |
|---|---|---|---|
Gemini 3 Pro Image | Dynamic lighting, complex action, environmental nuance | Up to 14 mixed reference images (up to 5 human refs) | Can struggle with literal, rigid text placement on background signs |
GPT Image 2 | Literal prompt adherence, multi-reference indexing, prop inserts | Multi-image referencing via edit endpoints | Can produce a slightly glossy, hyper-clean commercial finish |
Recraft (V4 / Vector) | Graphic novel boards, animated previs, icon/prop design, SVG boards | Reusable | Less suited for photorealistic cinematic film grain |
Rather than locking your entire pipeline into a single vendor, route each shot to the engine best suited for the task. In JammyJar, you can switch between Gemini, GPT Image, and Recraft within the same workspace, applying your saved style themes across different underlying models without losing your prompt history or asset library.
Frequently asked questions
What is an AI storyboard generator? An AI storyboard generator is a tool or workflow that uses generative image models to produce sequential visual panels illustrating camera angles, character staging, and scene composition for film, video, or animation preproduction.
How do I keep backgrounds consistent across an AI storyboard? Generate a clean, empty "location plate" of your setting before creating individual shots. Attach that master plate as a visual reference for every subsequent panel taking place in that environment, and maintain a fixed lighting description in every prompt.
Do I need a dedicated storyboard tool like Boords or Firefly? Dedicated tools offer valuable collaboration features like frame-level client commenting and PDF exports. However, for direct visual generation, using a multi-model workspace with a structured shot-card method gives you tighter control over reference images and model choice.
How many reference images should I use per shot? Use two to three primary reference images per panel: one master character anchor, one master location plate, and optionally the previous frame to verify screen direction and eyelines.
What breaks first in an AI-generated storyboard sequence? Lighting direction and background geometry break far faster than character faces. If unanchored by a location plate and a locked lighting block, background furniture, window placements, and shadow angles will shift between cuts.
Build your next board frame by frame
Take your current script and set aside the impulse to start generating scene action immediately. Step into the empty set first.
Generate your master location plates for your primary scenes. Write down your gaffer-style lighting block with specific clock positions and color temperatures. Assemble your character anchor sheets. When you load your first shot card, you won't be hoping the model remembers the room; you will be handing it the blueprint.
Pick one three-shot dialogue sequence from your next project. Build the location plate, set the lighting lock, enforce the 180-degree rule across the cuts, and run the panels through your workspace. Once you control the space, the storyboards finally hold together.