Scene prompts are the broadest and most complex prompt type. Where a portrait is one face and a token is one silhouette, a scene is a full environment with depth, lighting, atmosphere, and narrative potential. Getting a scene prompt right means your players see the tavern, the dungeon, or the battlefield before you describe it. Getting it wrong means a generic landscape that could be anywhere. This guide covers how to prompt for scenes that feel like places your party could actually visit.
Use scene prompts when you need an establishing shot or location illustration — a tavern exterior, a dungeon corridor, a mountain pass, a throne room, a forest clearing. Scenes set the stage for encounters; they are the "where" before the "who" and the "what."
Do not use scene prompts when a character is the focus. A dragon perched on a mountain is a monster prompt, not a scene. A tavern with a bard performing is a scene with a secondary character. If the central figure is more important than the environment, use the character or monster prompt type and add a simple background.
The practical rule: if removing all characters from the image would still leave a useful illustration, it is a scene. If removing the characters would leave an empty room, it is a character portrait with a background.
The most important decision in a scene prompt is the depth layering. A flat scene — "a forest" — produces a pretty picture with no spatial logic. A layered scene — "a forest clearing, fallen log in the foreground, mossy boulders midground, ancient trees fading into mist background" — produces an image with a sense of place and scale. Depth layers are what make a scene prompt produce an illustration you can set an encounter in, rather than a wallpaper you can only look at.
Lighting and time of day is the second critical decision. "A dungeon corridor" looks completely different depending on whether you say "lit by torchlight with warm shadows" or "dimly illuminated by scattered bioluminescent fungi." The time of day also determines the emotional register — dawn conveys hope, dusk conveys closure, midnight conveys danger, noon conveys safety. Never leave lighting unspecified unless you want the model to choose randomly.
Presence of characters is the third axis. A scene with no characters is an establishing shot — pure atmosphere. A scene with tiny silhouetted figures conveys scale and activity without focusing on any individual. A scene with a single prominent figure crosses into "portrait with background" territory. Decide whether characters are present before writing the prompt; adding them as an afterthought usually produces awkward scale mismatches.
Weather and atmosphere is often overlooked but transforms the mood. Rain, fog, snow, smoke, dust — these particles fill the air between the viewer and the scene, adding depth and mood for free. A dry scene reads as clinical; a weather-affected scene reads as lived-in. Always specify at least one atmospheric element.
a tavern, inside, wooden tables, fireplace, warm, fantasy art
wide establishing shot of a medieval tavern interior, rough-hewn timber beams overhead, stone hearth with a crackling fire on the left wall casting warm orange light, scattered wooden tables with carved chairs and half-empty tankards in the midground, a worn oak bar counter along the back wall with bottles on a shelf behind it, staircase ascending to a shadowy second floor on the right, amber candlelight from iron chandeliers, hearth smoke curling into the rafters, a hunched figure silhouetted at the far corner table, dark fantasy illustration, atmospheric lighting, lived-in detail
The "before" prompt describes a generic tavern. The model has seen a thousand taverns and will produce the average of all of them. The "after" prompt specifies foreground (tables), midground (bar counter), background (staircase), lighting (hearth fire, chandeliers), atmosphere (smoke in rafters), and a single character presence (silhouette) — six specific visual decisions that make this tavern unique.
ancient ruins in a jungle, overgrown, mysterious, dramatic lighting, detailed
wide establishing shot of crumbling stone ruins in a tropical jungle, moss-covered broken columns and a collapsed archway in the foreground, a vine-choked stone altar as the midground focal point, dense canopy and shafts of golden afternoon light filtering through the trees in the background, hanging vines and ferns draped over every surface, thin mist rising from the warm stone, a tiny explorer figure standing before the altar for scale, emerald and gold color palette, fantasy environment illustration, volumetric light, lush atmosphere
The "before" prompt relies on mood words ("mysterious," "dramatic") that are too abstract for the model to render specifically. The "after" prompt translates mood into concrete visual elements — "thin mist rising" for mystery, "shafts of golden afternoon light" for drama, "tiny explorer figure" for scale. Every mood keyword should have a corresponding concrete visual description.
Midjourney is the strongest model for scene art — its painterly style naturally produces atmospheric environments. Use --ar 16:9 or --ar 21:9 for cinematic wide shots. Add --v 6.1 or later for the best spatial coherence. Midjourney's main weakness is text rendering — if your scene includes signs, books, or inscriptions, they will be garbled. Describe text elements as "worn sign with illegible text" to avoid garbled output.
DALL-E 3 produces cleaner, more geometric scene art — closer to a concept art board than an oil painting. This is ideal for battle maps and architectural reference but less atmospheric for session handouts. DALL-E 3 handles text better than Midjourney but still struggles with more than two or three words. Use DALL-E 3 when you need a clean, readable environment layout; use Midjourney when you need mood and atmosphere.
Stable Diffusion is the fastest model for scene generation and produces the most detailed output when guided correctly. Structure your prompt with spatial tags: wide shot, environment, [location], foreground [element], midground [element], background [element], [lighting], [weather], [atmosphere], fantasy landscape, highly detailed. Use the negative prompt: character portrait, face, close-up, blurry, low contrast, flat colors. ControlNet with a depth map or edge map can enforce perspective consistency for repeated views of the same location.
The template mandates depth layers before style. Character presence is a separate decision, not a default. The "lived-in detail" keyword at the end pushes the model away from empty-room output. As with other templates, order matters — foreground and midground come before atmosphere and style.