A character portrait is the most common prompt type for DND players and dungeon masters. This guide covers when a portrait is the right choice, how to compose one that actually produces usable art, and what mistakes turn a promising idea into a blurry mess.
A character portrait prompt is the right tool when you need a face-forward image of a specific character — a player character avatar, an NPC who appears repeatedly in your campaign, or a villain you want to show your table. Portraits emphasize the character's identity: facial features, expression, costume details, and the lighting that sets their mood.
Portraits are the wrong choice when you need to show the character in action or in context. If you want your rogue mid-backstab, a full-body or scene prompt will serve you better. If you need a token for a virtual tabletop, the token prompt type produces top-down silhouettes that portraits cannot replicate.
The sweet spot for a portrait is any situation where recognizing the character matters more than what they are doing. Session-zero character sheets, recurring NPC handouts, campaign wiki avatars — these are all portrait territory. One-shots where the NPC appears once and never again? A brief description at the table may be more efficient than generating art.
The most important decision in a portrait prompt is the cropping. Head-and-shoulders framing keeps the focus on the face and expression but sacrifices armor and weapon detail. Three-quarter framing shows more of the body but reduces the emotional punch of the face. Full-body shots under a "portrait" prompt type often produce stiff, neutral poses because the model tries to fit everything in — that is what the full-body type is for.
Lighting direction is the second key decision. Dramatic side lighting (Rembrandt lighting) adds mood but can hide part of the face. Even, soft lighting shows detail but may look flat. A slight warm key light from above-left with a cool fill from the right is one reliable starting point for a fantasy torchlight mood. Rim light and backlight can work as accents, but pair them with enough front fill to keep the face readable.
Background complexity is the third axis. A simple dark background isolates the character and keeps the model focused on costume and face. An environment background adds story context but invites the model to spend tokens on scenery instead of the character. If you include a background, keep it to one or two nouns: "stone wall," "forest clearing," "throne room." Avoid full scene descriptions inside a portrait prompt.
a tiefling warlock with red skin and horns, wearing dark robes, holding a staff, in a dark place, fantasy art
head-and-shoulders portrait of a tiefling warlock, crimson skin, curved horns sweeping back from the brow, amber eyes catching firelight, wearing tattered velvet robes over black leather, a gnarled ironwood staff rests against the shoulder, warm torchlight from upper left, dark stone crypt wall behind, painterly fantasy illustration, expressive face, detailed skin texture
The "before" prompt is vague about framing, lighting, and background. The model has to guess what "dark place" means, and it will often guess wrong — producing a dark, muddy image. The "after" prompt specifies the cropping, the exact lighting setup, the background as a single element, and adds texture keywords ("detailed skin texture," "painterly") that push the model toward higher-quality output.
elf ranger, green cloak, bow, forest background, detailed, high quality
three-quarter portrait of a wood elf ranger, sun-weathered copper skin, pointed ears emerging from braided auburn hair, sharp green eyes scanning the treeline, wearing a hooded moss-green cloak over studded leather armor, a longbow slung across the back, dappled forest light filtering through canopy, simple backdrop of ancient oak bark, digital fantasy painting, warm natural lighting, focused expression
The "before" prompt is a keyword list. Models treat keyword lists as a buffet — they sample from each term without prioritizing any. The "after" prompt is a sentence that flows from the character outward, giving the model a clear visual hierarchy: face first, then costume, then environment, then style.
Midjourney v6+ responds well to natural-language descriptions. Add --ar 3:4 for portrait aspect ratio — the default 1:1 square crops heads or wastes space. Use --style raw if you want less of the Midjourney "painterly gloss" and more direct adherence to your description. Avoid comma-separated keyword dumps; Midjourney v6 prefers complete sentences.
ChatGPT image generation and DALL-E generally respond well to clear natural-language layout instructions. For a head-and-shoulders result, specify the crop and explain which body parts should remain outside the frame. Content handling can change over time, so describe the visual purpose and context plainly rather than trying to work around a particular filter.
Stable Diffusion responds strongly to comma-separated tags, which is the opposite of Midjourney's preference. Lead with the strongest visual tags: 1girl, tiefling, portrait, crimson skin, curved horns, dark robes, torchlit. Add quality tags like masterpiece, best quality, highly detailed early in the prompt. Use a negative prompt to exclude common SD artifacts: lowres, bad anatomy, extra fingers, blurry, watermark.
Fill the relevant brackets with concrete details and remove any bracket you do not need. Models vary in how they weight prompt order, but placing framing and identity early creates a clearer hierarchy and makes the prompt easier to revise.