Tokens are the most utilitarian prompt type in DND art generation. They are not meant to be beautiful — they are meant to be readable. A good token lets your players tell which figure on the map is the goblin, which is the paladin, and which is the warlord at 400% zoom on a 1080p screen. This guide covers how to prompt for tokens that actually work on a virtual tabletop.
Use token prompts whenever you need a top-down or 3/4 overhead view of a character or creature for placement on a virtual tabletop like Roll20, Foundry VTT, or MapTool. Tokens are the visual shorthand that identifies figures on a battle map — they are small, centered, and designed to be recognizable at 40-80 pixels.
Do not use token prompts when you want to show a character in full detail. A token prompt produces a simple, centered silhouette — it cannot show a character's face, expression, or complex costume layers with any fidelity. If you need a portrait for a character sheet, use the portrait prompt type. If you need a battle map illustration, use the scene prompt type.
The practical test: if the image will be displayed at less than 100 pixels wide on a grid, it is a token. If it will be displayed larger than that, it is a portrait, full-body, or scene.
The defining constraint of a token is size. At 64x64 pixels (a standard Roll20 grid cell), fine details vanish. The prompt must emphasize features that survive extreme downsampling: silhouette shape, color blocking, and a single distinctive element (a hat, horns, a cape) that makes the figure identifiable from adjacent tokens.
View angle is the first and most critical decision. "Top-down view" produces a bird's-eye silhouette that works perfectly on grid maps. "3/4 overhead view" adds a slight angle that reveals the face and front of the costume at the cost of some grid alignment. "Isometric view" produces a 3D-style token that looks great in isolation but aligns poorly with flat 2D maps. Pick top-down for standard VTT use; pick 3/4 if you are building tokens for a specific map style that uses angled art.
Background handling is the second critical decision. "Transparent background" is ideal for VTT tokens because it allows the map to show through. "Simple dark background" works if your VTT platform does not support transparency or if you prefer tokens with a solid backing circle. "No background" is ambiguous — some models interpret this as "white background," others as "no background elements." Specify exactly what you want.
Level of detail is the third axis. More detail in the prompt produces more detail in the output, but detail below the pixel threshold of your VTT is wasted. A token that will be 64px wide does not benefit from "intricate filigree on the breastplate" — that detail becomes noise at display size. Prioritize shape and color over texture and fine detail.
a dwarf fighter token for roll20, top down, with a shield and axe, fantasy style
top-down view token of a dwarf fighter, centered single figure on transparent background, squat broad silhouette, round shield on the left arm painted with a copper hammer emblem, battleaxe held vertically in the right hand, iron helmet with a short nose-guard, reddish-brown beard framing the chin, 1:1 square framing, clean readable outline, flat color blocking, VTT token art, no shadow beneath the figure
The "before" prompt uses the word "token" which is ambiguous — the model may produce a coin-like object instead of a character figure. The "after" prompt specifies the exact view angle, background, and the key silhouette-defining features (shield, axe, helmet shape) that make the dwarf identifiable at small size. "Flat color blocking" tells the model to prioritize bold color areas over fine gradients, which produces tokens that read clearly at 64px.
dragonborn paladin token, white scales, glowing sword, holy symbol
3/4 overhead view token of a dragonborn paladin, centered figure on transparent background, white-scaled draconic head with a short snout and brow ridges, heavy plate armor with a golden sunburst holy symbol on the chest, a glowing longsword held pointing forward, sweeping tail curling behind, bright warm lighting from above, bold silhouette, clean edges, VTT token illustration, 1:1 aspect ratio
The "before" prompt is a keyword list that omits the view angle and background entirely. The model has to guess how to present a "dragonborn paladin" — it might produce a full portrait, a side view, or a creature illustration. The "after" prompt locks in the view, background, and silhouette, ensuring the output is usable as a token.
Midjourney struggles with top-down views because its training data heavily favors portraits and landscapes. Use --ar 1:1 for square tokens. Add --style raw to reduce Midjourney's stylistic gloss, which makes tokens harder to read at small sizes. Reinforce the top-down instruction by saying "bird's-eye overhead view" instead of just "top-down" — Midjourney responds better to evocative descriptions than technical terms.
ChatGPT image generation or DALL-E can be a practical choice for VTT tokens when instruction-following is more important than painterly atmosphere. Be explicit: "The image must be a top-down overhead view showing a single centered figure on a plain background." Background transparency support varies, so plan to remove a solid background in an image editor when necessary.
Stable Diffusion can produce excellent tokens with the right tag structure. Lead with: top-down view, centered figure, transparent background, VTT token, 1:1 framing. Use the negative prompt aggressively: side view, portrait, face visible, detailed background, shadow, gradient, photorealistic. These negatives block the model's tendency to produce portrait-style output. ControlNet with a depth map can enforce the overhead angle if your setup supports it.
The template prioritizes view angle, silhouette, and background before any character detail. This order ensures the model processes the structural constraints first. Character-specific details (race, class, weapon) come in the middle. Style and formatting instructions anchor the end.