Concept

Text-to-Image Prompting Basics

The anatomy of an image prompt — subject, medium, style, composition, lighting — and an iteration workflow across 2026's tools.

A good image prompt has an anatomy, and naming its parts is how you get repeatable results instead of lucky ones. The core slots: subject (what or who, with concrete detail), medium (photo, oil painting, 3D render, watercolor), style (an era, movement, or aesthetic), composition (shot type, angle, framing), and lighting (soft, golden hour, neon, rim light). 'A dog' is a coin flip; 'a wet golden retriever puppy, close-up portrait, shallow depth of field, soft window light, photograph' specifies enough that the model's choices land near your intent. Start every prompt by filling these slots deliberately, even roughly, then tighten.

The 2026 tool landscape splits by strength, and prompt style follows the tool. Midjourney v7 leans aesthetic and stylized, and rewards evocative phrasing plus its own parameters. OpenAI's gpt-image-1 (the engine behind image generation in ChatGPT) and Google's Imagen 4 and native Gemini image generation are strong at prompt adherence and legible text-in-image, and take plain, literal descriptions well. Flux, the open-weight family from Black Forest Labs, gives you local control and fine-tuning. Match the prompt to the model: literal and structured for the adherence-focused engines, evocative and stylistic for Midjourney. The same words do not produce the same image across tools.

Image prompting is iteration, not one-shot. Write a base prompt, generate a batch, then change one variable at a time — swap the lighting, then the lens, then the style — so you learn what each term actually does in that model. Keep the seed fixed when a tool exposes it to isolate the effect of a single change; vary the seed when you want fresh compositions. Save the prompts that work, because you are building a personal library of terms that reliably do what you mean. Treat the first generation as a sketch that tells you which slot to tighten next, not as a verdict on the whole idea.

Two habits sharpen results fast. First, be concrete about what matters and silent about what doesn't — over-specifying every detail can fight the model, while naming the two or three things you actually care about (the subject's expression, the palette, the mood) leaves useful room where you don't. Second, describe what you want, not what you don't: positive description is what these models are built to follow, and exclusions are a separate, weaker tool covered next lesson. When a result is close but wrong in one way, change the single word governing that aspect rather than rewriting the whole prompt and losing what already worked.

Check your understanding
Q1. The same prompt gives a painterly, stylized result in Midjourney but a flat, literal one in Imagen. Why?
Q2. Your generated portrait is close, but the lighting is wrong. What is the best next move?
· Score 100% on the quiz.