Lesson Outcome
Before generating, it is easy to collect a folder of beautiful images and call it a moodboard. Then one shot contains one character, the next contains a distant relative, and the object suddenly changes shape. The problem is usually not the number of images. They simply have no roles.
By the end of this lesson, you will have a compact set of anchors for a face, outfit, object, location, lighting, color, composition, or movement. Every important property will have one primary authority, and you will know what to take from it and what to leave out on purpose.
Key point: a reference is useful not because it is beautiful, but because it holds one specific property of the scene.
A Reference Is More Than a Pretty Picture
Words such as “premium,” “cinematic,” and “dark” sound clear until two people picture completely different shots. To one person, “premium” is a black background and cool light. To another, it is gold, smoke, and three chandeliers nobody invited.
A reference becomes operational when you can finish the sentence: “I am taking exactly this from it.” A portrait may control the face and hairstyle; a jacket photo, the fabric and silhouette; an office frame, the geometry of the room; and an advertising frame, only the side light and empty space on the right.
If two images show the face differently, appoint one primary source in advance. Otherwise, the model may sincerely try to follow both and produce a third person who was never invited to the project.
Four Roles That Should Not Be Mixed
References differ by function, not by beauty.
- A model reference is material you actually provide to the selected generation route when it can accept that input. It can help preserve a face, object, outfit, location, or another property.
- An author reference is an example you study yourself for lighting, shot size, composition, material, or an editing technique. You may never upload it to a generator.
- A control frame is an exact start, middle, or end state when the chosen route can genuinely use it.
- A movement example is a video or sequence from which you read a path, speed, pause, or camera behavior.
One image may help several lines of thinking. In one specific generation, though, give it one narrow job. That way you do not ask one file to be a face, a room, an advertising style, and a camera instruction at once.
Turn a Moodboard into Rules
A moodboard is a collection of examples used to find a visual direction. Its job is not to fill a board beautifully. Its job is to turn an impression into visible decisions you can check in the next frame.
Imagine a perfume bottle in black water. Instead of “I want it premium and atmospheric,” you can see cool side lighting, a glossy black surface, a clean close-up, empty space on the right, and slow ripples on the water. That is no longer a mood in fog; it is a set of observable qualities.
For every useful example, record only what changes your shot:
- lighting: where it comes from, whether it is hard or soft, cool or warm;
- color: two or three dominant colors;
- composition: where the main subject sits and how much empty space surrounds it;
- shot size and angle: what the viewer notices first;
- material: for example, glass, water, metal, fabric, or smoke;
- movement: what changes and at what speed;
- exclusion: what kind of result definitely does not fit.
You are not copying someone else’s bottle or scene. It is like showing a photo to a hairstylist and saying, “Not this head; this volume and shape.” Take a transferable decision, not someone else’s work as a whole.
The task determines the number of examples. One frame may be enough for a lighting direction. A recurring character may need several clean angles. There is no universal number of images or videos: extra material does not automatically create stronger control.
Character, Object, and Location: Three Separate Anchors
Character
If a character appears in several shots, collect clean images where the face, hairstyle, clothing, and silhouette are readable. A regular suit, a mask, and damaged clothing are not one averaged identity; they are separate states. Do not mix them without a reason.
Object, Product, or Logo
Show anything that must stay exact separately and cleanly: its shape, material, required sides, without accidental text or distracting objects. An incorrect logo is not an “interpretation.” It is an incorrect logo.
Location
The location should support the action. A beautiful background with cables, screens, and smoke can look expensive while hiding the character, breaking the geometry, and leaving no room to move. Lock stable landmarks, overall lighting, depth, and a clear action area.
When Materials Conflict or Are Missing
Do not try to settle a conflict with a longer prompt. A prompt is not required to become a family therapist for two incompatible pictures.
When two references conflict, do one of three things: choose a primary source for the disputed property, narrow the second source’s role, or remove it from this task. If material is missing, record the next action: find an official source, photograph it yourself, extract a frame from a video, or prepare a test image. Separately note the author, rights, or usage restriction when they are known for the selected material.
Practice: Build a Reference Pack
Take one future scene. First name only the required properties: for example, face, outfit, object, location, and lighting. Then assign one primary example to each property. Copy this form for every material.
File or link:
Role:
Use for:
Do not use for:
Preserve:
Risk if the material is missing:
Used in scene or segment:
Author, rights, or usage restriction:Check Before the Next Frame
The set is ready when every required property has one primary example or a concrete plan to obtain it; roles do not duplicate or contradict one another; you have recorded visible qualities; character, object, and location are not mixed into an accidental source; and you know what the selected generation route will actually receive and what remains guidance for you alone.
Takeaway: you now have a set of anchors with roles, not a pile of attractive images. The next frame can rely on face, object, location, and lighting separately, without having to guess everything at once.