Most image-AI briefs fail before the first image is generated. They do not fail in the tool, they fail in the instruction. A brief written in generic e-commerce vocabulary ('product photo, white background, good lighting') gives the AI style information, not garment information. And without garment information, the AI fills the gaps its own way.

This is what a typical brief looks like, the one anyone would write without thinking twice:

It is a reasonable instruction if you are talking to a photographer who already knows the physical garment, because they have it in front of them and work out the rest on their own. But an image AI does not have the garment in front of it. It does not know whether 'red' is a matte burgundy or a bright crimson, it does not know whether the dress falls fluid or has structure, it does not know what trims it carries or where they go. It fills those gaps with the statistically most likely option, which is almost never yours.

Professional judgement is not writing a longer brief. It is writing one with the variables that really change the result: fabric, drape, trims, colour with a reference and framing with a visual reference.

Fabric: it is not enough to name the garment, you have to name how the material behaves. A crêpe does not move like a knit, a denim does not move like a silk. If the AI does not know which fabric it is, it picks a default, and that default rarely matches yours.

Drape: fluid, structured, flared, close to the body. It is the difference between a garment that hangs under its own weight and one that holds shape by construction. Without this variable, the AI tends to soften any garment towards the same generic fold.

Trims: buttons, zippers, buckles, visible or concealed closures. They are small but they are where a generic image shows first: invented trims, badly placed or missing altogether.

Colour with a reference: 'red' is not an instruction, it is a range. A Pantone reference, a photographed fabric swatch, or at least a comparative description ('burgundy, not crimson') narrows that range to something the AI can get right.

Visual framing reference: an example of the composition you want (angle, distance, type of shot) does more for the result than two paragraphs describing it in words.

With those five variables, the same brief is rewritten like this:

The difference between the two briefs is not length, it is what kind of information they contain. The first describes a photo. The second describes a garment. And an image AI has to be given the garment, not the photo, because it builds the photo itself from what you tell it about the garment.

This does not remove the need to iterate. Even with a brief written with judgement, the first generation can miss some detail and will need adjusting. But the difference between iterating twice and iterating ten times is, almost always, the quality of the starting brief.

Correct product image featuring a model.

Accurate product image; preserves the garment's color, design, and drape.

Incorrect product image; it does not match the garment.