The most common way to write an AI image prompt is to enter one sentence such as “a penguin using a computer” and wait for a random result. If you are lucky, it works. If not, you regenerate seven or eight times. In my experience, structuring the prompt raises the success rate from about 30% to 70%. This article explains the method I use in practice.

The Four-Layer Prompt Structure

Split the prompt into four blocks. Each block answers one question:

Layer 1: Subject. What should be drawn?

This is the most basic layer. Describe the main character, scene, and action. The more specific, the better. “A penguin” and “a small penguin wearing an orange scarf, sitting at a desk with an open laptop in front of it” produce completely different results.

Layer 2: Style. What style?

Watercolor, 3D render, pixel art, colored pencil, Japanese illustration, minimal line art. Style decides the overall “feel” of the image. Colored pencil and flat illustration are relatively safe choices that do not look too AI-generated.

Layer 3: Composition. How should it be arranged?

Camera angle (top-down, eye level, low angle), subject placement (center, left third), whitespace position (empty space on the right for text), aspect ratio (16:9 banner, 1:1 square).

Layer 4: Constraints. What should be avoided?

Many people ignore this layer, but it is very effective for output control. “No text,” “no yellow beak,” “no oversaturated colors,” “no photorealistic style.”

Four-layer prompt structure

Practical Gemini/ChatGPT Prompt Examples

These are formats I have actually used in Gemini.

Example 1: Blog Cover Image

Subject: A small penguin sitting at a desk with three monitors showing different AI tool interfaces
Style: Colored-pencil style, soft warm tones, slightly hand-drawn
Composition: 16:9 landscape, penguin in the left third, whitespace on the right for a headline
Constraints: No photorealism, no excessively sharp edges, no yellow pointed beak (the beak is round and orange)

Example 2: Social Image

Subject: A small penguin holding a magnifying glass and looking at a glowing block of code
Style: Flat illustration with distinct color blocks and subtle texture
Composition: 1:1 square, centered subject, simple background
Constraints: No 3D effect, no gradient background, use one pale background color

Example 3: Tutorial Step Diagram

Subject: A simple flowchart with a microphone icon on the left, an AI-processing gear icon in the center, and a subtitle icon on the right, connected by arrows
Style: Clean line illustration in dark blue and orange
Composition: 16:9 landscape, three elements spaced evenly
Constraints: No photorealistic imagery, no extra decoration, and use English if any text appears

What these examples share is a clear structure, with each part on its own line. Gemini understands this format well in natural language. It does not need the -- parameter syntax used by Midjourney.

More Scenario Prompts You Can Copy Directly

The three examples above are tool-oriented. The following are the scenarios Penchan switches between most often in real work.

Article Cover Image (Blog, Newsletter, Press Release)

Scenario: Main image for a blog article, newsletter, or press release. Usually 16:9, with space on the right for a title. Best tools: Gemini/ChatGPT (first choice, strongest instruction following), Midjourney (after translating into English) How to use: Fill in the topic and title keywords, then paste into the Gemini chat window.

Subject: Three notebooks scattered on a desk, a steaming cup of coffee, and an open laptop showing a simple text editor
Style: Watercolor, soft morning light, subtle paper texture
Composition: 16:9 landscape, objects concentrated on the left, blank space on the right for a headline
Color palette: Warm beige background with pale brown and light blue, low overall saturation
Topic keyword: [Enter a topic, for example: morning writing habits]
Exclude: Text, logos, 3D effects, excessively sharp edges, and highly saturated color blocks

Penchan tip: A blog cover should echo the page’s main color. In practice, upload an existing cover first and tell Gemini to “refer to this image’s color tone.” Consistency improves a lot.

Social Post Image (IG, Threads, X)

Scenario: Square or 4:5 vertical image for a short post. It needs to catch attention and stop the scroll. Best tools: Gemini, ChatGPT, Midjourney How to use: Choose the ratio by platform: 1:1 for X and Threads, 4:5 for IG and Facebook.

Subject: A simple visual metaphor for [post topic, for example: information anxiety]
Style: Flat illustration with distinct color blocks and slight hand-drawn irregularity
Composition: 1:1 square, subject centered slightly above the middle, lower third left empty for overlaid text
Color palette: Low-saturation muted colors, mainly deep blue-gray with a touch of warm orange
Mood: Quiet with a little humor, like a friend mentioning something small
Exclude: Text, facial close-ups, highly saturated neon, gradient backgrounds, and 3D rendering

Penchan tip: The biggest risk for social images is looking too similar to everyone else. Fix a color palette, such as deep blue-gray plus warm orange, and apply it to every post. Over time, followers will recognize the image as yours.

Product Promo Image (E-commerce, Crowdfunding)

Scenario: Context image for an e-commerce product page or crowdfunding page. It should make people want to buy without looking like stock material. Best tools: Gemini/ChatGPT (first choice, can upload product photo as reference), Midjourney (for atmosphere images) How to use: Always upload a real product photo before using this prompt.

Subject: Use the uploaded product as reference and place it in an everyday setting, such as a desk on a weekend afternoon beside an open book and a cup of tea
Style: Lifestyle photography, natural light, shallow depth of field
Composition: 4:5 portrait, product centered in the lower third, upper background slightly blurred
Lighting: Side light from the upper right, creating a soft shadow on the product
Mood: Slow, quiet, and lived-in, like a candid moment
Exclude: Plastic texture, excessive smoothness, artificial-looking people, handshake or suit-based business scenes, and fabricated product details
Important: The product's appearance, color, and logo must exactly match the uploaded image and must not be changed

Penchan tip: The last line, “do not change product appearance,” is important. Gemini sometimes helpfully “beautifies” a product, but then the output differs from the real product by one shade and clients get angry.

Character Illustration (Avoiding AI Faces)

Scenario: A blog illustration needs a person. AI-generated faces often have unnatural eyes and teeth. Best tools: Gemini, ChatGPT, Midjourney How to use: The key is avoiding front-facing close-ups and using back views or side faces.

Subject: A person sitting at a desk by a window, viewed from behind or from the side, with a book and pen nearby
Style: Hand-drawn colored pencil, visible paper texture, slightly uneven lines
Composition: 16:9 landscape, person in the left third, no front-facing facial features
Angle: A 45-degree view from above and behind, showing the back of the head and shoulders, face turned toward the window
Color palette: Warm orange afternoon light with pale green, low saturation
Exclude: Front-facing faces, close-ups of teeth, direct eye contact with the camera, plastic-looking skin, and perfect facial features

Penchan tip: If the prompt contains words like “front-facing” or “close-up,” AI easily draws a strange face. Use descriptions like “back view,” “45-degree side face,” or “only up to the shoulders,” and it almost never goes wrong. If you really need a face, use real photo material or shoot it yourself.

Information Diagrams (Flowcharts, Comparison Diagrams)

Scenario: An article needs a simple diagram to explain a flow or comparison. This is not a formal infographic. Best tools: Gemini/ChatGPT (can draw simple line diagrams), manual Figma work (most stable; AI-generated text is often blurry) How to use: If the diagram contains text, ask AI to draw only the graphics and add the text manually in Figma.

Subject: A simple three-step flowchart with three rounded rectangles arranged left to right and connected by arrows
Elements:
  First box: A sheet-of-paper icon representing input data
  Second box: A gear combined with an AI chip, representing processing
  Third box: A speech-bubble icon representing output
Style: Minimal line art with consistent stroke width and no fill, or pale fills only
Composition: 16:9 landscape, three boxes evenly spaced, ample background whitespace
Color palette: Pure white background #FFFFFF, dark gray lines #2D3748, a small amount of pale blue #90CDF4 as the accent
Exclude: Any text in any language, 3D effects, gradients, shadows, and unnecessary decoration

Penchan tip: The final line, “no text of any kind,” is the key. AI-generated text is almost always blurry or wrong. It is better to leave the image empty and add clean Chinese text in Figma. This saves an entire retry round.

Reference Images: The Key to Consistency

Pure text prompts have a ceiling: AI can only guess the image in your head. Reference images can close that gap significantly.

The practical method is to upload the image directly to Gemini, then tell it: “Refer to this image’s style and character design, then generate the following content.”

This is especially useful for solving character consistency. For example, the brand penguin has an orange rounded beak, but real penguins in AI training data mostly have yellow pointed beaks. If you only emphasize “orange rounded beak” in text, the model often gets pulled back to the yellow pointed beak. Once you attach a reference image, the error rate drops noticeably.

Before and after prompt optimization

How to Reduce the AI Look

AI-generated images have an “AI look” that people can recognize at a glance: high saturation, overly smooth textures, edges that are unnaturally sharp, lighting that is too perfect, gradients. There are several ways to reduce it:

Specify a textured style. Colored pencil, watercolor, pastel, crayon. These styles naturally include irregular strokes and textures, so they look less AI-generated than 3D render styles.

Lower saturation. Add “soft tones,” “low saturation,” or “muted colors” to the prompt. AI’s default colors tend to be highly saturated. Once you pull them down, the whole image feels much more comfortable.

Add a little imperfection. “Slightly hand-drawn,” “edges should not be too sharp,” “natural lighting, not over-HDR.” These small instructions make the final image feel less overly clean.

Avoid styles AI is best at. Hyperrealistic portraits, sci-fi scenes, 3D product renders. These are AI’s comfort zones, and the result often looks obviously AI-made. Imperfect styles like colored pencil and hand-drawn illustration tend to have much less AI look.

I use a colored-pencil style for almost all Penchan brand images for a simple reason: it is the least likely to look AI-generated at first glance.

Pitfall: The Penguin Beak Story

This pitfall deserves its own section because it shows a fundamental limitation of AI image generation.

The brand penguin has an orange rounded beak. A very simple feature, but AI keeps drawing it wrong.

The first instinct was that the prompt was not clear enough, so the line the penguin has an brown rounded beak, NOT yellow, NOT pointy was added. It improved things, but errors still appeared occasionally.

The likely reason is that the model applies the common appearance of a penguin, including a yellow pointed beak. No matter how strongly the prompt emphasizes the difference, the model is occasionally pulled back to that pattern.

The final solution was to combine reference images with text constraints. Attach one reference image with the correct beak, and also write “orange rounded beak” explicitly in the prompt. Only after using both did the success rate stabilize.

Lesson: AI output is strongly tied to training data. When what you want differs from common patterns in the training data, text description alone is not enough. Give it a visual reference.

Prompt Writing Differences by Tool

Comparison itemGemini (Nano Banana Pro / Nano Banana 2)Latest MidjourneyChatGPT built-in (ChatGPT Images)
LanguageChinese and English both workEnglish recommended for direct prompts; the website’s Conversational mode accepts ChineseChinese works through automatic conversational translation
FormatNatural language, no special syntaxNeeds parameters like --ar, --styleNatural language, conversational
Negative constraintsDirectly write “do not XX”Use --no parameterDirectly write “do not XX”
Reference imagesUpload an image with a text descriptionUpload a reference on the website; on Discord, put the image URL at the beginning of the promptAttach an image in the ChatGPT conversation
Style controlDescribe the style in text--style raw plus style keywordsDescribe in text, weaker control
Learning curveLowHighLow

For model-version differences, also see Gemini Free vs Pro differences.

Complete Image Generation Workflow

The flow from idea to finished image:

  1. First decide the purpose and placement of the image
  2. Write the prompt with the four-layer structure (subject, style, composition, constraints)
  3. If a brand character is involved, attach a reference image
  4. Generate 3-4 images and choose the closest one
  5. If none are right, adjust the weakest layer in the prompt and generate again
  6. After choosing, do final tweaks in Figma (add text, adjust colors, crop)

The whole process takes about 5-15 minutes per image. A new scene takes longer the first time because it needs more rounds to find the right direction.

Penchan’s experience

I first encountered AI image generation during the early Midjourney Discord era. Later, I moved my primary workflow to Gemini and ChatGPT for a simple reason: they follow Chinese instructions well, allow direct reference-image uploads, and keep brand characters much more consistent than text descriptions alone. I also tested Canva’s AI image generation for a while, but its gradient handling and overall texture did not suit my work, so I did not return to it.

“Colored pencil + no gradients” is the fixed foundation for my brand images. The default high-saturation, gradient-heavy, 3D texture of AI images is too easy to recognize at a glance. Colored pencil adds hand-drawn texture and irregularity, making it less likely to fall into that artificial look.

Building a prompt library is another habit I developed over the past few years. Whenever I find a useful instruction structure, I save it. The next time I need a similar image, changing a few words is much faster than starting from zero. The pen-pings series is how I organize and share these frequently used prompts.

Prompting has no finish line. Every time a tool version changes, methods that worked before may stop working, and different models produce different results. In the long run, the key to consistently producing usable images is building your own instruction library and iterating with tool versions, not clinging to one “god prompt.”

Further Reading


Compiled by Penna | Penchan