blog/use-cases·Sep 7, 2026·6 min·by the eroq team
YouTube thumbnails with AI images — contrast, one subject, space
How to render a thumbnail that works at phone size — widescreen, high contrast, one subject, and negative space where your title will go later.
A thumbnail is a picture that has to survive being looked at for a fifth of a second, at the size of a postage stamp, next to nine other pictures trying to do the same thing. Almost every rule that follows comes out of that one constraint. Here is how to render one in the Image studio, and the one thing you will have to do somewhere else.
Say it up front — eroq does not set type
There is no text tool, no title layer, no caption field and no font picker anywhere in the product. If you write "with the words BIG MISTAKE in yellow" into an image prompt, you will get something that looks like lettering from a distance and falls apart the moment anyone reads it. Rendered text is texture, not typography.
So the workflow splits cleanly in two: eroq renders the picture, and your editor sets the type. Anything will do the second half — Photoshop, Figma, Canva, Affinity. What matters is that you render the picture knowing the type is coming, which is what the rest of this article is about.
Put the words in the negative prompt rather than the positive one:
text, lettering, captions, subtitles, watermark, logo, signature, ui overlay
That stops the model decorating your clean background with unreadable scribble.
Widescreen, and nothing else
Set the format to 16:9 before you write the prompt. It renders at 1344 × 768 on all three image models, which is the widest native frame available, and it is the shape every video platform shows a thumbnail in.
Do not render a square and crop it later. A square composition with its top and bottom sliced off gives you a centered subject in a letterbox, which is exactly the thumbnail everyone scrolls past. A native 16:9 render places the subject off-center and gives you the horizontal room the format was asking for. The aspect ratios guide has the longer argument and the other four formats.
One subject, one idea
At thumbnail size the viewer resolves roughly three shapes. That is the whole budget.
- One subject. A face, an object, a silhouette. Not a face and a car and a chart.
- One action or state. Reaching, falling, glowing, broken. A verb reads at small size; a noun list does not.
- Big. Fill a third to a half of the frame with the subject. What looks aggressively oversized on your monitor looks correct on a phone.
Write the prompt that way too. "Close-up of a cracked ceramic mug on a dark counter, steam still rising" is a thumbnail. "A cozy kitchen scene with coffee, pastries, plants and morning light" is a stock photo that will disappear into the grid.
Contrast is most of the job
The reason thumbnails converge on the same look is not fashion, it is physics: at small sizes only luminance separation survives. Three ways to ask for it.
- Separate the subject from the ground by value, not by color. Bright subject on a dark ground, or dark silhouette against a bright one. Say it in the prompt: "lit subject against a deep unlit background".
- Name a hard light. Hard studio flash, a single rim light, hard low sun, neon glow as the key. Soft even lighting is the enemy of a thumbnail; it makes everything the same brightness and the image turns to mush.
- Limit the palette. Two dominant colors and one accent. "Deep teal background, warm amber key light" beats any adjective about vibrancy.
If a render looks great full screen and vanishes when you zoom out, it is a contrast problem nine times out of ten. Shrink the preview to a thumbnail-sized square on your own screen before judging any candidate.
Prompt for the empty space
This is the trick that makes AI thumbnails work with type. Do not render a full composition and then hunt for somewhere to put the title. Ask for the hole in the picture up front, as part of the framing.
Phrases that do this reliably: "subject in the right third, flat dark background across the left half", "generous negative space above the subject", "low horizon with empty sky in the upper two thirds", "clean uncluttered area on the left for copy".
Then decide where your type lives and stay consistent across the channel — one side, every time — so you can reuse the same layout file for every video and only swap the render.
Close-up of a weathered mechanic's hand holding a single bright red bolt against a deep unlit workshop background, subject filling the right third of a wide frame, hard rim light from the upper right, flat dark space across the left half, grease on the knuckles, amber and near-black palette, tense and deliberate.
Widescreen, one subject, one action, hard light, two colors, and a hole on the left with your title's name on it.
Which model, and what it costs
Image One for anything photographic — faces, hands, objects, real-world scenes. Image Anime for illustrated and character-led channels, where a consistent line style becomes channel identity. Image Art when the thumbnail is a mood rather than an object.
Every image is 10 credits on every model, and batches run 1 to 4. So a batch of four is 40 credits, and the honest workflow is two batches: four takes to find the composition, then four more on the winner with one clause changed. Eighty credits for a thumbnail you actually chose beats ten for the first thing that came out.
Every render lands in your library with its recipe, so once a thumbnail performs you can Remix it — same framing, same light, new subject — and the channel gets a visual grammar instead of a scrapbook. For the rest of the pipeline, the faceless YouTube workflow covers script, narration and footage, and the faceless YouTube page has the credit arithmetic per video.
The rule you cannot prompt around
Thumbnails are where people are most tempted to put a famous face. You cannot. Real identifiable people without their consent are out of bounds — no likenesses, no impersonation, no "in the style of a photograph of" a named person. It is an account-level rule, not a stylistic preference, and it is enforced.
Invented people are fine, including as a recurring face you keep across a channel. Save one as a Character with reference photos and @mention it in every thumbnail prompt, and you get a consistent host without hiring one.
FAQ
Can eroq put the title text on my thumbnail?
No. There is no text, title or caption tool, and lettering that appears inside a render is unreliable texture rather than real typography. Render the picture in eroq with deliberate empty space, then set the type in your own editor.
What size do thumbnails come out at?
16:9 renders at 1344 × 768 on all three image models, which is the widest native format available. There is no upscaler, so that is the file you work with — pick 16:9 at generation time rather than cropping a square afterwards.
How do I stop the model from adding fake text to the image?
Put "text, lettering, captions, watermark, logo" in the negative prompt, and avoid describing signs, screens or packaging in the positive prompt. The negative prompt entry explains how that list interacts with the rest of your wording.
Can I use a celebrity's face in a thumbnail?
No. Real identifiable people without consent are prohibited across the whole platform, thumbnails included. Build a recurring invented face instead — save it as a Character with reference photos and @mention it so it stays the same person across every video.
Render wide, light it hard, leave the hole for the title — open the Image studio and batch four.
Make this with the models behind the post — start with 50 free credits , or browse every engine and its price .