blog/tutorials·Jul 28, 2026·4 min·by the eroq team
Image generation for character products — prompts, consistency, cost
Prompting a diffusion engine well, keeping a character recognizable across generations, and budgeting image features at 10 credits a render.
Chat retains; images convert. A character the user can see — in new scenes, new outfits, new moods — is the single strongest premium mechanic we know in this category. Here is how to build the feature well against eroq-krea2.
Speaking diffusion
Krea2 is a diffusion engine, and diffusion has a dialect. It reads both prose and comma-separated tags; what it rewards is specificity in the right order:
subject and identity, action or pose, outfit, setting, lighting, camera, style
{
"model": "eroq-krea2",
"prompt": "young woman with silver bob and amber eyes, leaning on a workbench, grease-stained overalls, spaceship hangar, warm practical lights, 35mm, shallow depth of field",
"negative_prompt": "blurry, extra fingers, watermark, text"
}
Three rules that do most of the work:
- Front-load identity. Tokens earlier in the prompt weigh more. The character's defining traits — hair, eyes, build — come first, every time.
- Always send a negative prompt.
blurry, extra fingers, watermark, textis the floor; add whatever your style must exclude. It is a free quality lever. - One scene per prompt. Diffusion averages competing instructions into mush. "In the hangar AND at the beach" produces neither.
cfg_scale (prompt adherence) defaults sensibly; raise it toward 9–12 when the engine takes too much creative liberty, drop toward 4–6 when results look overcooked.
The consistency problem
The hard problem in character imagery is that a diffusion model invents a new face every call. Prompting alone never fully solves it — what does most of the work is a reference photo. Store your cast as Characters with their reference images attached, then pass the character in elements as char:<id> and the same face comes back render after render. Prompt discipline carries the rest:
- Fix a canonical description. Write one 15–25 token identity block per character and prepend it verbatim to every image prompt. Word-for-word stability matters; synonyms drift the face.
- Fix the style block too. A consistent rendering style ("35mm, soft grain") makes faces read as the same person even when features wobble.
- Seed galleries, don't stream them. For a character's public gallery, generate in batches, curate the on-model results, discard the rest. At 10 credits a render, a curated 12-image gallery costs about $1.60 including rejects — price the feature, not the attempt.
Refused or empty generations refund automatically, and a prompt the engine declines returns an explicit content_blocked code rather than a silent failure — so your retry logic can tell "rephrase" from "try again".
Where images fit the product
The pattern that converts, in order of effort:
- The reveal — a one-time "see them" moment early in a relationship. One image, massive activation effect.
- Scene stills — user-triggered "show me this moment" during chat. Charge your users per render; your cost is a known 10 credits.
- The gallery — curated, drip-released, subscription-gated. Batch-generated off-peak.
Each maps to one POST /v1/images/generations call — the reference covers the parameters, including batch for 1–4 renders in a call and aspect for portrait or landscape frames. Outputs come back inline as base64 for you to store, and every render is also saved to your Library with the recipe that made it, so a character's best takes stay findable and re-runnable later. Pass private: true when you would rather eroq kept nothing: no file, no library entry, and the prompt shows as stars in the usage ledger.
Budget table
| Feature | Renders/user/mo | Credits | ~Cost |
|---|---|---|---|
| Reveal (once) | 1 | 10 | $0.10 |
| Scene stills | 12 | 120 | $1.20 |
| Gallery drops (curated 3:1) | 16 | 160 | $1.60 |
Against a $10–15/month subscription, imagery lands comfortably inside margin while being the most visible thing the subscription buys.
FAQ
How do I keep the same face across AI-generated images?
Attach reference photos to a Character and pass it into the call as char:<id>; the engine holds the face instead of inventing a new one each time. Back that up with a canonical 15–25 token identity block, repeated word for word, and a fixed style block. Expect to curate — generate a few, keep the on-model ones, at 10 credits a render.
Which image model should a character product use?
All three cost 10 credits and share the same parameters. Image One is the photoreal flagship and the right default for selfies and scene stills. Image Anime handles manga and illustration; Image Art is painterly and follows long descriptive prompts more faithfully.
What should go in a negative prompt?
blurry, extra fingers, watermark, text is the floor for character work, plus whatever your style must exclude. It costs nothing and it is the cheapest quality lever available. Pair it with cfg_scale — raise toward 9–12 when the engine takes too much liberty, drop toward 4–6 when results look overcooked.
What content is allowed for my characters?
Fiction under the written acceptable-use policy — dark and intense themes included — as long as every character is 18 or older. Anything involving minors, real identifiable people without consent and anything illegal are banned absolutely; those requests return an explicit content_blocked code and are never charged, so your retry logic can tell "rephrase" from "try again".
That combination — cheap to run, premium to perceive — is why imagery is the first feature we tell builders to add after chat.
CharactersMake it in eroq
Cast your own in Characters
Everything above happens in the studio: same engines, same settings, rendered on your click.
50 free credits · no card · private mode on every account