blog/developers·Aug 14, 2026·4 min·by the eroq team
Let your users send pictures — vision in roleplay chat
Image input is the highest-retention feature per credit in character products. The UX patterns, the one-request implementation, and your moderation duties.
Text roleplay has a one-way intimacy problem: the character describes her world, the user can only describe his back. Vision input closes the loop — he sends a photo of his desk, his dog, his dinner, and she reacts to his actual day. In production character apps, that reaction moment is one of the strongest retention events per credit spent.
The feature, mechanically
Chat completions accept image parts on user turns: content becomes an array mixing text and image_url entries (https URL or data URI, up to 2 images per request). Each attached image adds 2 credits on top of the completion — a photo-reaction turn on RP mini costs 3 credits total, about three cents.
{
"model": "eroq-rp-mini",
"context": "Mira: sardonic starship mechanic. Relationship: three weeks in.",
"messages": [{
"role": "user",
"content": [
{ "type": "text", "text": "look what i built today" },
{ "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,…" } }
]
}]
}
The engine looks at the image and the character responds in character — not with an image caption. That distinction is the product: "That's a beautiful golden retriever" is a vision demo; "wipes grease off her hands okay, the dog is cuter than you. What's his name?" is a relationship.
UX patterns that earn their credits
- The reaction moment. A camera button in the chat composer, full stop. Users discover it themselves; the first "she noticed the details" reply is the hook.
- Prompted shares. Have the character ask — "show me where you're sitting right now" — at natural beats. Prompted photos convert 3–4× better than a passive button, and they pace your vision spend.
- Memory callbacks. Fold what she saw into your rolling summary ("his desk faces a window; the dog is Biscuit"). A callback two days later — "how's Biscuit?" — is the cheapest wow in the category, and it costs zero extra credits because it's just context.
- Cap it visibly. Two images per request is the API ceiling; a per-day allowance in your free tier keeps the unit economics boring. Vision is cheap per event, not free at scale — budget it like every other flat price.
The responsibilities that stay yours
User-submitted images are user-generated content, and the duties that come with UGC don't transfer to your model provider:
- Scan uploads on your side (CSAM detection against industry hash lists at minimum) before the API call. eroq's acceptable-use policy bans minors and non-consensual real-person content absolutely — requests outside it return
content_blockedand are never charged — but detection tooling on the upload path is your legal surface, not a nice-to-have. - Age-gate the feature with the rest of your product. Photo exchange belongs behind the same age gate and policy as the roleplay itself.
- Store nothing you don't need. Pass data URIs through and keep only what the product requires (the summary line, not the photo). RP+ and RP mini are eroq's own models and retain nothing upstream. If the character sends pictures back, that's image generation — and those you can host durably with
store: true.
FAQ
How many images can a user attach to one message?
Up to 2 per request, as https URLs or data URIs, on user turns only. Each one adds 2 credits on top of the completion, so a photo turn on RP mini costs 3 credits all-in and the same turn on RP+ costs 5.
Do both roleplay models support vision?
Yes — RP+ and RP mini both accept image parts, at the same +2 credits per image. Since the surcharge is identical, route photo turns to mini unless the picture is the subject of the reply, and keep the flagship for the beats that carry weight.
Does the character actually see the photo, or just guess?
The engine looks at the image and answers in character, which is the whole product difference. A vision demo captions the picture; a companion reacts to it — noticing the mess on the desk rather than listing the objects on it. Fold what she saw into your rolling summary and the callback two days later costs no extra credits.
Who is responsible for moderating what users upload?
You are. Scan uploads on your side before the API call — CSAM detection against industry hash lists at minimum — and age-gate the feature with the rest of your product. eroq's acceptable-use policy bans minors and non-consensual real-person content absolutely, and out-of-policy requests return content_blocked without a charge, but detection on the upload path is your legal surface, not ours.
Ship it in an afternoon
Vision input is the rare feature that's one composer button, one array change in an existing call, and no new infrastructure. A key and the 50 free credits cover the whole test: send the API a photo of your own desk and watch the character notice the coffee cups. That reaction is the feature — the rest is product discipline.
Make it in eroq
Build it on the API
An API key takes a minute. The same engines, the same prices, async jobs and webhooks.
50 free credits · no card · private mode on every account