A discount on premium plans when you sign upGet your discount
eroq

blog/developers·Aug 12, 2026·4 min·by the eroq team

Streaming AI responses over SSE — a practical guide

Server-sent events end to end — parsing chunk frames, relaying streams through your own backend, and the buffering bugs that only appear in production.


Streaming is the difference between a character who is present and a spinner. The mechanics of SSE are simple; the bugs are all in the details this guide covers.

The wire format

With stream: true, /v1/chat/completions answers text/event-stream. Each frame is a data: line and a blank line; deltas arrive in the standard chunk shape:

data: {"id":"cmpl_…","object":"chat.completion.chunk","choices":[{"delta":{"content":"That "}}]}

data: {"id":"cmpl_…","object":"chat.completion.chunk","choices":[{"delta":{"content":"noise"}}]}

data: {"object":"chat.completion.usage","usage":{"credits_spent":3,"credits_remaining":997}}

data: [DONE]

Two eroq-specific frames worth knowing: the usage event just before [DONE] carries the meter, and an error event ({"error":{…}}) replaces the crash you would otherwise have to infer from a dropped connection. If the stream errors before any content arrived, the call has already refunded itself.

Parsing without the classic bug

The classic bug: treating every network chunk as a complete frame. TCP does not respect your line breaks — a frame can arrive split across reads. Buffer, split on \n\n, and keep the remainder:

const reader = res.body.getReader()
const decoder = new TextDecoder()
let buffer = ''
while (true) {
  const { done, value } = await reader.read()
  if (done) break
  buffer += decoder.decode(value, { stream: true })   // stream: true matters for UTF-8
  const frames = buffer.split('\n\n')
  buffer = frames.pop() ?? ''                          // last piece may be incomplete
  for (const frame of frames) {
    const data = frame.replace(/^data: /, '')
    if (data === '[DONE]') continue
    const parsed = JSON.parse(data)
    if (parsed.error) throw new Error(parsed.error.message)
    const delta = parsed.choices?.[0]?.delta?.content
    if (delta) render(delta)
  }
}

The { stream: true } on decoder.decode is not decoration: without it, a multi-byte character split across chunks becomes mojibake. Emoji-heavy roleplay finds this bug within the hour.

Relaying through your backend

Never ship your API key to a browser — relay the stream through the backend that already owns your key and your character sheets. The relay is thin, but two details make or break it:

// Node/Express-style relay
app.post('/chat', async (req, res) => {
  res.setHeader('Content-Type', 'text/event-stream')
  res.setHeader('Cache-Control', 'no-cache')
  res.flushHeaders()                                   // 1. headers out immediately

  const upstream = await fetch('https://eroq.ai/v1/chat/completions', { /* … */ })
  for await (const chunk of upstream.body) {
    res.write(chunk)                                   // 2. relay bytes, re-frame nothing
  }
  res.end()
})
  1. Flush headers immediately, or your reverse proxy may buffer the whole response and deliver it at once — streaming that arrives as a block. On nginx, also set X-Accel-Buffering: no.
  2. Relay bytes verbatim. Parsing and re-serializing frames in the relay doubles your bug surface for zero value. Parse on the client, where you render.

UX details that separate good from great

  • Render on a small timer (30–50ms), not per-delta — per-token DOM writes jank on mobile.
  • Show the first token fast, then let it flow. Perceived latency lives almost entirely in time-to-first-token.
  • On mid-stream failure, keep the partial text and offer a retry affordance. A half-reply that stays beats a reply that vanishes — and since interrupted streams past first output were still generated, keeping the text respects what was paid for.
  • Let users abort. Closing your relay's response should close the upstream request; an abandoned stream you keep consuming is money spent rendering to nobody.

FAQ

How do I parse SSE chunks from the eroq API correctly?

Buffer the bytes, split on \n\n, and keep the trailing fragment for the next read — a frame can arrive split across network chunks. Pass { stream: true } to TextDecoder.decode so multi-byte characters survive a split, and skip the [DONE] sentinel before parsing JSON.

Does streaming cost more than a normal completion?

No. Credits are charged per completion, not per token, so a streamed reply and a buffered one both cost the same 1 credit on RP mini or 3 on RP+ — whether the reply runs to 50 tokens or the 1,200-token cap. There is no pricing reason to leave stream: false.

What is the usage event at the end of the stream?

A frame just before data: [DONE] carrying credits_spent and credits_remaining, so you can update a balance display without a second API call. An error frame shaped {"error":{…}} replaces the dropped connection you would otherwise have to infer, and a stream that fails before any content has already refunded itself.

Why does my stream arrive all at once instead of token by token?

Almost always a proxy buffering the response. Flush your headers before the first write, and on nginx set X-Accel-Buffering: no. The second usual suspect is re-serializing frames inside your relay — relay the bytes verbatim and parse on the client, where you render.

Flat pricing has one more consequence here: streaming costs exactly what buffering costs — the same 1 or 3 credits. There is no reason not to stream, which is why every example in the docs does.

Tagsstreamingssechat api

Make it in eroq

Build it on the API

An API key takes a minute. The same engines, the same prices, async jobs and webhooks.

50 free credits · no card · private mode on every account