stream parameter to true in your request. The model will then stream the response to the client in chunks, rather than returning the entire response at once.
Here is an example of how to stream a response, and process it:
Additional information
For SSE (Server-Sent Events) streams, OpenRouter occasionally sends comments to prevent connection timeouts. These comments look like:data: fields, and buffering for you:
eventsource-parser
data: events with an error field. See Handling Errors During Streaming below.
The generation ID is returned in the X-Generation-Id response header for all endpoints (chat completions, completions, responses, and messages), which can be useful for debugging and correlating requests.
Some SSE client implementations might not parse the payload according to spec, which leads to an uncaught error when you JSON.stringify the non-JSON payloads. We recommend the following clients:
The final usage chunk (Chat Completions)
On the Chat Completions endpoint (/api/v1/chat/completions), every stream ends with an extra chunk that carries the usage object for the request, sent just before the [DONE] message. OpenAI’s spec emits this chunk with an empty choices array, but many clients crash when accessing choices[0].delta on it, so OpenRouter intentionally deviates: the usage chunk contains one choice with a content-free delta that repeats the finish_reason (and native_finish_reason) of the stream.
finish_reason appears twice: once on the last content-bearing chunk and again on the usage chunk. Clients that validate streams should treat the usage chunk as an accounting frame rather than a second terminal event.
This shape is specific to Chat Completions. Other endpoints follow their own specs: the Responses API (/api/v1/responses) reports usage in the response.completed event, and the Messages API (/api/v1/messages) reports it in the message_delta event before message_stop.
Stream cancellation
Streaming requests can be cancelled by aborting the connection. For supported providers, this immediately stops model processing and billing.Provider Support
Provider Support
Supported
- OpenAI, Azure, Anthropic
- Fireworks, Mancer, Recursal
- AnyScale, Lepton, OctoAI
- Novita, DeepInfra, Together
- Cohere, Hyperbolic, Infermatic
- Avian, XAI, Cloudflare
- SFCompute, Nineteen, Liquid
- Friendli, Chutes, DeepSeek
- AWS Bedrock, Groq, Modal
- Google, Google AI Studio, Minimax
- HuggingFace, Replicate, Perplexity
- Mistral, AI21, Featherless
- Lynn, Lambda, Reflection
- SambaNova, Inflection, ZeroOneAI
- AionLabs, Alibaba, Nebius
- Kluster, Targon, InferenceNet
Handling errors during streaming
OpenRouter handles errors differently depending on when they occur during the streaming process:Errors before the response is committed
If an error occurs before OpenRouter has committed the response, you get a standard JSON error response with the appropriate HTTP status code. That covers failures raised before the request reaches a provider, and provider failures visible at connection time such as a connection error or a non-2xx upstream status.- 400: Bad Request (invalid parameters)
- 401: Unauthorized (invalid API key)
- 402: Payment Required (insufficient credits)
- 429: Too Many Requests (rate limited)
- 502: Bad Gateway (provider error)
- 503: Service Unavailable (no available providers)
Errors after the response is committed (mid-stream)
Once the provider has returned response headers, the200 OK status is committed even if no token has been produced yet. Any error after that point arrives as an SSE event rather than as an HTTP status:
- The error appears at the top level alongside standard response fields (id, object, created, etc.)
- A
choicesarray is included withfinish_reason: "error"to properly terminate the stream - The HTTP status remains 200 OK since headers were already sent
- The stream is terminated after this unified error event
- The error can be the first and only event in the stream, so treat a
200carrying anerrorchunk with no content as a failure, not a success
Code examples
Here’s how to properly handle both types of errors in your streaming implementation:API-specific behavior
Different API endpoints may handle streaming errors slightly differently:- OpenAI Chat Completions API: Returns
ErrorResponsedirectly if no chunks were processed, or includes error information in the response if some chunks were processed - OpenAI Responses API: May transform certain error codes (like
context_length_exceeded) into a successful response withfinish_reason: "length"instead of treating them as errors