Streaming lets an application display output as it arrives rather than waiting for the complete response. For a supported chat model, set stream: true on a chat completions request. Gatepx returns a stream of events that your application must read until completion or failure.
Read text deltas
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.gatepx.ai/v1",
apiKey: process.env.GATEPX_API_KEY,
maxRetries: 0,
});
let text = "";
try {
const stream = await client.chat.completions.create({
model: process.env.GATEPX_MODEL,
messages: [{ role: "user", content: "Explain streaming in three sentences." }],
stream: true,
});
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content ?? "";
text += delta;
process.stdout.write(delta);
}
console.log("\nGeneration completed.");
} catch (error) {
console.error("Generation interrupted; treat received text as partial.");
throw error;
}This example runs on a server and prints text to its terminal. If a browser needs live output, have your backend forward the stream to it. Keep the Gatepx API key on the backend rather than constructing an authenticated Gatepx client in the browser.
Not every chunk contains text
A stream may include events with an empty choices array or a delta without content. Optional access and an empty-string fallback prevent those events from breaking a text renderer. Tool calls and other structured fields require their own handling; they are not ordinary text deltas.
The SDK reads the event framing for you. If you use raw HTTP instead, parse the documented server-sent event format rather than treating each received network packet as one complete JSON response. Packet boundaries and event boundaries are different.
Distinguish partial output from success
A connection can fail after some text has arrived. Your interface should retain the distinction between a completed answer and a partial answer. Do not save partial output as a successful result just because the initial HTTP connection opened.
A retry is a new generation. It can produce different text and may incur another charge. Do not append a restarted generation to the old partial output as if the two were one continuous response. Decide whether to offer an explicit retry, discard the partial result, or show it with an interrupted status.
Test the interface, not just the first token
- Verify that the application handles both non-text chunks and a normal completed stream.
- Check how cancellation or a dropped connection is shown to the reader.
- Make sure a second attempt creates a separate result rather than silently duplicating partial text.
- Check request and usage records without logging the API key or sensitive prompt content.
Streaming changes delivery timing, not the need to evaluate a model or review its pricing. The streaming documentation is the current API reference, and the troubleshooting checklist covers the diagnostic information to retain when a generation fails.