Build a Streaming LLM Chat App in Python with Reflex
Build a custom Python chat app with Reflex, async LLM streaming, conversation history, retryable errors, and responsive styling. Includes a complete runnable example.

Reflex is a Python framework for building streaming chat applications with a custom browser interface, Python event handlers, and per-session conversation state. You can combine chat with document search, model inputs, charts, approval forms, and other application pages. The UI is composed from components, so its layout is not limited to a fixed chatbot template.
This guide builds one complete text chat app using the OpenAI Responses API. It includes async streaming, a bounded conversation history, duplicate-submission protection, and recovery after a failed request. Document retrieval, persistent accounts, and deployment policies are extensions described below; they are not implemented by this sample.
What Reflex adds to a Python LLM integration
The model provider generates text. Reflex connects your Python application logic to an interactive web interface: input controls, streaming message display, loading states, errors, navigation, and styling. An event handler can call a hosted model, a local inference service, or a retrieval pipeline.
| Workflow | How to build it with Reflex |
|---|---|
| Streaming chat | An async event consumes provider deltas and yields UI updates. |
| Custom chat layout | Compose message cards, sidebars, forms, and responsive layouts using components and CSS properties. |
| Document question answering | Call a retrieval service in Python, then display the answer and source links in your own layout. |
| Model and image interfaces | Combine input forms, file upload, previews, and prediction outputs in the same application. |
| Analytics beside chat | Add charts, tables, and filters that share the relevant application state. |
| A specialized React control | Wrap an existing React component and expose its props and events to Python. |
See the AI applications guide, model and media examples, and linked XY charts tutorial for those adjacent workflows.
Set up the app
Follow the Reflex installation guide to create a blank app. In its project directory, install the provider SDK:
Set OPENAI_API_KEY in your backend environment using your shell or deployment secret settings. Do not put the key in a component, a State variable, or a committed source file. The app starts without a key and explains the missing configuration when you try to send a message.
The example defaults to gpt-5.4, the model used by this guide. Set OPENAI_MODEL to another Responses-compatible model available to your project if needed. Provider requests incur the provider's charges; Reflex does not supply a model subscription.
The example styles standard HTML elements with Tailwind classes. In rxconfig.py, keep your generated app name and include rx.plugins.TailwindV4Plugin() in the plugins list. If your blank template already includes it, keep that entry. See the Tailwind setup guide for the complete configuration.
Complete streaming chat app
Replace the blank app module with the following code. The provider request runs outside the state lock. Each short async with self section updates only this session's UI state.
ExpandCollapse
Run the app:
Send a short question, then a follow-up that refers to the answer. Text should appear progressively; the composer is disabled while a request runs. Clear the conversation to start over. A missing key produces a configuration message. A failed or incomplete provider request removes the partial turn, restores your question, and enables retry.
This is a text-only example. It renders generated content as text rather than executable HTML. It does not implement tool calling, attachments, persistent storage, or a Stop button.
How streaming and conversation history work
AsyncOpenAI makes the provider request asynchronously. The Responses streaming protocol emits typed events: text deltas update the assistant card, and the completion event confirms that the answer finished. An unexpected end, failure, refusal, timeout, or incomplete result is not silently saved as a successful answer.
The event uses a Reflex background task. State changes take short locks; network waits happen outside those locks. yield sends the updated UI state to the browser. A server-side processing guard rejects overlapping submissions even if a second event arrives before the button becomes disabled. This does not make CPU-heavy inference nonblocking: run that work in an appropriate worker or model service.
The sample retains the latest eight completed question-and-answer pairs and includes them with the next question. Older turns are discarded after a successful response. This is bounded recent context, not permanent memory. Input characters and output tokens have limits, and each request has a two-minute overall deadline. The output token budget includes reasoning tokens; an incomplete response can require a different budget or shorter request. Character limits are not exact token accounting.
The SDK's automatic retries are disabled here so a failed operation returns control to the user. A production application can use a deliberate retry policy, but should distinguish retryable failures from invalid configuration and account for repeated-request cost. Clearing this UI does not erase provider records. store=False disables storing the response for later API retrieval; it is not a promise of zero data retention. Consult the provider's data controls.
Customize the interface beyond a chat template
Change the component tree independently of the provider call. Place the conversation beside a document viewer, put settings in a sidebar, add a source panel beneath each answer, or insert an approval form before executing an action. Use CSS styling and responsive props for your layout. Integrate a specialized editor or viewer through React component wrapping.
Reflex also supports the data-app and model-demo workflows associated with Streamlit, Dash, Gradio, and NiceGUI. Start with the pandas data app, XY cross-filtering tutorial, and model/media interfaces. Compare framework tradeoffs in the framework comparison hub.
Streaming reduces the time before users can read part of an answer; it does not guarantee faster model generation or prove that one framework is universally fastest. Measure first visible text, update latency, total response time, and concurrent-session behavior for your workload. For high-frequency streams, batch deltas before updating state to avoid excessive messages and rendering. See performance and execution.
Add document retrieval and production services
For document question answering, retrieve authorized passages before calling the model, include those passages as context, and render source links with the answer. Retrieval needs its own ingestion, chunking, search, permission checks, and evaluation. The sample above does not implement those steps or guarantee that generated answers are grounded. The AI applications guide explains how to organize the broader workflow.
For another model provider, replace the SDK client, request parameters, stream parsing, and error handling. Keep the Reflex layout and the event-to-state pattern. Provider APIs are not necessarily drop-in replacements, especially for tool calls, conversation state, and multimodal input.
For deployment, add identity and authorization, per-user usage limits, persistent conversation ownership if needed, and observability around provider failures and latency. In-memory session state is not a durable chat database. A background event is not a job queue that survives restarts. Test two independent sessions and your expected concurrency before treating this sample as a production service. Continue with the chat tutorial and deployment guide.
Written by
Frequently asked questions
Can Reflex build a streaming Python chatbot?
Yes. A Reflex async event can consume a model provider's stream, update per-session state, and yield incremental text to the browser. The complete example above includes conversation history, duplicate-submission protection, and retryable error handling.
Can I customize the chat UI beyond a standard chat box?
Yes. Compose your own message cards, navigation, sidebars, document viewers, charts, forms, and result panels. Reflex supports CSS styling, responsive layouts, and wrapping React components for controls outside its existing library.
Can I use Reflex for RAG and other model providers?
Yes. Call retrieval and inference services from Python and display sources and results in your UI. You must implement the retrieval and permission logic, and adapt the provider-specific request, stream events, and errors. The text chat sample does not implement RAG.
Does the example remember the entire conversation?
No. It keeps the latest eight completed turns in the current session and sends them with the next request. Add durable storage and an explicit context strategy when you need longer histories, cross-device access, or conversations that survive restarts.
Is Reflex always faster than other Python UI frameworks?
There is no universal speed ranking. This example avoids blocking provider I/O and keeps state locks short. Model latency, update frequency, payload size, frontend complexity, and concurrent users still affect performance; measure the actual application.


