Back to all articles
FastAPI & Python12 min readMarch 10, 2026

Streaming AI & LLM Responses with FastAPI Server-Sent Events (SSE) into Flutter

A step-by-step implementation guide to real-time AI token streaming: building async SSE generators in FastAPI and rendering token-by-token typewriter animations in Flutter.

Bayajit Islam

Written by Bayajit Islam

Freelance Flutter & Backend Developer • Dhaka, Bangladesh

Streaming AI & LLM Responses with FastAPI Server-Sent Events (SSE) into Flutter
When OpenAI launched ChatGPT, it fundamentally redefined user expectations for AI-driven mobile interfaces. Waiting 15 to 30 seconds in complete silence for a Large Language Model to generate an entire 500-word response creates a terrible user experience that feels unresponsive and broken. Users expect immediate visual feedback: seeing tokens stream onto the mobile screen word-by-word with typewriter fluidity the instant they tap send. In this comprehensive technical tutorial, I demonstrate how to build an end-to-end AI token streaming pipeline using FastAPI Server-Sent Events (SSE) on the backend and reactive stream consumers in Flutter.

1. Why Server-Sent Events (SSE) Beat WebSockets for LLMs

When developers consider streaming data to a mobile app, they often default to WebSockets. However, WebSockets are bidirectional, require stateful connection handshakes, complicate load balancer stickiness, and do not play nicely with HTTP/2 multiplexing.

For LLM generation, the communication flow is fundamentally unidirectional: the mobile client sends a prompt once via HTTP POST, and the server streams back tokens continuously until complete. Server-Sent Events (SSE) operate over standard HTTP/1.1 or HTTP/2, require zero custom protocols, integrate naturally with existing authentication headers, and auto-reconnect if the cellular connection hiccups.

Because SSE uses standard HTTP, it traverses corporate firewalls and mobile proxy filters without special port configurations, ensuring rock-solid connectivity across diverse cellular networks.

Server-Sent Events (SSE) operate over standard HTTP using Content-Type: text/event-stream. They bypass the complexity of stateful WebSocket servers while providing sub-millisecond token delivery to mobile clients.

2. Building the Asynchronous Token Generator in FastAPI

In FastAPI, streaming is achieved using Starlette's StreamingResponse combined with an asynchronous Python generator. We stream tokens from an LLM provider (such as OpenAI, Anthropic, or a locally hosted HuggingFace model via Ollama) and format them into the standard SSE event format: 'data: {json} '.

By consuming the LLM stream asynchronously via an async for loop, FastAPI yields each token chunk to the network socket immediately without accumulating tokens in server memory.

FastAPI StreamingResponse yielding formatted SSE token payloads
from fastapi import FastAPI, Depends
from fastapi.responses import StreamingResponse
from pydantic import BaseModel
import asyncio
import json

app = FastAPI()

class ChatPrompt(BaseModel):
    message: str

async def ai_token_generator(prompt: str):
    # Simulated token stream from an LLM engine
    mock_tokens = ["Flutter", " and", " FastAPI", " deliver", " unprecedented", " mobile", " AI", " speed."]
    for token in mock_tokens:
        await asyncio.sleep(0.08)  # Realistic token generation interval
        payload = json.dumps({"token": token, "done": False})
        yield f"data: {payload}

"
    
    yield f"data: {json.dumps({'token': '', 'done': True})}

"

@app.post("/v1/chat/stream")
async def stream_chat_response(body: ChatPrompt):
    return StreamingResponse(
        ai_token_generator(body.message),
        media_type="text/event-stream",
        headers={
            "Cache-Control": "no-cache",
            "Connection": "keep-alive",
            "X-Accel-Buffering": "no",  # Disables Nginx/Caddy proxy buffering
        },
    )

3. Consuming SSE in Flutter with Dio and StreamTransformers

On the Flutter side, handling SSE requires consuming the HTTP response body as a raw byte stream rather than decoding a completed JSON string. We use Dio with ResponseType.stream.

By piping the response stream through utf8.decoder and LineSplitter(), we extract each 'data: ' line as it arrives over the network, parse the token JSON, and append it to our reactive state (BLoC or Riverpod). The UI widget listens to this state and rebuilds smoothly at 60fps as new tokens arrive.

To prevent UI micro-stutters during high-speed token generation (such as 100+ tokens per second), debounce state notifications slightly (e.g. 16ms window) so that Flutter paints at a smooth 60fps frame rate without unnecessary re-layout thrashing.

Flutter Dio stream pipeline extracting SSE tokens for real-time UI rendering
Future<void> streamAiResponse(String prompt, Function(String) onToken) async {
  final dio = Dio();
  final response = await dio.post<ResponseBody>(
    'https://api.bayajitislam.com/v1/chat/stream',
    data: {'message': prompt},
    options: Options(responseType: ResponseType.stream),
  );

  response.data!.stream
      .cast<List<int>>()
      .transform(utf8.decoder)
      .transform(const LineSplitter())
      .listen((line) {
        if (line.startsWith('data: ')) {
          final jsonStr = line.substring(6);
          final data = jsonDecode(jsonStr);
          if (data['done'] == false) {
            onToken(data['token'] as String);
          }
        }
      });
}

4. Managing Stream State with Riverpod or BLoC

In production Flutter applications, streaming tokens should not be stored in naive StatefulWidget setState variables. Doing so causes full-screen rebuilds on every single token, degrading frame rates.

Instead, encapsulate the streaming state inside a Riverpod AsyncNotifier or BLoC. As each token arrives from the Dio stream, append it to a StringBuffer and emit an updated state to the UI. Wrap only the active message bubble inside a Consumer widget to isolate rebuilds strictly to that specific chat item.

5. Overcoming Reverse Proxy Buffering Gotchas

A notorious issue in production AI streaming occurs when reverse proxies (Nginx, Caddy, Cloudflare) buffer incoming server chunks until they reach 4KB before flushing them to the mobile device. This ruins the real-time effect, causing tokens to arrive in awkward batches.

Always include the header 'X-Accel-Buffering: no' in your FastAPI response, and ensure your proxy configuration disables proxy_buffering for SSE routes.

Additionally, configure your cloud load balancer timeouts to allow long-running streaming connections (e.g. 120 seconds) so that extensive reasoning models do not encounter gateway 504 errors mid-generation.

Final Thoughts

Real-time streaming transforms AI from a sluggish novelty into an interactive and tactile conversation. By combining FastAPI's async StreamingResponse with Flutter's reactive stream architecture, you deliver the gold standard in mobile AI user experience.

Key Takeaways

  • Server-Sent Events (SSE) provide lightweight, unidirectional streaming over standard HTTP.
  • Always disable proxy buffering with 'X-Accel-Buffering: no' for instant token delivery.
  • Use Dio ResponseType.stream and LineSplitter in Flutter to parse incoming token chunks.
  • Implement cancel tokens to stop upstream LLM generation when users interrupt.
  • Render real-time tokens using flutter_markdown for rich formatted mobile chat displays.

Frequently Asked Questions

Does SSE work reliably over cellular networks on mobile?

Yes. SSE operates over standard HTTP/2 and HTTP/1.1 connections. Modern mobile OS network stacks handle SSE effortlessly, and Flutter's LineSplitter pipeline processes incoming chunks without dropping packets.

How do I secure an SSE streaming endpoint?

Because the request is initiated with a standard HTTP POST or GET, you pass your standard Bearer JWT authorization token in the request headers, exactly like any other secure FastAPI endpoint.

Can I cancel an in-progress LLM generation stream from Flutter?

Yes. When the user taps a 'Stop Generating' button, call responseStream.cancel() or cancelToken.cancel() in Dio. FastAPI detects the broken client socket and immediately terminates the upstream LLM generator, saving API token costs.

How do you handle Markdown rendering while tokens are actively streaming?

Use package:flutter_markdown with an incremental string buffer. Flutter's layout engine can incrementally parse Markdown headers, bullet points, and code syntax blocks without crashing on unclosed tags.

How do you track token usage and billing for streaming requests?

Accumulate generated tokens on the server during the generator loop. In the generator finally block or completion event, send a telemetry event to your database or billing service recording the total input and output token count.

Tags:FastAPIAILLMSSEFlutterStreaming
Bayajit Islam

Need an AI Mobile App or Scalable Backend?

I'm Bayajit Islam, an AI Mobile App Developer with 2+ years of hands-on experience architecting cross-platform apps for iOS, Android & Desktop with Flutter, paired with high-performance Python & FastAPI backends, streaming LLMs, and DevOps.