1. Why Server-Sent Events (SSE) Beat WebSockets for LLMs
When developers consider streaming data to a mobile app, they often default to WebSockets. However, WebSockets are bidirectional, require stateful connection handshakes, complicate load balancer stickiness, and do not play nicely with HTTP/2 multiplexing.
For LLM generation, the communication flow is fundamentally unidirectional: the mobile client sends a prompt once via HTTP POST, and the server streams back tokens continuously until complete. Server-Sent Events (SSE) operate over standard HTTP/1.1 or HTTP/2, require zero custom protocols, integrate naturally with existing authentication headers, and auto-reconnect if the cellular connection hiccups.
Because SSE uses standard HTTP, it traverses corporate firewalls and mobile proxy filters without special port configurations, ensuring rock-solid connectivity across diverse cellular networks.
2. Building the Asynchronous Token Generator in FastAPI
In FastAPI, streaming is achieved using Starlette's StreamingResponse combined with an asynchronous Python generator. We stream tokens from an LLM provider (such as OpenAI, Anthropic, or a locally hosted HuggingFace model via Ollama) and format them into the standard SSE event format: 'data: {json} '.
By consuming the LLM stream asynchronously via an async for loop, FastAPI yields each token chunk to the network socket immediately without accumulating tokens in server memory.
from fastapi import FastAPI, Depends
from fastapi.responses import StreamingResponse
from pydantic import BaseModel
import asyncio
import json
app = FastAPI()
class ChatPrompt(BaseModel):
message: str
async def ai_token_generator(prompt: str):
# Simulated token stream from an LLM engine
mock_tokens = ["Flutter", " and", " FastAPI", " deliver", " unprecedented", " mobile", " AI", " speed."]
for token in mock_tokens:
await asyncio.sleep(0.08) # Realistic token generation interval
payload = json.dumps({"token": token, "done": False})
yield f"data: {payload}
"
yield f"data: {json.dumps({'token': '', 'done': True})}
"
@app.post("/v1/chat/stream")
async def stream_chat_response(body: ChatPrompt):
return StreamingResponse(
ai_token_generator(body.message),
media_type="text/event-stream",
headers={
"Cache-Control": "no-cache",
"Connection": "keep-alive",
"X-Accel-Buffering": "no", # Disables Nginx/Caddy proxy buffering
},
)3. Consuming SSE in Flutter with Dio and StreamTransformers
On the Flutter side, handling SSE requires consuming the HTTP response body as a raw byte stream rather than decoding a completed JSON string. We use Dio with ResponseType.stream.
By piping the response stream through utf8.decoder and LineSplitter(), we extract each 'data: ' line as it arrives over the network, parse the token JSON, and append it to our reactive state (BLoC or Riverpod). The UI widget listens to this state and rebuilds smoothly at 60fps as new tokens arrive.
To prevent UI micro-stutters during high-speed token generation (such as 100+ tokens per second), debounce state notifications slightly (e.g. 16ms window) so that Flutter paints at a smooth 60fps frame rate without unnecessary re-layout thrashing.
Future<void> streamAiResponse(String prompt, Function(String) onToken) async {
final dio = Dio();
final response = await dio.post<ResponseBody>(
'https://api.bayajitislam.com/v1/chat/stream',
data: {'message': prompt},
options: Options(responseType: ResponseType.stream),
);
response.data!.stream
.cast<List<int>>()
.transform(utf8.decoder)
.transform(const LineSplitter())
.listen((line) {
if (line.startsWith('data: ')) {
final jsonStr = line.substring(6);
final data = jsonDecode(jsonStr);
if (data['done'] == false) {
onToken(data['token'] as String);
}
}
});
}4. Managing Stream State with Riverpod or BLoC
In production Flutter applications, streaming tokens should not be stored in naive StatefulWidget setState variables. Doing so causes full-screen rebuilds on every single token, degrading frame rates.
Instead, encapsulate the streaming state inside a Riverpod AsyncNotifier or BLoC. As each token arrives from the Dio stream, append it to a StringBuffer and emit an updated state to the UI. Wrap only the active message bubble inside a Consumer widget to isolate rebuilds strictly to that specific chat item.
5. Overcoming Reverse Proxy Buffering Gotchas
A notorious issue in production AI streaming occurs when reverse proxies (Nginx, Caddy, Cloudflare) buffer incoming server chunks until they reach 4KB before flushing them to the mobile device. This ruins the real-time effect, causing tokens to arrive in awkward batches.
Always include the header 'X-Accel-Buffering: no' in your FastAPI response, and ensure your proxy configuration disables proxy_buffering for SSE routes.
Additionally, configure your cloud load balancer timeouts to allow long-running streaming connections (e.g. 120 seconds) so that extensive reasoning models do not encounter gateway 504 errors mid-generation.
Final Thoughts
Real-time streaming transforms AI from a sluggish novelty into an interactive and tactile conversation. By combining FastAPI's async StreamingResponse with Flutter's reactive stream architecture, you deliver the gold standard in mobile AI user experience.
Key Takeaways
- Server-Sent Events (SSE) provide lightweight, unidirectional streaming over standard HTTP.
- Always disable proxy buffering with 'X-Accel-Buffering: no' for instant token delivery.
- Use Dio ResponseType.stream and LineSplitter in Flutter to parse incoming token chunks.
- Implement cancel tokens to stop upstream LLM generation when users interrupt.
- Render real-time tokens using flutter_markdown for rich formatted mobile chat displays.
Frequently Asked Questions
Does SSE work reliably over cellular networks on mobile?
Yes. SSE operates over standard HTTP/2 and HTTP/1.1 connections. Modern mobile OS network stacks handle SSE effortlessly, and Flutter's LineSplitter pipeline processes incoming chunks without dropping packets.
How do I secure an SSE streaming endpoint?
Because the request is initiated with a standard HTTP POST or GET, you pass your standard Bearer JWT authorization token in the request headers, exactly like any other secure FastAPI endpoint.
Can I cancel an in-progress LLM generation stream from Flutter?
Yes. When the user taps a 'Stop Generating' button, call responseStream.cancel() or cancelToken.cancel() in Dio. FastAPI detects the broken client socket and immediately terminates the upstream LLM generator, saving API token costs.
How do you handle Markdown rendering while tokens are actively streaming?
Use package:flutter_markdown with an incremental string buffer. Flutter's layout engine can incrementally parse Markdown headers, bullet points, and code syntax blocks without crashing on unclosed tags.
How do you track token usage and billing for streaming requests?
Accumulate generated tokens on the server during the generator loop. In the generator finally block or completion event, send a telemetry event to your database or billing service recording the total input and output token count.



