pavelsavara/blazor-multi-node

Make a Blazor Interactive Server circuit survive an app-pod restart in Kubernetes: resume the same circuit on another pod via Redis (HybridCache + Data Protection). Local kind cluster, .NET 10.

★ 0Forks 0C#GitHub ↗Compare

README

Blazor Server circuit persistence across pods

A small, complete example showing how to make a Blazor Interactive Server app survive an app-pod restart in Kubernetes: the browser reconnects to a different pod and resumes the same circuit — counter and demo state intact — instead of starting over.

It runs as 2 app replicas on a local kind cluster with a Redis pod. The page shows which pod is serving you; delete that pod and watch the state survive on another node.

If you know Blazor and the Kubernetes basics but have never wired up circuit persistence, this repo is the missing how-to. The five ingredients below are the whole trick.


The problem

Blazor Interactive Server keeps each user's UI state in a circuit — an in-memory object graph on one server, kept alive by a SignalR (WebSocket) connection. That's great until the server goes away. In Kubernetes, pods are restarted constantly (rolling updates, node drains, evictions, autoscaling). By default, when the pod holding your circuit dies:

  • the circuit (and everything in it) is gone, and
  • the browser reconnects to another pod that has never heard of your circuit, so Blazor builds a brand-new one — your counter resets, your session restarts.

Blazor has a built-in answer: circuit persistence. When a circuit disconnects, its declared state can be serialized to a distributed cache and resumed on whatever server the browser lands on next. You just have to configure it. This repo configures it end-to-end.


The five ingredients

Everything app-side lives in Program.cs and Home.razor; the cluster side is in deploy/k8s/app.yaml.

1. Mark the state you want to survive

Only declaratively persisted state is captured when a circuit pauses. Use [PersistentState] on the properties that must survive. Plain fields are lost.

// Home.razor
[PersistentState] public int Count { get; set; }
[PersistentState] public string? MyPersistedDemoState { get; set; }
[PersistentState] public DateTime CreatedAt { get; set; }

(The pod name on the page is deliberately not persisted — it's read live, so it changes after a restart and proves you landed on a different pod.)

2. Give Blazor a distributed store for circuits

This is the core switch. Register a HybridCache backed by a Redis IDistributedCache. Simply having a HybridCache makes Blazor auto-wire CircuitOptions.HybridPersistenceCache, which replaces the default in-memory circuit-persistence provider with the Redis-backed one.

// Program.cs
builder.Services.AddStackExchangeRedisCache(o => o.Configuration = redisConnectionString);
builder.Services.AddHybridCache();   // <-- this is what turns on cross-server persistence

3. Share Data Protection keys across pods

Persisted circuit state is encrypted with ASP.NET Core Data Protection. Each pod generates its own keys by default, so pod B can't decrypt what pod A wrote. Point every pod at the same keys in Redis and pin a fixed application name.

// Program.cs
builder.Services.AddDataProtection()
    .PersistKeysToStackExchangeRedis(redis, "blazor-multi-node:DataProtection-Keys")
    .SetApplicationName("blazor-multi-node");   // must be identical on every pod

At-rest encryption (optional). The key ring is stored in Redis unencrypted here (Data Protection logs No XML encryptor configured...) — fine for a local demo. To encrypt it at rest, add .ProtectKeysWithCertificate(cert). That's orthogonal to the cross-node decryption above, but it adds a requirement: every pod needs the same certificate, including its private key (each pod unwraps the key ring on startup), so distribute one PFX to all replicas via a Kubernetes Secret. See the commented block in Program.cs.

4. Make sure the state actually reaches Redis before the pod dies

This is the part most people miss. There is no shutdown hook that persists a live circuit. State is written to Redis only when the disconnected-circuit retention timer fires. So the sequence on SIGTERM (pod delete) has to be: connection closes → circuit becomes "disconnected" → retention timer fires → persist to Redis — all before the process exits.

Three settings make that window exist:

// Program.cs
builder.Services.Configure<CircuitOptions>(o =>
{
    o.DisconnectedCircuitRetentionPeriod = TimeSpan.FromSeconds(1);   // persist quickly after disconnect
});

builder.Services.Configure<HostOptions>(o =>
{
    o.ServicesStopConcurrently = true;   // so the grace delay overlaps Kestrel closing connections
    o.ShutdownTimeout = TimeSpan.FromSeconds(20);
});
builder.Services.AddHostedService<ShutdownGraceService>();  // delays shutdown ~10s so the persist completes
# deploy/k8s/app.yaml
terminationGracePeriodSeconds: 30   # give Kubernetes room for the grace window above

See ShutdownGraceService.cs for the (tiny) hosted service. Without ServicesStopConcurrently = true, the default sequential stop runs the delay before Kestrel closes connections, so it wouldn't cover the persist.

5. Route SignalR correctly — sticky during a session, reroute on pod death

SignalR needs session affinity: the negotiate request and the follow-up WebSocket must reach the same pod, or you get 404 No Connection with that ID and the circuit never even establishes. Use ClientIP affinity.

# deploy/k8s/app.yaml (Service)
sessionAffinity: ClientIP

This does not prevent the cross-node resume: when you delete the pinned pod, its endpoint is removed and the affinity entry is dropped, so kube-proxy sends the browser's reconnect to the surviving pod — which then resumes the circuit from Redis. The client-side reconnect timing is nudged in app.js so Redis is populated before the first resume attempt.

Common mistake: leaving sessionAffinity: None "so reconnects spread to another pod." That breaks SignalR for normal operation. Affinity + endpoint removal already gives you the different-pod reconnect on a restart.


Prerequisites

  • Docker (the image is built and the cluster runs inside Docker — no local .NET SDK needed).
  • bash. The documented path uses WSL on Windows; macOS/Linux work the same way.
  • kind and kubectl — the first script installs them if missing.

The app targets .NET 10; it is built inside the mcr.microsoft.com/dotnet/sdk:10.0 image (deploy/Dockerfile).


Run it

cd scripts

bash 00-all.sh       # create cluster + build/load image + deploy  (or run 01, 02, 03 separately)
bash 04-status.sh    # show the 2 app pods and which node each is on

Open http://localhost:8080.

See it survive a restart

  1. The page shows "Served by pod: blazor-app-…" — the pod holding your circuit. Note the Counter and My persisted demo state.

  2. Click Increment counter a few times.

  3. Delete that pod (the page prints the exact command, or use the helper script):

    bash 05-kill-active-pod.sh blazor-app-xxxxxxxxxx-yyyyy
  4. Back in the browser, the reconnect overlay appears briefly, then:

    • Counter and My persisted demo state are unchanged (restored from Redis),
    • "Served by pod" shows a different pod — you resumed on the other node.

Inspect what's in Redis at any time:

bash redis-inspect.sh

You'll usually see only the Data Protection keys. The per-circuit key is short-lived by design: RestoreCircuitAsync deletes it the moment a pod resumes the circuit, so it's gone right after a successful resume.

Tear everything down:

bash 99-teardown.sh

How it works, end to end

[browser] --SignalR--> [pod A]  circuit + state in memory
   |  kubectl delete pod A  (SIGTERM)
   |     pod A: connection closes -> circuit "disconnected" -> 1s timer ->
   |             encrypt + write circuit state to Redis (grace window keeps the process alive)
   |  affinity entry for pod A dropped (endpoint removed)
   v
[browser] --reconnect--> [pod B]  no circuit in memory ->
             resume: read state from Redis -> decrypt (shared DP keys) -> circuit restored

Each ingredient maps to one failure you'd otherwise hit:

Without it Symptom
[PersistentState] (1) field values not captured → reset after restart
HybridCache over Redis (2) provider stays in-memory → nothing to resume from
Shared DP keys + app name (3) pod B can't decrypt pod A's state → resume silently fails
Grace window (4) process exits before the persist → Redis empty → fresh circuit
ClientIP affinity (5) SignalR 404 No Connection with that ID → circuit never connects

Pods vs. nodes, and production notes

  • This demo restarts a pod, which is the realistic and frequent event (rolling updates, drains, evictions). Surviving a full node reboot uses the same mechanism — as long as Redis survives.
  • Redis is the single point of truth here, and a single pod. For production you'd run Redis with persistence and/or replication (or a managed Redis), so circuit state and Data Protection keys outlive a node loss.
  • The timing values (DisconnectedCircuitRetentionPeriod, the ~10s grace delay, terminationGracePeriodSeconds) are tuned for a snappy demo. Tune them for your shutdown profile; if a counter ever resets after a restart, the persist didn't finish in time — lengthen the grace window or the first reconnect interval in app.js.

Layout

src/BlazorMultiNode/      the Blazor Server app
  Program.cs              ingredients 2-5 (Redis, HybridCache, Data Protection, CircuitOptions, HostOptions)
  ShutdownGraceService.cs the graceful-shutdown delay (ingredient 4)
  Components/Pages/Home.razor   the [PersistentState] demo page (ingredient 1)
  wwwroot/app.js          reconnect tuning for the resume
deploy/
  Dockerfile              multi-stage build (sdk:10.0 -> aspnet:10.0)
  kind-cluster.yaml       1 control-plane + 2 workers, NodePort 30080 -> localhost:8080
  k8s/                    namespace, redis, app (Deployment + Service with ClientIP affinity)
scripts/                  bash helpers (run from WSL): 00-all, 01..05, 99-teardown, redis-inspect

License

MIT.

Contributors

pavelsavara

Issues