mdevolde/multi_tenant_stream_processing

Master thesis doc and operatos.

★ 0Forks 0C++GitHub ↗Compare

README

multi_tenant_stream_processing

This repository contains the implementation used for a master thesis on secure multi-tenant stream processing. The system combines MQTT, Apache Flink, and an SGX-backed processing backend in order to compare plain stream processing with enclave-backed confidential processing.

Repository structure

  • producer: synthetic MQTT producer for transport and hospital data
  • raw_flux_consummer: Flink job that consumes MQTT streams and either processes them directly or relays them to the SGX backend
  • legacy_sgx_operator: local SGX host and enclave implementation
  • benchmarks: latency, throughput, and capacity benchmarks

Each submodule has its own README with module-specific details.

High-level architecture

Plain path

  1. the producer publishes clear-text MQTT payloads
  2. the Flink job consumes those payloads
  3. Flink applies:
    • a single-stream SQL transformation
    • or a plain SQL join between two streams
  4. Flink serializes the result rows to JSON
  5. Flink publishes the results to MQTT

SGX path

  1. the producer retrieves the enclave public key through MQTT control topics
  2. the producer provisions a tenant stream key
  3. the producer encrypts business payloads before publication
  4. the Flink job consumes encrypted MQTT envelopes
  5. the Flink job relays control and data-plane requests to the local SGX host
  6. the SGX host forwards calls into the enclave
  7. the enclave decrypts and processes the data
  8. the resulting JSON records are returned through the host and Flink back to MQTT

Main workloads

Single-stream workload

The default single-stream query is:

SELECT
  'passenger' AS type,
  station_id,
  count_in,
  count_out,
  occupancy_pct,
  raw_json AS data
FROM typed_input
WHERE occupancy_pct >= 0.25

This workload is used to benchmark a simple projection and filter pipeline.

Join workload

The current benchmark join correlates passengers and trips on station identifiers, conceptually equivalent to:

SELECT
  passengers.station_id AS station_id,
  passengers.count_in AS count_in,
  passengers.count_out AS count_out,
  passengers.occupancy_pct AS occupancy_pct,
  passengers.source_ts_ms AS passengers_source_ts_ms,
  trips.source_ts_ms AS trips_source_ts_ms,
  trips.fare_class AS fare_class,
  trips.station_out AS station_out
FROM passengers
JOIN trips
ON passengers.station_id = trips.station_in

In plain mode, this join is executed by Flink SQL. In SGX mode, it is executed through an explicit two-stream join configuration in the backend.

Main components

Producer

The producer:

  • generates synthetic transport and hospital events
  • publishes MQTT payloads at a configurable tick interval
  • optionally adds padding for benchmark experiments
  • optionally encrypts payloads for SGX scenarios

See producer/README.md.

Flink job

The Flink module:

  • reads configuration from environment variables
  • creates MQTT-backed input streams
  • supports plain single-stream and plain join execution
  • supports enclave-backed control and data relays

See raw_flux_consummer/README.md.

SGX backend

The SGX backend:

  • exposes a small local HTTP API
  • manages tenant configuration and stream-key provisioning
  • executes restricted single-stream SQL inside the enclave
  • executes an explicit two-stream join inside the enclave

See legacy_sgx_operator/README.md.

Benchmark tooling

The benchmark package supports:

  • fixed-window latency measurements
  • fixed-window throughput measurements
  • load-ramp capacity benchmarks
  • scenario matrices across plain and SGX modes

See benchmarks/README.md.

Running the stack

Base stack

The base stack starts:

  • MQTT
  • Flink JobManager
  • Flink TaskManagers
  • the Flink job submitter
  • the default producer

Command:

docker compose -f docker-compose.yaml up -d

SGX-backed stack

The SGX overlay adds:

  • Flink connectivity to the host SGX backend
  • SGX-specific runtime configuration
  • producer-side encrypted publishing

Command:

docker compose -f docker-compose.yaml -f docker-compose.legacy-host.yaml up -d

For SGX runs, the local host backend must also be available. A typical launch looks like:

LEGACY_SGX_BIND_ADDR=0.0.0.0 \
LEGACY_SGX_PORT=3000 \
LEGACY_SGX_ENCLAVE_PATH=/path/to/enclave.signed.so \
/path/to/legacy_sgx_host

Benchmarks

The benchmark suite can compare:

  • plain vs sgx
  • simple vs join
  • different payload padding sizes
  • fixed-load runs vs capacity runs

Examples:

python3 benchmarks/benchmark_latency.py --mode plain --workload simple
python3 benchmarks/benchmark_throughput.py --mode sgx --workload join
python3 benchmarks/run_benchmark_matrix.py --runs 5 --payload-padding-values 0,8192
python3 benchmarks/run_capacity_matrix.py --payload-padding-values 0,8192

Notes

  • The SGX backend is a legacy SGX implementation and depends on the SGX stack available on the host machine.
  • The repository contains benchmarking code, because this is the purpose of our thesis.

Contributors

mdevolde

Issues