This repository contains the implementation used for a master thesis on secure multi-tenant stream processing. The system combines MQTT, Apache Flink, and an SGX-backed processing backend in order to compare plain stream processing with enclave-backed confidential processing.
- producer: synthetic MQTT producer for transport and hospital data
- raw_flux_consummer: Flink job that consumes MQTT streams and either processes them directly or relays them to the SGX backend
- legacy_sgx_operator: local SGX host and enclave implementation
- benchmarks: latency, throughput, and capacity benchmarks
Each submodule has its own README with module-specific details.
- the producer publishes clear-text MQTT payloads
- the Flink job consumes those payloads
- Flink applies:
- a single-stream SQL transformation
- or a plain SQL join between two streams
- Flink serializes the result rows to JSON
- Flink publishes the results to MQTT
- the producer retrieves the enclave public key through MQTT control topics
- the producer provisions a tenant stream key
- the producer encrypts business payloads before publication
- the Flink job consumes encrypted MQTT envelopes
- the Flink job relays control and data-plane requests to the local SGX host
- the SGX host forwards calls into the enclave
- the enclave decrypts and processes the data
- the resulting JSON records are returned through the host and Flink back to MQTT
The default single-stream query is:
SELECT
'passenger' AS type,
station_id,
count_in,
count_out,
occupancy_pct,
raw_json AS data
FROM typed_input
WHERE occupancy_pct >= 0.25This workload is used to benchmark a simple projection and filter pipeline.
The current benchmark join correlates passengers and trips on station
identifiers, conceptually equivalent to:
SELECT
passengers.station_id AS station_id,
passengers.count_in AS count_in,
passengers.count_out AS count_out,
passengers.occupancy_pct AS occupancy_pct,
passengers.source_ts_ms AS passengers_source_ts_ms,
trips.source_ts_ms AS trips_source_ts_ms,
trips.fare_class AS fare_class,
trips.station_out AS station_out
FROM passengers
JOIN trips
ON passengers.station_id = trips.station_inIn plain mode, this join is executed by Flink SQL. In SGX mode, it is executed through an explicit two-stream join configuration in the backend.
The producer:
- generates synthetic transport and hospital events
- publishes MQTT payloads at a configurable tick interval
- optionally adds padding for benchmark experiments
- optionally encrypts payloads for SGX scenarios
See producer/README.md.
The Flink module:
- reads configuration from environment variables
- creates MQTT-backed input streams
- supports plain single-stream and plain join execution
- supports enclave-backed control and data relays
See raw_flux_consummer/README.md.
The SGX backend:
- exposes a small local HTTP API
- manages tenant configuration and stream-key provisioning
- executes restricted single-stream SQL inside the enclave
- executes an explicit two-stream join inside the enclave
See legacy_sgx_operator/README.md.
The benchmark package supports:
- fixed-window latency measurements
- fixed-window throughput measurements
- load-ramp capacity benchmarks
- scenario matrices across plain and SGX modes
See benchmarks/README.md.
The base stack starts:
- MQTT
- Flink JobManager
- Flink TaskManagers
- the Flink job submitter
- the default producer
Command:
docker compose -f docker-compose.yaml up -dThe SGX overlay adds:
- Flink connectivity to the host SGX backend
- SGX-specific runtime configuration
- producer-side encrypted publishing
Command:
docker compose -f docker-compose.yaml -f docker-compose.legacy-host.yaml up -dFor SGX runs, the local host backend must also be available. A typical launch looks like:
LEGACY_SGX_BIND_ADDR=0.0.0.0 \
LEGACY_SGX_PORT=3000 \
LEGACY_SGX_ENCLAVE_PATH=/path/to/enclave.signed.so \
/path/to/legacy_sgx_hostThe benchmark suite can compare:
plainvssgxsimplevsjoin- different payload padding sizes
- fixed-load runs vs capacity runs
Examples:
python3 benchmarks/benchmark_latency.py --mode plain --workload simple
python3 benchmarks/benchmark_throughput.py --mode sgx --workload join
python3 benchmarks/run_benchmark_matrix.py --runs 5 --payload-padding-values 0,8192
python3 benchmarks/run_capacity_matrix.py --payload-padding-values 0,8192- The SGX backend is a legacy SGX implementation and depends on the SGX stack available on the host machine.
- The repository contains benchmarking code, because this is the purpose of our thesis.