A small, deliberately explainable secure remote device management / agent orchestration demo, built as a backend engineering exercise.
Three independent applications talk to each other over HTTPS: a management server, a device agent, and an administrator CLI. An administrator lists connected agents, sends a command to one of them, and reads back what it returned.
Status: official Steps 1, 2 and 3 are complete. Registration, multiple agents, job assignment, polling delivery, result collection, an admin API behind a bearer token, a CLI that drives all of it — TLS underneath the whole thing, with a private demo CA that agents and administrators verify the server against — and commands, results and application logs persisted to PostgreSQL running as its own service. Generic commands with an internal queue (Step 4) come next — see Roadmap.
- Python 3.14 (developed and tested on 3.14.4; 3.11+ should work)
- Docker (or Docker Compose) — the database runs as its own service. See External persistence (PostgreSQL).
python3 -m venv .venv
.venv/bin/pip install -r requirements.txtGenerate the demo PKI first. Certificates are never committed, so a fresh checkout has none:
.venv/bin/python scripts/generate_certs.pyThis writes certs/{ca.crt,ca.key,server.crt,server.key}. See
Secure transport for what each file is and who may hold it.
Start the database. It is a separate service, so it comes up before the application and stays up across restarts of it:
cp .env.example .env # then set POSTGRES_PASSWORD and ADMIN_TOKEN
docker compose up -d postgresTerminal 1 — the management server. It serves HTTPS only and refuses to start if the
certificate, the key, ADMIN_TOKEN or the database is missing:
ADMIN_TOKEN=dev-admin-token \
SERVER_CERT_PATH=certs/server.crt \
SERVER_KEY_PATH=certs/server.key \
.venv/bin/python -m server.runTLS enabled
Listening on https://localhost:8443
Certificate: certs/server.crt
Server 0.1.0 starting. Persistence: postgresql at postgresql+psycopg://agent_manager:***@localhost:5432/agent_manager
Terminal 2 — a device agent (start as many as you like, each with its own identity):
export SERVER_URL=https://localhost:8443 CA_CERT_PATH=certs/ca.crt
.venv/bin/python -m agent.client
DEV_AGENT_ID=agent-002 .venv/bin/python -m agent.clientTerminal 3 — the administrator:
export SERVER_URL=https://localhost:8443 CA_CERT_PATH=certs/ca.crt
export ADMIN_TOKEN=dev-admin-token
.venv/bin/python -m admin.cli agentsThose defaults are already the built-in ones, so with a .env in place the exports are optional.
Interactive API docs: https://localhost:8443/docs — your browser will warn about the certificate,
because a private CA is exactly what a browser is supposed to distrust. Run the tests with
.venv/bin/pytest -q.
$ python -m admin.cli agents
ID HOSTNAME OS VERSION STATUS
778623bf lugal Linux 0.1.0 online
d160c496 lugal Linux 0.1.0 online
bb862fb7 lugal Linux 0.1.0 online
Send echo to one agent and wait for the answer. Agents may be named by any unique prefix:
$ python -m admin.cli echo 778623bf "Hello World" --wait
Command queued for 778623bf.
Job ID: 20d3ad6c-e25c-4e71-82c1-481f58a46602
Status: completed
Result:
Hello World
Job ID: 20d3ad6c-e25c-4e71-82c1-481f58a46602
Agent: 778623bf-36bb-41b5-b47e-95ec8657183a
Operation: echo
Arguments: Hello World
Created: 2026-08-11 17:26:33 UTC
Delivered: 2026-08-11 17:26:33 UTC
Completed: 2026-08-11 17:26:33 UTC
Without --wait the command returns immediately and the result is read separately:
$ python -m admin.cli result 20d3ad6c
Status: completed
Result:
Hello World
The command reached only that agent:
$ python -m admin.cli jobs 778623bf
JOB OPERATION STATUS CREATED COMPLETED
20d3ad6c echo completed 2026-08-11 17:26:33 UTC 2026-08-11 17:26:33 UTC
$ python -m admin.cli jobs d160c496
No commands have been sent to this agent.
Stop an agent gracefully. It reports the result before exiting:
$ python -m admin.cli kill 778623bf --wait
Status: completed
Result:
agent shutting down
# in the agent's terminal
Received command 8576764b: kill
Reported completed for command 8576764b
Kill command received. Shutting down.
Agent stopped.
About thirty seconds later it is reported offline, while the others carry on:
$ python -m admin.cli agents
ID HOSTNAME OS VERSION STATUS
778623bf lugal Linux 0.1.0 offline
d160c496 lugal Linux 0.1.0 online
bb862fb7 lugal Linux 0.1.0 online
| command | what it does |
|---|---|
health |
check the server is reachable (needs no token) |
agents |
list agents and their status |
agent <id> |
show one agent |
echo <id> "<text>" [--wait] |
send an echo command |
kill <id> [--wait] |
ask an agent to shut down gracefully |
jobs <id> |
an agent's command history |
result <job> |
a command's result |
Only echo and kill are exposed, which is what Step 1 calls for. The agent's local allowlist is
larger, and the server's is the authority on what may be assigned.
Admin CLI (admin/cli.py)
|
| HTTPS + Authorization: Bearer $ADMIN_TOKEN
v
+-------------------------------------+
| Management Server |
| FastAPI |
| |
| api/agent.py api/admin.py | <- transport / HTTP
| \ / |
| security/identity.py | <- who is calling?
| \ / |
| services/{agent,job}_service.py | <- business rules
| | |
| models.py / database.py | <- persistence
+-------------------------------------+
| ^
SQLite | HTTPS (mutual TLS next)
|
+-----------------+
| Device Agent |
+-----------------+
Two independent trust relationships meet at this server:
| Caller | Authenticated by | Status |
|---|---|---|
| Device agent | AgentIdentity — a dev header today, a client certificate under mutual TLS |
in place |
| Administrator | Authorization: Bearer $ADMIN_TOKEN |
in place |
They are separate credentials on separate routers — holding one grants nothing on the other. The
agent API never consults the admin token, and the admin API never looks at AgentIdentity.
The exercise talks about connected clients. This implementation polls over HTTP and holds no persistent socket per agent — each request is its own short-lived connection. So:
Connection state here is application-level presence, not an open socket. Registration and keep-alive update
last_seen;statusis derived from it at read time. An agent isonlinewhen it has contacted the server withinAGENT_OFFLINE_AFTER_SECONDS(default 30).
The server manages many concurrently-running agents without keeping a connection open for any of them, which is what makes this design cheap to scale and trivial to restart on either side. It is worth saying out loud rather than letting the word "connected" imply a socket that does not exist.
Admin CLI Server Device Agent
| | |
|-- POST .../jobs ------>| job: queued |
| echo "Hello World" | |
| |<-- POST /jobs/claim --------| every 2s
| | job: delivered ---------->|
| | | execute_operation()
| | | (allowlist only)
| |<-- POST /jobs/{id}/result --|
| | job: completed |
|-- GET /jobs/{id} ----->| |
|<-- "Hello World" ------| |
Claiming is a POST, not a GET: it moves the job to delivered and stamps delivered_at, and a
GET that mutates state is neither cacheable nor safely retryable. Two concurrent claims cannot both
win — the transition is a conditional UPDATE ... WHERE status = 'queued' whose row count decides.
Why polling. One code path, no connection state to manage, trivially testable, and every step is visible in the server log. WebSockets would later replace only the delivery step; the job table, the statuses and the result endpoint would stay exactly as they are.
No agent endpoint takes an agent identifier. There is nothing to pass, and therefore nothing to pass wrongly:
X-Dev-Agent-Id (today) client certificate (mutual TLS)
\ /
v v
get_agent_identity()
|
AgentIdentity(source, stable_id)
|
find_by_identity() -> the Agent row
|
SELECT ... FROM jobs WHERE agent_pk = <that row> AND status = 'queued'
Addressing agents by name is an admin-side concept (/api/admin/agents/{agent_id}/jobs), because
managing a fleet means choosing which device to act on. An agent's scope comes from who it is.
The same rule governs results: a job is looked up scoped to the calling agent, so knowing another
agent's job_id is not enough to write to it. That case returns 404, not 403 — a "forbidden"
would confirm the job exists and belongs to someone else, turning job_id into an oracle for
mapping the fleet.
queued ──(agent claims)──> delivered ──(agent reports)──> completed
\─> failed
That is the entire state machine. Reporting twice on a job is a 409; results are recorded once.
The single most important design decision. Authentication is resolved at the transport edge and handed to the rest of the application as a plain value object:
TLS / transport <- SSL objects, peer certificates live here
|
v
CertIdentityProvider <- the ONLY code that parses X.509
|
v
AgentIdentity <- plain frozen dataclass of strings
|
v
API / services / persistence <- unaware that TLS exists at all
Nothing below AgentIdentity ever sees a Request, an SSLObject or an x509.Certificate.
| Step | Provider | stable_id |
subject |
|---|---|---|---|
| 1 | DevIdentityProvider |
dev:agent-001 |
None |
| 2+ | CertIdentityProvider |
mtls:<sha256-fingerprint> |
CN=agent-001,O=Demo |
Registration, job claiming and result submission all sit behind this one dependency, so swapping the provider changes who an agent is without changing what any of them do.
Today an agent identifies itself with an X-Dev-Agent-Id header (falling back to the hostname in
the request body). Anyone can send any value. It is a stable local-development identifier and
nothing more — it is not a credential, not a secret, and not a security boundary. It exists purely
so the registration, command and result logic can be built and tested before the PKI is wired into
identity. Mutual TLS replaces the provider with certificate-derived identity; the API and service
contracts do not change.
TLS does not fix this. It stops the header being read on the wire and proves to the agent which
server it reached, but the handshake authenticates the server to the client, not the reverse.
Encrypting a claim does not verify it: any process holding ca.crt can open a valid TLS session and
then declare itself any agent it likes.
Authorization: Bearer $ADMIN_TOKEN establishes that the caller is the administrator. TLS
establishes that the server is the real one, and keeps the token from being read in transit. Two
separate properties, and Step 2 supplies the one that was missing:
Step 1 Bearer <token> over HTTP -> authenticated, NOT confidential
now Bearer <token> over HTTPS -> authenticated AND confidential
What TLS does not do is authorize anyone. A completed handshake grants nothing — a wrong token
over a perfectly valid TLS session still returns 401. The token is compared with
secrets.compare_digest (a plain == returns early on the first differing byte, which leaks the
secret one character at a time through timing), is never logged, and never appears in a response or
an error message.
ADMIN_TOKEN is required — the server refuses to start without it, and there is no default,
because a shipped default credential is precisely how a demo secret reaches production.
A managed device: it starts up, describes itself, registers with the management server, learns the
agent_id the server assigned it, and keeps that registration alive.
agent/ is an independent application. It imports nothing from server/ — the HTTP API is the
only contract between them, so the two halves could run on different machines or live in separate
repositories. Its response model deliberately tolerates unknown fields (the mirror image of the
server's strict request validation), so a server that starts returning more cannot break agents
already deployed.
agent/
├── config.py # environment-driven settings; server URL + CA trust anchor
├── system_info.py # device facts — no networking, no job execution
├── client.py # AgentClient (httpx.AsyncClient) + lifecycle + entry point
└── operations.py # the safe-operation allowlist
system_info.py sits underneath both consumers, so registration never depends on the
job-execution module:
system_info.py
/ \
v v
client.py operations.py
| variable | default | purpose |
|---|---|---|
SERVER_URL |
https://localhost:8443 |
HTTPS; plain http:// is only for local experiments |
CA_CERT_PATH |
unset | required for HTTPS — the CA that verifies the server certificate |
DEV_AGENT_ID |
agent-001 |
development identity — not a credential |
AGENT_VERSION |
0.1.0 |
reported as metadata |
JOB_POLL_INTERVAL_SECONDS |
2 |
how often to ask for work |
KEEP_ALIVE_INTERVAL_SECONDS |
10 |
how often to report being alive (offline after ~30s) |
CONNECT_TIMEOUT_SECONDS |
5 |
explicit — nothing waits forever |
REQUEST_TIMEOUT_SECONDS |
10 |
explicit — nothing waits forever |
CLIENT_CERT_PATH / CLIENT_KEY_PATH |
unset | mutual TLS: prove this agent's identity |
AGENT_LOG_LEVEL |
INFO |
DEBUG also shows every HTTP request |
The two cadences are separate because they answer different questions: how often a device checks for work bounds how long an administrator waits, while how often it says it is alive only needs to beat the offline threshold. Sharing one timer would mean either a sluggish demo or five times the database writes.
every JOB_POLL_INTERVAL_SECONDS:
job = claim_next_job()
while job: # drain, don't sleep between commands
result = execute_operation(...)
submit_result(...)
if operation == "kill": stop
keep_alive_if_due() # a backlog must not starve it
job = claim_next_job()
keep_alive_if_due()
Work is drained rather than taken one command per interval, and the keep-alive check runs inside that loop — an agent busy working must never be reported offline. Step 4 replaces this with an internal queue and concurrent execution; a sequential loop is easier to follow and to trust for now.
The agent sends X-Dev-Agent-Id: <DEV_AGENT_ID> on every request. Any process can send any value;
it carries no secret and proves nothing. It exists only so the server has a stable key to file this
agent under while the demo PKI does not exist. Do not treat it as a credential.
Under mutual TLS identity moves into the handshake: the agent presents a client certificate, and the server derives identity from the validated certificate rather than from a header it was handed.
The server has no heartbeat endpoint yet, so keep_alive() re-calls the idempotent
POST /api/agent/register, which refreshes last_seen and keeps the agent reported as online.
It is a separate, clearly-labelled method so nothing comes to depend on registration doubling as a
heartbeat:
now keep_alive() -> POST /api/agent/register (this shim)
later keep_alive() -> POST /api/agent/heartbeat
It also verifies that the same development identity keeps mapping to the same agent_id, and logs
an error if it ever changes — that would mean the server filed a duplicate record.
Restarting an agent re-registers it: the server returns the same agent_id rather than creating
a second row. The client treats 201 Created and 200 OK identically — whether a record already
existed is the server's business, not the device's.
agent/operations.py holds the allowlist: ping, get_system_info, get_time, echo, sleep
and kill. Of these, the server currently permits only echo and kill to be assigned.
Dispatch is a lookup in a literal dictionary of already-imported functions, so a job's operation
string is a dictionary key, never a lookup path — there is no route from remote input to an
arbitrary callable. No shell, no subprocess, no eval/exec, no dynamic import. Arguments are
validated by a strict model per operation, and sleep is capped at 30 seconds so a malformed job
cannot wedge the agent. A command outside the allowlist becomes a failed result, never a crash.
The kill handler kills nothing: it is a pure function returning
{"message": "agent shutting down"}. The lifecycle stops, by recognising the operation after the
result has been reported. A handler that called sys.exit would be untestable, would terminate
before its result could be reported, and would put process control inside the code that executes
remote input.
It never touches other processes, the machine, or the filesystem — and it takes no arguments, so there is no pid or target to supply. Shutdown does not depend on the server accepting the final result: one bounded attempt, then the agent exits regardless, because an operator killing an agent during an outage must not find it still running.
All transport security is assembled in one function, build_ssl_context():
context = ssl.create_default_context(cafile=CA_CERT_PATH)
context.minimum_version = ssl.TLSVersion.TLSv1_2
context.load_cert_chain(CLIENT_CERT_PATH, CLIENT_KEY_PATH) # mutual TLS, later
httpx.AsyncClient(base_url=SERVER_URL, verify=context)create_default_context(cafile=...) replaces the trust anchors rather than adding to them, so the
agent trusts our demo CA and nothing else — a public CA has no business vouching for this server.
Verification is left entirely to the TLS stack: chain, validity and SAN. No hand-rolled certificate
checks, and never verify=False. An https:// URL without CA_CERT_PATH is refused at startup,
because the alternative is an opaque handshake failure several layers down.
See Secure transport for the full trust model.
The third independent application. It imports neither server nor agent, reads no database, and
knows nothing about SQLAlchemy — the HTTP API is the whole contract, so it could be installed on
another machine or split into its own repository unchanged.
admin/
├── config.py # environment-driven settings; server URL + CA trust anchor
├── client.py # AdminClient (sync httpx.Client) -> returns data, never prints
├── formatting.py # data -> strings; no HTTP, no I/O
└── cli.py # argparse, exit codes
Synchronous on purpose: the CLI is one request and an exit, so there is nothing to overlap. The agent is async because it runs a loop forever — a different problem. Transport and presentation are separated so either can change without the other.
| variable | default | purpose |
|---|---|---|
SERVER_URL |
https://localhost:8443 |
HTTPS; plain http:// is only for local experiments |
CA_CERT_PATH |
unset | required for HTTPS — the CA that verifies the server certificate |
ADMIN_TOKEN |
unset | required for every command except health; TLS does not replace it |
CONNECT_TIMEOUT_SECONDS / REQUEST_TIMEOUT_SECONDS |
5 / 10 |
explicit |
ADMIN_LOG_LEVEL |
INFO |
DEBUG also shows every HTTP request |
Expected failures print one line and exit non-zero, never a traceback, and the token appears in no message:
$ python -m admin.cli agents
Cannot reach the server: unable to connect to management server at https://localhost:8443
$ ADMIN_TOKEN=wrong python -m admin.cli agents
Authentication failed (check ADMIN_TOKEN): the server rejected the administrator token
Step 2 of the exercise asks for a secure encryption protocol between client and server. This implementation does not define one. It uses TLS.
That is the whole design decision, and it is deliberate. Hand-rolling a protocol — AES-encrypting JSON bodies, RSA-wrapping content keys, inventing a nonce scheme, bolting on an HMAC — produces something that looks cryptographic and is almost certainly broken. Padding oracles, nonce reuse, unauthenticated ciphertext, no replay protection, no forward secrecy, no key rotation and no downgrade resistance are the normal outcomes, and none of them are visible in a passing test suite. TLS is the standard answer to this exact problem, has been attacked by everyone for thirty years, and arrives already implemented in OpenSSL.
So the application protocol is unchanged and the security lives underneath it:
Application HTTP + JSON <- identical to Step 1, byte for byte
|
Transport TLS 1.3 <- session keys, confidentiality, integrity
|
Wire TCP
Not one route, schema, service or database call changed. Uvicorn terminates TLS and FastAPI sees
ordinary HTTP requests, which is why the entire Step 1 test suite still runs unmodified through
TestClient.
Demo Root CA ca.key the CA's private signing key
| ca.crt the public trust anchor
+-- signs --> server.crt SAN: DNS:localhost, IP:127.0.0.1
server.key the server's private key
Agent --trusts--> ca.crt --verifies--> server.crt
Admin --trusts--> ca.crt --verifies--> server.crt
| file | what it is | who holds it |
|---|---|---|
ca.crt |
the trust anchor: clients believe certificates it signed | agents, admins, freely distributable |
ca.key |
signs certificates | scripts/generate_certs.py only — the server never loads it |
server.crt |
authenticates the management server during the handshake | the server; sent to every client, public |
server.key |
proves the server owns server.crt |
the server host only, never distributed |
A private CA rather than a self-signed certificate: a self-signed cert has to be distributed to and pinned by every client, and replacing it means touching every client. With a CA, the anchor is distributed once and server certificates can be reissued underneath it.
server.crt does not encrypt anything. It authenticates the server — proving possession of
server.key for a certificate chaining to the CA. The handshake then derives ephemeral session
keys, and those protect the traffic. The certificate's job is answering "who am I talking to",
not "how is this scrambled".
The certificate carries DNS:localhost, IP:127.0.0.1 in its Subject Alternative Name. Clients match
the address they dialled against those entries and ignore the Common Name entirely — CN-based
matching was deprecated for good reason. Without this check, any certificate the CA ever issued
would be accepted for any host, so a party holding a legitimately-issued certificate for one name
could impersonate any other. Verified against the running server:
$ openssl s_client -connect 127.0.0.1:8443 -CAfile certs/ca.crt -verify_hostname localhost
Verify return code: 0 (ok)
$ openssl s_client -connect 127.0.0.1:8443 -CAfile certs/ca.crt \
-verify_hostname not-in-the-san.example
verify error:num=62:hostname mismatch
Verify return code: 62 (hostname mismatch)
Same server, same CA, same open socket — only the name differs.
| ✅ Confidentiality | nobody on the network can read commands, results, or the admin token |
| ✅ Integrity | nobody can alter a command or a result in flight |
| ✅ Server authentication | clients know they reached the real server, not an impostor |
| ❌ Client authentication | the handshake is one-directional — the client proves nothing |
That last row is the important one, and it has two consequences.
ADMIN_TOKEN is still required. TLS and the token answer different questions:
TLS "is this the real server?" confidentiality + integrity + server identity
ADMIN_TOKEN "is this caller an admin?" application-level authorization
A completed handshake grants no authorization at all. Anyone holding ca.crt — a public file — can
open a perfectly valid TLS session to this server; what they cannot do is call /api/admin without
the token. Dropping the token because "it's encrypted now" would confuse a confidentiality control
with an authentication one and leave the admin API open to anyone who can reach the port:
$ ADMIN_TOKEN=wrong-token python -m admin.cli agents
Authentication failed (check ADMIN_TOKEN): the server rejected the administrator token
$ curl --cacert certs/ca.crt -H "Authorization: Bearer wrong-token" \
https://localhost:8443/api/admin/agents
{"detail":"Administrator authentication required."} # HTTP 401 — TLS succeeded
X-Dev-Agent-Id is still not authentication. TLS now hides the header from passive observation
and stops the agent being redirected to an impostor server. It does not make the claim true.
Encrypting an assertion does not verify it, and any process with ca.crt can complete a valid
handshake and then declare itself any agent it likes. Only mutual TLS closes this.
One parameter is set by hand — the minimum version — and only because it cannot otherwise be guaranteed:
context.minimum_version = ssl.TLSVersion.TLSv1_2A stock SSLContext reports minimum_version as the MINIMUM_SUPPORTED sentinel, which defers to
whatever the host's OpenSSL configuration permits — and that varies by distribution. Naming the
floor turns "we use secure defaults" from a claim into something the test suite asserts.
Nothing else is tuned: no cipher list, no curve list, no version ceiling. OpenSSL's own defaults are better maintained than anything hand-written here, and leaving the ceiling open means TLS 1.3 is negotiated whenever both ends support it, which in practice is always:
$ openssl s_client -connect localhost:8443 -CAfile certs/ca.crt
Protocol : TLSv1.3
Cipher : TLS_AES_256_GCM_SHA384
Verify return code: 0 (ok)
python -m server.run serves HTTPS or it does not serve:
python -m server.run
|
v
validate TLS configuration
|
+---+---+
valid invalid
| |
v v
HTTPS exit 1, with the reason
There is no TLS_ENABLED flag, no autodetection that quietly degrades when a variable is
misspelled, and no plaintext fallback. An operator whose TLS is broken must find out immediately,
not by later discovering that traffic they believed was encrypted never was. Four failures are
checked before anything binds a port — unset variable, missing file, unreadable file, and a
certificate and key that do not match:
$ python -m server.run
Cannot start: SERVER_CERT_PATH is not set, so TLS cannot be configured -- it is the server
certificate presented to agents and administrators.
The management server serves HTTPS only; it will not fall back to plaintext.
...
$ SERVER_CERT_PATH=certs/ca.crt SERVER_KEY_PATH=certs/server.key python -m server.run
Cannot start: TLS certificate and key do not form a usable pair:
certificate certs/ca.crt
key certs/server.key
OpenSSL reported: [X509: KEY_VALUES_MISMATCH] key values mismatch (_ssl.c:4181)
The mismatch check is real, not a guess: the pair is loaded into a throwaway SSLContext, exactly
what Uvicorn does moments later.
.gitignore blocks *.key, *.pem and the whole of certs/ except .gitkeep, so generated
material cannot be committed by accident:
$ git ls-files | grep -E '\.key$|\.crt$'
(nothing)
$ git ls-files certs/
certs/.gitkeep
Certificates are regenerated per checkout rather than shared. Keys are written 0600 at creation
time via os.open, so they are never briefly world-readable, and the generator prints paths, SANs
and fingerprints — never key material.
Step 2 authenticates the server to its clients. The remaining half authenticates clients to the server:
today mutual TLS
----- ----------
Agent Agent
| verifies server cert | verifies server cert
v | presents agent cert
Server v
Server
| verifies agent cert
v
CertIdentityProvider -> AgentIdentity
The architecture is already shaped for it. AgentIdentity, identity_source, identity_stable_id,
certificate_fingerprint and certificate_subject all exist and are persisted today; swapping
DevIdentityProvider for CertIdentityProvider behind get_identity_provider() changes who an
agent is without touching registration, job claiming, result submission, the service layer or the
database — that is what the identity seam is for.
It is deferred rather than dismissed because it is not trivial: Uvicorn exposes no stable interface for handing the validated peer certificate to an ASGI application (the ASGI TLS extension remains a draft it does not implement), so it needs work at the protocol layer. Shipping working TLS first was the better trade.
Step 3 asks for commands, results and logs to live in an external database. Here "external" means what it says: a database service running as its own process, started and stopped independently of the application.
Admin CLI ──HTTPS+token──┐
v
FastAPI server ── SQLAlchemy ──> ┌──────────────┐
^ │ PostgreSQL │ separate process
Agent ────HTTPS──────────┘ │ (container) │
├──────────────┤
│ agents │
│ jobs │ commands + results
│ log_events │ operational logs
└──────────────┘
The two halves start and stop separately, which is the demonstration:
docker compose up -d postgres # the database
.venv/bin/python -m server.run # the application, on the hostdocker-compose.yml contains PostgreSQL and nothing else. The server, agent and admin CLI stay on
the host: containerising them would hide the separation behind a single up, and the point is that
the database outlives any of them.
Persistence was already behind SQLAlchemy, and server/database.py was already the only module that
knew which database was in use. So the move was a URL, a driver, and a startup check:
DATABASE_URL |
sqlite:///./agent_manager.db → postgresql+psycopg://… |
requirements.txt |
psycopg[binary] |
server/database.py |
pool_pre_ping for a networked database, plus check_database_connection() |
No model, service, route or repository changed to accommodate PostgreSQL. The column types were
already portable — JSON rather than JSONB, String(36) rather than UUID, Python-side defaults
rather than server_default — so the same create_all() builds the same schema on either engine.
The one genuine difference is that PostgreSQL returns timezone-aware datetimes where SQLite returns
naive ones, and ensure_utc() already normalised both.
These answer different questions and are deliberately kept apart.
Admin creates command ──> jobs row: queued
|
agent claims v
delivered
|
agent reports v
┌───────────┴───────────┐
completed + result failed + error
jobs is the source of truth for what was commanded and what came back — one row per command,
carrying job_id, the owning agent (agent_pk, the surrogate key behind the public agent_id),
operation, arguments, status, created_at, delivered_at, completed_at, result and
error. The command is committed when the administrator creates it;
the result when the agent reports it. Reconstructing history from log lines instead would turn an
auditable state machine into text that has to be parsed to be believed.
log_events is operational logging: what the server was doing, and when. Nothing in it is required
to answer "what did we run and what came back".
Integration happens at the logging layer, not at the call sites, so every logger.info already in
the codebase is persisted without being rewritten:
logger.info("Job %s created", job_id)
|
├──> StreamHandler ─────────────────────────> console (synchronous)
|
└──> QueueHandler ──> queue ──> QueueListener thread ──> INSERT INTO log_events
QueueHandler.emit does no I/O — it puts the record on an in-memory queue and returns — so the
request that logged never waits for the database. The queue is bounded (10,000 records): if the
database is slow or gone, records are dropped rather than buffered until the process runs out of
memory.
Two failure modes get explicit attention in server/logging_setup.py:
- Recursion. A handler that logs while handling a record feeds itself, and writing to a database
makes that real, because SQLAlchemy and psycopg log. Three independent guards: a thread-local
re-entrancy flag, a filter rejecting
sqlalchemy/psycopg/server.logging_setup, and the fact that persistence failures never re-enter logging at all. - A dead database. Failures go to stderr via
handleError, outside the logging system. The record is lost from the database and kept on the console — a logging outage must not become an application outage.
Only server.* and uvicorn.error are persisted. uvicorn.access is excluded deliberately: an
idle agent polls every two seconds, so access logging alone would add ~30 rows per minute per agent
and bury the events an operator came to read. The same reasoning applies to the keep-alive, which
reuses the registration endpoint — a genuine re-registration is logged, a routine ping is not.
What is never written. No tokens, no Authorization headers, no key material, no tracebacks, no
command arguments or results. The fields on LogEvent are an allowlist rather than a copy of the
record, and the database password is masked (postgresql+psycopg://agent_manager:***@…) everywhere
it is displayed. A failed admin authentication is logged — repeated failures are worth seeing — but
never the token that was presented, not even a prefix: the whole reason require_admin compares in
constant time is to avoid leaking it a fragment at a time.
PostgreSQL unreachable ──> ConfigurationError, exit 1
╳───> fall back to local SQLite
There is no fallback, and that is a deliberate refusal. A server that quietly downgraded to a local
file would look healthy while writing the fleet's command history somewhere nobody is looking, and
the operator would find out only when they went looking for an audit trail that was never there.
The check runs in server/run.py before anything binds a port, so the message reads as the
configuration problem it is rather than as a crash:
Cannot start: Cannot reach the database at postgresql+psycopg://agent_manager:***@localhost:5432/agent_manager.
The server will not start without it, and it will not fall back to a local SQLite file [...]
docker compose up -d postgres
| runtime | schema management | |
|---|---|---|
Real runs, the demo, tests/integration/ |
PostgreSQL | create_all() |
Unit tests (tests/) |
throwaway SQLite per test | create_all() |
This is not a hedge. SQLAlchemy is the persistence abstraction, so the models, services and routes
under test are the same either way, and an empty file-backed database per test costs milliseconds
and needs no running service. What the unit tests deliberately do not prove is that the system
uses an external database — that is what the PostgreSQL end-to-end run and tests/integration/
are for.
- Migrations.
create_all()creates missing tables and does nothing about changing one. A real deployment would use Alembic so schema changes are versioned, reviewable and reversible. - Logs to an observability stack. A relational table is queryable and good enough at this size, but production logging belongs somewhere built for retention, indexing and alerting.
- Credentials and connections. A managed secret rather than
.env, TLS on the database connection, and a pool sized against the real workload.
.venv/bin/pytest -q # 293 tests, no services required| suite | what it covers |
|---|---|
tests/test_*.py |
server: registration, admin API, auth boundary, job flow, isolation |
tests/test_persistence.py |
what actually reaches the database for a command and its result |
tests/test_log_events.py |
the logging pipeline: mapping, filtering, failure, secrets |
tests/integration/ |
the same guarantees against a real PostgreSQL — opt-in |
tests/agent/ |
agent: operations allowlist, client, serving loop, kill semantics |
tests/admin/ |
CLI: client, formatting, commands, exit codes, independence |
tests/test_certs.py |
the generated PKI: SANs, CA:FALSE, EKU, chain, key permissions |
tests/test_server_tls_config.py |
every way TLS can be misconfigured fails startup |
tests/test_client_tls.py |
both clients verify; nothing in the source disables verification |
tests/test_tls_integration.py |
real handshakes against a real Uvicorn TLS server |
Agent and admin tests run entirely through httpx.MockTransport — no server, no socket — so each
application is tested as the independent thing it is. Independence is also asserted directly: a test
imports each package in a subprocess and checks that nothing named server or agent appears in
sys.modules.
The TLS integration test is the exception, deliberately: TestClient never opens a socket and so
can say nothing about transport security. It starts python -m server.run in a subprocess on a free
port with a throwaway PKI, and asserts what actually distinguishes a secure transport — the correct
CA succeeds, an unrelated CA fails, a name outside the SAN fails, a plaintext client cannot reach
the API, the negotiated version is TLS 1.3, and a wrong admin token gets 401 over a successful
handshake. The server is terminated in a finally, so a failing test cannot leak a process holding
a port.
Two of these tests are written against specific failure modes rather than "an error occurred". The
untrusted-CA cases unwrap the exception chain and assert SSLCertVerificationError with
unable to get local issuer certificate, because asserting on httpx.TransportError alone would
also pass if the server were simply down. The SAN test verifies the name over an already-open,
known-good socket, because the obvious version — dialling a second loopback address — passes
whether or not hostname checking works at all: the server binds one address, so that connection
fails at TCP before TLS is ever consulted.
The PostgreSQL suite is skipped unless you point it at a database, so pytest stays fast and needs
nothing running:
docker compose up -d postgres
POSTGRES_TEST_URL='postgresql+psycopg://agent_manager:...@localhost:5432/agent_manager' \
.venv/bin/pytest tests/integration -qIt is not the API suite re-pointed at another engine — SQLAlchemy already makes those tests
engine-independent, and running them twice would double the runtime to re-prove what the abstraction
guarantees. It covers only what SQLite cannot answer for: that create_all() emits valid PostgreSQL
DDL, that JSON arguments and results survive a real driver as structures, that timezone-aware
timestamps read back correctly, and that rows are visible to a connection that did not write them.
Each test builds and drops its own schema, so it can run against the demo database without
disturbing it.
| Step | Contents | Status |
|---|---|---|
| 1 | Client/server connectivity, multiple agents, admin CLI, echo + kill, results | ✅ done |
| 2 | Encryption: TLS, demo PKI, server authentication | ✅ done |
| 2b | Mutual TLS: client certificates, certificate-derived identity | planned |
| 3 | Microservices: external database for logs and command history | ✅ done |
| 4 | Generic commands, async server, heartbeat, internal agent queue, scale | planned |
| 5 | Scale and functionality testing | planned |
- Agent identity is self-declared.
X-Dev-Agent-Idproves nothing. TLS keeps it off the wire in cleartext and authenticates the server, but the client is not authenticated at all — mutual TLS is what closes this. - The demo CA is a demo CA.
ca.keysits unencrypted next to the certificates it signed. A real deployment keeps the signing key offline or in an HSM, and issues short-lived certificates. - No certificate revocation. Nothing consults a CRL or OCSP; a compromised certificate stays valid until it expires. A revocation denylist is a natural companion to mutual TLS.
create_all()instead of migrations. It creates missing tables and does nothing about changing one. Alembic is the production answer; see External persistence.- Logs live in a relational table. Queryable and adequate at this size, but production logging belongs in a stack built for retention, indexing and alerting.
- A log record can be dropped under sustained database failure. The queue is bounded on purpose:
keeping every record would trade a logging outage for an application outage. Dropped records are
still on the console, and
jobs— the auditable record — is written synchronously and never affected. - Sequential command execution. One command at a time per agent, by design until Step 4.
- A dead agent still reads
onlinefor up toAGENT_OFFLINE_AFTER_SECONDS. That is inherent to presence-based status, and the alternative — a stored flag — is worse, because a killed or partitioned agent never gets to writeofflineat all. - A registration body with no
hostnameand noX-Dev-Agent-Idreturns 400, not the 422 an invalid body would normally get: identity is resolved before body validation, and the development provider useshostnameas its fallback. This coupling disappears with mutual TLS, when identity comes from the certificate and never from the body.
server/ # the management server
├── main.py # create_app(), lifespan, router wiring -- transport-agnostic
├── run.py # the executable entry point; TLS enforced, no fallback
├── config.py # environment-driven settings, incl. require_tls_material()
├── database.py # engine, session factory, get_db <-- the only DB-aware module
├── logging_setup.py # console + queued background writer to log_events
├── models.py # Agent, Job, LogEvent
├── schemas.py # pydantic request/response models
├── security/
│ ├── identity.py # AgentIdentity + providers <-- the seam
│ └── admin_auth.py # bearer-token dependency
├── api/
│ ├── health.py
│ ├── agent.py # register, claim, submit result
│ ├── admin.py # agents, jobs, results
│ └── serialization.py
└── services/
├── agent_service.py
└── job_service.py
agent/ # the device agent -- imports nothing from server/
├── config.py
├── system_info.py
├── client.py
└── operations.py
admin/ # the CLI -- imports neither server/ nor agent/
├── config.py
├── client.py
├── formatting.py
└── cli.py
scripts/
└── generate_certs.py # the demo PKI; the only code that touches ca.key
docker-compose.yml # the external database service -- PostgreSQL, nothing else
certs/ # generated demo PKI (git-ignored; only .gitkeep is tracked)
tests/
├── test_*.py # server, persistence, logging, certificates, TLS
├── integration/ # against a real PostgreSQL (opt-in)
├── agent/ # device agent
└── admin/ # CLI