TechNapoleon/server-client-admin

★ 0Forks 0PythonGitHub ↗Compare

README

Agent Manager Demo

A small, deliberately explainable secure remote device management / agent orchestration demo, built as a backend engineering exercise.

Three independent applications talk to each other over HTTPS: a management server, a device agent, and an administrator CLI. An administrator lists connected agents, sends a command to one of them, and reads back what it returned.

Status: official Steps 1, 2 and 3 are complete. Registration, multiple agents, job assignment, polling delivery, result collection, an admin API behind a bearer token, a CLI that drives all of it — TLS underneath the whole thing, with a private demo CA that agents and administrators verify the server against — and commands, results and application logs persisted to PostgreSQL running as its own service. Generic commands with an internal queue (Step 4) come next — see Roadmap.

Requirements

  • Python 3.14 (developed and tested on 3.14.4; 3.11+ should work)
  • Docker (or Docker Compose) — the database runs as its own service. See External persistence (PostgreSQL).

Quick start

python3 -m venv .venv
.venv/bin/pip install -r requirements.txt

Generate the demo PKI first. Certificates are never committed, so a fresh checkout has none:

.venv/bin/python scripts/generate_certs.py

This writes certs/{ca.crt,ca.key,server.crt,server.key}. See Secure transport for what each file is and who may hold it.

Start the database. It is a separate service, so it comes up before the application and stays up across restarts of it:

cp .env.example .env      # then set POSTGRES_PASSWORD and ADMIN_TOKEN
docker compose up -d postgres

Terminal 1 — the management server. It serves HTTPS only and refuses to start if the certificate, the key, ADMIN_TOKEN or the database is missing:

ADMIN_TOKEN=dev-admin-token \
SERVER_CERT_PATH=certs/server.crt \
SERVER_KEY_PATH=certs/server.key \
.venv/bin/python -m server.run
TLS enabled
Listening on https://localhost:8443
Certificate: certs/server.crt
Server 0.1.0 starting. Persistence: postgresql at postgresql+psycopg://agent_manager:***@localhost:5432/agent_manager

Terminal 2 — a device agent (start as many as you like, each with its own identity):

export SERVER_URL=https://localhost:8443 CA_CERT_PATH=certs/ca.crt

.venv/bin/python -m agent.client
DEV_AGENT_ID=agent-002 .venv/bin/python -m agent.client

Terminal 3 — the administrator:

export SERVER_URL=https://localhost:8443 CA_CERT_PATH=certs/ca.crt
export ADMIN_TOKEN=dev-admin-token

.venv/bin/python -m admin.cli agents

Those defaults are already the built-in ones, so with a .env in place the exports are optional.

Interactive API docs: https://localhost:8443/docs — your browser will warn about the certificate, because a private CA is exactly what a browser is supposed to distrust. Run the tests with .venv/bin/pytest -q.

The Step 1 workflow

$ python -m admin.cli agents
ID        HOSTNAME  OS     VERSION  STATUS
778623bf  lugal     Linux  0.1.0    online
d160c496  lugal     Linux  0.1.0    online
bb862fb7  lugal     Linux  0.1.0    online

Send echo to one agent and wait for the answer. Agents may be named by any unique prefix:

$ python -m admin.cli echo 778623bf "Hello World" --wait
Command queued for 778623bf.
Job ID: 20d3ad6c-e25c-4e71-82c1-481f58a46602

Status: completed
Result:
  Hello World

Job ID:    20d3ad6c-e25c-4e71-82c1-481f58a46602
Agent:     778623bf-36bb-41b5-b47e-95ec8657183a
Operation: echo
Arguments: Hello World
Created:   2026-08-11 17:26:33 UTC
Delivered: 2026-08-11 17:26:33 UTC
Completed: 2026-08-11 17:26:33 UTC

Without --wait the command returns immediately and the result is read separately:

$ python -m admin.cli result 20d3ad6c
Status: completed
Result:
  Hello World

The command reached only that agent:

$ python -m admin.cli jobs 778623bf
JOB       OPERATION  STATUS     CREATED                  COMPLETED
20d3ad6c  echo       completed  2026-08-11 17:26:33 UTC  2026-08-11 17:26:33 UTC

$ python -m admin.cli jobs d160c496
No commands have been sent to this agent.

Stop an agent gracefully. It reports the result before exiting:

$ python -m admin.cli kill 778623bf --wait
Status: completed
Result:
  agent shutting down
# in the agent's terminal
Received command 8576764b: kill
Reported completed for command 8576764b
Kill command received. Shutting down.
Agent stopped.

About thirty seconds later it is reported offline, while the others carry on:

$ python -m admin.cli agents
ID        HOSTNAME  OS     VERSION  STATUS
778623bf  lugal     Linux  0.1.0    offline
d160c496  lugal     Linux  0.1.0    online
bb862fb7  lugal     Linux  0.1.0    online

CLI reference

command what it does
health check the server is reachable (needs no token)
agents list agents and their status
agent <id> show one agent
echo <id> "<text>" [--wait] send an echo command
kill <id> [--wait] ask an agent to shut down gracefully
jobs <id> an agent's command history
result <job> a command's result

Only echo and kill are exposed, which is what Step 1 calls for. The agent's local allowlist is larger, and the server's is the authority on what may be assigned.

Architecture

                          Admin CLI  (admin/cli.py)
                              |
                              |  HTTPS  +  Authorization: Bearer $ADMIN_TOKEN
                              v
              +-------------------------------------+
              |        Management Server            |
              |            FastAPI                  |
              |                                     |
              |  api/agent.py     api/admin.py      |   <- transport / HTTP
              |        \             /              |
              |     security/identity.py            |   <- who is calling?
              |        \             /              |
              |  services/{agent,job}_service.py     |   <- business rules
              |             |                       |
              |        models.py / database.py      |   <- persistence
              +-------------------------------------+
                     |                     ^
                  SQLite                   |  HTTPS  (mutual TLS next)
                                           |
                                  +-----------------+
                                  |  Device Agent   |
                                  +-----------------+

Two independent trust relationships meet at this server:

Caller Authenticated by Status
Device agent AgentIdentity — a dev header today, a client certificate under mutual TLS in place
Administrator Authorization: Bearer $ADMIN_TOKEN in place

They are separate credentials on separate routers — holding one grants nothing on the other. The agent API never consults the admin token, and the admin API never looks at AgentIdentity.

What "connected" means here

The exercise talks about connected clients. This implementation polls over HTTP and holds no persistent socket per agent — each request is its own short-lived connection. So:

Connection state here is application-level presence, not an open socket. Registration and keep-alive update last_seen; status is derived from it at read time. An agent is online when it has contacted the server within AGENT_OFFLINE_AFTER_SECONDS (default 30).

The server manages many concurrently-running agents without keeping a connection open for any of them, which is what makes this design cheap to scale and trivial to restart on either side. It is worth saying out loud rather than letting the word "connected" imply a socket that does not exist.

The command path

Admin CLI                 Server                      Device Agent
    |                        |                             |
    |-- POST .../jobs ------>|  job: queued                |
    |    echo "Hello World"  |                             |
    |                        |<-- POST /jobs/claim --------|  every 2s
    |                        |    job: delivered ---------->|
    |                        |                             |  execute_operation()
    |                        |                             |  (allowlist only)
    |                        |<-- POST /jobs/{id}/result --|
    |                        |    job: completed           |
    |-- GET /jobs/{id} ----->|                             |
    |<-- "Hello World" ------|                             |

Claiming is a POST, not a GET: it moves the job to delivered and stamps delivered_at, and a GET that mutates state is neither cacheable nor safely retryable. Two concurrent claims cannot both win — the transition is a conditional UPDATE ... WHERE status = 'queued' whose row count decides.

Why polling. One code path, no connection state to manage, trivially testable, and every step is visible in the server log. WebSockets would later replace only the delivery step; the job table, the statuses and the result endpoint would stay exactly as they are.

Why Agent B can never receive Agent A's command

No agent endpoint takes an agent identifier. There is nothing to pass, and therefore nothing to pass wrongly:

X-Dev-Agent-Id (today)      client certificate (mutual TLS)
         \                          /
          v                        v
              get_agent_identity()
                     |
        AgentIdentity(source, stable_id)
                     |
        find_by_identity()  ->  the Agent row
                     |
   SELECT ... FROM jobs WHERE agent_pk = <that row> AND status = 'queued'

Addressing agents by name is an admin-side concept (/api/admin/agents/{agent_id}/jobs), because managing a fleet means choosing which device to act on. An agent's scope comes from who it is.

The same rule governs results: a job is looked up scoped to the calling agent, so knowing another agent's job_id is not enough to write to it. That case returns 404, not 403 — a "forbidden" would confirm the job exists and belongs to someone else, turning job_id into an oracle for mapping the fleet.

Job lifecycle

queued ──(agent claims)──> delivered ──(agent reports)──> completed
                                                      \─> failed

That is the entire state machine. Reporting twice on a job is a 409; results are recorded once.

The identity seam

The single most important design decision. Authentication is resolved at the transport edge and handed to the rest of the application as a plain value object:

        TLS / transport            <- SSL objects, peer certificates live here
              |
              v
       CertIdentityProvider        <- the ONLY code that parses X.509
              |
              v
         AgentIdentity             <- plain frozen dataclass of strings
              |
              v
   API / services / persistence    <- unaware that TLS exists at all

Nothing below AgentIdentity ever sees a Request, an SSLObject or an x509.Certificate.

Step Provider stable_id subject
1 DevIdentityProvider dev:agent-001 None
2+ CertIdentityProvider mtls:<sha256-fingerprint> CN=agent-001,O=Demo

Registration, job claiming and result submission all sit behind this one dependency, so swapping the provider changes who an agent is without changing what any of them do.

⚠️ The development identity is NOT authentication

Today an agent identifies itself with an X-Dev-Agent-Id header (falling back to the hostname in the request body). Anyone can send any value. It is a stable local-development identifier and nothing more — it is not a credential, not a secret, and not a security boundary. It exists purely so the registration, command and result logic can be built and tested before the PKI is wired into identity. Mutual TLS replaces the provider with certificate-derived identity; the API and service contracts do not change.

TLS does not fix this. It stops the header being read on the wire and proves to the agent which server it reached, but the handshake authenticates the server to the client, not the reverse. Encrypting a claim does not verify it: any process holding ca.crt can open a valid TLS session and then declare itself any agent it likes.

The admin token — still required, now also confidential

Authorization: Bearer $ADMIN_TOKEN establishes that the caller is the administrator. TLS establishes that the server is the real one, and keeps the token from being read in transit. Two separate properties, and Step 2 supplies the one that was missing:

Step 1    Bearer <token> over HTTP    ->  authenticated, NOT confidential
now       Bearer <token> over HTTPS   ->  authenticated AND confidential

What TLS does not do is authorize anyone. A completed handshake grants nothing — a wrong token over a perfectly valid TLS session still returns 401. The token is compared with secrets.compare_digest (a plain == returns early on the first differing byte, which leaks the secret one character at a time through timing), is never logged, and never appears in a response or an error message.

ADMIN_TOKEN is required — the server refuses to start without it, and there is no default, because a shipped default credential is precisely how a demo secret reaches production.

The device agent (agent/)

A managed device: it starts up, describes itself, registers with the management server, learns the agent_id the server assigned it, and keeps that registration alive.

agent/ is an independent application. It imports nothing from server/ — the HTTP API is the only contract between them, so the two halves could run on different machines or live in separate repositories. Its response model deliberately tolerates unknown fields (the mirror image of the server's strict request validation), so a server that starts returning more cannot break agents already deployed.

agent/
├── config.py       # environment-driven settings; server URL + CA trust anchor
├── system_info.py  # device facts — no networking, no job execution
├── client.py       # AgentClient (httpx.AsyncClient) + lifecycle + entry point
└── operations.py   # the safe-operation allowlist

system_info.py sits underneath both consumers, so registration never depends on the job-execution module:

              system_info.py
              /            \
             v              v
        client.py        operations.py

Configuration

variable default purpose
SERVER_URL https://localhost:8443 HTTPS; plain http:// is only for local experiments
CA_CERT_PATH unset required for HTTPS — the CA that verifies the server certificate
DEV_AGENT_ID agent-001 development identity — not a credential
AGENT_VERSION 0.1.0 reported as metadata
JOB_POLL_INTERVAL_SECONDS 2 how often to ask for work
KEEP_ALIVE_INTERVAL_SECONDS 10 how often to report being alive (offline after ~30s)
CONNECT_TIMEOUT_SECONDS 5 explicit — nothing waits forever
REQUEST_TIMEOUT_SECONDS 10 explicit — nothing waits forever
CLIENT_CERT_PATH / CLIENT_KEY_PATH unset mutual TLS: prove this agent's identity
AGENT_LOG_LEVEL INFO DEBUG also shows every HTTP request

The two cadences are separate because they answer different questions: how often a device checks for work bounds how long an administrator waits, while how often it says it is alive only needs to beat the offline threshold. Sharing one timer would mean either a sluggish demo or five times the database writes.

The serving loop

every JOB_POLL_INTERVAL_SECONDS:
    job = claim_next_job()
    while job:                        # drain, don't sleep between commands
        result = execute_operation(...)
        submit_result(...)
        if operation == "kill": stop
        keep_alive_if_due()           # a backlog must not starve it
        job = claim_next_job()
    keep_alive_if_due()

Work is drained rather than taken one command per interval, and the keep-alive check runs inside that loop — an agent busy working must never be reported offline. Step 4 replaces this with an internal queue and concurrent execution; a sequential loop is easier to follow and to trust for now.

⚠️ X-Dev-Agent-Id is temporary development identity, NOT authentication

The agent sends X-Dev-Agent-Id: <DEV_AGENT_ID> on every request. Any process can send any value; it carries no secret and proves nothing. It exists only so the server has a stable key to file this agent under while the demo PKI does not exist. Do not treat it as a credential.

Under mutual TLS identity moves into the handshake: the agent presents a client certificate, and the server derives identity from the validated certificate rather than from a header it was handed.

Keep-alive is a deliberate temporary shim

The server has no heartbeat endpoint yet, so keep_alive() re-calls the idempotent POST /api/agent/register, which refreshes last_seen and keeps the agent reported as online. It is a separate, clearly-labelled method so nothing comes to depend on registration doubling as a heartbeat:

now       keep_alive() -> POST /api/agent/register     (this shim)
later     keep_alive() -> POST /api/agent/heartbeat

It also verifies that the same development identity keeps mapping to the same agent_id, and logs an error if it ever changes — that would mean the server filed a duplicate record.

Restarting an agent re-registers it: the server returns the same agent_id rather than creating a second row. The client treats 201 Created and 200 OK identically — whether a record already existed is the server's business, not the device's.

Safe operations

agent/operations.py holds the allowlist: ping, get_system_info, get_time, echo, sleep and kill. Of these, the server currently permits only echo and kill to be assigned.

Dispatch is a lookup in a literal dictionary of already-imported functions, so a job's operation string is a dictionary key, never a lookup path — there is no route from remote input to an arbitrary callable. No shell, no subprocess, no eval/exec, no dynamic import. Arguments are validated by a strict model per operation, and sleep is capped at 30 seconds so a malformed job cannot wedge the agent. A command outside the allowlist becomes a failed result, never a crash.

Kill means "stop this agent process"

The kill handler kills nothing: it is a pure function returning {"message": "agent shutting down"}. The lifecycle stops, by recognising the operation after the result has been reported. A handler that called sys.exit would be untestable, would terminate before its result could be reported, and would put process control inside the code that executes remote input.

It never touches other processes, the machine, or the filesystem — and it takes no arguments, so there is no pid or target to supply. Shutdown does not depend on the server accepting the final result: one bounded attempt, then the agent exits regardless, because an operator killing an agent during an outage must not find it still running.

How the agent verifies the server

All transport security is assembled in one function, build_ssl_context():

context = ssl.create_default_context(cafile=CA_CERT_PATH)
context.minimum_version = ssl.TLSVersion.TLSv1_2
context.load_cert_chain(CLIENT_CERT_PATH, CLIENT_KEY_PATH)  # mutual TLS, later
httpx.AsyncClient(base_url=SERVER_URL, verify=context)

create_default_context(cafile=...) replaces the trust anchors rather than adding to them, so the agent trusts our demo CA and nothing else — a public CA has no business vouching for this server. Verification is left entirely to the TLS stack: chain, validity and SAN. No hand-rolled certificate checks, and never verify=False. An https:// URL without CA_CERT_PATH is refused at startup, because the alternative is an opaque handshake failure several layers down.

See Secure transport for the full trust model.

The administrator CLI (admin/)

The third independent application. It imports neither server nor agent, reads no database, and knows nothing about SQLAlchemy — the HTTP API is the whole contract, so it could be installed on another machine or split into its own repository unchanged.

admin/
├── config.py      # environment-driven settings; server URL + CA trust anchor
├── client.py      # AdminClient (sync httpx.Client) -> returns data, never prints
├── formatting.py  # data -> strings; no HTTP, no I/O
└── cli.py         # argparse, exit codes

Synchronous on purpose: the CLI is one request and an exit, so there is nothing to overlap. The agent is async because it runs a loop forever — a different problem. Transport and presentation are separated so either can change without the other.

variable default purpose
SERVER_URL https://localhost:8443 HTTPS; plain http:// is only for local experiments
CA_CERT_PATH unset required for HTTPS — the CA that verifies the server certificate
ADMIN_TOKEN unset required for every command except health; TLS does not replace it
CONNECT_TIMEOUT_SECONDS / REQUEST_TIMEOUT_SECONDS 5 / 10 explicit
ADMIN_LOG_LEVEL INFO DEBUG also shows every HTTP request

Expected failures print one line and exit non-zero, never a traceback, and the token appears in no message:

$ python -m admin.cli agents
Cannot reach the server: unable to connect to management server at https://localhost:8443

$ ADMIN_TOKEN=wrong python -m admin.cli agents
Authentication failed (check ADMIN_TOKEN): the server rejected the administrator token

Secure transport (TLS)

Step 2 of the exercise asks for a secure encryption protocol between client and server. This implementation does not define one. It uses TLS.

That is the whole design decision, and it is deliberate. Hand-rolling a protocol — AES-encrypting JSON bodies, RSA-wrapping content keys, inventing a nonce scheme, bolting on an HMAC — produces something that looks cryptographic and is almost certainly broken. Padding oracles, nonce reuse, unauthenticated ciphertext, no replay protection, no forward secrecy, no key rotation and no downgrade resistance are the normal outcomes, and none of them are visible in a passing test suite. TLS is the standard answer to this exact problem, has been attacked by everyone for thirty years, and arrives already implemented in OpenSSL.

So the application protocol is unchanged and the security lives underneath it:

Application   HTTP + JSON        <- identical to Step 1, byte for byte
                  |
Transport     TLS 1.3            <- session keys, confidentiality, integrity
                  |
Wire          TCP

Not one route, schema, service or database call changed. Uvicorn terminates TLS and FastAPI sees ordinary HTTP requests, which is why the entire Step 1 test suite still runs unmodified through TestClient.

The private demo CA

        Demo Root CA          ca.key   the CA's private signing key
              |               ca.crt   the public trust anchor
              +-- signs -->   server.crt   SAN: DNS:localhost, IP:127.0.0.1
                              server.key   the server's private key

Agent  --trusts--> ca.crt --verifies--> server.crt
Admin  --trusts--> ca.crt --verifies--> server.crt
file what it is who holds it
ca.crt the trust anchor: clients believe certificates it signed agents, admins, freely distributable
ca.key signs certificates scripts/generate_certs.py only — the server never loads it
server.crt authenticates the management server during the handshake the server; sent to every client, public
server.key proves the server owns server.crt the server host only, never distributed

A private CA rather than a self-signed certificate: a self-signed cert has to be distributed to and pinned by every client, and replacing it means touching every client. With a CA, the anchor is distributed once and server certificates can be reissued underneath it.

server.crt does not encrypt anything. It authenticates the server — proving possession of server.key for a certificate chaining to the CA. The handshake then derives ephemeral session keys, and those protect the traffic. The certificate's job is answering "who am I talking to", not "how is this scrambled".

Why SAN validation matters

The certificate carries DNS:localhost, IP:127.0.0.1 in its Subject Alternative Name. Clients match the address they dialled against those entries and ignore the Common Name entirely — CN-based matching was deprecated for good reason. Without this check, any certificate the CA ever issued would be accepted for any host, so a party holding a legitimately-issued certificate for one name could impersonate any other. Verified against the running server:

$ openssl s_client -connect 127.0.0.1:8443 -CAfile certs/ca.crt -verify_hostname localhost
Verify return code: 0 (ok)

$ openssl s_client -connect 127.0.0.1:8443 -CAfile certs/ca.crt \
      -verify_hostname not-in-the-san.example
verify error:num=62:hostname mismatch
Verify return code: 62 (hostname mismatch)

Same server, same CA, same open socket — only the name differs.

What TLS does and does not protect

✅ Confidentiality nobody on the network can read commands, results, or the admin token
✅ Integrity nobody can alter a command or a result in flight
✅ Server authentication clients know they reached the real server, not an impostor
❌ Client authentication the handshake is one-directional — the client proves nothing

That last row is the important one, and it has two consequences.

ADMIN_TOKEN is still required. TLS and the token answer different questions:

TLS          "is this the real server?"      confidentiality + integrity + server identity
ADMIN_TOKEN  "is this caller an admin?"      application-level authorization

A completed handshake grants no authorization at all. Anyone holding ca.crt — a public file — can open a perfectly valid TLS session to this server; what they cannot do is call /api/admin without the token. Dropping the token because "it's encrypted now" would confuse a confidentiality control with an authentication one and leave the admin API open to anyone who can reach the port:

$ ADMIN_TOKEN=wrong-token python -m admin.cli agents
Authentication failed (check ADMIN_TOKEN): the server rejected the administrator token

$ curl --cacert certs/ca.crt -H "Authorization: Bearer wrong-token" \
       https://localhost:8443/api/admin/agents
{"detail":"Administrator authentication required."}     # HTTP 401 — TLS succeeded

X-Dev-Agent-Id is still not authentication. TLS now hides the header from passive observation and stops the agent being redirected to an impostor server. It does not make the claim true. Encrypting an assertion does not verify it, and any process with ca.crt can complete a valid handshake and then declare itself any agent it likes. Only mutual TLS closes this.

TLS version and cipher policy

One parameter is set by hand — the minimum version — and only because it cannot otherwise be guaranteed:

context.minimum_version = ssl.TLSVersion.TLSv1_2

A stock SSLContext reports minimum_version as the MINIMUM_SUPPORTED sentinel, which defers to whatever the host's OpenSSL configuration permits — and that varies by distribution. Naming the floor turns "we use secure defaults" from a claim into something the test suite asserts.

Nothing else is tuned: no cipher list, no curve list, no version ceiling. OpenSSL's own defaults are better maintained than anything hand-written here, and leaving the ceiling open means TLS 1.3 is negotiated whenever both ends support it, which in practice is always:

$ openssl s_client -connect localhost:8443 -CAfile certs/ca.crt
Protocol  : TLSv1.3
Cipher    : TLS_AES_256_GCM_SHA384
Verify return code: 0 (ok)

Fail fast, never fall back

python -m server.run serves HTTPS or it does not serve:

python -m server.run
      |
      v
validate TLS configuration
      |
  +---+---+
valid   invalid
  |       |
  v       v
HTTPS   exit 1, with the reason

There is no TLS_ENABLED flag, no autodetection that quietly degrades when a variable is misspelled, and no plaintext fallback. An operator whose TLS is broken must find out immediately, not by later discovering that traffic they believed was encrypted never was. Four failures are checked before anything binds a port — unset variable, missing file, unreadable file, and a certificate and key that do not match:

$ python -m server.run
Cannot start: SERVER_CERT_PATH is not set, so TLS cannot be configured -- it is the server
certificate presented to agents and administrators.
The management server serves HTTPS only; it will not fall back to plaintext.
...

$ SERVER_CERT_PATH=certs/ca.crt SERVER_KEY_PATH=certs/server.key python -m server.run
Cannot start: TLS certificate and key do not form a usable pair:
    certificate  certs/ca.crt
    key          certs/server.key

OpenSSL reported: [X509: KEY_VALUES_MISMATCH] key values mismatch (_ssl.c:4181)

The mismatch check is real, not a guess: the pair is loaded into a throwaway SSLContext, exactly what Uvicorn does moments later.

Private keys never reach git

.gitignore blocks *.key, *.pem and the whole of certs/ except .gitkeep, so generated material cannot be committed by accident:

$ git ls-files | grep -E '\.key$|\.crt$'
(nothing)

$ git ls-files certs/
certs/.gitkeep

Certificates are regenerated per checkout rather than shared. Keys are written 0600 at creation time via os.open, so they are never briefly world-readable, and the generator prints paths, SANs and fingerprints — never key material.

Next: mutual TLS

Step 2 authenticates the server to its clients. The remaining half authenticates clients to the server:

today                         mutual TLS
-----                         ----------
Agent                         Agent
  | verifies server cert        | verifies server cert
  v                             | presents agent cert
Server                          v
                              Server
                                | verifies agent cert
                                v
                              CertIdentityProvider  ->  AgentIdentity

The architecture is already shaped for it. AgentIdentity, identity_source, identity_stable_id, certificate_fingerprint and certificate_subject all exist and are persisted today; swapping DevIdentityProvider for CertIdentityProvider behind get_identity_provider() changes who an agent is without touching registration, job claiming, result submission, the service layer or the database — that is what the identity seam is for.

It is deferred rather than dismissed because it is not trivial: Uvicorn exposes no stable interface for handing the validated peer certificate to an ASGI application (the ASGI TLS extension remains a draft it does not implement), so it needs work at the protocol layer. Shipping working TLS first was the better trade.

External persistence (PostgreSQL)

Step 3 asks for commands, results and logs to live in an external database. Here "external" means what it says: a database service running as its own process, started and stopped independently of the application.

Admin CLI ──HTTPS+token──┐
                         v
                  FastAPI server  ── SQLAlchemy ──> ┌──────────────┐
                         ^                          │  PostgreSQL  │  separate process
Agent ────HTTPS──────────┘                          │  (container) │
                                                    ├──────────────┤
                                                    │ agents       │
                                                    │ jobs         │  commands + results
                                                    │ log_events   │  operational logs
                                                    └──────────────┘

The two halves start and stop separately, which is the demonstration:

docker compose up -d postgres     # the database
.venv/bin/python -m server.run    # the application, on the host

docker-compose.yml contains PostgreSQL and nothing else. The server, agent and admin CLI stay on the host: containerising them would hide the separation behind a single up, and the point is that the database outlives any of them.

How little had to change

Persistence was already behind SQLAlchemy, and server/database.py was already the only module that knew which database was in use. So the move was a URL, a driver, and a startup check:

DATABASE_URL sqlite:///./agent_manager.db → postgresql+psycopg://…
requirements.txt psycopg[binary]
server/database.py pool_pre_ping for a networked database, plus check_database_connection()

No model, service, route or repository changed to accommodate PostgreSQL. The column types were already portable — JSON rather than JSONB, String(36) rather than UUID, Python-side defaults rather than server_default — so the same create_all() builds the same schema on either engine. The one genuine difference is that PostgreSQL returns timezone-aware datetimes where SQLite returns naive ones, and ensure_utc() already normalised both.

jobs holds commands and results; log_events holds logs

These answer different questions and are deliberately kept apart.

Admin creates command ──> jobs row: queued
                                      |
                          agent claims v
                                   delivered
                                      |
                          agent reports v
                          ┌───────────┴───────────┐
                    completed + result       failed + error

jobs is the source of truth for what was commanded and what came back — one row per command, carrying job_id, the owning agent (agent_pk, the surrogate key behind the public agent_id), operation, arguments, status, created_at, delivered_at, completed_at, result and error. The command is committed when the administrator creates it; the result when the agent reports it. Reconstructing history from log lines instead would turn an auditable state machine into text that has to be parsed to be believed.

log_events is operational logging: what the server was doing, and when. Nothing in it is required to answer "what did we run and what came back".

Logs reach the database without slowing requests

Integration happens at the logging layer, not at the call sites, so every logger.info already in the codebase is persisted without being rewritten:

logger.info("Job %s created", job_id)
      |
      ├──> StreamHandler ─────────────────────────> console   (synchronous)
      |
      └──> QueueHandler ──> queue ──> QueueListener thread ──> INSERT INTO log_events

QueueHandler.emit does no I/O — it puts the record on an in-memory queue and returns — so the request that logged never waits for the database. The queue is bounded (10,000 records): if the database is slow or gone, records are dropped rather than buffered until the process runs out of memory.

Two failure modes get explicit attention in server/logging_setup.py:

  • Recursion. A handler that logs while handling a record feeds itself, and writing to a database makes that real, because SQLAlchemy and psycopg log. Three independent guards: a thread-local re-entrancy flag, a filter rejecting sqlalchemy/psycopg/server.logging_setup, and the fact that persistence failures never re-enter logging at all.
  • A dead database. Failures go to stderr via handleError, outside the logging system. The record is lost from the database and kept on the console — a logging outage must not become an application outage.

Only server.* and uvicorn.error are persisted. uvicorn.access is excluded deliberately: an idle agent polls every two seconds, so access logging alone would add ~30 rows per minute per agent and bury the events an operator came to read. The same reasoning applies to the keep-alive, which reuses the registration endpoint — a genuine re-registration is logged, a routine ping is not.

What is never written. No tokens, no Authorization headers, no key material, no tracebacks, no command arguments or results. The fields on LogEvent are an allowlist rather than a copy of the record, and the database password is masked (postgresql+psycopg://agent_manager:***@…) everywhere it is displayed. A failed admin authentication is logged — repeated failures are worth seeing — but never the token that was presented, not even a prefix: the whole reason require_admin compares in constant time is to avoid leaking it a fragment at a time.

If the database is missing, the server does not start

PostgreSQL unreachable ──> ConfigurationError, exit 1
                     ╳───> fall back to local SQLite

There is no fallback, and that is a deliberate refusal. A server that quietly downgraded to a local file would look healthy while writing the fleet's command history somewhere nobody is looking, and the operator would find out only when they went looking for an audit trail that was never there. The check runs in server/run.py before anything binds a port, so the message reads as the configuration problem it is rather than as a crash:

Cannot start: Cannot reach the database at postgresql+psycopg://agent_manager:***@localhost:5432/agent_manager.
The server will not start without it, and it will not fall back to a local SQLite file [...]

    docker compose up -d postgres

SQLite is still the right tool for unit tests

runtime schema management
Real runs, the demo, tests/integration/ PostgreSQL create_all()
Unit tests (tests/) throwaway SQLite per test create_all()

This is not a hedge. SQLAlchemy is the persistence abstraction, so the models, services and routes under test are the same either way, and an empty file-backed database per test costs milliseconds and needs no running service. What the unit tests deliberately do not prove is that the system uses an external database — that is what the PostgreSQL end-to-end run and tests/integration/ are for.

Production would do three things differently

  • Migrations. create_all() creates missing tables and does nothing about changing one. A real deployment would use Alembic so schema changes are versioned, reviewable and reversible.
  • Logs to an observability stack. A relational table is queryable and good enough at this size, but production logging belongs somewhere built for retention, indexing and alerting.
  • Credentials and connections. A managed secret rather than .env, TLS on the database connection, and a pool sized against the real workload.

Testing

.venv/bin/pytest -q     # 293 tests, no services required
suite what it covers
tests/test_*.py server: registration, admin API, auth boundary, job flow, isolation
tests/test_persistence.py what actually reaches the database for a command and its result
tests/test_log_events.py the logging pipeline: mapping, filtering, failure, secrets
tests/integration/ the same guarantees against a real PostgreSQL — opt-in
tests/agent/ agent: operations allowlist, client, serving loop, kill semantics
tests/admin/ CLI: client, formatting, commands, exit codes, independence
tests/test_certs.py the generated PKI: SANs, CA:FALSE, EKU, chain, key permissions
tests/test_server_tls_config.py every way TLS can be misconfigured fails startup
tests/test_client_tls.py both clients verify; nothing in the source disables verification
tests/test_tls_integration.py real handshakes against a real Uvicorn TLS server

Agent and admin tests run entirely through httpx.MockTransport — no server, no socket — so each application is tested as the independent thing it is. Independence is also asserted directly: a test imports each package in a subprocess and checks that nothing named server or agent appears in sys.modules.

The TLS integration test is the exception, deliberately: TestClient never opens a socket and so can say nothing about transport security. It starts python -m server.run in a subprocess on a free port with a throwaway PKI, and asserts what actually distinguishes a secure transport — the correct CA succeeds, an unrelated CA fails, a name outside the SAN fails, a plaintext client cannot reach the API, the negotiated version is TLS 1.3, and a wrong admin token gets 401 over a successful handshake. The server is terminated in a finally, so a failing test cannot leak a process holding a port.

Two of these tests are written against specific failure modes rather than "an error occurred". The untrusted-CA cases unwrap the exception chain and assert SSLCertVerificationError with unable to get local issuer certificate, because asserting on httpx.TransportError alone would also pass if the server were simply down. The SAN test verifies the name over an already-open, known-good socket, because the obvious version — dialling a second loopback address — passes whether or not hostname checking works at all: the server binds one address, so that connection fails at TCP before TLS is ever consulted.

The PostgreSQL suite is skipped unless you point it at a database, so pytest stays fast and needs nothing running:

docker compose up -d postgres
POSTGRES_TEST_URL='postgresql+psycopg://agent_manager:...@localhost:5432/agent_manager' \
    .venv/bin/pytest tests/integration -q

It is not the API suite re-pointed at another engine — SQLAlchemy already makes those tests engine-independent, and running them twice would double the runtime to re-prove what the abstraction guarantees. It covers only what SQLite cannot answer for: that create_all() emits valid PostgreSQL DDL, that JSON arguments and results survive a real driver as structures, that timezone-aware timestamps read back correctly, and that rows are visible to a connection that did not write them. Each test builds and drops its own schema, so it can run against the demo database without disturbing it.

Roadmap

Step Contents Status
1 Client/server connectivity, multiple agents, admin CLI, echo + kill, results ✅ done
2 Encryption: TLS, demo PKI, server authentication ✅ done
2b Mutual TLS: client certificates, certificate-derived identity planned
3 Microservices: external database for logs and command history ✅ done
4 Generic commands, async server, heartbeat, internal agent queue, scale planned
5 Scale and functionality testing planned

Known limitations

  • Agent identity is self-declared. X-Dev-Agent-Id proves nothing. TLS keeps it off the wire in cleartext and authenticates the server, but the client is not authenticated at all — mutual TLS is what closes this.
  • The demo CA is a demo CA. ca.key sits unencrypted next to the certificates it signed. A real deployment keeps the signing key offline or in an HSM, and issues short-lived certificates.
  • No certificate revocation. Nothing consults a CRL or OCSP; a compromised certificate stays valid until it expires. A revocation denylist is a natural companion to mutual TLS.
  • create_all() instead of migrations. It creates missing tables and does nothing about changing one. Alembic is the production answer; see External persistence.
  • Logs live in a relational table. Queryable and adequate at this size, but production logging belongs in a stack built for retention, indexing and alerting.
  • A log record can be dropped under sustained database failure. The queue is bounded on purpose: keeping every record would trade a logging outage for an application outage. Dropped records are still on the console, and jobs — the auditable record — is written synchronously and never affected.
  • Sequential command execution. One command at a time per agent, by design until Step 4.
  • A dead agent still reads online for up to AGENT_OFFLINE_AFTER_SECONDS. That is inherent to presence-based status, and the alternative — a stored flag — is worse, because a killed or partitioned agent never gets to write offline at all.
  • A registration body with no hostname and no X-Dev-Agent-Id returns 400, not the 422 an invalid body would normally get: identity is resolved before body validation, and the development provider uses hostname as its fallback. This coupling disappears with mutual TLS, when identity comes from the certificate and never from the body.

Project layout

server/                # the management server
├── main.py            # create_app(), lifespan, router wiring -- transport-agnostic
├── run.py             # the executable entry point; TLS enforced, no fallback
├── config.py          # environment-driven settings, incl. require_tls_material()
├── database.py        # engine, session factory, get_db  <-- the only DB-aware module
├── logging_setup.py   # console + queued background writer to log_events
├── models.py          # Agent, Job, LogEvent
├── schemas.py         # pydantic request/response models
├── security/
│   ├── identity.py    # AgentIdentity + providers   <-- the seam
│   └── admin_auth.py  # bearer-token dependency
├── api/
│   ├── health.py
│   ├── agent.py       # register, claim, submit result
│   ├── admin.py       # agents, jobs, results
│   └── serialization.py
└── services/
    ├── agent_service.py
    └── job_service.py

agent/                 # the device agent -- imports nothing from server/
├── config.py
├── system_info.py
├── client.py
└── operations.py

admin/                 # the CLI -- imports neither server/ nor agent/
├── config.py
├── client.py
├── formatting.py
└── cli.py

scripts/
└── generate_certs.py  # the demo PKI; the only code that touches ca.key

docker-compose.yml     # the external database service -- PostgreSQL, nothing else
certs/                 # generated demo PKI (git-ignored; only .gitkeep is tracked)
tests/
├── test_*.py          # server, persistence, logging, certificates, TLS
├── integration/       # against a real PostgreSQL (opt-in)
├── agent/             # device agent
└── admin/             # CLI

Contributors

TechNapoleon

Issues