A small two-part system for diagnosing CI build failures without wading through raw Jenkins console logs.
- Log Proxy (
server/) — a Flask server that proxies and caches Jenkins build logs on disk, so any given build's log is only ever downloaded from Jenkins once. - Log Proxy Client (
client/) — a CLI that downloads a build's log from the proxy and deduplicates it before printing, producing a smaller, more readable log for a human (or a downstream system such as alerting or an LLM).
Jenkins (ci.jenkins.io)
^
| one-time download per build_id
|
Log Proxy (Flask, :8080) <---- disk cache (cache/{build_id}.log)
^
| GET/HEAD /logs/{build_id}?offset=&limit=
|
Log Proxy Client (CLI)
|
v
deduplicated log, printed to stdout
The proxy and client are coupled through exactly one HTTP contract. The client has no knowledge that Jenkins exists; if the proxy were ever backed by a different log source, the client would not change.
log_cache.py— domain logic: given a build ID, return the path to its cached log, downloading from Jenkins on a cache miss. Knows nothing about HTTP or Flask.app.py— Flask routes. Translateslog_cache's exceptions into HTTP status codes and handles query-parameter parsing/validation.
proxy_client.py— thin HTTP client for the log proxy's API. Knows nothing about Jenkins.dedupe.py— two deduplication passes: strip leading Jenkins timestamps, then collapse consecutive duplicate lines.cli.py— entry point wiring the two together.
.
├── Makefile
├── README.md
├── REFLECTION.md
├── requirements.txt
├── server/
│ ├── __init__.py
│ ├── app.py
│ └── log_cache.py
├── client/
│ ├── __init__.py
│ ├── cli.py
│ ├── dedupe.py
│ └── proxy_client.py
└── tests/
├── test_app.py
├── test_log_cache.py
├── test_dedupe.py
└── test_proxy_client.py
- Python 3.9+
make
make build
Creates a virtualenv at .venv/ and installs flask, requests, and pytest.
make run-server
Starts the log proxy on http://localhost:8080. Logs are cached under ./cache/, created on first use.
make run-client ARGS="<build_id>"
For example:
make run-client ARGS="500"
Downloads and deduplicates the console log for build 500 from the proxy (default http://localhost:8080; override with --server <url>), printing the result to stdout.
make test
Runs the full pytest suite. Tests avoid the network and the real Jenkins server: log_cache/proxy_client tests mock requests.get, and app tests seed the on-disk cache directly rather than mocking Flask internals.
Returns Content-Length for the full cached log. Ignores offset/limit if supplied.
Optional query params:
offset(default0)limit(default: remainder of file)
Returns the requested byte slice with Content-Length and Content-Type: text/plain; charset=utf-8.
| Condition | Status |
|---|---|
| Success | 200 |
| Unknown build (Jenkins 404) | 404 |
Non-integer build_id in URL |
404 |
| Jenkins unreachable or errored | 502 |
Non-integer or negative offset/limit |
400 |
offset beyond the log's length |
416 |
- Timestamp format:
ci.jenkins.iocurrently returns 403 for anonymous requests (a change from the "public infrastructure" the assignment describes), so a live console log couldn't be pulled to confirm the exact format in use today. This assumes Jenkins Timestamper's ISO-8601 format (2024-01-15T13:45:23.456Z message) — the most common Timestamper configuration on real Jenkins pipelines. If the real format differs,strip_timestampsdegrades gracefully to a no-op rather than corrupting output. - Second dedup technique: consecutive identical lines are collapsed into
line (xN). Non-adjacent repeats (the same line reappearing later in the file) are intentionally left alone — see Future Improvements. - HEAD ignores
offset/limit: it always reports the size of the whole cached log, since it answers "how big is this resource," not "how big would this particular slice be." offsetequal to the file's length returns an empty 200 body rather than a 416; only strictly past the end of the file counts as an invalid range.- No cache invalidation: correct only because the assignment guarantees builds are finished and logs are immutable by the time they're requested.
- No lock around concurrent first-time downloads of the same uncached build ID — two simultaneous cold requests could both download. Harmless (the same immutable bytes get written either way) but wasteful; left out deliberately rather than adding concurrency-control complexity the assignment doesn't ask for.
- A shared, distributed cache (e.g., S3 or Redis) instead of local disk, if the proxy ever needed to run as more than one instance.
- A lock per build ID to avoid duplicate concurrent downloads on a cache miss.
- Non-adjacent duplicate detection (e.g., the same stack trace appearing more than once, not just back-to-back).
- Structured (JSON) error responses, if a downstream system other than a human ever needed to parse proxy errors.
- Recorded fixture logs for an integration-style test exercising the full client → proxy → Jenkins path.