vellankikoti/cloudbees-sse-assignment

★ 0Forks 0PythonGitHub ↗Compare

README

Jenkins Log Proxy & Dedup Client

A small two-part system for diagnosing CI build failures without wading through raw Jenkins console logs.

Overview

  • Log Proxy (server/) — a Flask server that proxies and caches Jenkins build logs on disk, so any given build's log is only ever downloaded from Jenkins once.
  • Log Proxy Client (client/) — a CLI that downloads a build's log from the proxy and deduplicates it before printing, producing a smaller, more readable log for a human (or a downstream system such as alerting or an LLM).

Architecture

Jenkins (ci.jenkins.io)
        ^
        | one-time download per build_id
        |
  Log Proxy (Flask, :8080) <---- disk cache (cache/{build_id}.log)
        ^
        | GET/HEAD /logs/{build_id}?offset=&limit=
        |
  Log Proxy Client (CLI)
        |
        v
  deduplicated log, printed to stdout

The proxy and client are coupled through exactly one HTTP contract. The client has no knowledge that Jenkins exists; if the proxy were ever backed by a different log source, the client would not change.

Server (server/)

  • log_cache.py — domain logic: given a build ID, return the path to its cached log, downloading from Jenkins on a cache miss. Knows nothing about HTTP or Flask.
  • app.py — Flask routes. Translates log_cache's exceptions into HTTP status codes and handles query-parameter parsing/validation.

Client (client/)

  • proxy_client.py — thin HTTP client for the log proxy's API. Knows nothing about Jenkins.
  • dedupe.py — two deduplication passes: strip leading Jenkins timestamps, then collapse consecutive duplicate lines.
  • cli.py — entry point wiring the two together.

Project Structure

.
├── Makefile
├── README.md
├── REFLECTION.md
├── requirements.txt
├── server/
│   ├── __init__.py
│   ├── app.py
│   └── log_cache.py
├── client/
│   ├── __init__.py
│   ├── cli.py
│   ├── dedupe.py
│   └── proxy_client.py
└── tests/
    ├── test_app.py
    ├── test_log_cache.py
    ├── test_dedupe.py
    └── test_proxy_client.py

Requirements

  • Python 3.9+
  • make

Installation

make build

Creates a virtualenv at .venv/ and installs flask, requests, and pytest.

Running the Server

make run-server

Starts the log proxy on http://localhost:8080. Logs are cached under ./cache/, created on first use.

Running the Client

make run-client ARGS="<build_id>"

For example:

make run-client ARGS="500"

Downloads and deduplicates the console log for build 500 from the proxy (default http://localhost:8080; override with --server <url>), printing the result to stdout.

Running Tests

make test

Runs the full pytest suite. Tests avoid the network and the real Jenkins server: log_cache/proxy_client tests mock requests.get, and app tests seed the on-disk cache directly rather than mocking Flask internals.

API

HEAD /logs/{build_id}

Returns Content-Length for the full cached log. Ignores offset/limit if supplied.

GET /logs/{build_id}

Optional query params:

  • offset (default 0)
  • limit (default: remainder of file)

Returns the requested byte slice with Content-Length and Content-Type: text/plain; charset=utf-8.

Condition Status
Success 200
Unknown build (Jenkins 404) 404
Non-integer build_id in URL 404
Jenkins unreachable or errored 502
Non-integer or negative offset/limit 400
offset beyond the log's length 416

Assumptions

  • Timestamp format: ci.jenkins.io currently returns 403 for anonymous requests (a change from the "public infrastructure" the assignment describes), so a live console log couldn't be pulled to confirm the exact format in use today. This assumes Jenkins Timestamper's ISO-8601 format (2024-01-15T13:45:23.456Z message) — the most common Timestamper configuration on real Jenkins pipelines. If the real format differs, strip_timestamps degrades gracefully to a no-op rather than corrupting output.
  • Second dedup technique: consecutive identical lines are collapsed into line (xN). Non-adjacent repeats (the same line reappearing later in the file) are intentionally left alone — see Future Improvements.
  • HEAD ignores offset/limit: it always reports the size of the whole cached log, since it answers "how big is this resource," not "how big would this particular slice be."
  • offset equal to the file's length returns an empty 200 body rather than a 416; only strictly past the end of the file counts as an invalid range.
  • No cache invalidation: correct only because the assignment guarantees builds are finished and logs are immutable by the time they're requested.
  • No lock around concurrent first-time downloads of the same uncached build ID — two simultaneous cold requests could both download. Harmless (the same immutable bytes get written either way) but wasteful; left out deliberately rather than adding concurrency-control complexity the assignment doesn't ask for.

Future Improvements

  • A shared, distributed cache (e.g., S3 or Redis) instead of local disk, if the proxy ever needed to run as more than one instance.
  • A lock per build ID to avoid duplicate concurrent downloads on a cache miss.
  • Non-adjacent duplicate detection (e.g., the same stack trace appearing more than once, not just back-to-back).
  • Structured (JSON) error responses, if a downstream system other than a human ever needed to parse proxy errors.
  • Recorded fixture logs for an integration-style test exercising the full client → proxy → Jenkins path.

Contributors

vellankikoti

Issues