AbraaoAlves/LETS_GO

★ 0Forks 0ShellGitHub ↗Compare

README

LETS_GO: Lab Experiment Track System for Governance and Operations

Project Summary

The system is a Laboratory Experiment Tracking System (LETS) to manage laboratory governance and operations. Created to solve this challange (PROBLEM.md)

Architecture

See ARCHITECTURE.md for the runtime map, domain graph, and database invariant map used to guide implementation.

Quick Start

Start the stack and explore (watch mode)

docker compose up

Flyway applies the Postgres migrations, then run_pipeline.sh loads seed data, imports valid CSV fixtures, and proves the invalid fixtures are rejected by database constraints/triggers. Postgres stays up afterward so you can connect and run queries. Stop and clean up with docker compose down -v.

Run the pipeline once as a test (no watch)

docker compose run --rm pipeline_tester; docker compose down -v

Same migrations and pipeline, but it runs once and tears everything down on exit. Test results print directly to the terminal without container-log prefixes — handy for a quick pass/fail check.

Port override

If local port 5432 is busy, keep the same internal database wiring and publish a different host port by prefixing either command:

POSTGRES_PORT=55432 docker compose up

Connect with:

docker compose exec postgres psql -U postgres -d lab

Useful demo queries:

-- Project collaborators (M:N)
SELECT p.title, array_agg(r.name ORDER BY r.name) AS collaborators
FROM projects p
JOIN project_researchers pr ON pr.project_id = p.id
JOIN researchers r ON r.id = pr.researcher_id
GROUP BY p.title
ORDER BY p.title;

-- Experiment ancestor chain
WITH RECURSIVE chain AS (
  SELECT id, title, predecessor_experiment_id, 0 AS depth
  FROM experiments
  WHERE title = 'Gamma Planning Run'
  UNION ALL
  SELECT e.id, e.title, e.predecessor_experiment_id, c.depth + 1
  FROM experiments e
  JOIN chain c ON c.predecessor_experiment_id = e.id
)
SELECT depth, title FROM chain ORDER BY depth;

-- Samples used by each experiment, derived from measurements
SELECT e.title, array_agg(DISTINCT s.sample_code) FILTER (WHERE s.id IS NOT NULL) AS samples
FROM experiments e
LEFT JOIN measurements m ON m.experiment_id = e.id
LEFT JOIN samples s ON s.id = m.sample_id
GROUP BY e.title
ORDER BY e.title;

Schema Overview

The schema has seven tables: researchers, projects, project_researchers, samples, experiments, measurement_types, and measurements. IDs are generated by Postgres; CSV fixtures resolve relationships through stable keys documented in docs/csv-ingestion.md.

Open Questions & Assumptions

Full detail, including impact and database enforcement for each assumption, lives in QUESTIONS_ASSUMPTIONS.md. The short version:

  • Runtime boundary: this is a CSV ingestion workflow for lab exports, not a web API or application service. Durable invariants live in Postgres constraints, foreign keys, checks, and triggers.
  • Measurements: one measurements table stores typed JSONB payloads classified by measurement_type. Known core shapes are database-checked; new types can be registered as data before their shape is hardened.
  • Project lifecycle: completed and cancelled projects freeze new descendant experiments and measurements; planning and active remain writable.
  • Experiment lineage: follow-ups are traceability links only. They do not copy samples, hypotheses, or status, and cross-project follow-ups are allowed.
  • Sample usage: a measurement row is the record that an experiment used a sample. There is no separate experiment_samples table until the lab confirms sample usage must be pre-registered or tracked without a measurement.

Questions I would clarify with the lab before building this further:

  • Are samples inventory-managed physical resources with quantities, depletion, disposal, or chain-of-custody history?
  • Must an experiment pre-register samples before any measurements are recorded?
  • Are measurement corrections routine enough to require append-only versioning instead of strict update/delete blocking?
  • Do researcher roles drive approvals or access control, or are they only descriptive metadata?
  • What do real machine CSV exports look like, and do they require durable staging plus row-level error reports?

Decisions & Trade-offs

0. Runtime context: CSV ingestion workflow

Based on A0, I assume the first runtime context is a CSV ingestion workflow for lab exports, not a full web API or long-running application service.

This keeps the solution focused on what the challenge asks for: a Postgres data model, Docker startup, and seed data that prove the model works. The domain rules should live as close to the data as possible through tables, foreign keys, unique constraints, enums, and check constraints. Any CSV-specific parsing or row-shape validation can live in the ingestion step when that workflow is implemented.

Below is the rationale behind this choice and the trade-offs considered.


  • Web API Microservice Approach : Build an HTTP API around the model, with request handlers, DTOs, service classes, authentication concerns, and validation in the application layer. Why it was rejected: The prompt asks for a database model that can be started with Docker and inspected live. Adding an API now would create more code to explain without proving the core model better.

  • Full Domain Application Layer Approach : Build a DDD-style application layer with use cases, repositories, and schema validators before the database is exercised. Why it was rejected: The current ambiguity is mostly about domain invariants, not transport or orchestration. Capturing those invariants in the schema and documentation gives faster value and keeps the interview surface smaller.

  • [Chosen] CSV Ingestion Workflow Approach : Treat the lab's current spreadsheet/export world as the input boundary. CSV rows are parsed and normalized before insert, while Postgres remains the source of truth and final validation guard for relationships and durable constraints. See docs/csv-ingestion.md for the flat-CSV → JSONB transform seam that bridges this decision with the JSONB measurement model in Decision 1.

Advantages

  • Simple Delivery Path: Matches the required Docker + Postgres + seed-data deliverable without adding an unrequested API surface.
  • Pragmatic Migration from Spreadsheets: Fits the lab's current reality: they are already working from spreadsheet-like systems and need a clean target model.
  • Centralized Data Integrity: Core rules are enforced by the database, so future import scripts, CLIs, or APIs inherit the same constraints.

Accepted Trade-offs & Risks :

  • Batch-Oriented Feedback: CSV ingestion usually reports errors after a file is processed, not interactively while a researcher is typing.
  • Importer Responsibility: CSV header mapping, required columns, type coercion, and row-level error messages still need to be handled outside the database.
  • Limited Product Surface: This does not yet solve user workflows such as login, approval, dashboards, or real-time experiment editing.
Mitigation Strategies:

Keep the database strict on durable invariants, document the expected CSV contract near the ingestion script when it exists, and add a small staging/error-report table only if bad production files become a real workflow problem.

0.1. CSV shape: simplified contract over raw lab exports

Real laboratory machines often produce noisy, denormalized flat CSV files: one row can mix project, experiment, sample, researcher, and measurement columns, and bad rows may contain malformed values or inconsistent names. I am aware that a production-grade lab importer would need a broader ETL/staging workflow for that raw shape.

For this challenge, I intentionally use a simpler CSV contract that is close to the target database model. The goal is to prove the relational design and database invariants, not to build a general-purpose CSV cleaning product.

Below is the rationale behind this choice and the trade-offs considered.


  • Raw Machine Export ETL Approach : Accept the exact denormalized files emitted by lab equipment and split them into the target tables during ingestion, including header mapping, row quarantine, deduplication, idempotency, and malformed-line recovery. Why it was deferred: Those are real production concerns, but they require real source files and operational error-handling requirements. Building that now would bury the database model under importer behavior that the challenge does not ask for.

  • SQL Seed-Only Approach : Skip CSV ingestion entirely and prove the schema only with hand-written INSERT statements. Why it was rejected: It would make the Docker demo smaller, but it would not exercise the CSV boundary established in A0.

  • [Chosen] Simplified Contract CSV Approach : Use explicit, predictable CSV fixtures that already follow the domain vocabulary and can be loaded through temporary staging tables plus SQL transforms. Postgres remains the source of truth and final validation guard through foreign keys, unique constraints, checks, and triggers.

Advantages

  • Honest Challenge Scope: Keeps attention on the schema, invariants, Docker startup, and seed/demo data requested by the problem.
  • Executable Without Importer Sprawl: A reviewer can run the pipeline and inspect the database without reading a custom ETL framework first.
  • Still Compatible With A0: CSV remains the input boundary, while durable rules stay enforced by the database rather than an absent application layer.

Accepted Trade-offs & Risks :

  • Not a Production Raw-Export Importer: The solution does not claim to handle arbitrary machine exports, malformed delimiters, duplicate names, or row-level quarantine.
  • Upstream Normalization Assumed: The CSV contract assumes stable identifiers or already-separated entity rows before data reaches the database ingestion step.
  • Simpler Error Feedback: Failed rows surface as Postgres cast, FK, unique, check, or trigger errors, not as polished per-row business messages.
Mitigation Strategies:

Treat the simplified CSV shape as a documented contract for this challenge. If real raw lab exports become part of the requirement, add a small durable staging/error-report layer or a separate pre-processing step that converts raw files into this contract before COPY. Do not relax the database constraints; they remain the final guard under A0.

0.2. Tooling: Flyway schema + Bash CSV pipeline

The implementation splits schema history from executable ingestion tests: Flyway owns migrations, while run_pipeline.sh owns seed loading, CSV imports, and acceptance assertions.


  • One application tool does migration plus ingestion : Build a CLI or service that migrates, imports, and validates. Why it was rejected: It adds an application layer the challenge does not need and cuts against A0.

  • Flyway does schema and seed/test data : Put demo data and CSV assertions into migrations. Why it was rejected: It couples disposable fixtures to schema history and gives no clean place for invalid-row assertions.

  • Plain Postgres init scripts : Rely on /docker-entrypoint-initdb.d. Why it was rejected: It only runs on a fresh volume and is not explicit versioned migration history.

  • [Chosen] Flyway + run_pipeline.sh : Flyway applies versioned DDL. The pipeline imports CSV fixtures through TEMP tables and lets Postgres accept or reject each case.

Advantages

  • Separation of Concerns: Schema migrations are durable history; CSV fixtures are executable proof data.
  • TDD Shape: The pipeline proves every documented invariant at the database boundary.
  • Single Command: docker compose up still builds the schema and leaves a running database.

Accepted Trade-offs & Risks :

  • Two Moving Parts: Compose ordering has to run Postgres, then Flyway, then the tester.
  • Bash Harness: This is not a production importer or a full test framework.
  • Fixture-Level Feedback: Errors surface as raw Postgres enum, FK, unique, CHECK, cast, or trigger failures.
Mitigation Strategies:

Keep the harness small and assertion-driven. Promote it to a real importer/test runner only when real CSV error reporting becomes part of the product requirement.

1. Measurements: The JSONB approach

My first design decisions in this system revolves around how to store Measurements. The requirements state that measurements can take several forms and that new kinds of measurements are added occasionally as the lab adopts new techniques.

To solve this flexibility requirement, I chose to implement a JSONB column (value) within a unified measurements table to store the polymorphic payload, rather than relying on strict relational alternatives.

Below is the rationale behind this choice and the trade-offs considered.


  • Classic Relation Approach : A single measurements table containing optional columns for every possible type: numeric_value, unit, categorical_value, and text_value. Why it was rejected: While this maintains strict database-level data typing, it fails to support the requirement that new measurement types will be added over time. Every time the lab adopts a new technique, it would require executing a migration (ALTER TABLE) to add new columns. This introduces unnecessary operational friction, potential table locks in production, and leads to a heavily sparse table filled with NULL values.

  • Entity-Attribute-Value (EAV) Approach : Splitting the data into measurement_attributes (metadata like name, type, unit) and measurement_values (rows mapping a measurement instance to an attribute and its value). Why it was rejected: The EAV pattern is notorious for complicating queries. Fetching a complete experiment profile with multiple measurements requires complex, multi-layered JOIN operations or pivoting data in application memory. It also degrades indexing performance and obfuscates the data structure for future developers.

  • [Chosen] JSONB Polymorphism Approach : By utilizing PostgreSQL's native JSONB data type, the table structure remains lean and completely decoupled from the specific domain of the scientific technique being introduced.


Advantages:

  • Future-Proof Extensibility: The lab can introduce a complex, nested measurement type tomorrow (e.g., a genomic sequence slice or an array of multidimensional sensor readings) without requiring a single database migration.
  • Performance: JSONB stores data in a decomposed binary format. This allows us to inject GIN (Generalized Inverted Index) indexes directly on JSON keys, ensuring that queries targeting specific internal properties remain fast.
  • Operational Simplicity: Avoids the complexity of managing multiple joined tables or running risky structural schema updates on growing production databases.

Accepted Trade-offs & Risks :

  • No type safety from the column definition: Unlike a typed column, a JSONB column does not, on its own, guarantee that a numeric measurement holds a valid number. That guarantee has to be added explicitly.

  • Known JSONB shapes are still database-enforced: A JSONB column does not infer the expected shape from measurement_type by itself, but Postgres can enforce that contract with a CHECK using jsonb_typeof. Under A0, the database is the last guard: a numeric measurement inserted as a JSON string is rejected before it becomes durable data.

Mitigation Strategies:

Enforce the payload contract with a Postgres CHECK on value that branches on measurement_type: assert the shape and jsonb_typeof of the known core types (a numeric reading must be a JSON number, a unit must be a JSON string), while staying permissive for not-yet-seen types so adopting a new technique still needs no migration. The CSV importer then relies on the database to reject malformed rows. Promoting a brand-new shape to a validated core type is an additive CHECK migration — cheap, no table rewrite. See A1 for the full reasoning, and docs/csv-ingestion.md for the concrete CHECK and transform SQL.

2. Sample usage tracking: FK-only vs. experiment_samples join table

Based on B4, the decision is whether to track which samples an experiment uses as a first-class relationship, or to derive it from the measurements that reference those samples.


  • experiment_samples join table approach: An explicit (experiment_id, sample_id) association is recorded before any measurement is taken. Enforcing that measurement.sample_id belongs to the experiment becomes a trigger check against this table. Why it was deferred: It introduces a new first-class entity the problem statement never mentions, requires maintaining a pre-registration step in the CSV ingestion workflow, and enforcement still needs a cross-table trigger (not a simple CHECK) — adding write cost and complexity before a real workflow has proven the need.

  • [Chosen] FK-only approach: measurement.sample_id FK → samples.id (nullable). The measurement row is the record that an experiment used a sample. No join table, no cross-table trigger.


Accepted Trade-offs & Risks:

  • Zero-measurement samples are invisible: A sample placed into an experiment but destroyed or consumed before any measurement is collected will not appear in the experiment's audit trail. This is a real GxP traceability gap.
  • Semantic overloading: "Sample used" is treated as synonymous with "sample that generated a measurement." Control samples and reagents that are inputs to the experiment but do not produce individual measurement rows are also invisible.
  • No scope guard: The FK accepts any valid samples.id in the system. A CSV typo that references a sample from an unrelated project passes silently.
  • Query cost at scale: Deriving sample usage via SELECT DISTINCT sample_id FROM measurements WHERE experiment_id = X becomes a bottleneck on high-frequency telemetry tables.
Mitigation Strategies:

If D1 confirms that samples can be consumed without generating measurements, introduce an experiment_samples associative table at that point — it is an additive migration, not a redesign. Until then, the FK-only model keeps CSV ingestion simple and avoids a join table the current scope does not justify.

3. Experiment lineage cycle detection

Based on C2, the decision is whether to enforce acyclicity in the experiment predecessor graph at the database level, and if so, which mechanism to use.


  • Recursive CTE trigger: On every INSERT or UPDATE to predecessor_experiment_id, walk the entire ancestor chain with a recursive CTE and reject if a cycle is detected. Full cycle detection, O(depth) per write, imposing read locks and full-tree traversal. Why it was rejected: The normal write path in a CSV ingestion workflow is append-only — retroactive lineage rewiring via UPDATE is an operational anomaly, not a routine operation. Paying the recursive traversal cost on every insert to guard against a mutation the workflow never performs is not justified.

  • Generation counter: A generation integer column enforces generation = predecessor.generation + 1 via a BEFORE INSERT trigger with a single predecessor lookup (O(1), not recursive), capped by a CHECK constraint. Any row that would create a cycle violates the monotonicity invariant. Why it was deferred: Under concurrent inserts targeting the same predecessor, the single-row lookup creates row-level lock contention — a real cost for a guard against UPDATE-based lineage rewiring that the append-only ingestion workflow does not perform.

  • [Chosen] CHECK-only deferral: The CHECK (predecessor_experiment_id <> id) from B2 blocks the one-hop self-reference case. Multi-hop cycle detection is deferred.


Accepted Trade-offs & Risks:

  • Silent multi-hop cycles under UPDATE: An UPDATE that rewires predecessor_experiment_id retroactively (e.g., changing B's predecessor from A to C when A→C already exists) creates a multi-hop cycle the database will not catch. If the generation counter escape hatch is also active, the lineage becomes inconsistent with corrupted generation values.
  • Generation counter's own concurrency cost: If the generation counter is activated, its BEFORE INSERT trigger introduces row-level lock contention on the predecessor row under concurrent inserts — a cost that is O(1) but not zero.
Mitigation Strategies:

The append-only nature of CSV ingestion over historical lab logs is the load-bearing assumption. If the workflow evolves to allow retroactive lineage edits, activate the generation counter (BEFORE INSERT trigger + CHECK (generation < MAX_DEPTH)) as the O(1) incremental guard. Reserve the recursive CTE trigger only if generation depth enforcement proves insufficient for the cycle patterns that emerge in practice.

Contributors

AbraaoAlves

Issues