The system is a Laboratory Experiment Tracking System (LETS) to manage laboratory governance and operations. Created to solve this challange (PROBLEM.md)
See ARCHITECTURE.md for the runtime map, domain graph, and database invariant map used to guide implementation.
docker compose upFlyway applies the Postgres migrations, then run_pipeline.sh loads seed data, imports valid
CSV fixtures, and proves the invalid fixtures are rejected by database constraints/triggers.
Postgres stays up afterward so you can connect and run queries. Stop and clean up with
docker compose down -v.
docker compose run --rm pipeline_tester; docker compose down -vSame migrations and pipeline, but it runs once and tears everything down on exit. Test results print directly to the terminal without container-log prefixes — handy for a quick pass/fail check.
If local port 5432 is busy, keep the same internal database wiring and publish a different host
port by prefixing either command:
POSTGRES_PORT=55432 docker compose upConnect with:
docker compose exec postgres psql -U postgres -d labUseful demo queries:
-- Project collaborators (M:N)
SELECT p.title, array_agg(r.name ORDER BY r.name) AS collaborators
FROM projects p
JOIN project_researchers pr ON pr.project_id = p.id
JOIN researchers r ON r.id = pr.researcher_id
GROUP BY p.title
ORDER BY p.title;
-- Experiment ancestor chain
WITH RECURSIVE chain AS (
SELECT id, title, predecessor_experiment_id, 0 AS depth
FROM experiments
WHERE title = 'Gamma Planning Run'
UNION ALL
SELECT e.id, e.title, e.predecessor_experiment_id, c.depth + 1
FROM experiments e
JOIN chain c ON c.predecessor_experiment_id = e.id
)
SELECT depth, title FROM chain ORDER BY depth;
-- Samples used by each experiment, derived from measurements
SELECT e.title, array_agg(DISTINCT s.sample_code) FILTER (WHERE s.id IS NOT NULL) AS samples
FROM experiments e
LEFT JOIN measurements m ON m.experiment_id = e.id
LEFT JOIN samples s ON s.id = m.sample_id
GROUP BY e.title
ORDER BY e.title;The schema has seven tables: researchers, projects, project_researchers, samples,
experiments, measurement_types, and measurements. IDs are generated by Postgres; CSV
fixtures resolve relationships through stable keys documented in
docs/csv-ingestion.md.
Full detail, including impact and database enforcement for each assumption, lives in QUESTIONS_ASSUMPTIONS.md. The short version:
- Runtime boundary: this is a CSV ingestion workflow for lab exports, not a web API or application service. Durable invariants live in Postgres constraints, foreign keys, checks, and triggers.
- Measurements: one
measurementstable stores typedJSONBpayloads classified bymeasurement_type. Known core shapes are database-checked; new types can be registered as data before their shape is hardened. - Project lifecycle:
completedandcancelledprojects freeze new descendant experiments and measurements;planningandactiveremain writable. - Experiment lineage: follow-ups are traceability links only. They do not copy samples, hypotheses, or status, and cross-project follow-ups are allowed.
- Sample usage: a measurement row is the record that an experiment used a sample. There is no
separate
experiment_samplestable until the lab confirms sample usage must be pre-registered or tracked without a measurement.
Questions I would clarify with the lab before building this further:
- Are samples inventory-managed physical resources with quantities, depletion, disposal, or chain-of-custody history?
- Must an experiment pre-register samples before any measurements are recorded?
- Are measurement corrections routine enough to require append-only versioning instead of strict update/delete blocking?
- Do researcher roles drive approvals or access control, or are they only descriptive metadata?
- What do real machine CSV exports look like, and do they require durable staging plus row-level error reports?
Based on A0, I assume the first runtime context is a CSV ingestion workflow for lab exports, not a full web API or long-running application service.
This keeps the solution focused on what the challenge asks for: a Postgres data model, Docker startup, and seed data that prove the model works. The domain rules should live as close to the data as possible through tables, foreign keys, unique constraints, enums, and check constraints. Any CSV-specific parsing or row-shape validation can live in the ingestion step when that workflow is implemented.
Below is the rationale behind this choice and the trade-offs considered.
-
Web API Microservice Approach : Build an HTTP API around the model, with request handlers, DTOs, service classes, authentication concerns, and validation in the application layer. Why it was rejected: The prompt asks for a database model that can be started with Docker and inspected live. Adding an API now would create more code to explain without proving the core model better.
-
Full Domain Application Layer Approach : Build a DDD-style application layer with use cases, repositories, and schema validators before the database is exercised. Why it was rejected: The current ambiguity is mostly about domain invariants, not transport or orchestration. Capturing those invariants in the schema and documentation gives faster value and keeps the interview surface smaller.
-
[Chosen] CSV Ingestion Workflow Approach : Treat the lab's current spreadsheet/export world as the input boundary. CSV rows are parsed and normalized before insert, while Postgres remains the source of truth and final validation guard for relationships and durable constraints. See docs/csv-ingestion.md for the flat-CSV → JSONB transform seam that bridges this decision with the JSONB measurement model in Decision 1.
- Simple Delivery Path: Matches the required Docker + Postgres + seed-data deliverable without adding an unrequested API surface.
- Pragmatic Migration from Spreadsheets: Fits the lab's current reality: they are already working from spreadsheet-like systems and need a clean target model.
- Centralized Data Integrity: Core rules are enforced by the database, so future import scripts, CLIs, or APIs inherit the same constraints.
- Batch-Oriented Feedback: CSV ingestion usually reports errors after a file is processed, not interactively while a researcher is typing.
- Importer Responsibility: CSV header mapping, required columns, type coercion, and row-level error messages still need to be handled outside the database.
- Limited Product Surface: This does not yet solve user workflows such as login, approval, dashboards, or real-time experiment editing.
Keep the database strict on durable invariants, document the expected CSV contract near the ingestion script when it exists, and add a small staging/error-report table only if bad production files become a real workflow problem.
Real laboratory machines often produce noisy, denormalized flat CSV files: one row can mix project, experiment, sample, researcher, and measurement columns, and bad rows may contain malformed values or inconsistent names. I am aware that a production-grade lab importer would need a broader ETL/staging workflow for that raw shape.
For this challenge, I intentionally use a simpler CSV contract that is close to the target database model. The goal is to prove the relational design and database invariants, not to build a general-purpose CSV cleaning product.
Below is the rationale behind this choice and the trade-offs considered.
-
Raw Machine Export ETL Approach : Accept the exact denormalized files emitted by lab equipment and split them into the target tables during ingestion, including header mapping, row quarantine, deduplication, idempotency, and malformed-line recovery. Why it was deferred: Those are real production concerns, but they require real source files and operational error-handling requirements. Building that now would bury the database model under importer behavior that the challenge does not ask for.
-
SQL Seed-Only Approach : Skip CSV ingestion entirely and prove the schema only with hand-written
INSERTstatements. Why it was rejected: It would make the Docker demo smaller, but it would not exercise the CSV boundary established in A0. -
[Chosen] Simplified Contract CSV Approach : Use explicit, predictable CSV fixtures that already follow the domain vocabulary and can be loaded through temporary staging tables plus SQL transforms. Postgres remains the source of truth and final validation guard through foreign keys, unique constraints, checks, and triggers.
- Honest Challenge Scope: Keeps attention on the schema, invariants, Docker startup, and seed/demo data requested by the problem.
- Executable Without Importer Sprawl: A reviewer can run the pipeline and inspect the database without reading a custom ETL framework first.
- Still Compatible With A0: CSV remains the input boundary, while durable rules stay enforced by the database rather than an absent application layer.
- Not a Production Raw-Export Importer: The solution does not claim to handle arbitrary machine exports, malformed delimiters, duplicate names, or row-level quarantine.
- Upstream Normalization Assumed: The CSV contract assumes stable identifiers or already-separated entity rows before data reaches the database ingestion step.
- Simpler Error Feedback: Failed rows surface as Postgres cast, FK, unique, check, or trigger errors, not as polished per-row business messages.
Treat the simplified CSV shape as a documented contract for this challenge. If real raw lab exports become part of the requirement, add a small durable staging/error-report layer or a separate pre-processing step that converts raw files into this contract before COPY. Do not relax the database constraints; they remain the final guard under A0.
The implementation splits schema history from executable ingestion tests: Flyway owns migrations,
while run_pipeline.sh owns seed loading, CSV imports, and acceptance assertions.
-
One application tool does migration plus ingestion : Build a CLI or service that migrates, imports, and validates. Why it was rejected: It adds an application layer the challenge does not need and cuts against A0.
-
Flyway does schema and seed/test data : Put demo data and CSV assertions into migrations. Why it was rejected: It couples disposable fixtures to schema history and gives no clean place for invalid-row assertions.
-
Plain Postgres init scripts : Rely on
/docker-entrypoint-initdb.d. Why it was rejected: It only runs on a fresh volume and is not explicit versioned migration history. -
[Chosen] Flyway +
run_pipeline.sh: Flyway applies versioned DDL. The pipeline imports CSV fixtures through TEMP tables and lets Postgres accept or reject each case.
- Separation of Concerns: Schema migrations are durable history; CSV fixtures are executable proof data.
- TDD Shape: The pipeline proves every documented invariant at the database boundary.
- Single Command:
docker compose upstill builds the schema and leaves a running database.
- Two Moving Parts: Compose ordering has to run Postgres, then Flyway, then the tester.
- Bash Harness: This is not a production importer or a full test framework.
- Fixture-Level Feedback: Errors surface as raw Postgres enum, FK, unique, CHECK, cast, or trigger failures.
Keep the harness small and assertion-driven. Promote it to a real importer/test runner only when real CSV error reporting becomes part of the product requirement.
My first design decisions in this system revolves around how to store Measurements. The requirements state that measurements can take several forms and that new kinds of measurements are added occasionally as the lab adopts new techniques.
To solve this flexibility requirement, I chose to implement a JSONB column (value) within a unified measurements table to store the polymorphic payload, rather than relying on strict relational alternatives.
Below is the rationale behind this choice and the trade-offs considered.
-
Classic Relation Approach : A single
measurementstable containing optional columns for every possible type:numeric_value,unit,categorical_value, andtext_value. Why it was rejected: While this maintains strict database-level data typing, it fails to support the requirement that new measurement types will be added over time. Every time the lab adopts a new technique, it would require executing a migration (ALTER TABLE) to add new columns. This introduces unnecessary operational friction, potential table locks in production, and leads to a heavily sparse table filled withNULLvalues. -
Entity-Attribute-Value (EAV) Approach : Splitting the data into
measurement_attributes(metadata like name, type, unit) andmeasurement_values(rows mapping a measurement instance to an attribute and its value). Why it was rejected: The EAV pattern is notorious for complicating queries. Fetching a complete experiment profile with multiple measurements requires complex, multi-layeredJOINoperations or pivoting data in application memory. It also degrades indexing performance and obfuscates the data structure for future developers. -
[Chosen] JSONB Polymorphism Approach : By utilizing PostgreSQL's native
JSONBdata type, the table structure remains lean and completely decoupled from the specific domain of the scientific technique being introduced.
- Future-Proof Extensibility: The lab can introduce a complex, nested measurement type tomorrow (e.g., a genomic sequence slice or an array of multidimensional sensor readings) without requiring a single database migration.
- Performance: JSONB stores data in a decomposed binary format. This allows us to inject GIN (Generalized Inverted Index) indexes directly on JSON keys, ensuring that queries targeting specific internal properties remain fast.
- Operational Simplicity: Avoids the complexity of managing multiple joined tables or running risky structural schema updates on growing production databases.
-
No type safety from the column definition: Unlike a typed column, a
JSONBcolumn does not, on its own, guarantee that a numeric measurement holds a valid number. That guarantee has to be added explicitly. -
Known JSONB shapes are still database-enforced: A
JSONBcolumn does not infer the expected shape frommeasurement_typeby itself, but Postgres can enforce that contract with aCHECKusingjsonb_typeof. Under A0, the database is the last guard: a numeric measurement inserted as a JSON string is rejected before it becomes durable data.
Enforce the payload contract with a Postgres CHECK on value that branches on measurement_type: assert the shape and jsonb_typeof of the known core types (a numeric reading must be a JSON number, a unit must be a JSON string), while staying permissive for not-yet-seen types so adopting a new technique still needs no migration. The CSV importer then relies on the database to reject malformed rows. Promoting a brand-new shape to a validated core type is an additive CHECK migration — cheap, no table rewrite. See A1 for the full reasoning, and docs/csv-ingestion.md for the concrete CHECK and transform SQL.
Based on B4, the decision is whether to track which samples an experiment uses as a first-class relationship, or to derive it from the measurements that reference those samples.
-
experiment_samplesjoin table approach: An explicit(experiment_id, sample_id)association is recorded before any measurement is taken. Enforcing thatmeasurement.sample_idbelongs to the experiment becomes a trigger check against this table. Why it was deferred: It introduces a new first-class entity the problem statement never mentions, requires maintaining a pre-registration step in the CSV ingestion workflow, and enforcement still needs a cross-table trigger (not a simpleCHECK) — adding write cost and complexity before a real workflow has proven the need. -
[Chosen] FK-only approach:
measurement.sample_id FK → samples.id(nullable). The measurement row is the record that an experiment used a sample. No join table, no cross-table trigger.
- Zero-measurement samples are invisible: A sample placed into an experiment but destroyed or consumed before any measurement is collected will not appear in the experiment's audit trail. This is a real GxP traceability gap.
- Semantic overloading: "Sample used" is treated as synonymous with "sample that generated a measurement." Control samples and reagents that are inputs to the experiment but do not produce individual measurement rows are also invisible.
- No scope guard: The FK accepts any valid
samples.idin the system. A CSV typo that references a sample from an unrelated project passes silently. - Query cost at scale: Deriving sample usage via
SELECT DISTINCT sample_id FROM measurements WHERE experiment_id = Xbecomes a bottleneck on high-frequency telemetry tables.
If D1 confirms that samples can be consumed without generating measurements, introduce an experiment_samples associative table at that point — it is an additive migration, not a redesign. Until then, the FK-only model keeps CSV ingestion simple and avoids a join table the current scope does not justify.
Based on C2, the decision is whether to enforce acyclicity in the experiment predecessor graph at the database level, and if so, which mechanism to use.
-
Recursive CTE trigger: On every
INSERTorUPDATEtopredecessor_experiment_id, walk the entire ancestor chain with a recursive CTE and reject if a cycle is detected. Full cycle detection, O(depth) per write, imposing read locks and full-tree traversal. Why it was rejected: The normal write path in a CSV ingestion workflow is append-only — retroactive lineage rewiring viaUPDATEis an operational anomaly, not a routine operation. Paying the recursive traversal cost on every insert to guard against a mutation the workflow never performs is not justified. -
Generation counter: A
generationinteger column enforcesgeneration = predecessor.generation + 1via aBEFORE INSERTtrigger with a single predecessor lookup (O(1), not recursive), capped by aCHECKconstraint. Any row that would create a cycle violates the monotonicity invariant. Why it was deferred: Under concurrent inserts targeting the same predecessor, the single-row lookup creates row-level lock contention — a real cost for a guard against UPDATE-based lineage rewiring that the append-only ingestion workflow does not perform. -
[Chosen]
CHECK-only deferral: TheCHECK (predecessor_experiment_id <> id)from B2 blocks the one-hop self-reference case. Multi-hop cycle detection is deferred.
- Silent multi-hop cycles under
UPDATE: AnUPDATEthat rewirespredecessor_experiment_idretroactively (e.g., changing B's predecessor from A to C when A→C already exists) creates a multi-hop cycle the database will not catch. If the generation counter escape hatch is also active, the lineage becomes inconsistent with corrupted generation values. - Generation counter's own concurrency cost: If the generation counter is activated, its
BEFORE INSERTtrigger introduces row-level lock contention on the predecessor row under concurrent inserts — a cost that is O(1) but not zero.
The append-only nature of CSV ingestion over historical lab logs is the load-bearing assumption. If the workflow evolves to allow retroactive lineage edits, activate the generation counter (BEFORE INSERT trigger + CHECK (generation < MAX_DEPTH)) as the O(1) incremental guard. Reserve the recursive CTE trigger only if generation depth enforcement proves insufficient for the cycle patterns that emerge in practice.