Barbhuiya12/PINN

★ 0Forks 0PythonGitHub ↗Compare

README

PINN — Physics-Informed Neural Networks of the Saint-Venant Equations

A reproduction study and laboratory course based on:

Feng, D., Tan, Z., & He, Q. Z. (2023). Physics-informed neural networks of the Saint-Venant equations for downscaling a large-scale river model. Water Resources Research, 59, e2022WR033168. doi:10.1029/2022WR033168

The work is organised in four stages — data exploration, loss function analysis, training, and validation — carried out across six independent river reaches.

Everything needed to run it is contained in this repository, including the reference simulations and the authors' trained weights.

Maintained by Siddik Barbhuiya


What this is

A large-scale river model gives you water depth at coarse resolution. A few gauges give you depth at a few points along a reach. Can you recover the full high-resolution field in between — not by interpolating, but by requiring that the answer satisfies the Saint-Venant equations?

That is what a physics-informed neural network does here. A small network maps position and time to depth and velocity:

input (x, t)  ──►  tanh network  ──►  output (u, h)

It is trained on almost nothing — one snapshot profile, the two boundaries, and two to five interior gauges, typically 3–9% of the reference field. The rest it must infer from the governing equations, which are imposed as a penalty on the loss using automatic differentiation.

There is no mesh and no time-stepping. The trained model is a continuous function you can query at any (x, t), on or off the grid.


Quick start

git clone https://github.com/Barbhuiya12/PINN.git
cd PINN

bash SETUP.sh                 # builds an isolated conda env, ~5 min
conda activate pinn_sve

python phase4_validation/validate.py --case 3

Expected output:

CASE 3   variant: case3   weights: paper
PINN               eps_h = 6.6354e-03    RMSE = 0.054    NSE = 0.9997
                   paper: eps_h = 6.635e-03, RMSE = 0.054  (deviation 0.0%)

deviation 0.0% means your environment is correct and the published study reproduces on your machine. Do this before anything else.

The code is TensorFlow 1.15 / Python 3.7, deliberately. It depends on tf.contrib.opt.ScipyOptimizerInterface, which was deleted in TensorFlow 2 and has no drop-in replacement. Do not try to port it — see SETUP.md.


Course structure

The work is divided into four stages. Every script takes --case N, and each stage folder contains its own README with full instructions and a statement of what is to be submitted.

Phase Folder Question Runtime
1. Data exploration phase1_data_exploration/ What is the model shown, and what must it invent? seconds
2. Loss function analysis phase2_loss_analysis/ How do the Saint-Venant equations become a loss? seconds
3. Training phase3_training/ Can you reproduce the published model yourself? minutes–hours
4. Validation phase4_validation/ Is it any good, and compared to what? seconds

The authors' trained weights are included, so Phases 1, 2 and 4 all run on day one. You do not have to wait for training.

Phase 2 is the core. It contains the only code you write from scratch — the Saint-Venant residuals — with an automatic checker that reports a difference of exactly zero when you get them right.


The six cases

One reach each. They are fully independent, so no two people can hand in the same numbers. Full descriptions in docs/02_case_assignments.md.

Case Reach What it demonstrates Published ε_h
1 analytical, flat bed moving wet/dry boundary 3.075 × 10⁻³
2 analytical, sloping bed RK4 reference, tidal forcing 3.685 × 10⁻²
3 HEC-RAS sub-reach the cleanest success 6.635 × 10⁻³
4 same reach, 10× longer where a standard PINN fails, and Fourier features fix it 3.093 × 10⁻¹ → 7.176 × 10⁻²
5 full reach PINN vs linear interpolation 9.876 × 10⁻²
6 full reach, 2/3/4/5 gauges how many gauges do you actually need? 1.396 × 10⁻¹ → 8.142 × 10⁻²

Repository layout

PINN/
├── README.md                  this file
├── SETUP.md / SETUP.sh        environment, with every version pinned and justified
├── environment.yml
│
├── docs/
│   ├── 01_theory.md           the equations, and how they become a loss
│   ├── 02_case_assignments.md the six reaches, in detail
│   ├── 03_code_map.md         what every file does
│   ├── 04_deliverables.md     what to hand in, and how it is marked
│   └── 05_known_issues.md     quirks and bugs in the released code — READ THIS
│
├── engine/                    the authors' PINN. Do not edit.
│   ├── SVE_module_dynamic.py           depth only, Cases 1–2
│   ├── SVE_module_dynamic_h.py         depth + velocity, Cases 3–5
│   └── SVE_module_dynamic_h_mff_ts.py  + Fourier features, Cases 4–6
│
├── common/                    shared course code
│   ├── cases.py               the case registry — start reading here
│   ├── dataprep.py            reference data pipeline
│   ├── modelbuild.py          adapter over the three engine signatures
│   └── metrics.py             ε_h, RMSE, NSE, interpolation baseline
│
├── data/
│   ├── HEC-RAS/case{3,4,5,6}/ reference simulations (HDF5, 11 MB)
│   └── saved_model/           the weights that produced the published figures
│
├── phase1_data_exploration/
├── phase2_loss_analysis/
├── phase3_training/
├── phase4_validation/
│
└── results/                   everything you generate lands here

Reproduction status

All results below were produced on CPU with the code in this repository.

From the authors' weights — every PINN value in Tables 3 and 4 reproduces exactly, to the last published digit, across all six cases.

From scratch — the released package contains no training script, and its L-BFGS-B routine crashes on a freshly initialised model. Both are worked around in phase3_training/train.py. Results:

Case Published ε_h Retrained ε_h Deviation
3 6.635 × 10⁻³ 6.495 × 10⁻³ 2.1% (slightly better)
5 9.876 × 10⁻² 1.133 × 10⁻¹ 14.7%
2 3.685 × 10⁻² 4.714 × 10⁻² 27.9%
4 standard 3.093 × 10⁻¹ 4.827 × 10⁻¹ 56.1%
4 Fourier 7.176 × 10⁻² 3.314 × 10⁻¹ factor of 4.6

Case 3 took 170 seconds. Case 4's Fourier variant does not reproduce — most likely because the training length is nowhere recorded and the count recoverable from the shipped files is a lower bound.

Reproducing the published number is not the goal of this project. Running the experiment and giving an honest, quantified account of what you got is.


What the code does that the paper does not say

Eleven items are documented in docs/05_known_issues.md. Read it early — it will save you an afternoon debugging something already known. The most consequential:

  • The HEC-RAS data is in feet; the momentum residual is written in SI. The paper reports the domain as 305 m and 914 m; the code reads 1000 and 3000. g = 9.81 m/s² should be 32.2 ft/s², and Manning's US-unit factor of 1.49 is absent. ε_h is a ratio and so is unit-free, which is why nothing catches it.
  • The "observation noise" is set to SNR = 500 dB, i.e. σ ≈ 10⁻²⁵. The observations are exact. Robustness to realistic observation error is never tested anywhere in the study.
  • The continuity residual drops a term from the conservative form, so mass is not strictly conserved.
  • Manning's n is hard-coded inside the momentum function; the value passed in is ignored.
  • The interpolation baseline in Table 4 uses a different norm from the PINN column, because np.linalg.norm silently changes meaning between 1-D and 2-D input.

None of these overturn the paper's conclusions. All of them change how you should read it. Finding a twelfth is the most valuable thing you can do in this project.


Ground rules

  • Do not edit engine/. If you think it is wrong, that is a finding for your report, not a patch.
  • Write only into results/. Never into data/.
  • Every claim about your case needs a number you generated. "The error is small" is worth nothing.
  • Report what happened, not what should have happened.

Citation

Cite the original paper. This repository is teaching material built around someone else's research; the science is theirs.

Feng, D., Tan, Z., & He, Q. Z. (2023). Physics-informed neural networks of the Saint-Venant equations for downscaling a large-scale river model. Water Resources Research, 59, e2022WR033168. https://doi.org/10.1029/2022WR033168

@article{Feng2023PINNSVE,
  author  = {Feng, Dongyu and Tan, Zeli and He, QiZhi},
  title   = {Physics-Informed Neural Networks of the {Saint-Venant} Equations
             for Downscaling a Large-Scale River Model},
  journal = {Water Resources Research},
  volume  = {59},
  number  = {2},
  pages   = {e2022WR033168},
  year    = {2023},
  doi     = {10.1029/2022WR033168}
}

The code and reference data:

Feng, D. (2022). PINN SVE data assimilation [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.7118168

@dataset{Feng2022PINNSVEcode,
  author    = {Feng, Dongyu},
  title     = {{PINN} {SVE} data assimilation},
  publisher = {Zenodo},
  year      = {2022},
  doi       = {10.5281/zenodo.7118168}
}

Credits and licence

Part Author
engine/, data/ Feng, Tan & He — from the Zenodo release above
common/, phase1..4/, docs/, this README Siddik Barbhuiya

engine/ differs from the authors' release by one line per file — see SETUP.md.

The authors' Data Availability Statement reads: "The data and code to reproduce the results and figures are publicly available at the Zenodo repository." They are redistributed here on that basis, with attribution. The paper itself is © 2023 American Geophysical Union.

The course material added by this repository is released under CC BY 4.0. Full terms for both parts are set out in LICENSE; citation metadata is in CITATION.cff.

Correspondence on the original study: Z. Tan, [email protected]

Contributors

Barbhuiya12

Issues