A Nextflow DSL2 pipeline for GPU-accelerated basecalling of Oxford Nanopore POD5 files using Dorado. Designed as an upstream workflow that generates basecalled BAM and FASTQ files for use in downstream analyses.
This workflow:
- Downloads the specified Dorado basecalling model
- Builds a minimap2 index of the reference genome
- Basecalls each POD5 file in parallel (GPU-accelerated)
- Merges per-POD5 BAM files by sample
- Converts merged BAMs to compressed FASTQ
- Nextflow ≥ 22.10.0
- Singularity (container runtime; Docker not supported by default profiles)
All bioinformatics tools (Dorado, minimap2, samtools, pigz) are provided via the container oras://ghcr.io/shians/dorado-container:1.1.1-singularity.
| Process | CPUs | Memory | Time limit |
|---|---|---|---|
| Basecalling (GPU) | 16 | 64 GB | 24h |
| Genome indexing | 8 | 32 GB | 12h |
| BAM merging | 8 | 32 GB | 12h |
| BAM → FASTQ | 4 | 16 GB | 4h |
An NVIDIA GPU (A30 or compatible) is required for basecalling. High-performance storage is recommended for POD5 files (1–5 GB per file).
A tab-separated file mapping sample identifiers to directories containing POD5 files.
Format (pod5_sheet.tsv):
sample_id path
control_rep1 /data/nanopore/pod5/control_replicate1
control_rep2 /data/nanopore/pod5/control_replicate2
case_rep1 /data/nanopore/pod5/case_replicate1_flowcell1
case_rep1 /data/nanopore/pod5/case_replicate1_flowcell2sample_id: Identifier for each sample (used to name output files). Multiple rows with the samesample_idare merged into a single output file, which is useful when one sample was sequenced across multiple flow cells or run folders.path: Absolute path to a directory containing.pod5files (searched recursively)
A FASTA file used for alignment during basecalling (via minimap2). Gzipped FASTA (.fa.gz, .fasta.gz) is supported.
A Dorado model string. The workflow supports base models and modification models:
| Type | Example |
|---|---|
| Base model | [email protected] |
| With modifications | [email protected]_5mCG_5hmCG@v2 |
Supported modification codes: 5mCG_5hmCG, 5mC, 6mA, 5mC_5hmC, 4mC_5mC
Run nextflow run shians/dorado_workflow --help to list all valid model strings (requires Nextflow ≥ 25.10.0).
| Parameter | Required | Default | Description |
|---|---|---|---|
--pod5_sheet |
Yes | — | Path to POD5 sample sheet TSV |
--reference_genome |
Yes | — | Path to reference genome FASTA |
--basecall_model |
Yes | — | Dorado basecalling model string |
--dna |
Yes* | false |
DNA sequencing mode (uses lr:hq minimap2 preset) |
--cdna |
Yes* | false |
cDNA sequencing mode (uses splice:hq minimap2 preset) |
--output_dir |
No | "output" |
Root directory for output files |
--publish_merged_bams |
No | true |
Save merged BAM files to output |
--publish_fastq |
No | true |
Save converted FASTQ files to output |
*Exactly one of --dna or --cdna must be specified.
| Profile | Executor | Use Case |
|---|---|---|
singularity |
Local | Local machine with GPU and Singularity |
gpu_slurm |
SLURM | Generic SLURM cluster with GPU |
gpu_wehi |
SLURM | WEHI cluster (A30 GPU) |
Resource limits are defined by process labels. To override them, create a custom.config file and pass it with -c custom.config.
| Label | Used by | Default CPUs | Default Memory | Default Time |
|---|---|---|---|---|
large |
Basecalling | 16 | 64 GB | 24h |
medium |
Genome indexing, BAM merging | 8 | 32 GB | 12h |
small |
Model download, BAM → FASTQ | 4 | 16 GB | 4h |
Example custom.config:
process {
withLabel: large {
cpus = 24
memory = '128.GB'
time = '48h'
}
withLabel: medium {
cpus = 16
memory = '64.GB'
time = '24h'
}
}Pass the custom config alongside a profile:
nextflow run shians/dorado_workflow \
-profile gpu_wehi \
-c custom.config \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]" \
--dnanextflow run shians/dorado_workflow \
-profile gpu_wehi \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]" \
--dnanextflow run shians/dorado_workflow \
-profile gpu_wehi \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]" \
--cdnanextflow run shians/dorado_workflow \
-profile gpu_wehi \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]_5mCG_5hmCG@v2" \
--dnanextflow run shians/dorado_workflow \
-profile gpu_wehi \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]" \
--dna \
-resumenextflow run shians/dorado_workflow \
-profile gpu_wehi \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]" \
--dna \
--output_dir /scratch/project/basecallednextflow run shians/dorado_workflow \
-profile gpu_wehi \
--pod5_sheet pod5_sheet.tsv \
--reference_genome /path/to/genome.fa \
--basecall_model "[email protected]" \
--dna \
--publish_fastq falseoutput/ # Controlled by --output_dir
├── basecalled/ # Merged BAM files (one per sample)
│ ├── sample1.bam
│ └── sample2.bam
├── fastq/ # Compressed FASTQ files (one per sample)
│ ├── sample1.fastq.gz
│ └── sample2.fastq.gz
└── logs/
└── runtime_reports/
├── main_timeline.html # Interactive execution timeline
├── main_report.html # Execution summary report
└── main_trace.txt # Detailed task trace log
BAM files include alignment tags from minimap2, making them suitable for downstream tools that use alignment information (e.g., splice site analysis, modification calling).