intial version

This commit is contained in:
Tom Kasper
2026-09-17 22:16:42 +01:00
parent e23ed24fa9
commit 08467102c8
19 changed files with 311 additions and 0 deletions
+24
View File
@@ -1,2 +1,26 @@
# eDNA_Stream_E01
## Scope
- Establish whether squiggle-based taxonomy is feasible on the small computational budget that is available
## Dataset:
- NO-MISS bacterial isolates, v10.4.1 chemistry on PromethION flow cells (https://epi2me.nanoporetech.com/nomiss_96bc_p2i_sup_2026/)
- Downsampling at the pod5 level intended for compute/time reasons -> the full dataset is 293 pod5 / 1.6 TB of raw data
## Dependencies:
- dorado >= v 2.0.0
- aws CLI
- conda/mamba (mamba recommended)
- python env with:
- python 3.14
- snakemake
- python-dotenv
## Optional:
- .env file with NCBI API key -> slightly speeds up reference fasta download
## Running:
```bash
cd workflow
snakemake --cores <cores> --resources gpu=<n_gpus>
```