10b0a29d2b60e9b451ec0e13c7f7a3c50305d6af
eDNA_Stream_E01
Scope
- Establish whether squiggle-based taxonomy is feasible on the small computational budget that is available
Dataset:
- NO-MISS bacterial isolates, v10.4.1 chemistry on PromethION flow cells (https://epi2me.nanoporetech.com/nomiss_96bc_p2i_sup_2026/)
- Downsampling at the pod5 level intended for compute/time reasons -> the full dataset is 293 pod5 / 1.6 TB of raw data
Dependencies:
- dorado >= v 2.0.0
- aws CLI
- conda/mamba (mamba recommended)
- python env with:
- python 3.14
- snakemake
- python-dotenv
- custom-models package: live in a separate repo (
git@git.tk-ai.eu:Tom/eDNA_Stream_E01_custom_models.git) included here as thecustom_modelsgit submodule; the workflow pip-installs it from that checkout (seeworkflow/envs/torch_*.yaml)- clone with:
git clone --recurse-submodules ...(orgit submodule update --initon an existing clone) - update to a newer version:
git submodule update --remote custom_models, then commit the new pinned commit
- clone with:
- .env file with NCBI API key -> slightly speeds up reference fasta download
Running:
cd workflow
snakemake --cores <cores> --resources gpu=<n_gpus> --use-conda
Notes:
- Reads aligned to the FP traps by minimap are excluded for training -> esp. the E. coli reads are highly represented and should not be there, but are stable even when only using stringent alignment
Languages
Python
100%