30 lines
1.4 KiB
Markdown
30 lines
1.4 KiB
Markdown
# eDNA_Stream_E01
|
|
|
|
## Scope
|
|
- Establish whether squiggle-based taxonomy is feasible on the small computational budget that is available
|
|
|
|
## Dataset:
|
|
- NO-MISS bacterial isolates, v10.4.1 chemistry on PromethION flow cells (https://epi2me.nanoporetech.com/nomiss_96bc_p2i_sup_2026/)
|
|
- Downsampling at the pod5 level intended for compute/time reasons -> the full dataset is 293 pod5 / 1.6 TB of raw data
|
|
|
|
## Dependencies:
|
|
- dorado >= v 2.0.0
|
|
- aws CLI
|
|
- conda/mamba (mamba recommended)
|
|
- python env with:
|
|
- python 3.14
|
|
- snakemake
|
|
- python-dotenv
|
|
- custom-models package: live in a separate repo (`git@git.tk-ai.eu:Tom/eDNA_Stream_E01_custom_models.git`) included here as the `custom_models` git submodule; the workflow pip-installs it from that checkout (see `workflow/envs/torch_*.yaml`)
|
|
- clone with: `git clone --recurse-submodules ...` (or `git submodule update --init` on an existing clone)
|
|
- update to a newer version: `git submodule update --remote custom_models`, then commit the new pinned commit
|
|
- .env file with NCBI API key -> slightly speeds up reference fasta download
|
|
|
|
## Running:
|
|
```bash
|
|
cd workflow
|
|
snakemake --cores <cores> --resources gpu=<n_gpus> --use-conda
|
|
```
|
|
|
|
## Notes:
|
|
- Reads aligned to the FP traps by minimap are excluded for training -> esp. the E. coli reads are highly represented and should not be there, but are stable even when only using stringent alignment |