eDNA_Stream_E01

Scope

  • Establish whether squiggle-based taxonomy is feasible on the small computational budget that is available

Dataset:

Dependencies:

  • dorado >= v 2.0.0
  • aws CLI
  • conda/mamba (mamba recommended)
  • python env with:
    • python 3.14
    • snakemake
    • python-dotenv
  • custom-models package: live in a separate repo (git@git.tk-ai.eu:Tom/eDNA_Stream_E01_custom_models.git) included here as the custom_models git submodule; the workflow pip-installs it from that checkout (see workflow/envs/torch_*.yaml)
    • clone with: git clone --recurse-submodules ... (or git submodule update --init on an existing clone)
    • update to a newer version: git submodule update --remote custom_models, then commit the new pinned commit
  • .env file with NCBI API key -> slightly speeds up reference fasta download

Running:

cd workflow
snakemake --cores <cores> --resources gpu=<n_gpus> --use-conda

Notes:

  • Reads aligned to the FP traps by minimap are excluded for training -> esp. the E. coli reads are highly represented and should not be there, but are stable even when only using stringent alignment
S
Description
No description provided
Readme
9.7 MiB
Languages
Python 100%