Metagenomics

September 11, 2023 ยท View on GitHub

As it currently stands, we do not recommend plassembler for metagenomic sequences. This is because of their high diversity, leading to difficulties in recovering chromosome-length contigs for bacteria. Additionally, Unicycler (a core dependency of Plassembler) is not recommended for metagenomes.

However, we anticipate that as sequencing becomes more accurate and cheaper, it will be increasingly possible to assemble plasmids using a plassembler like approach from metagenomes - it's a work in progress.

So as a test, we tried assembling the ZYMO HMW DNA Standard dataset from this paper, under ENA accession PRJEB48692. This mock community contains 7 bacteria and 1 fungus isolate. Notably, this dataset had extremely had deep (all bacterial chromosomes >100x coverage) and long (N50 > 20kbp) reads, so is unlikely to reflect your real-world metagenomic data as of 2023.

Get Data

# installation
mamba create -n fastq-dl fastq-dl
conda activate fastq-dl

# downloads all the read sets
fastq-dl PRJEB48692	

conda deactivate

Run Plassembler

We decided to use -m 10000, because we figures that smalll plasmids would be missed by Flye anyway, and wanted complete chromosome assemblies, and a -c 500000. We used 32 threads on 16 cores and allocated 80 GB of RAM.

plassembler run -d Plassembler_DB -l ERR7287988.fastq.gz -1 ERR7255689_1.fastq.gz -2 ERR7255689_2.fastq.gz \
-f -t 32 -q 10 -o zymo_R10.4_flye -m 10000 -c 500000

plassembler took around 8 hours (wall clock) to finish and excitingly we assembled all 7 bacterial chromosomes using Flye (unsurprising!) along with the 5 plasmids indicated in the ground truth (1 E. coli 100kbp, 1 S. enterica 49kbp and 3 small S. aureus plasmids (6, 2 and 2 kbp)) with genome fraction 100% from QUAST.

So in theory plassembler might work on metagenomes, but I would caution against using it, for now.

Contigs 34, 61, 87, 101 and 109 match what was found in the ground truth.

contiglengthmean_depth_shortcircularityPLSDB_hitACC_NUCCOREDescription_NUCCOREplasmid_copy_number_shortplasmid_copy_number_long
34110007336.5circularYesNZ_CP061531.1Escherichia coli strain WEM25 plasmid p1, complete sequence1.671.37
6149661357.16not_circularYesNZ_CP012345.2Salmonella enterica subsp. enterica serovar Choleraesuis str. ATCC 10708 plasmid pCFSAN000679_01, complete sequence1.783.17
83962839.11not_circularYesNZ_CP069918.1Klebsiella oxytoca strain FDAARGOS_1334 plasmid unnamed70.190
8763679554.23circularYesNZ_CP013628.1Staphylococcus aureus strain RIVM4293 plasmid pRIVM4293, complete sequence.47.5412.67
91535514.61not_circularYesNZ_CP068597.1Paenibacillus sonchi strain LMG 24727 plasmid unnamed2, complete sequence0.070
93501013.32not_circularYesNZ_CP068597.1Paenibacillus sonchi strain LMG 24727 plasmid unnamed2, complete sequence0.070
10129932561.92circularYesNZ_MH785226.1Staphylococcus aureus strain ph1 plasmid pRIVM1295-2, complete sequence12.751.65
10327891018.1not_circularYesCP048737.1Enterobacter sp. T2 plasmid unnamed1, complete sequence5.074.75
10626671045.66not_circularYesNZ_CP066061.1Actinomyces oris strain FDAARGOS_1051 plasmid unnamed5.24.79
108233716.17not_circularYesNZ_CP069918.1Klebsiella oxytoca strain FDAARGOS_1334 plasmid unnamed70.080
10922162188.82circularYesNZ_CP013624.1Staphylococcus aureus strain RIVM1076 plasmid pRIVM1076, complete sequence.10.891.08
14110491015.41not_circularYesNZ_CP066061.1Actinomyces oris strain FDAARGOS_1051 plasmid unnamed5.054.77
145830930.78not_circularYesNZ_CP066061.1Actinomyces oris strain FDAARGOS_1051 plasmid unnamed