Section 3 of 4
3 Results
Che-Chun Chen, Hsin-Yun Lu, and Ying-Ning Ho · about 3 minutes
3.1 Cross-platform performance and resource efficiency
To validate the system’s stability and versatility, we processed a raw Pod5 dataset containing 1 million reads on both native Linux environments and Windows systems utilizing the Windows Subsystem for Linux 2 (WSL2). Detailed hardware specifications for the testing environments are provided in Table S1, available at supplementary material Bioinformatics Advances online.
Resource monitoring analysis identified basecalling and AI inference as the primary consumers of GPU memory. Notably, while Machine A exhibited significantly higher peak VRAM usage during basecalling compared to the AI inference step, VRAM consumption was comparable between these two stages across the other test configurations (Table S2, available at supplementary material Bioinformatics Advances online). These findings allow us to conclude that integrating the AI module does not impose hardware requirements exceeding those already mandated by standard Nanopore basecalling. Benchmarking confirms that the workflow efficiently processes standard MinION amplicon runs on consumer-grade laptops equipped with NVIDIA GPUs (a minimum of 6 GB VRAM is required, while 8 GB or more is recommended for optimal AI inference).
3.2 Benchmarking taxonomic accuracy using mock communities
To rigorously evaluate taxonomic classification accuracy, we analyzed three independent DNA samples extracted from the ZymoBIOMICS® Microbial Community Standard (D6300). We compared the performance of microbiONT v1.0.0 (utilizing the Emu-based custom workflow) against the widely used EPI2ME Desktop v5.3.0 platform (utilizing the Minimap2-based workflow) by calculating the coefficient of determination (_R_2) and the root mean square error (RMSE) between the observed and expected relative abundances of the bacterial genera.
Benchmarking results revealed a consistent performance advantage for microbiONT. Across three replicates, microbiONT demonstrated higher classification precision, achieving _R_2 values from 0.76 to 0.82 and RMSE values ranging from 2.41 to 2.81. In contrast, the EPI2ME workflow exhibited significantly higher deviation, with poor model fit with _R_2 values strictly below 0.5 (0.35–0.40) and RMSE values ranging between 4.42 and 4.63 (Fig. S4, available at supplementary material Bioinformatics Advances online). A paired Student’s t-test confirmed that microbiONT’s improvements in both predictive error (RMSE) and goodness of fit (_R_2) are statistically significant (_P _< 0.05). A detailed genus-level analysis identified the primary source of this discrepancy. The EPI2ME workflow showed substantial misclassification, particularly within the Escherichia genus and an inflated assignment of reads to the others category. This tendency to misclassify reads into non-target groups suggests a limitation in the standard EPI2ME classification strategy for this specific mock community. Conversely, microbiONT’s integrated workflow, which leverages the Emu algorithm alongside customized parameter optimization, effectively minimized these off-target assignments, resulting in a community profile that more closely mirrored the theoretical composition.
While full-length 16S/18S sequencing theoretically provides species-level resolution, our evaluation reveals that classification accuracy at this taxonomic rank is highly volatile and heavily dependent on the composition of the reference database. As demonstrated in our supplementary evaluation (Table S5, available at supplementary material Bioinformatics Advances online), relying exclusively on standard public databases (e.g. unmodified NCBI) results in severe taxonomic deviations, including extreme false-negative rates and inflated non-target assignments. Although microbiONT can achieve accurate species-level profiling when supplemented with specific reference sequences (Fig. S5, available at supplementary material Bioinformatics Advances online), standard databases frequently lack the necessary completeness to resolve taxa with high sequence homology reliably. Therefore, given the significant discrepancies introduced by database variations, we strongly recommend utilizing the genus level as the primary and most robust basis for standard taxonomic interpretation.