The Sequence Read Archive (SRA) is the primary archive of high-throughput sequencing data hosted by the National Institutes of Health (NIH). This collection contains genome sequences from MPXV, the Monkeypox virus, deposited into the SRA. The SRA represents the largest publicly available repository of raw MPXV sequencing data. Where possible, raw sequence data were processed by DNAstack through a unified bioinformatics pipeline to produce genome assemblies and variant calls. Methodology: SRA-formatted data was converted to the standard FASTQ format using the sra-toolkit (https://github.com/ncbi/sra-tools). FASTQ files were aligned to the MPXV reference genome (https://www.ncbi.nlm.nih.gov/nuccore/NC_063383) to produce alignment files (BAM format), which were then used to call variants using the Genome Analysis Toolkit (GATK, https://github.com/broadinstitute/gatk), stored as variant call format (VCF) files. Reads were also assembled into MPXV genomes and genomic regions using iVar (https://github.com/andersen-lab/ivar), which were then assigned to MPXV lineages using Nextclade (https://github.com/nextstrain/nextclade). Data Location: Global Data Source: NIH SRA...
This collection is openly accessible.
NCBI SRA Monkeypox Genomes is published on Viral AI.