The Sequence Read Archive (SRA) is the primary archive of high-throughput sequencing data hosted by the National Institutes of Health (NIH). This collection contains genome sequences from SARS-CoV-2 deposited into the SRA. The SRA represents the largest publicly available repository of raw SARS-CoV-2 sequencing data. Where possible raw sequence data were processed by DNAstack through a unified bioinformatics pipeline to produce genome assemblies and variant calls. Methodology: SRA-formatted data was converted to the standard FASTQ format using the sra-toolkit (https://github.com/ncbi/sra-tools). FASTQ files were aligned to the SARS-CoV-2 reference genome (https://www.ncbi.nlm.nih.gov/nuccore/MN908947) to produce alignment files (BAM format), which were then used to call variants, stored as variant call format (VCF) files. Reads were also assembled into SARS-CoV-2 genomes and genomic regions, which were assigned to SARS-CoV-2 lineages using Pangolin (https://github.com/cov-lineages/pangolin). The full bioinformatics processing pipeline written in WDL can be found on Dockstore (https://dockstore.org/workflows/github.com/DNAstack/covid-processing-pipeline/covid-19-varcal:master)...
This collection is openly accessible.
NCBI SRA SARS-CoV-2 Genomes is published on Viral AI.