We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

q-bio.OT

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Quantitative Biology > Other Quantitative Biology

Title: Evolutionary trajectory and origin of SARS-CoV-2

Authors: Anyou Wang
Abstract: Traditionally alignment-based phylogenetics faces challenges to uncover the evolutionary trajectory of SARS-CoV-2. This study develops a novel alignment-free system and reveals the evolutionary trajectory of SARS-CoV-2 from more than one million of genome sequences. This new system contains Fr\'echet distance(Fr) and artificial recurrent neural network. Fr computes the dissimilarity between variant and reference genome, which is decomposed into 84 features (4 single nucleotides, 16 dinucleotides and 64 codons). Recurrent neural network predicts and forecasts time-series Fr trajectory, inferring SARS-CoV-2 evolutionary trajectory and origin. Generally SARS-CoV-2 genome mutates rapidly via deletion during COVID-19 pandemic. Among single nucleotides, C mutates fast but T changes slowly. C-prefix dinucleotide (e.g. CG and CT) also loses dramatically during evolution. Similarly, the virus genome also deletes several codons prefixed by C (e.g. CCT) but gains several T and A prefix codons (e.g. TTA and ATT) during its evolution. Interestingly, codon CCT and CT centrally control the entire SARS-CoV-2 genome, and their evolutionary trajectories fit COVID-19 cases spike. Therefore C-prefix feature trajectory marks SARS-CoV-2 evolution. This study further identifies total 34 SARS-Co-2 variants, which can be classified into 3 groups, slight mutation group, middle level deletion, and high deletion. The slight deletion group and the high deletion group have low infection capacity. The middle deletion group gradually deletes their genome with a certain rhythm trajectory, corresponding to the pandemic peaks, which causes most of the global COVID-19 cases. Mink is the origin of SARS-Co-2, and the origin path follows this order: mink, cat, tiger, mouse, bat and pangolin. Together, this mink-origin SARS-Co-2 evolves with C-driven rhythm deletions to infect humans.
Comments: 15 pages, 9 figures
Subjects: Other Quantitative Biology (q-bio.OT)
Cite as: arXiv:2110.07696 [q-bio.OT]
  (or arXiv:2110.07696v1 [q-bio.OT] for this version)

Submission history

From: Anyou Wang [view email]
[v1] Thu, 14 Oct 2021 20:10:20 GMT (6036kb)
[v2] Mon, 18 Jul 2022 18:34:06 GMT (4402kb)
[v3] Sat, 23 Mar 2024 17:25:37 GMT (4460kb)

Link back to: arXiv, form interface, contact.