We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.LG

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Machine Learning

Title: Fast differentiable DNA and protein sequence optimization for molecular design

Abstract: Designing DNA and protein sequences with improved function has the potential to greatly accelerate synthetic biology. Machine learning models that accurately predict biological fitness from sequence are becoming a powerful tool for molecular design. Activation maximization offers a simple design strategy for differentiable models: one-hot coded sequences are first approximated by a continuous representation which is then iteratively optimized with respect to the predictor oracle by gradient ascent. While elegant, this method suffers from vanishing gradients and may cause predictor pathologies leading to poor convergence. Here, we build on a previously proposed straight-through approximation method to optimize through discrete sequence samples. By normalizing nucleotide logits across positions and introducing an adaptive entropy variable, we remove bottlenecks arising from overly large or skewed sampling parameters. The resulting algorithm, which we call Fast SeqProp, achieves up to 100-fold faster convergence compared to previous versions of activation maximization and finds improved fitness optima for many applications. We demonstrate Fast SeqProp by designing DNA and protein sequences for six deep learning predictors, including a protein structure predictor.
Comments: All code available at this http URL; Moved example sequences from Suppl to new Figure 2, Added new benchmark comparison to Section 4.3, Moved some technical comparisons to Suppl, Added new Methods section
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Journal reference: BMC Bioinformatics 22 (2021) 1-20
DOI: 10.1186/s12859-021-04437-5
Cite as: arXiv:2005.11275 [cs.LG]
  (or arXiv:2005.11275v2 [cs.LG] for this version)

Submission history

From: Johannes Linder [view email]
[v1] Fri, 22 May 2020 17:03:55 GMT (1228kb,D)
[v2] Sun, 20 Dec 2020 22:44:01 GMT (2418kb,D)

Link back to: arXiv, form interface, contact.