Current browse context:
q-bio
Change to browse by:
References & Citations
Computer Science > Machine Learning
Title: Amortized Tree Generation for Bottom-up Synthesis Planning and Synthesizable Molecular Design
(Submitted on 12 Oct 2021 (v1), last revised 12 Mar 2022 (this version, v2))
Abstract: Molecular design and synthesis planning are two critical steps in the process of molecular discovery that we propose to formulate as a single shared task of conditional synthetic pathway generation. We report an amortized approach to generate synthetic pathways as a Markov decision process conditioned on a target molecular embedding. This approach allows us to conduct synthesis planning in a bottom-up manner and design synthesizable molecules by decoding from optimized conditional codes, demonstrating the potential to solve both problems of design and synthesis simultaneously. The approach leverages neural networks to probabilistically model the synthetic trees, one reaction step at a time, according to reactivity rules encoded in a discrete action space of reaction templates. We train these networks on hundreds of thousands of artificial pathways generated from a pool of purchasable compounds and a list of expert-curated templates. We validate our method with (a) the recovery of molecules using conditional generation, (b) the identification of synthesizable structural analogs, and (c) the optimization of molecular structures given oracle functions relevant to drug discovery.
Submission history
From: Wenhao Gao [view email][v1] Tue, 12 Oct 2021 22:43:25 GMT (11435kb,D)
[v2] Sat, 12 Mar 2022 19:18:25 GMT (11486kb,D)
Link back to: arXiv, form interface, contact.