Substructure-based neural machine translation for retrosynthetic prediction

Authors
Ucak, Umit V.Kang, TaekKo, JunsuLee, Juyong
Issue Date
2021-01-11
Publisher
BMC
Citation
JOURNAL OF CHEMINFORMATICS, v.13, no.1
Abstract
With the rapid improvement of machine translation approaches, neural machine translation has started to play an important role in retrosynthesis planning, which finds reasonable synthetic pathways for a target molecule. Previous studies showed that utilizing the sequence-to-sequence frameworks of neural machine translation is a promising approach to tackle the retrosynthetic planning problem. In this work, we recast the retrosynthetic planning problem as a language translation problem using a template-free sequence-to-sequence model. The model is trained in an end-to-end and a fully data-driven fashion. Unlike previous models translating the SMILES strings of reactants and products, we introduced a new way of representing a chemical reaction based on molecular fragments. It is demonstrated that the new approach yields better prediction results than current state-of-the-art computational methods. The new approach resolves the major drawbacks of existing retrosynthetic methods such as generating invalid SMILES strings. Specifically, our approach predicts highly similar reactant molecules with an accuracy of 57.7%. In addition, our method yields more robust predictions than existing methods.
Keywords
COMPUTER-ASSISTED DESIGN; AIDED SYNTHESIS DESIGN; ORGANIC-CHEMISTRY; TRANSFORMER; LANGUAGE; OUTCOMES; SYSTEM; MODEL; TOOL; COMPUTER-ASSISTED DESIGN; AIDED SYNTHESIS DESIGN; ORGANIC-CHEMISTRY; TRANSFORMER; LANGUAGE; OUTCOMES; SYSTEM; MODEL; TOOL; Retrosynthesis planning; Neural machine translation; Seq-to-seq; Attention
ISSN
1758-2946
URI
https://pubs.kist.re.kr/handle/201004/117538
DOI
10.1186/s13321-020-00482-z
Appears in Collections:
KIST Article > 2021
Files in This Item:
There are no files associated with this item.
Export
RIS (EndNote)
XLS (Excel)
XML

qrcode

Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

BROWSE