Generative chemistry is revolutionizing drug discovery, using AI to efficiently explore entirely new areas of chemical space in the search for novel molecules. CADD solution Flare’s generative chemistry integration layer MolGenAI, combines the power of AI with Flare’s robust physics-based molecule ideation and evaluation.
What is MolGenAI?
MolGenAI is a generative AI tool that is based on AstraZeneca’s REINVENT4 technology,1 essentially running this algorithm in a user-friendly interface within Flare™ and combining it with Flare’s radial plot filtering. The algorithm uses models based on known molecules in the ChEMBL database, representing the chemical space as a probability distribution. It then samples this distribution to output structures as SMILES strings. The model, or prior, can be retrained on a user’s input structures and new molecules will more closely represent the input structures. A typical MolGenAI workflow is shown in Figure 1.

Ideally, you would have 100+ ligands of a congeneric series from your project, but if this is not available, Spark™ or Hit Expander can be used to generate the necessary training set. Less than 100 compounds tends to output more unfavorable results. MolGenAI will then retrain the default prior model on your training set and output results that are similar in 2D structure to your training set. In Flare, the MolGenAI window also allows you to pre-filter your results by Radial Plot score to optimize the properties of your results. In this article, we will give an example of how to approach a typical drug discovery project for covalent inhibitors.
How can MolGenAI be used for covalent inhibitor design?
Experiment set-up
To illustrate the utility of MolGenAI for inhibitor design, we will use inhibitor 1 that targets the serine protease prolyl oligopeptidase (POP).2 Because there are very few published analogs of this compound, the Hit Expander tool in Flare was used to generate a 100+ analogs to retrain the default prior model (Figure 2).

While there is some precedent for nitriles to act as covalent inhibitors of this target, observed by co-crystallization of nitrile inhibitors,3 it’s more likely that the nitrile-serine bond is quite weak and behaves more as a non-covalent inhibitor, evidenced by short residence times.4 In any case, the covalent warhead can be modified in the post-processing of the results.
Once the training set is ready, the Radial Plot score can be optimized for the ligand series to pre-filter MolGenAI results. A minimum Radial Plot score can be applied to the search to avoid compounds with unfavorable properties, such as low molecular weight fragments. In this experiment, the default settings were used for the following five properties: MW, #Atoms, SlogP, TPSA, and Flexibility. A minimum Radial Plot score of 0.9 was applied to the search.
Results
The MolGenAI experiment yielded some interesting results. Of course, many analogs of compound 1 were obtained that only changed a small portion of the molecule, such as the benzyl group, since the core structure was kept consistent. These results resembled those of a Spark bioisostere replacement experiment.
Many results, however, changed multiple aspects of the molecule and would not have been found with a Spark search. Notable results from the MolGenAI experiment include compounds 2, 3, and 4 (Figure 3).

While compound 2 is technically just an R-group replacement of the benzyl group with an acetyl group, this compound would likely not have been found with Spark, as an acetyl group is not a good bioisostere of a benzyl group. Acetyl groups are smaller and have different electrostatics, and there are many more benzyl-like fragments in the Spark databases that are more likely to be a closer bioisosteric match than an acetyl group.
Compounds 3 and 4 are quite structurally different, however. For compound 3, the algorithm has removed the benzyl group as well as the pyrrolidinone portion of the scaffold and has added halogens at the two ortho positions. While this compound is not reported as an inhibitor of POP, compounds 5 and 6 are. These two compounds come from an SAR study of boronic acid POP inhibitors. While compound 5 is structurally similar (replacing the Cl in compound 3 with a F) but not very active, compound 6 from the same series is quite potent.2
Even more structurally dissimilar to compound 1 is compound 4, for which the algorithm has effectively just excised the aromatic portion of the scaffold. While this result is not a known POP inhibitor, compound 7 which only differs by a carbonyl on the benzyl group is reported and is quite potent.5 Compound 8, or Z-Pro-Prolinal (ZPP), is also quite similar to compound 4, replacing the nitrile with an aldehyde and the benzyl group with a Cbz group. This inhibitor is very potent2 and has been co-crystallized with the protein.6
With MolGenAI, we have obtained results for three separate ligand series, all with striking similarities to known POP inhibitors. These series were all discovered independently over a period of several years, while MolGenAI was able to find them in a 20-minute experiment.
Physics-based and AI solutions for compound idea generation
Cresset offers physics-based and AI-based solutions to efficiently explore chemical space for new compound idea generation. Spark targets a user-selected region for bioisostere replacement and scaffold hopping – starting from fragments of real molecules. In this way, Spark enables you to generate diverse, non-obvious ideas that retain the characteristics of active molecules.
As we can see in the example above, MolGenAI can go a step further than the fragment databases supplied with Spark (e.g., eMolecules and ChEMBL). If your training set contains 100+ compounds with the same molecular scaffold but only varies the R-group or Hit Expander additions, the algorithm will do its best to maintain the common substructure in the results. However, you will still get more varied results, namely (1) R-groups that are not bioisosterically similar by Spark’s standards and that vary in size, and (2) different yet structurally similar scaffolds in addition to the varied R-groups. This tool can therefore be used to simultaneously search for similar scaffolds and R-groups.
On the other hand, if you have 500+ compounds in the training set, such as from a Spark R-group replacement experiment, MolGenAI is more likely to behave like a Spark search, keeping the common structure constant and varying the R-groups. In this way, you can search for R-groups that are not bioisosteres of that of one starting compound. Future Flare development will include SparkAI, which will combine generative AI with Spark, keeping the scaffold constant while you search for R-groups, or vice-versa.
Can I use MolGenAI with Cresset 3D scoring during the experiment?
Currently, MolGenAI uses transfer learning to find fragments in the ChEMBL database with similar SMILES strings to the molecules in your training set, using a 2D-based search. Cresset field points can be used in post-processing of the results for analysis of electrostatics. For example, the generated analogs can be aligned to a reference ligand in a ligand-based alignment or docked to a protein active site. In the upcoming version of Flare, you will be able to use Cresset field points and structure-based docking during the transfer learning to output the most promising compounds.
Conclusions
In this article, we have illustrated that MolGenAI’s generative AI function can successfully be applied to the discovery of covalent inhibitors, even without sufficient activity data to retrain the algorithm. Utilizing Hit Expander to generate a theoretical training set based on one known hit, we obtained analogs of two separate active series in one simple experiment. With effective post-processing of results, MolGenAI can be a powerful tool to generate ideas for new or backup ligand series.
References
- Loeffler, H.H., He, J., Tibo, A. et al., Reinvent 4: Modern AI–driven generative molecule design. J. Cheminform. 2024, 16, 20. https://doi.org/10.1186/s13321-024-00812-5
- Plescia, J., Dufresne, C., Janmamode, N. et al., Discovery of covalent prolyl oligopeptidase boronic ester inhibitors. Eur. J. Med. Chem. 2020, 185, 111783.
- Kaszuba, K., Róg, T., Danne, R. et al. Molecular dynamics, crystallography and mutagenesis studies on the substrate gating mechanism of prolyl oligopeptidase. Biochimie 2012, 94, 6, 1398-1411. https://doi.org/10.1016/j.biochi.2012.03.012
- Plescia, J., De Cesco, S., Patrascu, M., et al., Integrated Synthetic, Biophysical, and Computational Investigations of Covalent Inhibitors of Prolyl Oligopeptidase and Fibroblast Activation Protein α. J. Med. Chem. 2019, 62, 17, 7874-7884. https://doi.org/10.1021/acs.jmedchem.9b00642
- Jarho, E. M., Wallén, E. A. A., Christiaans, J. A. M. et al., Dicarboxylic Acid Azacycle l-Prolyl-pyrrolidine Amides as Prolyl Oligopeptidase Inhibitors and Three-Dimensional Quantitative Structure−Activity Relationship of the Enzyme−Inhibitor Interactions. J. Med. Chem. 2005, 48, 4772-4782. https://doi.org/10.1021/jm0500020
- Fülöp, V., Böcskei, Z., Polgár L., Prolyl Oligopeptidase. Cell 1998, 94, 161-170.