Latest advances in FEP: We’re Going to Need a Bigger Drug Discovery Toolbox…

SHARE

Introduction

Time flies when you are having fun! I can’t believe that it has been over five years since writing the blog post – Free Energy Perturbation (FEP): Another technique in the drug discovery toolbox – which introduced Free Energy Perturbation (FEP) and some of the broader concepts needed when running such calculations. Fast forward to today and I thought it would be good to discuss FEP again, because a lot has happened in those five years.

Free Energy Perturbation remains a very exciting endeavor and as the technique becomes more reliable and predictive, it has the increased potential to help science-based industries such as pharmaceutical, biotechnology and agrochemicals to become more efficient. This is because it will allow them to move away from the traditional expensive exploratory ‘lab-based’ approach and replace it with more efficient in silico prediction simulations.

FEP continues to grow, and as our understanding of the technique develops, then so does the approach to running more-accurate simulations. Let’s look at some of the advances in FEP that we have seen over the past five years.

Lambda window selection

One of the challenges of setting up FEP perturbation maps is deciding how many lambda windows should be calculated for each link. Calculating too few lambda windows potentially leads to poor results being generated, while in contrast, calculating too many wastes valuable GPU time. Five years ago, I was ‘guessing’ the number of lambda windows required based on the complexity of the transformation taking place (i.e. the type and numbers of atoms being changed). However, this guess was regularly wrong which resulted in the link having to be recalculated. However, by using ‘short’ exploratory calculations to give an automated ‘educated guess’ to the numbers of lambda windows required, it is now possible to potentially reduce this wasteful (and frustrating) experience: not only does it potentially keep GPU costs down, but it removes the unnecessary pressure on the scientist to guess correctly, hopefully giving them a greater peace of mind.

Schematic showing how the automatic lambda scheduling algorithm calculates the number of lambda windows required

Figure 1. The automatic lambda scheduling algorithm significantly reduces the amount of guess work required when defining the numbers of lambda windows to be calculated for each transformation.

Force fields

At the center of any FEP calculation is how the system is described and modeled. Getting this right is essential for generating reliable simulation results. If the molecular system consists of ‘standard residues’ and ligand atoms within the limits of the force field description, then the molecular description will be good enough to give some potentially excellent results. But nature and life rarely behave like that! Significant errors are often obtained – especially for ligands – where the underlying force field is unable to accurately describe the system. These are usually attributed to the poor description of torsion angles. One solution available is to run quantum mechanics calculations to generate some improved parameters for specific torsions. This improvement in the description should in turn result in more accurate FEP simulations.

Plot showing torsion angles described by the force field and torsion angles calculated using quantum mechanics

Figure 2. Some ligands have torsions which are not described particularly well by the selected force field. It is often possible to use QM calculations to modify torsion parameters so that these torsions have more appropriate behavior.

Generating a broad set of parameters that can accurately model how diverse sets of ligands interact with their biological macromolecular environments is essential. Looking back five years, Cresset was on the verge of joining the Open Force Field Initiative – a body of academic and commercial scientists brought together to develop (initially) an accurate ligand force field that could be used in conjunction with a macromolecular force field such as AMBER. Several improvements to the OpenFF have been achieved over these five years and more are to come…

But this still raises an issue – how the ligand-based force field and macromolecular force field interact with each other (or not!). In recent years, there has been an increased need to model covalent inhibitors within their binding site environment, and because there are no parameters to connect the two worlds together, it is often difficult to model the covalent systems correctly. However, work is ongoing within the industry to continue to improve the description of ligands and proteins within a single force field; we are also developing ideas on how to reliably model such covalent systems in the future.

Charges

In my original blog post, I mentioned that successfully modeling charge changes in Relative Binding Free Energy (RBFE) studies were problematic and recommended that all molecules should have the same formal charge. This is not always possible, and removing charged ligands from the dataset removes valuable information. By introducing a counterion to neutralize the charged ligand now gives us a way to retain the same formal charge across the perturbation map where the formal charges of the ligands differ. Perturbations involving charged ligands do indeed give results that are potentially less reliable, but we have found it is possible to maximize the reliability of the result by running longer simulations when compared to a non-charged transformation.

Drug design software Flare showing options to change simulation length in an FEP calculation.

Figure 3. By increasing the simulation length for the lambda windows in a charge change transformation, it is possible to reliably model charge changes in relative-binding free energy experiments.

Water

The position of water molecules in molecular simulations is crucial and this is especially true with FEP experiments. RBFE calculations can be susceptible to different hydration environments: If the ligand in the forward direction of a particular link has an inconsistent hydration environment compared to the starting ligand in the reverse direction, then this has the potential to result in the hysteresis of the ΔΔG calculation between the forward and reverse transformations. Ensuring that all the ligands are adequately hydrated in the FEP experiment is a critical step to getting the best results. It is possible to use techniques such as 3D-RISM and GIST to help understand where initial hydration is lacking in the system, while sampling techniques such as Grand Canonical Non-equilibrium Candidate Monte-Carlo (GCNCMC) techniques – which use Monte-Carlo steps to simultaneously add/remove water molecules – provide an opportunity of ensuring appropriate hydration of the ligands.

Targets

There are a lot of possible targets out there! An obvious choice is to just consider ‘easy’ targets which consist of water-soluble proteins with a few hundred amino acids, but what about more challenging targets? For example, targets located at a membrane often involve simulating tens-of-thousands of atoms requiring larger amounts of processor time; but patience is a virtue, and simulations can be set-up involving these large systems (such as GPCRs) to give some very good results. Having done these expensive calculations and understood what is possible in terms of result accuracy, it is now possible to experiment with truncating the system so that fewer atoms are modeled, potentially reducing simulation time while not significantly impacting on the quality of the result.

Image of P2Y1 protein in gray cartoon showing lipid membrane molecules in yellow

Figure 4. Proteins found within a lipid membrane (yellow molecules), such as P2Y1 (secondary structure cartoon, gray) can be accurately modeled using relative-binding free energy methods.

Active Learning FEP

One of the workflows which demonstrates further efficiency gains involves running FEP and 3D-QSAR methods together: FEP simulations can provide accurate binding predictions but are inherently slow, while more-rapid QSAR methods use ligand-based information to generate results for larger sets of molecules at the expense of being less accurate than FEP methods. When generating a large ensemble of virtual hits/designs using bioisostere replacement approaches (such as those available in Spark™) or virtual screening studies (such as those done with Blaze™), we can select a subset of these molecules to run the FEP calculation on and then use QSAR methods to rapidly predict the binding affinity of the remaining set based on the initial FEP result. Molecules in the larger set which look interesting are then added to the FEP set and are subsequently calculated, and the process is repeated until no improvement is obtained.

Figure 5. Schematic showing the use of Active Learning for a set of in silico designs.

Absolute FEP

One of the drawbacks of Relative Binding Free Energy (RBFE) is the limited scope of the ligand changes that can be modeled with the technique: it is not uncommon to be limited to a 10-atom change in a molecule pair. This is a reasonable limitation when FEP is being used to prioritize compounds in a lead-optimization stage of a pharma project, but it is very restrictive at the earlier hit identification phase, where exploration of a larger area of chemical space is done (for example, with virtual screening). With Absolute Binding Free Energy (ABFE), there is greater freedom when setting up each ligand: they can be calculated independently of one another.
Schematic of the free energy cycles used to calculate the binding free energy of molecules

Figure 6. The free energy cycle used to calculate the binding free energy of molecules. For both the bound and unbound states, the ligand is decoupled from its environment by first turning off the electrostatic interactions between the ligand and environment, followed by the van der Waals (vdW) parameters of the ligand atoms. The ligand is restrained within the binding site throughout the bound calculations.

An additional advantage of this approach is that you are not restricted to using the same protein structure for all the compounds. Protonation states of binding-site residues are likely to be partially defined by the ligand present. This means that different protein structures with different protonation states can be used depending on the ligand being studied.

In ABFE, it is often the case that there is an off-set error in the calculated results when compared with experimental free energy of binding. This is often because of the over-simplistic description of the binding process which doesn’t consider the changes taking place with the protein, such as protonation state changes on residues: ABFE is calculated using an energy cycle in which the unoccupied binding site and the ligand-bound binding site are a similar conformation. It can be envisaged that as the ligand approaches the protein and starts to interact, then the protein potentially changes conformation to accommodate the ligand. This protein conformational change (and any associated protonation state changes) in the most part is not accounted for in the ABFE calculation, potentially resulting in some residual error in the ΔG calculation when comparing with experiment.

Running ABFE calculations is more demanding than running RBFE calculations as such studies require longer amounts of time to equilibrate compared to RBFE experiments. Running RBFE calculations for a congeneric series of 10 ligands would probably take 100 GPU hours while the equivalent ABFE experiment would probably take 1000 GPU hours. But it is not just about GPU hours: to get the best out of RBFE often requires a lot of tinkering and testing by the scientists which is very demanding on time-sensitive projects. However, as the domain of applicability within our own Flare FEP offering has been developed our scientists have gained deeper insight into solving increasingly complex challenges. This has enabled them to effectively support our customers to get the best possible results for their project.

Reimagining the Future…

ABFE has enormous potential, if some of the simulation issues can be solved. Using my rose-tinted spectacles, it is possible to see how this could help the science-based industries to become more efficient. One potential drug discovery workflow where ABFE could be incredibly useful is in the reliable selection of hits from a virtual screening experiment. Such techniques explore significant areas of chemical space, where a structurally diverse set of compounds are purchased/made and tested to identify potential hits. The chemical space around the tested hits is then usually explored to see if there is SAR available around these hits. Rather than relying on expensive ‘purchasing, testing, and analyzing’ approaches, it is possible to see that ABFE could identify the initial diverse hits, followed by exploration of the immediate chemical space around the initial hits, using -for example – an active learning approach to allow efficient exploration of the immediate area around the identified hits. This would lead to a shortlist of compound designs earmarked for synthesis that are ideal for the target being studied.

Therefore, free energy perturbation methods have the potential to unleash the power of Digital Transformation – to explore chemical space in silico and then identify the molecules for synthesis that really matter for your project in the most efficient way.

Related Science Resources

Application of generative AI to design covalent inhibitors of prolyl oligopeptidase
Generative chemistry is revolutionizing drug discovery, using AI to efficiently explore entirely new areas of chemical space in the search for...
Flare™ V12 released: Significant improvements to FEP calculations and structure-based methods
We are pleased to announce the release of Flare V12, which introduces a range of new capabilities and enhancements aimed at improving the efficiency,...

Subscribe & Don't Miss Out

Receive our newsletter to be among the first to hear about product releases, case studies, opinion articles, events and more.