In this case study, we test the predictive abilities of Free Energy Perturbation (FEP) using computational methods available within Cresset’s software Flare™, to arrive at accurate binding free energies for new ligand suggestions withing the TYK2 system.
Introduction
Free energy perturbation (FEP) calculations offer the opportunity to accelerate ligand design by accurately ranking ligands within projects. In previous work1 we have demonstrated competitive results across literature benchmark FEP datasets. Here we show a case study where we have prospectively predicted the binding affinities of a set of compounds binding to tyrosine kinase 2 (TYK2) in a blind experiment, mimicking the process a project would undertake to prioritize new designs.
TYK2 is a member of the Janus kinase (JAK) family which contains four members; JAK1, JAK2, JAK3 and TYK2, each of which associates with a distinct set of receptors to induce intracellular signalling.2 TYK2 is known to associate with several cytokine receptors, and the inhibition of TYK2 kinase activity is linked to therapeutic strategies in various autoimmune diseases such as psoriasis and inflammatory bowel diseases (e.g. Crohn’s disease).3,4
In a 2013 paper, Liang et al2,3 presented a large congeneric series of ligands aimed at exploring two areas of binding within the TYK2 active site: the first explored the hinge region and the second investigated the linker region.2 A set of 16 ligands from the first paper is now commonly used as the TYK2 benchmark dataset.1,5 As these compounds originate from the same project and share the 4-aminopyridine benzamide core, this presents the opportunity to use the TYK2 system from our benchmark experiment for prospective calculations on additional compounds. We chose a set of 18 ligands from the paper exploring the linker region3 (i.e. R-group modifications are at the C4-position of the upper phenyl ring), as shown in Figure 1(B), to produce a blind predictive study with Flare FEP in production mode.

Methods
FEP experiment for TYK2 benchmark compounds
In this instance, we have already run a successful benchmark experiment for the TYK2 system,6. Here, I will give a brief recap of the preparation steps:
We began by carefully preparing the protein (PDB:4GIH), first with our automated Protein Prep tools in Flare and followed with a careful inspection of the active site. We then used Molecular Dynamics (MD) with Grand Canonical Nonequilibrium Candidate Monte Carlo (GCNCMC), to provide confidence that the model could capture the bioactive ligand-protein binding mode. Finally, we used 3D-RISM water stability analysis7 with the crystallographic ligand (0X5) in place, to add water sites so the system was appropriately hydrated going into a benchmark FEP calculation. For a more detailed explanation of our recommended protein preparation steps prior to a Flare FEP benchmark study, please refer to some of our other materials (A quick and accurate hit to lead process for the discovery of CDK9 inhibitors and FEP on membrane targets: capturing lipid exposed binding in the P2Y1 GPCR complex). Here I will give a brief recap of the protein preparation and subsequent FEP benchmark process:
Using the prepared protein, the benchmark dataset of 16 ligands1 underwent ligand alignment Conformational Hunt and Alignment against the crystallographic ligand, to get as tight an alignment as possible (Figure 2, A). These ligand conformations and prepared protein were then used to benchmark the performance of Flare FEP predictions vs the experimental binding affinities.

By using these careful preparation methods prior to our benchmark study, we obtained improved predictive affinities and enhanced statistical results (r2 = 0.8 and MUE = 0.50) with respect to those published by Wang et al.5 which eliminated the need for subsequent alterations and troubleshooting of any problematic perturbations. The graph in Figure 3 presents these benchmark results.

Prediction of the unknown hinge binding designs
With confidence in the performance of our benchmark experiment, 18 ligands from the hinge binder variants2 were identified, all contained the core present in the benchmark ligand set but explored different substituents at the C-4 position of the phenyl ring in the core (see Figure 1, B). The different substituents are detailed in Table 1; the binding affinities are also shown in this table although this information was not used until the assessment of the predictions. It is also worth noting the activity of these ligands ranged from 1.0 to 60.8nM (pIC50 of 9.0 to 7.2).
| Compound Title | R-group Substitution | TYK2 Ki (nM) |
| Liang_4 | -NH2 | 1.0 |
| Liang_5 | ![]() | 3.0 |
| Liang_6 | ![]() | 13.5 |
| Liang_7 | -OH | 3.1 |
| Liang_8 | ![]() | 7.8 |
| Liang_9 | ![]() | 5.4 |
| Liang_11 | ![]() | 21.8 |
| Liang_12 | ![]() | 10.3 |
| Liang_13 | ![]() | 4.4 |
| Liang_14 | ![]() | 3.3 |
| Liang_15 | ![]() | 60.8 |
| Liang_16 | ![]() | 35.1 |
| Liang_17 | ![]() | 7.8 |
| Liang_18 | ![]() | 30.1 |
| Liang_19 | ![]() | 1.8 |
| Liang_20 | ![]() | 9.5 |
| Liang_21 | ![]() | 3.4 |
| Liang_22 | ![]() | 10.9 |
Using both the 16 benchmark TYK2 compounds (with known activities, Figure 2, A) and the 18 ligands from the hinge binder dataset (input with no activity data, Figure 2, B), a production FEP perturbation map was automatically created. This map contained only one ligand with known activity, compound ejm_46 (from the benchmark series), where R=H (C-4 position of the phenyl ring) in Figure 1, B.
The network included 27 links between the 19 compounds, totalling 54 perturbations as each link was run in forward and reverse directions, ‘dual-way’, to obtain hysteresis for each link. Given the chemical similarity between the unknown compounds and the selected known compound in the perturbation network, the link scores are good and did not require any intermediates to bridge difficult perturbations. However, they still have a wide range of scores from 0.55 to 0.95 showing that some of these perturbations are non-trivial changes. The estimated calculation of the full perturbation network would take 265 GPU hours.
Results
As outlined at the start of this case study, we conducted this production run to rank the 18 “new”, hinge binder variants to aid in prioritizing designs for synthesis. The best predicted binder was compound Liang_6 with a relative binding affinity prediction of -13.1 kcal/mol (±1.3 kcal/mol), while the worst predicted binder was Liang_15 at -9.9 kcal/mol (±0.5 kcal/mol). In all RBFE calculations, we aim for an error within 2 kcal/mol. The predicted error helps us to understand how reliable our predictions are; the larger the error value, the higher the uncertainly. Thus, in our dataset, only one ligand had an error greater than this (Liang_16 with an error of ±2.3 kcal/mol), meaning we would have less confidence in this predicted value. The remaining ligands all had predicted errors within 2 kcal/mol. The top 10 predicted binders (those with the lowest ΔG values) are detailed in bold in Table 2 along with their associated errors. In this instance, we decided that compounds with a ΔG value equal to or less than -11.0 kcal/mol could be considered for further assay testing, providing 10 ligands to take forward. This ΔG value was chosen as the known active, ejm_46 (from the benchmark series) had a ΔG value of -11.3 kcal/mol, providing us with some overlap in binding energies to the known compound. It is worth noting, in Production mode of Flare FEP, the known active is chosen by the software based on which compound in a provided dataset best connects to the unknown ligands; it is not based on activity.
The eight remaining ligands, which we would not have taken forward for testing, had predicted binding affinities ranging between -9.9 to -10.8 kcal/mol. In a real project, a user may wish to take a ligand from this group of “worse” binders forward if it supported a hypothesis to address a further issue (e.g. solubility or ADME properties). As can be seen in Table 2, and from the range of predicted binding affinities quoted so far, there is an overlap in predictions within the middle of the dataset. Seven compounds fall between the narrow 0.6 kcal/mol range of -10.6 and -11.2 kcal/mol (Liang_8, 9, 12, 13, 14, 19 and 20). In such a case, if the user wished to take compounds forward from this portion of the dataset, they could assess their hypothesis and choose those which best address other concerns of the project (e.g. solubility), as the error values for these ligands overlap with the arbitrary ΔG cut off which we have applied.
| Ligand Title | Predicted ΔG (kcal/mol) | Predicted ranking | Predicted Activity Error ΔG (kcal/mol) |
| Liang_6 | -13.1 | 1 | 1.3 |
| Liang_5 | -12.4 | 2 | 1.4 |
| Liang_21 | -12.4 | 2 | 1.3 |
| Liang_4 | -12.2 | 4 | 1.1 |
| Liang_7 | -11.6 | 5 | 0.6 |
| Liang_11 | -11.5 | 5 | 0.5 |
| Liang_9 | -11.2 | 7 | 0.5 |
| Liang_8 | -11.1 | 8 | 0.5 |
| Liang_13 | -11.0 | 9 | 0.4 |
| Liang_19 | -11.0 | 9 | 0.8 |
| Liang_14 | -10.8 | 11 | 0.6 |
| Liang_12 | -10.8 | 11 | 0.7 |
| Liang_20 | -10.6 | 13 | 0.9 |
| Liang_16 | -10.4 | 14 | 2.3 |
| Liang_18 | -10.2 | 15 | 0.6 |
| Liang_22 | -10.2 | 15 | 1.5 |
| Liang_17 | -10.1 | 17 | 0.5 |
| Liang_15 | -9.9 | 18 | 0.5 |
We can now compare the ligands’ ranking order for experimental vs predictive ΔG values (Table 3). For the predicted top ten best binders, the experimental ranking is mostly in agreement with predicted ranking. Despite the more difficult perturbation, adding five heavy atoms to grow a pyrazole ring, Liang_21 has been well predicted and ranked highly, matching the experimental ranking as a better binder in the dataset. Two of the compounds have been over-predicted though, Liang_6 by 2.6 kcal/mol and Liang_11 by 1.1 kcal/mol and thus ranked more highly than seen experimentally. As previously stated, an error of 2 kcal/mol or less is expected within an RBFE calculation so Liang_11 is not outside what we would expect to see. However, due to the limited activity range, the compounds in the middle of the dataset only differ by 0.2 kcal/mol in some instances. Thus, as Liang_11 is overpredicted by 1.1 kcal/mol, this has taken its experimental ranking from number 15 in the dataset, to be ranked at number five in the predicted data. Liang_6 contained a sulfonyl group addition, whilst Liang_11 sees the perturbation to a tert-butyl group, both more complicated transformations which we can consider in more detail.
| Ligand Title | TYK2 Ki (nM) | Experimental ΔG (kcal/mol) | Experimental ranking | Predicted ΔG (kcal/mol) | Predicted ranking | Predicted Activity Error ΔG (kcal/mol) |
| Liang_6 | 13.5 | -10.7 | 14 | -13.1 | 1 | 1.3 |
| Liang_5 | 3 | -11.6 | 3 | -12.4 | 2 | 1.4 |
| Liang_21 | 3.4 | -11.6 | 3 | -12.4 | 2 | 1.3 |
| Liang_4 | 1 | -12.3 | 1 | -12.2 | 4 | 1.1 |
| Liang_7 | 3.1 | -11.6 | 3 | -11.6 | 5 | 0.6 |
| Liang_11 | 21.8 | -10.4 | 15 | -11.5 | 5 | 0.5 |
| Liang_9 | 5.4 | -11.3 | 8 | -11.2 | 7 | 0.5 |
| Liang_8 | 7.8 | -11.1 | 9 | -11.1 | 8 | 0.5 |
| Liang_13 | 4.4 | -11.4 | 7 | -11.0 | 9 | 0.4 |
| Liang_19 | 1.8 | -11.9 | 2 | -11.0 | 9 | 0.8 |
Sulfur-containing groups are known to be more challenging for FEP calculations. Problems with forcefields and sulfur can include: (i) sulfur tends to be asymmetric, and thus point charges are arguably insufficient, (ii) sulfur is sufficiently large that cheaper QM methods, which are acceptable for deriving the parameters for groups such as CHNO, can struggle with sulfur, and (iii) if a forcefield is trying to use only one term for all sulfurs, this can cause issues as SO2 is very different to SO and S alone. Such difficulties usually present themselves as large hysteresis in FEP cycles where a sulfur is added or removed. However, in this case, the hysteresis of the cycle containing Liang_6 was acceptable (±1.1 kcal/mol) and Liang_6 had an estimated error of 1.3 kcal/mol but, with further investigation using the “Convergence” analysis in Flare FEP, we could see there was under convergence in the bound state in both directions. Hence, if we were to add an ‘instance’ to the links associated to the Liang_6 compound, we would increase the simulation time in the bound state (from 4ns to 8ns) to allow for better convergence and hopefully improve the prediction of this ligand. An example of this is shown in Figure 4.

Overall, if we were to send the top ten predicted binders as ranked by Flare FEP (those with a predicted ΔG of -11 kcal/mol or lower), eight would have been active experimentally and only two were weaker than predicted. It is worth noting that Flare FEP captured the best-known binder in the top five compounds, Liang_4, and only narrowly missed one compound, Liang_14 (experimental ΔG -11.6 kcal/mol and predicted ΔG -10.8 kcal/mol), from the top ten. This ligand was ranked at 11 in the predictive study and had an error of ±0.7 kcal/mol, so would have been up to the user as to whether they included this compound for further testing. Importantly, Flare FEP also correctly predicted the worst binder of the dataset, Liang_15 (ΔGexp= -9.8 kcal/mol, ΔGpred= -9.9 kcal/mol ±0.5 kcal/mol), demonstrating that we can reduce the number of weak binders that would go through to assays unnecessarily.
It is worth noting the small dynamic range of only 1.2 kcal/mol between the best ten binders, and an overall activity range of 2.5 kcal/mol across the whole dataset, which is slightly below the recommended value (3.0 kcal/mol) if this were a benchmark series.8 In a good FEP study, you can expect to see an error between 1-2 kcal/mol, hence the need for a dynamic range larger than this when creating a benchmark. Therefore, despite the low correlation (r2), Flare FEP has correctly predicted eight of the top ten binders from the experimental assay with errors within the accepted energetic range.
We retrospectively plotted the experimental ΔG values of the 18 ligands2 against the predicted ΔG values of the Flare FEP production run, shown in Figure 5.

Whilst the statistical output may not suggest a strong correlation between the predicted and experimental ΔG values, the ranking of the best and worst binders within the dataset was correct. Setting the precision to capture the eight ligands which Flare FEP ranked in the top ten binders (-10.9 kcal/mol, see Figure 5) the top left confusion matrix shows eight true positives, four false negatives (one of which is the benchmark compound with known activity, ejm_46), five true negatives and two false positives, i.e. there is an 80% chance of getting a true result. As previously stated, there is significant overlap in the middle of the dynamic range, and some of these compounds are associated with larger error bars (such as ligand Liang_16, Figure 5). Large error bars (i.e. those greater than 2 kcal/mol) indicate that we should be less confident in the predictions for a ligand associated with them. However, most compounds at either extreme of the plot (i.e. the best and worst binders) are well predicted, with low error. The mean unassigned error (MUE) is very good at 0.59 kcal/mol.
Troubleshooting
Flare FEP offers multiple tools to troubleshoot problematic links in both benchmark and production mode studies, including the option to add link instances without needing to rerun the full FEP network, saving computational time. Of the 27 links, only one was not included in the results by Flare, as the hysteresis of the link was too high and other link instances had been added to the ligand in question, Liang_16. This link was between compounds Liang_9 and 16 perturbing from an alcohol group to a morpholine ring respectively, without an intermediate, shown in Figure 6. The automatic intermediate generator within Flare FEP suggested the compound Liang_13 was the most appropriate intermediate (where R = CH2NH2) and hence, did not include a secondary one upon network generation.
Liang_16 also had a high error range of ±2.5 kcal/mol. In a case such as this, (i.e. high error and hysteresis) the user may wish to take ligand Liang_16, as it could be very active or not active at all, and measure it experimentally as the prediction isn’t robust, but it may be chemistry interesting to the user. However, this is obviously project specific.

Within Flare FEP you can color links by their percentage contribution to total error. Here, we can quickly identify links which may benefit from an additional instance if they also show some hysteresis, such as the link between Liang_7 and 22 seen in Figure 7.
In the case of this experiment, we chose not to add further instances to any of the links within the perturbation map as this was a production run rather than a benchmark study where we may spend more time fine-tuning experimental setup. Here, we wanted to ascertain if the production mode was able to provide “good or bad” robust predictions to these compounds with the parameters used from the benchmark (in this instance, default settings with the use of GCNCMC during equilibration). We could also use the “Enhance Connectivity” tool in Flare FEP to potentially add further links to weaker areas of the map.

In a live discovery project, the 18 compounds would have been new designs based on the benchmark scaffold, exploring a new area of the active site, potentially with different protein-ligand contacts being made. Therefore, in relative FEP we assume, despite the changes in the new design, that the same binding mode and geometry is conserved from benchmark to production mode and hence, this may account for some hysteresis in the predictions of the production run. We wish to emphasize that FEP is to be used as part of a wider testing process and discovery workflow. For example, the user may suggest the top ten binders and one ligand with more ambiguous predicted activity from the middle of the dataset. These molecules would go into an assay to obtain their experimental activities and potentially take one or two for further crystallographic studies to determine any protein rearrangement, confirming what is seen in modelling. Even as part of a larger workflow, FEP still offers the opportunity to triage a large set of congeneric compounds, identifying a subset which should go onto further testing and thus, greatly reducing the time and money spent on synthesizing weak binders.
Furthermore, within Flare, the user could investigate the best and worst predicted binders of the dataset to model, and thus better understand why these compounds have different activities. For example, both GIST and 3D-RISM7,9 are water analysis tools in Flare which can provide more information as to the activity or contacts observed in an FEP calculation. To learn more about these techniques, please check out our webinar Understanding water behavior for enhancing drug design. Focusing on the worst binder of the dataset, Liang_15, we could transfer the equilibrated complex back to the main Flare project to learn why it is predicted as the worst binder, investigating the binding pose and different torsions/ conformations it may adopt (in an MD simulation for example). Such work can inform the next round of designs by providing further understanding why such a functional group did not improve activity. In the case of Liang_15, its reduced activity is likely due to the R-group (CH2N(CH3)2) being unable to make interactions in the water exposed region where the substitutions are occurring.
Conclusions
In this case study we show how FEP could be used to help prioritize designs in a project, focusing on intrinsic activity (or binding), reducing the number of weak compounds (or poor binders) made within a project. Such a technique can increase the progress of a project by aiding the prioritization of interesting compounds to move a project towards its activity goals. Of course, you can also use it to promote designs that attempt to address issues other than intrinsic target activity. For example, using such a method can suggest whether to explore a compound with particularly challenging chemistry, i.e. is the predicted activity/ binding good enough to push forward into synthesis, such as Liang_4 or, is it worth exploring why some ligands were worse binders to better inform future designs, such as Liang_15. Here we have shown a real use case of Flare FEP in a production run, suggested how such results can be analyzed and subsequently used in a larger lead optimization workflow.
References
- Kuhn M., et al. Assessment of Binding Affinity via Alchemical Free-Energy Calculations. J. Chem. Inf. Model. 2020, 60, 6, 3120–3130. https://doi.org/10.1021/acs.jcim.0c00165
- Liang J., et al. Lead Optimization of a 4‑Aminopyridine Benzamide Scaffold To Identify Potent, Selective, and Orally Bioavailable TYK2 Inhibitors, J. Med. Chem. 2013, 56, 4521−4536. https://dx.doi.org/10.1021/jm400266t
- Liang J., et al. Lead identification of novel and selective TYK2 inhibitors, European Journal of Medicinal Chemistry, 2013, 67, 175 – 187. https://doi.org/10.1016/j.ejmech.2013.03.070
- Ramakrisna C., et al. Tyrosine kinase 2 inhibitors in autoimmune diseases, Autoimmunity Reviews, 2024, 23 (11), 103649, ISSN 1568-9972, https://doi.org/10.1016/j.autrev.2024.103649.
- Wang L., et al. Accurate and Reliable Prediction of Relative Ligand Binding Potency in Prospective Drug Discovery by Way of a Modern Free-Energy Calculation Protocol and Force Field, J. Am. Chem. Soc. 2015, 137, 2695−2703. DOI: 10.1021/ja512751q
- Kuhn, M. et al.; Assessment of Binding Affinity via Alchemical Free-Energy Calculations, J. Chem. Inf. Model. 2020, 60, 6, 3120–3130.
- Luchko, T. et al.; Three-Dimensional Molecular Theory of Solvation Coupled with Molecular Dynamics in Amber, J. Chem. Theory Comput. 2010, 6 (3), 607–624
- Hahn, D. et al.; Best Practices for Constructing, Preparing, and Evaluating Protein-Ligand Binding Affinity Benchmarks [Article v1.0]. Living Journal of Computational Molecular Science. 2022, 4 (1), 1497. DOI: 10.33011/livecoms.4.1.1497
- Vinter, J.G. Extended electron distributions applied to the molecular mechanics of some intermolecular interactions. J. Computer-Aided Mol. Des. 1994, 8, 653–668. DOI: 10.1007/BF00124013.















