Ensure novel ideas for your project with the new Spark databases

SHARE

To accompany the release of Spark™ V10.6, the Spark fragment and reagent databases have been updated and are now available for download. Derived by fragmenting compounds and reagents from commercial sources and the literature, these database are a great source of novel ideas for your drug discovery projects, ensuring at the same time that the results found by Spark are always associated to real, synthetically accessible compounds.

Fragment databases

The Spark ‘Commercial’ databases in this release are derived from the eMolecules Screening Compounds. With more than 6 million fragments to search overall, they provide an excellent source of chemical diversity for your experiments.

The Spark ‘ChEMBL’ databases have also been updated. Based on release 26 of ChEMBL, they provide more than 1.5 million additional fragments to search, derived from chemical literature compounds.

Compounds in both original source collections are filtered to remove molecules containing potentially toxic or reactive groups before the creation of the databases. Each compound is then fragmented independently, breaking the bonds which connect to heteroatoms, carbonyls, thiocarbonyls and bonds to rings. Specific functional groups such as carboxylic acids, nitro groups and rings are not fragmented. The frequency with which a given fragment occurs is captured together with the number of bonds that were broken to disconnect the fragment from the parent molecule.

All resultant fragments are subject to molecular weight, number of H-bond acceptor/donor and rotatable bond limits. They are then sorted by frequency and labelled as shown in Table 1.

Table 1: Fragment databases sorted by frequency.

Spark categoryDatabaseTotal number of fragments (to nearest 1,000)Frequency
CommercialVery Common68,000Fragments which appear in more than 725 molecules
Common68,000Fragments which appear in 215-724 molecules
Less Common212,000Fragments which appear in 65-214 molecules
Rare280,000Fragments which appear in 25-64 molecules
Very Rare527,000Fragments which appear in 9-24 molecules
Extremely Rare534,000Fragments which appear in 5-8 molecules
Ultra Rare770,000Fragments which appear in 3-4 molecules
Doubleton*1,053,000Fragments which appear in 2 molecules
Singleton*2,526,000Fragments which appear in a single molecule
ChEMBLCommon232,000Fragments which appear in more than 12 molecules
Rare232,000Fragments which appear in 4-12 molecules
Very Rare382,000Fragments which appear in 2-3 molecules
Extremely Rare*641,000Fragments which appear in a single molecule

*Contact us for further details.

Typically we would recommend to install only the databases including fragments which appear at least 3-4 times in the original collections. The databases containing fragments seen with lower frequency (Singleton, Doubleton and ChEMBL Extremely Rare) are very large, and may contain fragments derived from unrealistic/wrong structures in the original collections. If you do wish to use these databases then please contact Cresset Support for download instructions.

The number of fragments in each database per connection point count (excluding the databases containing only singletons and doubletons) is shown in Figure 1.

Counts of fragments in Spark databases

Figure 1: Count of fragments in Spark ‘Commercial’ and ‘ChEMBL’  databases split by the number of connection points of each fragment.

The most common fragments in the ChEMBL and Commercial databases have a significant overlap (Table 2). However, comparing the rarer fragments from each database shows significantly less overlap, highlighting the different areas of chemical space each database occupies.

Table 2: Overlap of the most common fragments in the ChEMBL and Commercial databases.

% overlapVery CommonCommonLess CommonRareVery RareExtremely RareUltra RareDoubleton*Singleton*Unique
ChEMBL common17%14%13%8%8%4%4%3%5%24%
ChEMBL rare2%6%9%9%10%6%5%4%6%43%
ChEMBL very rare1%2%4%5%7%5%5%5%7%58%
ChEMBL extremely rare*0%1%2%3%5%4%4%4%9%68%

*Contact us for further details.

With more than 6.9 million unique fragments to search, the Spark fragment databases provide an extremely large source of novel bioisosteres for Spark projects, which can be further complemented by generating fragments from your corporate collection with the Spark Database Generator, a dedicated and user-friendly interface to custom database creation within Spark.

Reagent databases

Monthly updates of the Spark reagent databases, derived from the eMolecules building blocks using an enhanced set of rules for chemical transformation, are included in the Spark V10.6 release. The November update includes over 314,000 reagents with up-to-date availability information, to make it easy for you to order the reagents you require to synthesize your favorite Spark results.

 Total1-5051-100101-150151-200201-250
eMolecules_acidCO23,98334016,73213,3613,486
eMolecules_acid41,545432,81115,61817,9395,134
eMolecules_alcohol18,032111,4357,5217,1931,872
eMolecules_alcoholO19,63434686,7739,6662,724
eMolecules_aliphatic_halide8,808139243,6513,421799
eMolecules_alkyne2,851275051,420781118
eMolecules_aromatic_alcoholO8,6250441,9275,0231,631
eMolecules_aromatic_aminesN18,55701114,20710,5673,672
eMolecules_aromatic_halide40,110845113,59222,7623,297
eMolecules_boronic4,49601281,8942,093381
eMolecules_cyano15,118201,0865,6626,2832,067
eMolecules_isocyanateCO55502017028778
eMolecules_olefin3,273165241,4191,089225
eMolecules_primary_aliphatic_amine19,01661,3668,4957,7631,386
eMolecules_primary_aliphatic_amineN11,57103985,2345,101838
eMolecules_primary_aliphatic_halide6,875126272,8862,705645
eMolecules_primary_aromatic_amines23,35003256,58112,1714,273
eMolecules_reductive_amination22,12738186,55110,6834,072
eMolecules_secondary_aliphatic_amineN15,06112774,2708,4132,100
eMolecules_sulfonicacid5,066316022,2651,761407
eMolecules_sulfonicacidSO23,0750133021,5841,176
eMolecules_thiol721720633016414
eMolecules_thiolS1,9861385371,078332

In the Spark results table, the eMolecules IDs for your favorite reagents can be easily exported from Spark and used to purchase the compounds from the eMolecules building blocks database, as shown in the web clip How to use the eMolecules reagents databases in Spark and access ordering information for the result.

Crystallographic fragments database

Spark V10.6 also includes the new ‘COD’ database (Figure 2).  This contains more than 440K fragments in their crystallographic conformation, derived from the Crystallography Open Database and available for download to all Spark customers.

New COD databases

Figure 2: The new ‘COD’ database is available to all Spark customers and includes more than 440K fragments in their crystallographic conformation.

Create your own Spark databases

If you have access to large collections of proprietary chemistry or specialized reagents, or if you want to only consider fragments from reagents you have in stock, you can add value to your Spark experiments by creating your own custom databases.

These can be easily prepared using the Database Generator (Figure 3), a dedicated and user-friendly interface to custom database creation within Spark, or using the equivalent functionality from the command line.

Spark database generator

Figure 3: Use the Spark Database Generator to create your own fragments and reagent databases.

Conclusion

This new release of the fragment and reagent databases, combined with custom databases from corporate collections generated with the Spark Database Generator, will provide an outstanding range of bioisosteres for your project.

Spark and your project

Please contact us to update to the latest databases, to learn how to make the best use of the Spark Database Generator, or to find out how Spark can impact your project.

Subscribe & Don't Miss Out

Receive our newsletter to be among the first to hear about product releases, case studies, opinion articles, events and more.