Currently released Spark™ databases

The currently released databases for Spark are derived from commercially available screening compounds (eMolecules screening compounds1), literature reports (ChEMBL2), patents (SureChEMBL2), commercial reagents (eMolecules building blocks3), small molecule crystal structures (Crystallography Open Database4 and Cambridge Structural Database5), theoretical ring systems (VEHICLe6), degrader linkers from experimental sources (PROTAC-DB libary7 and Enamine Linkers for Linkerology8), and agrochemicals (from PubChem9).

About 30 million of conformationally enriched fragments are available from the Cresset released databases for Spark. Fragments in the Spark database have up to four connection points.

DatabaseTotal fragsFrags with 1 connectionFrags with 2 connectionsFrags with 3 connectionsFrags with 4 connectionsRings only frags
Agrochemical10,9762,2453,9733,2041,554235
COD448,71271,990135,465140,826100,43111,460
CSD1,217,378159,086343,554397,447317,29127,879
ChEMBL2,160,380452,482769,882617,913320,10332,094
Commercial6,535,1382,373,8122,334,3011,310,505516,52037,384
SureChEMBL18,962,4543,948,4396,126,3085,495,7653,391,942213,389
Linkers – PROTAC – DB1,39801,398006
Linkers – Enamine8,29108,2910022
VEHICLe170,46757,52861,58637,21714,13635,206

See below for fragments from Reagents.

Update to the latest version of all databases by following the instructions at Installing Spark databases.

Fragments from screening compounds

The Commercial Spark databases are based on eMolecules screening compounds and are split based on the frequency of occurrence of the fragments.

  • VeryCommon (309 MB) – fragments which appear in at least 725 molecules
  • Common (577 MB) – fragments which appear in 215-724 molecules
  • LessCommon (1.2 GB) – fragments which appear in 65-214 molecules
  • Rare (1.6 GB) – fragments which appear in 25-64 molecules
  • VeryRare (3.1 GB) – fragments which appear in 9-24 molecules
  • ExtremelyRare (3.3 GB) – fragments which appear in 5-8 molecules
  • UltraRare (4.9 GB) – fragments which appear in 3-4 molecules

In general, fragments from the VeryCommon or Common databases are more likely to be readily synthesizable as they appear in many different commercially available molecules. Fragments from the VeryRare, ExtremelyRare and UltraRare databases are more likely to be non-drug-like or hard to make. These databases have been filtered to remove potentially toxic or reactive fragments (such as alkyl halides or nitroso functionalities). In addition, all phosphorus-containing fragments have been removed as the calculation of fields on phosphorus-containing functional groups is still under development. See below for a detailed analysis of these databases.

Two optional databases are also available:

  • Doubleton (2 files, each of 3.3 GB) – fragments which appear in two molecules
  • Singleton (3 files, each of 8.2 GB) – fragments which appear in a single molecule
DatabaseTotal FragsFrags with 1 connectionFrags with 2 connectionsFrags with 3 connectionsFrags with 4 connectionsRings Only
VeryCommon66,61720,58028,66414,1373,2361,932
Common104,02633,10242,86922,3895,6661,431
LessCommon188,21054,01979,60842,08112,5021,850
Rare250,97470,752104,04157,63418,5471,946
VeryRare480,125137,352189,072113,15740,5443,090
ExtremelyRare493,723151,993187,182112,15442,3942,974
UltraRare734,206244,128268,729159,73461,6153,752
Doubleton957,535335,484338,416200,01483,6214,426
Singleton3,259,7221,326,4021,095,720589,205248,39515,983

Typically we would recommend to install only the databases including fragments which appear at least 3-4 times in the original collections. The databases containing fragments seen with lower frequency are very large, and may contain fragments derived from unrealistic/wrong structures in the original collections. Contact support if you wish to download these optional databases.

Fragments from ChEMBL

The current ChEMBL Spark databases are based on Release 34 of ChEMBL and are split based on the frequency of occurrence of the fragments.

  • ChEMBL_common (1.7 GB) – fragments which appear in more than 12 molecules
  • ChEMBL_rare (2.4 GB) – fragments which appear in 4-12 molecules
  • ChEMBL_veryrare (3.2 GB) – fragments which appear in 2-3 molecules

An optional database is also available (contact Cresset support to download this database):

  • ChEMBL_extremelyrare (6.7 GB) – fragments which appear once
DatabaseTotal FragsFrags with 1 connectionFrags with 2 connectionsFrags with 3 connectionsFrags with 4 connectionsRings Only
ChEMBL_common288,33554,717109,91785,37638,3257,122
ChEMBL_rare371,18574,342133,992108,75854,0936,006
ChEMBL_veryrare497,783101,953175,864144,55175,4156,850
ChEMBL_extremelyrare1,003,077221,470350,109279,228152,27012,116

Fragments from SureChEMBL

The SureChEMBL Spark databases are based on the complete SureChEMBL compound collection and are split based on the frequency of occurrence of the fragments.

  • SureChEMBL_verycommon (4.1 GB) – fragments which appear in at least 45 molecules
  • SureChEMBL_common (6.9 GB) – fragments which appear in 14-44 molecules
  • SureChEMBL_uncommon (7.1 GB) – fragments which appear in 8-13 molecules

Three optional databases are also available:

  • SureChEMBL_rare (9.2 GB) – fragments which appear in 5-7 molecules
  • SureChEMBL_veryrare (7.7 GB) – fragments which appear in 4 molecules
  • SureChEMBL_extremelyrare (9.5 GB) – fragments which appear in 3 molecules

Additional optional databases are also available on-demand (contact Cresset support to download these databases):

  • SureChEMBL_doubleton (2 files, each of 9.2 GB) – fragments which appear twice
  • SureChEMBL_singleton (3 files, each of 9.6 GB) – fragments which appear once
DatabaseTotal FragsFrags with 1 connectionFrags with 2 connectionsFrags with 3 connectionsFrags with 4 connectionsRings Only
SureChEMBl_verycommon681,782117,432246,422211,113106,81514,569
SureChEMBl_common1,066,501196,423367,648323,807178,62314,873
SureChEMBl_uncommon1,066,063207,262358,220317,503183,07812,642
SureChEMBl_rare1,338,922283,978449,151385,655220,13815,895
SureChEMBl_veryrare1,113,372237,491362,024320,679193,17811,099
SureChEMBl_extremelyrare1,345,030308,443442,005374,426220,15615,565
SureChEMBl_doubleton2,650,383627,519865,648725,721431,49529,051
SureChEMBl_singleton3,525,009671,1861,084,7711,055,570713,48235,322

Reagents

Spark Reagent Databases are derived from eMolecules building blocks using the Cresset reagent importer, which converts a file of usable reagents into the corresponding R-group. For example, to create the eMolecules_acid database, all the eMolecules building blocks containing a C(=O)OH or C(=O)Cl group were processed to add the R-group to the database.

Using databases derived from available reagents ensures that the results of your Spark experiment are tethered to molecules that are readily synthetically accessible. Monthly updates for these databases provide reliable availability information on the reagents that you wish to employ.

The current list of Spark Reagents databases includes 23 common chemical transformations.

Fragments from linker degraders

Spark Linkers Databases are derived from the PROTAC-DB library,7 from the Tingjun Hou group, and Linkers for Linkerology from Enamine.8 The linker protecting groups, like alcohols, were replaced by Te atoms and linker duplicates removed within a collection. The Linkers databases were created using the ‘Pre-labelled’ fragmentation mechanism from the list Te-lablled molecules. All fragments in the Linkers databases have two connection points. Because the fragments in the Linkers databases will be larger and more flexible than the fragments in the small molecule databases; during the database creation the threshold for maximum molecular weight, number of heavy atoms and rotatable bonds in a fragment was increased and set to 9999, 99, and 99 respectively. The maximum number of conformations for a fragment was also increased to 10000. The linkers in the MADE collection were split into ‘Aromatic’ and ‘Aliphatic’ based on the presence or absence of an aromatic ring substructure in the fragment.

  • PROTAC-DB (1.1 MB) – 1,398 linkers
  • Enamine Stock (154 MB) – 2,453 linkers
  • Enamine MADE Aromatic (174 MB) – 2,708 linkers
  • Enamine MADE Aliphatic (520 GB)– 3,130 linkers

Using databases derived from experimentally derived sources ensure that the results of your Spark experiment are tethered to molecules that are synthetically accessible.

Fragments from Agrochemicals

Spark Agrochemical database is derived from the ‘Agrochemical’ classification of the open chemistry database, PubChem.9 The records in the ‘Agrochemical’ classification relate to agrochemicals used in agriculture, including chemical fertilizers, pesticides and more. Prior fragmenting the PubChem compounds:

i) 1,3 diones were enumerated and converted to the enolate form,

ii) all procides were removed,

and iii) the following known, a-c, transformations were applied:

The Agrochemical database was created using the ‘Molecules’ fragmentation mechanism with default settings. Fragments in the Agrochemical database have up to four connection points, including ring-only fragments.

  • PubChem (56 MB) – 10,976 fragments

Using an agrochemical-focussed fragment database ensures that the results of your Spark experiment have a property profile suitable for applications in the agrochemical sector.

Fragments from small molecules crystal structures

These databases contain fragments in their crystallographic conformation, derived from small molecule crystal structures.

The Spark COD database contains fragments from the Crystallography Open Database. This database is available for download to all Spark customers.

  • COD (799 MB) – 448,712 fragments

The Spark CSD Fragment Database is derived from the Cambridge Structural Database (CSD). A valid CSD-System license is required for use of this database. If you do not already have a license, please contact CCDC for assistance.

  • CSD (2.4 GB) – 1,217,268 fragments

Theoretical rings

A collection of theoretical ring systems derived from the VEHICLe6 database.

  • VEHICLe (250 MB) – 170,467 ring fragments

Create your own database

Spark fragment and reagent databases provide an excellent source of new bioisosteres. However, if you have access to significant proprietary chemistry, to specialized reagents, or simply want to only consider fragments from reagents that you have in stock then creating your own custom databases will add value to your Spark experiments.

Custom databases can be easily created using the Database Generator, a dedicated and user-friendly interface to custom database creation within Spark, or using the equivalent functionality from the command line.

Contact Cresset support if you need assistance with the Spark Database Generator or require additional chemical transformations for your reagent database generation.

References

  1. https://www.emolecules.com/products/screening-compounds
  2. https://www.ebi.ac.uk/chembl/
  3. https://www.emolecules.com/products/building-blocks
  4. http://www.crystallography.net/cod/
  5. https://www.ccdc.cam.ac.uk/solutions/csd-system/components/csd/
  6. Pitt, W. R.; Parry, D. M.; Perry, B. G.; Groom, C. R. Heteroaromatic Rings of the Future. J. Med. Chem. 2009, 52 (9), 2952–2963 https://doi.org/10.1021/jm801513z
  7. http://cadd.zju.edu.cn/protacdb/
  8. https://enamine.net/building-blocks-mob/linkers-for-linkerology
  9. https://pubchem.ncbi.nlm.nih.gov/