The currently released databases for Spark are derived from commercially available screening compounds (eMolecules screening compounds1), literature reports (ChEMBL2), patents (SureChEMBL2), commercial reagents (eMolecules building blocks3), small molecule crystal structures (Crystallography Open Database4 and Cambridge Structural Database5), theoretical ring systems (VEHICLe6), degrader linkers from experimental sources (PROTAC-DB libary7 and Enamine Linkers for Linkerology8), and agrochemicals (from PubChem9).
About 30 million of conformationally enriched fragments are available from the Cresset released databases for Spark. Fragments in the Spark database have up to four connection points.
| Database | Total frags | Frags with 1 connection | Frags with 2 connections | Frags with 3 connections | Frags with 4 connections | Rings only frags |
|---|---|---|---|---|---|---|
| Agrochemical | 10,976 | 2,245 | 3,973 | 3,204 | 1,554 | 235 |
| COD | 448,712 | 71,990 | 135,465 | 140,826 | 100,431 | 11,460 |
| CSD | 1,217,378 | 159,086 | 343,554 | 397,447 | 317,291 | 27,879 |
| ChEMBL | 2,160,380 | 452,482 | 769,882 | 617,913 | 320,103 | 32,094 |
| Commercial | 6,535,138 | 2,373,812 | 2,334,301 | 1,310,505 | 516,520 | 37,384 |
| SureChEMBL | 18,962,454 | 3,948,439 | 6,126,308 | 5,495,765 | 3,391,942 | 213,389 |
| Linkers – PROTAC – DB | 1,398 | 0 | 1,398 | 0 | 0 | 6 |
| Linkers – Enamine | 8,291 | 0 | 8,291 | 0 | 0 | 22 |
| VEHICLe | 170,467 | 57,528 | 61,586 | 37,217 | 14,136 | 35,206 |
See below for fragments from Reagents.
Update to the latest version of all databases by following the instructions at Installing Spark databases.
Fragments from screening compounds
The Commercial Spark databases are based on eMolecules screening compounds and are split based on the frequency of occurrence of the fragments.
- VeryCommon (309 MB) – fragments which appear in at least 725 molecules
- Common (577 MB) – fragments which appear in 215-724 molecules
- LessCommon (1.2 GB) – fragments which appear in 65-214 molecules
- Rare (1.6 GB) – fragments which appear in 25-64 molecules
- VeryRare (3.1 GB) – fragments which appear in 9-24 molecules
- ExtremelyRare (3.3 GB) – fragments which appear in 5-8 molecules
- UltraRare (4.9 GB) – fragments which appear in 3-4 molecules
In general, fragments from the VeryCommon or Common databases are more likely to be readily synthesizable as they appear in many different commercially available molecules. Fragments from the VeryRare, ExtremelyRare and UltraRare databases are more likely to be non-drug-like or hard to make. These databases have been filtered to remove potentially toxic or reactive fragments (such as alkyl halides or nitroso functionalities). In addition, all phosphorus-containing fragments have been removed as the calculation of fields on phosphorus-containing functional groups is still under development. See below for a detailed analysis of these databases.
Two optional databases are also available:
- Doubleton (2 files, each of 3.3 GB) – fragments which appear in two molecules
- Singleton (3 files, each of 8.2 GB) – fragments which appear in a single molecule
| Database | Total Frags | Frags with 1 connection | Frags with 2 connections | Frags with 3 connections | Frags with 4 connections | Rings Only |
|---|---|---|---|---|---|---|
| VeryCommon | 66,617 | 20,580 | 28,664 | 14,137 | 3,236 | 1,932 |
| Common | 104,026 | 33,102 | 42,869 | 22,389 | 5,666 | 1,431 |
| LessCommon | 188,210 | 54,019 | 79,608 | 42,081 | 12,502 | 1,850 |
| Rare | 250,974 | 70,752 | 104,041 | 57,634 | 18,547 | 1,946 |
| VeryRare | 480,125 | 137,352 | 189,072 | 113,157 | 40,544 | 3,090 |
| ExtremelyRare | 493,723 | 151,993 | 187,182 | 112,154 | 42,394 | 2,974 |
| UltraRare | 734,206 | 244,128 | 268,729 | 159,734 | 61,615 | 3,752 |
| Doubleton | 957,535 | 335,484 | 338,416 | 200,014 | 83,621 | 4,426 |
| Singleton | 3,259,722 | 1,326,402 | 1,095,720 | 589,205 | 248,395 | 15,983 |
Typically we would recommend to install only the databases including fragments which appear at least 3-4 times in the original collections. The databases containing fragments seen with lower frequency are very large, and may contain fragments derived from unrealistic/wrong structures in the original collections. Contact support if you wish to download these optional databases.
Fragments from ChEMBL
The current ChEMBL Spark databases are based on Release 34 of ChEMBL and are split based on the frequency of occurrence of the fragments.
- ChEMBL_common (1.7 GB) – fragments which appear in more than 12 molecules
- ChEMBL_rare (2.4 GB) – fragments which appear in 4-12 molecules
- ChEMBL_veryrare (3.2 GB) – fragments which appear in 2-3 molecules
An optional database is also available (contact Cresset support to download this database):
- ChEMBL_extremelyrare (6.7 GB) – fragments which appear once
| Database | Total Frags | Frags with 1 connection | Frags with 2 connections | Frags with 3 connections | Frags with 4 connections | Rings Only |
|---|---|---|---|---|---|---|
| ChEMBL_common | 288,335 | 54,717 | 109,917 | 85,376 | 38,325 | 7,122 |
| ChEMBL_rare | 371,185 | 74,342 | 133,992 | 108,758 | 54,093 | 6,006 |
| ChEMBL_veryrare | 497,783 | 101,953 | 175,864 | 144,551 | 75,415 | 6,850 |
| ChEMBL_extremelyrare | 1,003,077 | 221,470 | 350,109 | 279,228 | 152,270 | 12,116 |
Fragments from SureChEMBL
The SureChEMBL Spark databases are based on the complete SureChEMBL compound collection and are split based on the frequency of occurrence of the fragments.
- SureChEMBL_verycommon (4.1 GB) – fragments which appear in at least 45 molecules
- SureChEMBL_common (6.9 GB) – fragments which appear in 14-44 molecules
- SureChEMBL_uncommon (7.1 GB) – fragments which appear in 8-13 molecules
Three optional databases are also available:
- SureChEMBL_rare (9.2 GB) – fragments which appear in 5-7 molecules
- SureChEMBL_veryrare (7.7 GB) – fragments which appear in 4 molecules
- SureChEMBL_extremelyrare (9.5 GB) – fragments which appear in 3 molecules
Additional optional databases are also available on-demand (contact Cresset support to download these databases):
- SureChEMBL_doubleton (2 files, each of 9.2 GB) – fragments which appear twice
- SureChEMBL_singleton (3 files, each of 9.6 GB) – fragments which appear once
| Database | Total Frags | Frags with 1 connection | Frags with 2 connections | Frags with 3 connections | Frags with 4 connections | Rings Only |
|---|---|---|---|---|---|---|
| SureChEMBl_verycommon | 681,782 | 117,432 | 246,422 | 211,113 | 106,815 | 14,569 |
| SureChEMBl_common | 1,066,501 | 196,423 | 367,648 | 323,807 | 178,623 | 14,873 |
| SureChEMBl_uncommon | 1,066,063 | 207,262 | 358,220 | 317,503 | 183,078 | 12,642 |
| SureChEMBl_rare | 1,338,922 | 283,978 | 449,151 | 385,655 | 220,138 | 15,895 |
| SureChEMBl_veryrare | 1,113,372 | 237,491 | 362,024 | 320,679 | 193,178 | 11,099 |
| SureChEMBl_extremelyrare | 1,345,030 | 308,443 | 442,005 | 374,426 | 220,156 | 15,565 |
| SureChEMBl_doubleton | 2,650,383 | 627,519 | 865,648 | 725,721 | 431,495 | 29,051 |
| SureChEMBl_singleton | 3,525,009 | 671,186 | 1,084,771 | 1,055,570 | 713,482 | 35,322 |
Reagents
Spark Reagent Databases are derived from eMolecules building blocks using the Cresset reagent importer, which converts a file of usable reagents into the corresponding R-group. For example, to create the eMolecules_acid database, all the eMolecules building blocks containing a C(=O)OH or C(=O)Cl group were processed to add the R-group to the database.
Using databases derived from available reagents ensures that the results of your Spark experiment are tethered to molecules that are readily synthetically accessible. Monthly updates for these databases provide reliable availability information on the reagents that you wish to employ.
The current list of Spark Reagents databases includes 23 common chemical transformations.
Fragments from linker degraders
Spark Linkers Databases are derived from the PROTAC-DB library,7 from the Tingjun Hou group, and Linkers for Linkerology from Enamine.8 The linker protecting groups, like alcohols, were replaced by Te atoms and linker duplicates removed within a collection. The Linkers databases were created using the ‘Pre-labelled’ fragmentation mechanism from the list Te-lablled molecules. All fragments in the Linkers databases have two connection points. Because the fragments in the Linkers databases will be larger and more flexible than the fragments in the small molecule databases; during the database creation the threshold for maximum molecular weight, number of heavy atoms and rotatable bonds in a fragment was increased and set to 9999, 99, and 99 respectively. The maximum number of conformations for a fragment was also increased to 10000. The linkers in the MADE collection were split into ‘Aromatic’ and ‘Aliphatic’ based on the presence or absence of an aromatic ring substructure in the fragment.
- PROTAC-DB (1.1 MB) – 1,398 linkers
- Enamine Stock (154 MB) – 2,453 linkers
- Enamine MADE Aromatic (174 MB) – 2,708 linkers
- Enamine MADE Aliphatic (520 GB)– 3,130 linkers
Using databases derived from experimentally derived sources ensure that the results of your Spark experiment are tethered to molecules that are synthetically accessible.
Fragments from Agrochemicals
Spark Agrochemical database is derived from the ‘Agrochemical’ classification of the open chemistry database, PubChem.9 The records in the ‘Agrochemical’ classification relate to agrochemicals used in agriculture, including chemical fertilizers, pesticides and more. Prior fragmenting the PubChem compounds:
i) 1,3 diones were enumerated and converted to the enolate form,

ii) all procides were removed,

and iii) the following known, a-c, transformations were applied:

The Agrochemical database was created using the ‘Molecules’ fragmentation mechanism with default settings. Fragments in the Agrochemical database have up to four connection points, including ring-only fragments.
- PubChem (56 MB) – 10,976 fragments
Using an agrochemical-focussed fragment database ensures that the results of your Spark experiment have a property profile suitable for applications in the agrochemical sector.
Fragments from small molecules crystal structures
These databases contain fragments in their crystallographic conformation, derived from small molecule crystal structures.
The Spark COD database contains fragments from the Crystallography Open Database. This database is available for download to all Spark customers.
- COD (799 MB) – 448,712 fragments
The Spark CSD Fragment Database is derived from the Cambridge Structural Database (CSD). A valid CSD-System license is required for use of this database. If you do not already have a license, please contact CCDC for assistance.
- CSD (2.4 GB) – 1,217,268 fragments
Theoretical rings
A collection of theoretical ring systems derived from the VEHICLe6 database.
- VEHICLe (250 MB) – 170,467 ring fragments
Create your own database
Spark fragment and reagent databases provide an excellent source of new bioisosteres. However, if you have access to significant proprietary chemistry, to specialized reagents, or simply want to only consider fragments from reagents that you have in stock then creating your own custom databases will add value to your Spark experiments.
Custom databases can be easily created using the Database Generator, a dedicated and user-friendly interface to custom database creation within Spark, or using the equivalent functionality from the command line.
Contact Cresset support if you need assistance with the Spark Database Generator or require additional chemical transformations for your reagent database generation.
References
- https://www.emolecules.com/products/screening-compounds
- https://www.ebi.ac.uk/chembl/
- https://www.emolecules.com/products/building-blocks
- http://www.crystallography.net/cod/
- https://www.ccdc.cam.ac.uk/solutions/csd-system/components/csd/
- Pitt, W. R.; Parry, D. M.; Perry, B. G.; Groom, C. R. Heteroaromatic Rings of the Future. J. Med. Chem. 2009, 52 (9), 2952–2963 https://doi.org/10.1021/jm801513z
- http://cadd.zju.edu.cn/protacdb/
- https://enamine.net/building-blocks-mob/linkers-for-linkerology
- https://pubchem.ncbi.nlm.nih.gov/