# Open Targets Platform

![](/files/-MYK6gyfoxFZaPp0Y9LX)

Welcome!

The [Open Targets Platform](https://platform.opentargets.org) is a comprehensive tool that supports systematic identification and prioritisation of potential therapeutic drug targets.

By integrating publicly available datasets including data generated by the [Open Targets consortium](https://www.opentargets.org/), the Platform builds and scores target-disease associations to assist in drug target identification and prioritisation. It also integrates relevant annotation information about targets, diseases/phenotypes, drugs, variants, studies, and credible sets as well as their most relevant relationships.

The Platform is a freely available resource that is actively maintained with quarterly updates. Our [data can be accessed](/data-access) through an intuitive web user interface, an API, and data downloads. Likewise, our pipeline and infrastructure codebases are open-source and can be used to create a self-hosted private instance of the Platform with custom data. For more information, please review our [Licence documentation](/licence) and, if you use our data and/or pipelines, please cite [our latest publication](/citation).

Check out [our blog](https://blog.opentargets.org) to learn more about the Platform and the Open Targets research programme.

You can also join the [Open Targets Community](https://community.opentargets.org) and follow us on:

* LinkedIn: [Open Targets](https://www.linkedin.com/company/open-targets/)
* Bluesky: [@opentargets.org](https://bsky.app/profile/opentargets.org)
* X: [@opentargets](https://twitter.com/opentargets)
* YouTube: [Open Targets](https://www.youtube.com/opentargets)

For additional help with the Open Targets Platform, or to report bugs, data issues, or submit a feature request, please post on the [Open Targets Community](https://community.opentargets.org/t/welcome-to-the-open-targets-community-forum/838), using the relevant categories and tags. If the request you would like to make has already been posted, please like the post to indicate you would like this to be prioritised.

{% embed url="<https://www.youtube.com/watch?v=cuutCYPiYuw>" %}


# Getting started

Identifying evidence implicating drug targets with diseases or phenotypes constitutes one of the pivotal challenges of the Open Targets Platform. Several steps are required to aggregate, integrate, validate, and score the collected information, in order to contextualise and weight the underlying evidence. The Platform provides a framework to allow the interactive or programmatic interrogation of target-disease evidence.

### Data model

The Open Targets data model focuses on five main entities:

* [Target](/target): understood as any candidate drug binding molecule
* [Disease or Phenotype](/disease-or-phenotype): including any disease indications, phenotypes, measurements, biological processes and other relevant traits.
* [Variant](/variant): DNA variation that has been associated with a disease, trait, or phenotype
* [Study](/study): source of evidence linking genetic variants to traits, diseases, and molecular phenotypes.
* [Drug](/drug): molecules that can act as medicinal products.

### Entity annotation

These entities are annotated with a variety of public data sources, as well as information about the relationships between them. A more detailed description of the annotations is provided throughout the documentation. Some examples are:

* [Target tractability assessment](/target/tractability)
* [Target safety](/target/safety)
* [Baseline expression](/target/baseline-expression)
* [Molecular interactions](/target/molecular-interactions)
* [Clinical signs and symptoms](/disease-or-phenotype/clinical-signs-and-symptoms)
* [Pharmacovigilance](/drug/pharmacovigilance)
* [Bibliography](/bibliography)
* [GWAS & functional genomics](/gentropy)

### Evidence generation and association scoring

A pivotal component of the Platform is the integration of potentially causal evidence linking targets and diseases. The definition of the Platform evidence as well as expanded documentation on each of the data sources are available in the [Target - disease evidence](/evidence) section.

In order to contextualise the information, all evidence referring to unique target-disease pairs are aggregated in the form of associations. Expanded documentation on how associations are built and scored is available in the [Target - disease associations](/associations) section.

### Applications and data access

To start addressing therapeutic hypotheses users can find alternative ways to interface with the data:

* [Web interface](/web-interface) at [platform.opentargets.org](http://platform.opentargets.org). The Platform web application provides a user interface to navigate data relevant for building therapeutic hypotheses. Starting from the home page search box, users can navigate to entity pages, associations pages or evidence pages.
* [Data access](/data-access). More complex hypotheses might require advance interfaces or programmatic access. The Platform aims to deliver a number of alternative ways to access data to support the most intensive queries.

More details about the open source codebase and development can be found in the [Platform infrastructure](/data-access/platform-infrastructure) section.

## Training materials

### Online tutorial

Please note we are in the process of updating our training content following the most recent updates (25.03).

{% embed url="<https://www.ebi.ac.uk/training/online/courses/open-targets-quick-tour/>" %}

### Webinar

{% embed url="<https://www.youtube.com/watch?v=VZQql6CnbSE>" %}


# Target

## Overview

A target in the Platform is understood as any naturally-occurring molecule that can be targeted by a medicinal product. EMBL-EBI's [Ensembl](https://www.ensembl.org) database is used as source for human targets in the Platform, with the Ensembl gene ID as the primary identifier.&#x20;

Criteria for target inclusion:

* Genes from all biotypes encoded in canonical chromosomes
* Genes in alternative assemblies encoding for a reviewed protein product.

This definition accounts for some of the complexities of human targets; targets are not only protein coding genes, RNAs or pseudogenes are also considered. However, the current definition has some potential drawbacks. Some drug targets are the result of interactions between genes (e.g. gene fusion) or proteins (e.g. protein complexes). The current target entity does not yet consider these cases, which will be addressed in the future.

{% embed url="<https://www.youtube.com/watch?v=iYRCGRGI5K4>" %}

## Target annotation data sources

| Annotation data                                                                                  | Data source                                                                                                                                                                                                              |
| ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| [Drugs and Clinical Candidates](/target/drugs)                                                   | [Open Targets](https://github.com/opentargets/clinical_mining)                                                                                                                                                           |
| Protein, positional, and structural information (ProtVista); Subcellular location; Gene Ontology | [UniProt](https://www.uniprot.org)                                                                                                                                                                                       |
| [Molecular interactions](/target/molecular-interactions)                                         | [Open Targets](/target/molecular-interactions)                                                                                                                                                                           |
| Pathways                                                                                         | [Reactome](https://reactome.org)                                                                                                                                                                                         |
| [Baseline expression](/target/baseline-expression)                                               | [Tabula Sapiens;](https://tabula-sapiens-portal.ds.czbiohub.org/) [GTEx](https://gtexportal.org/home/); [PRIDE](https://www.ebi.ac.uk/pride/archive/); [The Database of Immune Cells (DICE)](https://dice-database.org/) |
| Comparative genomics                                                                             | [Ensembl Compara](https://www.ensembl.org/info/docs/api/compara/index.html)                                                                                                                                              |
| Mouse phenotypes                                                                                 | [MGI](http://www.informatics.jax.org/phenotypes.shtml)                                                                                                                                                                   |
| Cancer hallmarks                                                                                 | [Cancer Gene Census](https://cancer.sanger.ac.uk/census)                                                                                                                                                                 |
| [Chemical Probes](/target/chemical-probes-and-teps)                                              | [Probes & Drugs](https://www.probes-drugs.org)                                                                                                                                                                           |
| [Bibliography](/bibliography)                                                                    | [Open Targets](/bibliography)                                                                                                                                                                                            |
| [Target tractability](/target/tractability)                                                      | [Open Targets](/target/tractability)                                                                                                                                                                                     |
| [Target safety](/target/safety)                                                                  | [ToxCast](https://www.epa.gov/chemical-research/toxicity-forecasting), [AOPWiki](https://aopwiki.org), [PharmGKB](https://www.pharmgkb.org/) and selected publications                                                   |
| [Target Enabling Packages (TEPs)](/target/chemical-probes-and-teps)                              | [SGC](https://www.thesgc.org/tep)                                                                                                                                                                                        |
| CRISPR-Cas9 cancer cell line dependency                                                          | [Project Score](https://score.depmap.sanger.ac.uk)                                                                                                                                                                       |
| [Core Gene Essentiality](https://platform-docs.opentargets.org/target/core-gene-essentiality)    | [DepMap Portal](https://depmap.org/portal/)                                                                                                                                                                              |
| [Pharmacogenetics](https://platform-docs.opentargets.org/target/pharmacogenetics)                | [ClinPGx](https://www.clinpgx.org/)                                                                                                                                                                                      |
| Genetic constraint                                                                               | [gnomAD](https://gnomad.broadinstitute.org)                                                                                                                                                                              |
| 95% Molecular QTL credible sets                                                                  | [Open Targets ](/gentropy/fine-mapping)                                                                                                                                                                                  |


# Drugs and Clinical Candidates

The **Drugs and Clinical Candidates** widget on the Target page is based on the **Clinical Target** dataset. Unlike the [Disease page equivalent](/disease-or-phenotype/drugs), which fixes a disease and lists drugs, this view fixes a target and surfaces the diseases for which connected drugs have clinical evidence.

The dataset is built by joining two sources: clinical reports, which link drugs to diseases, and drug mechanism of action data, which links drugs to targets. Each row in the widget represents a drug linked to both the selected target and disease. The **maximum stage** shown is the highest clinical development stage that the drug has reached across all its supporting indications.

```mermaid
flowchart TD
    T["<b>Target</b><br>e.g. BRAF"]
    D["<b>Drug</b><br>e.g. Vemurafenib"]
    DIS["<b>Disease</b><br>e.g. Melanoma"]

    T -->|"Mechanism of action"| D
    D -->|"Clinical report"| DIS

    subgraph inferred["Inferred target–disease"]
        T
        D
        DIS
    end

    style T fill:#dcd8f5,stroke:#7b6fc4,color:#3b2e8a
    style D fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style DIS fill:#fce8dc,stroke:#c97a50,color:#8b3a10
    style inferred fill: transparent, stroke:#ccc, color:#888
```

{% hint style="info" %}
For a full description of the **clinical stage categories** and their ranking, see [Clinical stage categories](/drug/clinical-report#clinical-stage-categories) in the Clinical Report page.
{% endhint %}


# Tractability

Data for assessing tractability with small molecule, antibody, and other clinical modalities

## Overview

To support target prioritisation, the Open Targets Platform includes tractability data that identifies key details, including if there is a binding site suitable for small molecule binding, an accessible epitope for antibody based therapy, relevant data for using Proteolysis Targeting Chimeras (PROTACs), or a compound in clinical trials with a modality other than small molecule or antibody.

The tractability data can assist in target prioritisation by identifying potential drug targets suitable for discovery pipelines and therapeutic modalities that are most likely to succeed. It also supports further investigation of targets for which there are no ligands or experimental structures or those targets outside a "druggable" target family but with strong genetic associations.

Our target tractability is based on a modified version of [Approaches to target tractability assessment – a practical perspective](https://pubs.rsc.org/en/content/articlelanding/2018/md/c7md00633k#!divAbstract) and [The PROTACtable genome](https://doi.org/10.1038/s41573-021-00245-x) and has workflows that generate tractability assessments for small molecule (SM), antibody (AB), Proteolysis Targeting Chimeras (PR), and other clinical (OC) modalities.

## Assessments

The tractability assessments displayed on the Platform's target profile pages is the result of an [open-source computational pipeline](https://github.com/melschneider/tractability_pipeline_v2/tree/master/ot_tractability_pipeline_v2) that performs *in silico* tractability assessments with small molecule, antibody, PROTAC, and other clinical modality workflows.

Data sources used in the pipeline include UniProt, HPA, PDBe, DrugEBIlity, ChEMBL, Pfam, InterPro, Complex Portal, DrugBank, Gene Ontology, and BioModels.

Assessments common to all modalities, ingested from ChEMBL, are:&#x20;

* **Approved Drug**: the target has clinical precedence with Phase IV drugs&#x20;
* **Advanced Clinical**: the target has clinical precedence with Phase II or III drugs
* **Phase 1 Clinical**: the target has clinical precedence with Phase I drugs.

We also include additional assessments specific to each modality.

### Small molecule

* **Structure with Ligand**: Target has been co-crystallised with a small molecule (source: [Protein Data Bank](https://www.rcsb.org/))
* **High-Quality Ligand**: Target with ligand(s) (PFI ≤ 7, SMART hits ≤ 2, scaffolds ≥ 2) (source: [ChEMBL](https://www.ebi.ac.uk/chembl/))
* **High-Quality Pocket**: Target has a DrugEBIlity score of ≥ 0.7 (source: [DrugEBIlity](http://chembl.github.io/drugebility-structure-based-component/))
* **Med-Quality Pocket**: Target has a DrugEBIlity score between 0 and 0.7 (source: [DrugEBIlity](http://chembl.github.io/drugebility-structure-based-component/))
* **Druggable Family**: Target is considered druggable as per [Finan et al](https://pubmed.ncbi.nlm.nih.gov/28356508/)’s Druggable Genome pipeline.

### Antibody

* **UniProt loc high conf:** High confidence that the subcellular location of the target is either plasma membrane, extracellular region/matrix, or secretion (source: [Uniprot](https://www.uniprot.org/))
* **GO CC high conf:** High confidence that the subcellular location of the target is either plasma membrane, extracellular region/matrix, or secretion (source: [Gene Ontology](http://geneontology.org/))
* **UniProt loc med conf:** Medium confidence that the subcellular location of the target is either plasma membrane, extracellular region/matrix, or secretion (source: [Uniprot](https://www.uniprot.org/))
* **UniProt SigP or TMHMM**: Target has a predicted signal peptide or trans-membrane regions, and not destined to organelles (source: Uniprot SigP, [TMHMM](https://services.healthtech.dtu.dk/service.php?TMHMM-2.0))
* **GO CC med conf:** Medium confidence that the subcellular location of the target is either plasma membrane, extracellular region/matrix, or secretion (source: [Gene Ontology](http://geneontology.org/))
* **Human Protein Atlas loc:** High confidence that the target is located in the Plasma membrane (source: [HPA](https://www.proteinatlas.org/))

### PROTAC

* **Literature**: Target mentioned in a set of manually curated PROTAC-related publications (source: [Europe PMC](http://europepmc.org/))
* **UniProt Ubiquitination:** Target tagged with the Uniprot keyword “Ubl conjugation \[KW-0832]”, which indicates that the protein has a ubiquitination site, based on evidence from the literature (source: [Uniprot](https://www.uniprot.org/))
* **Database Ubiquitination:** Target has reported ubiquitination sites in [PhosphoSitePlus](https://www.phosphosite.org/homeAction.action), [mUbiSiDa](http://reprod.njmu.edu.cn/cgi-bin/mubisida/mUbiSiDa.php) (2013), or [Kim et al. 2011](https://www.sciencedirect.com/science/article/pii/S1097276511006757)
* **Half-life Data:** Target has available half-life data (source: [Mathieson et al. 2018](https://www.nature.com/articles/s41467-018-03106-1))
* **Small Molecule Binder:** Target has a reported small-molecule ligand in ChEMBL with a measured activity of at least 10 μM in a target-based assay (source: [ChEMBL](https://www.ebi.ac.uk/chembl/))

## Computational pipeline and datasets

The data is available for download as part of the target core annotation from [our data downloads page](https://platform.opentargets.org/downloads).

Alternatively, you can also download the input TSV file with the per-target assessments via FTP. To access this file, visit [our FTP site](http://ftp.ebi.ac.uk/pub/databases/opentargets/platform/) and click on the release version (e.g. 21.04), followed by "input", "target", and "tractability". You can then download the `tractability` TSV file. Descriptions of the columns found in the input file can be found on the [pipeline README.md file](https://github.com/melschneider/tractability_pipeline_v2/blob/master/README.md).

## Publications

Brown KK, Hann MM, Lakdawala AS, Santos R, Thomas PJ, Todd K. **Approaches to target tractability assessment - a practical perspective**. Medchemcomm. 2018 Feb 14;9(4):606-613. doi: [10.1039/c7md00633k](https://doi.org/10.1039/c7md00633k). PMID: [30108951](https://pubmed.ncbi.nlm.nih.gov/30108951/); PMCID: [PMC6072525](https://europepmc.org/article/PMC/PMC6072525).

Schneider M, Radoux CJ, Hercules A, Ochoa D, Dunham I, Zalmas LP, Hessler G, Ruf S, Shanmugasundaram V, Hann MM, Thomas PJ, Queisser MA, Benowitz AB, Brown K, Leach AR. **The PROTACtable genome**. Nat Rev Drug Discov. 2021 Jul 20. doi: [10.1038/s41573-021-00245-x](https://doi.org/10.1038/s41573-021-00245-x). PMID: [34285415](https://pubmed.ncbi.nlm.nih.gov/34285415/).


# Safety

Manually-curated safety data to support target prioritisation

## Overview

Throughout the drug discovery and development process, target safety assessments help with understanding the role of the drug target in normal physiology and potential unintended adverse consequences and safety liabilities when modulating the target with a chemical compound or drug (Brennan, 2017).

To support target prioritisation, we have manually curated experimental data and insights from publications and other well-known sources of target safety and toxicity data, including the [ToxCast](https://www.epa.gov/chemical-research/toxicity-forecasting), [AOPWiki](https://aopwiki.org) and [ClinPGx](https://www.clinpgx.org/) (Open Targets downstream analysis of the toxicity datasets).

Safety data is available on the target profile page and can be accessed to provide a systematic view of potentially relevant target safety liabilities.

## **Computational pipelines and datasets**

Target safety datasets are mapped to the correct Ensembl gene ID and ingested during our initial pipeline steps to enrich the target annotation object. The data is available for download as part of the target core annotation from [our data download page](https://platform.opentargets.org/downloads).

## Publications

Bowes J, Brown AJ, Hamon J, Jarolimek W, Sridhar A, Waldron G, Whitebread S. **Reducing safety-related drug attrition: the use of in vitro pharmacological profiling**. Nat Rev Drug Discov. 2012 Dec;11(12):909-22. doi: [10.1038/nrd3845](https://doi.org/10.1038/nrd3845). PMID: [23197038](https://pubmed.ncbi.nlm.nih.gov/23197038/).

Brennan R.J. (2017) **Target Safety Assessment: Strategies and Resources**. In: Gautier JC. (eds) Drug Safety Evaluation. Methods in Molecular Biology, vol 1641. Humana Press, New York, NY. doi: [10.1007/978-1-4939-7172-5\_12](https://doi.org/10.1007/978-1-4939-7172-5_12)

Force T, Kolaja KL. **Cardiotoxicity of kinase inhibitors: the prediction and translation of preclinical models to clinical outcomes**. Nat Rev Drug Discov. 2011 Feb;10(2):111-26. doi: [10.1038/nrd3252](https://doi.org/10.1038/nrd3252). PMID: [21283106](https://pubmed.ncbi.nlm.nih.gov/21283106/).

Ann M. Richard, Richard S. Judson, Keith A. Houck, Christopher M. Grulke, Patra Volarath, Inthirany Thillainadarajah, Chihae Yang, James Rathman, Matthew T. Martin, John F. Wambaugh, Thomas B. Knudsen, Jayaram Kancherla, Kamel Mansouri, Grace Patlewicz, Antony J. Williams, Stephen B. Little, Kevin M. Crofton, and Russell S. Thomas. **ToxCast Chemical Landscape: Paving the Road to 21st Century Toxicology**. Chemical Research in Toxicology 2016 *29* (8), 1225-1251. doi: [10.1021/acs.chemrestox.6b00135](https://doi.org/10.1021/acs.chemrestox.6b00135). PMID: [27367298](https://pubmed.ncbi.nlm.nih.gov/27367298/).

Lamore SD, Ahlberg E, Boyer S, Lamb ML, Hortigon-Vinagre MP, Rodriguez V, Smith GL, Sagemark J, Carlsson L, Bates SM, Choy AL, Stålring J, Scott CW, Peters MF. **Deconvoluting Kinase Inhibitor Induced Cardiotoxicity**. Toxicol Sci. 2017 Jul 1;158(1):213-226. doi: [10.1093/toxsci/kfx082](https://doi.org/10.1093/toxsci/kfx082). PMID: [28453775](https://pubmed.ncbi.nlm.nih.gov/28453775/); PMCID: [PMC5837613](https://europepmc.org/article/PMC/PMC5837613).

Lynch JJ 3rd, Van Vleet TR, Mittelstadt SW, Blomme EAG. **Potential functional and pathological side effects related to off-target pharmacological activity**. J Pharmacol Toxicol Methods. 2017 Sep;87:108-126. doi: [10.1016/j.vascn.2017.02.020](https://doi.org/10.1016/j.vascn.2017.02.020). PMID: [28216264](https://pubmed.ncbi.nlm.nih.gov/28216264/).

Urban L, Whitebread S, Hamon J et al. **Screening for safety-relevant off-target affinities.** In: Polypharmacology in Drug Discovery. Peters JU (Ed.). John Wiley and Sons, NJ, USA (2012). doi: [10.1002/9781118098141.ch2](https://onlinelibrary.wiley.com/doi/abs/10.1002/9781118098141.ch2).


# Chemical probes & TEPs

## Overview

To support further assessments about the suitability of targets for a discovery pipeline, the Open Targets Platform integrates information from various sources on chemical probes and Target Enabling Packages (TEPs).

### Chemical probes

A chemical probe is a small molecule that can act as chemical modulator of a system, such as a cell or an organism, by reversibly binding to a biological target. Chemical probes are expected to have a minimum standard of high affinity, binding selectivity and efficacy. Chemical probes do not need to meet the same requirements in terms of pharmacokinetics, pharmacodynamics, and bioavailability as drugs.

### Target Enabling Packages

Target Enabling Packages (TEPs) are a critical mass of reagents and knowledge allowing for the rapid biochemical and chemical exploration of a given target.

As noted by the [Structural Genomics Consortium](https://www.thesgc.org/tep), each TEP contains:

* Protein production methods
* Biochemical/biophysical assays for activity, affinity
* Structures of the protein, potentially including wild type and disease mutant proteins; full-length or domains; protein-ligand complexes; structures of close homologues
* Initial chemical matter from a fragment or small molecule screen

Additional components of TEPs may be developed on a case-by-case basis, based upon reasonable scientific need, in collaboration with TEP target nominators:

* An antibody or nanobody
* Cell-based assay
* CRISPR knockout

## Data sources

### Probes & Drugs Portal

The Probes & Drugs Portal – <https://www.probes-drugs.org/home/> – is a public resource joining together focused libraries of bioactive compounds (probes, drugs, specific inhibitor sets etc.) with commercially available screening libraries. The purpose of the Portal is to reflect the current state of bioactive compound space and to enable its exploration from different points of view. It is intended to serve as a central hub in chemical biology research by providing a unique integration of the most relevant probes sources such as the [Chemical Probes Portal](https://www.chemicalprobes.org), [Open Science Probes](https://www.sgc-ffm.uni-frankfurt.de), and [ProbeMiner](https://probeminer.icr.ac.uk).

### Structural Genomics Consortium

The Structural Genomics Consortium (SGC) – [https://www.thesgc.org/](https://www.thesgc.org) – is a public-private partnership focused on accelerating drug discovery using open science. The SGC’s core research operations are funded by pharmaceutical companies, governments, and charities, all of which act as research partners and participate in the governance of the SGC.

## Computational pipelines and datasets

Chemical probes datasets are downloaded from each data source listed above, mapped to the correct Ensembl gene ID, and ingested during our initial pipeline steps. The data is available for download as part of the target core annotation from [our data download page](https://platform.opentargets.org/downloads).


# Baseline expression

June 2026 version

## Overview

Baseline (healthy, unstimulated) human RNA and protein expression data can help determine whether a target is expressed broadly across tissues and cell types or selectively in specific biological contexts. The presence of a target molecule in the tissue or cell type of interest is an important consideration throughout the drug discovery and development process.

The Open Targets Platform integrates baseline expression data generated using three complementary technologies:

* Single-cell RNA sequencing (scRNA-seq)
* Bulk RNA sequencing (RNA-seq)
* Mass spectrometry (MS)-based proteomics

Data are collected from four sources:

| **Data source**      | **Technology**               | **Coverage**                         |
| -------------------- | ---------------------------- | ------------------------------------ |
| **Tabula Sapiens**   | Single-cell RNA-seq          | 66 tissues, 180 annotated cell types |
| **GTEx (v10)**       | Bulk RNA-seq                 | 54 tissues                           |
| **DICE**             | Bulk RNA-seq                 | 12 flow-sorted immune cell types     |
| **PRIDE (OTAR3091)** | Mass spectrometry proteomics | 74 tissues                           |

***

## Data Sources

### Tabula Sapiens

[Tabula Sapiens v2](https://tabula-sapiens-portal.ds.czbiohub.org/) is a single-cell transcriptomic atlas containing over 1.1 million cells spanning multiple human organs and tissues. The atlas was generate

### Genotype-Tissue Expression (GTEx)

[Genotype-Tissue Expression (GTEx)](https://gtexportal.org/home/) aims to build a comprehensive public resource for studying tissue-specific gene expression and its relationship to genetic variation. The GTEx v10 dataset contains bulk RNA-seq data from 54 tissues collected from 946 human donors.

### Database of Immune Cells (DICE)

The [Database of Immune Cells (DICE)](https://dice-database.org/) provides gene expression profiles from 12 flow-sorted immune cell types derived from peripheral blood samples from 91 healthy donors. Open Targets uses DICE build 2.23.2022. The data were generated using bulk RNA sequencing.

### PRoteomics IDEntifications Database (PRIDE)

PRoteomics IDEntification Database ([PRIDE](https://www.ebi.ac.uk/pride/archive/)) is a public repository for proteomics data. As part of an Open Targets project, the PRIDE team reanalysed four human proteomics datasets to provide protein expression measurements across multiple human tissues.

The following datasets were included:

| **Dataset**                                                         | **Description**           |
| ------------------------------------------------------------------- | ------------------------- |
| [PXD000561](https://www.ebi.ac.uk/pride/archive/projects/PXD000561) | 21 tissues from 6 donors  |
| [PXD000865](https://www.ebi.ac.uk/pride/archive/projects/PXD000865) | 34 tissues from 1 donor   |
| [PXD020192](https://www.ebi.ac.uk/pride/archive/projects/PXD020192) | 46 tissues from 2 donors  |
| [PXD016999](https://www.ebi.ac.uk/pride/archive/projects/PXD016999) | 32 tissues from 14 donors |

## Computational pipelines and datasets

#### Initial data processing

**Single-cell RNA sequencing (scRNA-seq) data from Tabula Sapiens**

Single-cell RNA sequencing data for tissues and cell types was downloaded from the [Tabula Sapiens portal](https://tabula-sapiens-portal.ds.czbiohub.org/) as AnnData files containing raw count matrices with the associated metadata. The data were then processed into pseudobulked expression profiles for each cell type, each tissue, and each cell type within each tissue. The pseudobulked expression profiles are generated with a dSum approach as described in [Cuomo et al. 2021](https://doi.org/10.1186/s13059-021-02407-x). In short, in order to obtain a pseudobulked sample we sum the raw counts of each gene across all cells within a donor for a given annotation (tissue, cell type or cell type within tissue). This results in one pseudobulked expression sample per donor per annotation. The summed counts are then normalised using the Counts Per Million (CPM) approach.

The result is three pseudobulked expression datasets: cell type, tissue, and cell type within tissue.

**Bulk RNA sequencing (RNA-seq) data from GTEx and DICE**

Bulk RNA sequencing data for tissues was downloaded in transcripts per million (TPM) units from [GTEx](https://gtexportal.org/home/downloads/adult-gtex/bulk_tissue_expression) and for cell types from the [DICE portal](https://dice-database.org/downloads).

**Mass spectrometry (MS)-based proteomics data from PRIDE**

Baseline mass spectrometry–based proteomics data were compiled from selected public PRIDE datasets curated using SDRF proteomics metadata guidelines. Raw files in these datasets had been reanalysed centrally using the open source tools MaxQuant (version 2.7.0.0) according to predefined dataset selection and curation guidelines, and the resulting protein-level quantification tables with associated SDRF metadata were downloaded from the [PRIDE FTP ](https://www.ebi.ac.uk/pride/archive/)resource.

#### Generating unaggregated expression profiles

All baseline expression data are processed into a unified set of unaggregated, donor-level expression profiles. We harmonise metadata and map tissue and cell type annotations to standardised ontologies: [UBERON](https://www.ebi.ac.uk/ols4/ontologies/uberon) and the [Cell Ontology](https://www.ebi.ac.uk/ols4/ontologies/cl). If a donor has multiple samples for the same annotation, we calculate the mean expression to produce a single value for each donor-annotation pair. The data structure is standardised, but expression values retain their original units: TPM for bulk RNA-seq, pseudobulk CPM (sum\[counts]) for pseudobulked scRNA-seq, and PPB (iBAQ) for proteomics.

#### Generating aggregated expression profiles

We aggregate the donor-level data to produce a summary of expression for each target in each tissue or cell type. This includes:

* **Expression Quartiles**: We calculate the minimum, 25th percentile, median, 75th percentile, and maximum expression values across all donors for a given annotation.
* **Expression Distribution Score**: We compute a score representing the breadth of expression within a dataset, defined as the proportion of annotations in which the target is expressed above a threshold of 0.5 (unit) median expression. The threshold is chosen to minimise the impact of noise (e.g. from ambient RNA).
* **Expression Specificity Score**: We use the tool CELLEX ([Timshel et al. 2020](https://doi.org/10.7554/eLife.55851)) which combines four expression-specificity metrics into a single, ranked specificity score that reflects how specifically a gene is expressed in a given annotation relative to all other annotations in that dataset. Because ranking is done within each annotation, a gene expressed in a tissue expressing many genes must have stronger specificity to rank highly than it would in a tissue with fewer distinct markers. Consequently, the annotation with the highest expression is not always the one where the gene has the highest specificity score. For a given gene, specificity is denoted as “high” for annotations where the gene is scored amongst the top 25% of specifically expressed genes.

To enable comparisons across similar tissues and cell types from different data sources, we also assign a set of high-level, parental tissue and cell type categories to each annotation. The parental labels are derived from manual grouping of the ontology mappings, for example, the heart left ventricle from GTEx and cardiac atrium from Tabula Sapiens both share the parental label of heart. There are 37 parental tissue labels and 24 parental cell type labels.

## Data Availability

These datasets can be accessed via the [Platform web interface](https://platform.opentargets.org/target/ENSG00000133703)**,** [download page](https://platform.opentargets.org/downloads) and GraphQL [API](https://platform.opentargets.org/api).

<figure><img src="/files/UkNBsVzJcr9CowkExPEO" alt=""><figcaption></figcaption></figure>

<sub>Example of new baseline expression widget for</sub> <sub></sub><sub>**KRAS**</sub><sub>.</sub>

## Publications

* GTEx Consortium. ***The GTEx Consortium atlas of genetic regulatory effects across human tissues*****.** Science. 2020;369(6509):1318–1330.
* The Tabula Sapiens Consortium. ***The Tabula Sapiens: A multiple-organ, single-cell transcriptomic atlas of humans*****.** Science. 2022;376(6594).
* Ha B, et al. ***Database of Immune Cell eQTLs, Expression, Epigenomics*****.** The Journal of Immunology. 2019;202(1 Supplement):131.18.
* Perez-Riverol Y, et al. ***The PRIDE database at 20 years: 2025 update*****.** Nucleic Acids Research. 2025;53(D1)–D553.
* Tyanova S, et al. ***The MaxQuant computational platform for mass spectrometry-based shotgun proteomics*****.** Nature Protocols. 2016;11(12):2301–2319.
* Cuomo AS, et al. ***Optimizing expression quantitative trait locus mapping workflows for single-cell studies*****.** Genome Biology. 2021;22(1):240.
* Timshel P, et al. ***Genetic mapping of etiologic brain cell types for obesity*****.** eLife. 2020;9.


# Molecular interactions

Systematically capturing target - target interactions of different nature

## **Overview**

The Molecular Interactions data aggregates and integrates interaction evidence reported in several resources to provide a systematic view on potentially relevant drug targets. Each of the integrated resources captures relationships of different nature including physical binary interactions, enzymatic reactions, or functional relationships. The information here available aims to capture not only the topology of the interaction network, but also the supporting experimental evidence reported on each of the databases.

In order to maximise coverage, the network contains all reported binary relationships between gene products (proteins and RNAs). Although the main focus are interactions between human molecules, the data also includes additional interactions between human gene products and molecules encoded in the genome of infectious pathogens (viruses and bacteria).

## **Data sources**

### IntAct

IntAct – <http://www.ebi.ac.uk/intact> – is a freely available, open source database for molecular interaction data. IntAct contains physical interactions derived from literature curation or direct user submissions.

Interactions are scored using the MI score. Benefiting from the PSI-MI controlled vocabulary, the Intact MI score provides a normalised (0 to 1) score that weights how recurrently an interaction has been reported, together with the confidence of the experimental techniques reported. Note that a high scoring interaction can be due to high-confidence evidence, but also a social bias on studying certain proteins. Generally speaking, scores > 0.4 correspond to medium to high confidence interactions, although some good-quality high-throughput interactions might still be scored below that threshold. More info on MI score can be found in the [Intact documentation](https://www.ebi.ac.uk/intact/pages/faq/faq.xhtml).

Interactions are grouped by interaction detection method and interaction type. As a consequence, the same pair of interactors might be split into multiple entries if individual proteins are reported to have different biological roles.

For IntAct, please note:

* The network only contains human and selected pathogen data from IntAct
* The majority of interactions are not directional and not signed. However, there are a proportion of interactions where the biological role of the participants can be stated and directionality specified (e.g. enzymatic reactions)

### Reactome

Reactome – [https://reactome.org/](https://reactome.org) – is an open source, open access, manually curated and peer-reviewed pathway database.

For Reactome, please note:

* Only human-human interactions are provided
* Interactions are directional and signed, with biological roles assigned to each participant if possible
* Protein interactions in Reactome are inferred from pathways and complexes based on Reactome internal [method](https://github.com/reactome/interaction-exporter/wiki/Methods).

### SIGNOR

SIGNOR, the SIGnaling Network Open Resource – [https://signor.uniroma2.it/](https://signor.uniroma2.it) – contains signaling information published in the scientific literature, which is manually curated and stored in a structured format.

For SIGNOR, please note:

* SIGNOR only contains human data
* Interactions are directional and signed, with biological roles assigned to each participant
* The network pulls information from the SIGNOR relations file

### STRING

STRING – <https://string-db.org> – contains functionally interacting proteins. While most interactions in the other resources capture different types of physical interaction between molecules, functional interactions do not necessarily interact physically. Both direct (physical) and indirect (functional) associations are derived from computational predictions, from knowledge transfer between organisms, or from interactions aggregated from other (primary) databases.\
\
STRING interactions provide an overall combined\_score, as well as each of the pieces of information that compose this score. More information on STRING scoring can be found on their [documentation page](https://string-db.org/cgi/info).

## Computational pipeline and datasets

The multipartite network displayed in the Open Targets Platform is the result of post-processing the information stored in a Neo4j graph database (graphDB). The graphDB does not provide STRING information but contains ComplexPortal information on stable protein complexes as an additional data source. The information of the graphDB is then exported together with STRINGdb and mapped to the Open Targets Platform targets (Ensembl Gene IDs).

The resulting dataset as well as all intermediate files can be found in the Open Targets Platform Data Access section or the [Intact FTP](ftp://ftp.ebi.ac.uk/pub/databases/intact/various/ot_graphdb/current/).

## Publications

When using this data please remember to acknowledge the sources:

Orchard S, Ammari M, Aranda B, Breuza L, Briganti L, Broackes-Carter F, Campbell NH, Chavali G, Chen C, del-Toro N, Duesbury M, Dumousseau M, Galeota E, Hinz U, Iannuccelli M, Jagannathan S, Jimenez R, Khadake J, Lagreid A, Licata L, Lovering RC, Meldal B, Melidoni AN, Milagros M, Peluso D, Perfetto L, Porras P, Raghunath A, Ricard-Blum S, Roechert B, Stutz A, Tognolli M, van Roey K, Cesareni G, Hermjakob H. **The MIntAct project--IntAct as a common curation platform for 11 molecular interaction databases**. Nucleic Acids Res. 2014 Jan;42(Database issue):D358-63. doi: 10.1093/nar/gkt1115. Epub 2013 Nov 13. PMID: 24234451; PMCID: PMC3965093.

Jassal B, Matthews L, Viteri G, Gong C, Lorente P, Fabregat A, Sidiropoulos K, Cook J, Gillespie M, Haw R, Loney F, May B, Milacic M, Rothfels K, Sevilla C, Shamovsky V, Shorser S, Varusai T, Weiser J, Wu G, Stein L, Hermjakob H, D'Eustachio P. **The reactome pathway knowledgebase**. Nucleic acids research. 2020 Jan;48(D1) D498-D503. doi: 10.1093/nar/gkz1031. PubMed PMID: 31691815. PubMed Central PMCID: PMC7145712.

Licata L, Lo Surdo P, Iannuccelli M, Palma A, Micarelli E, Perfetto L, Peluso D, Calderone A, Castagnoli L, Cesareni G. **SIGNOR 2.0, the SIGnaling Network Open Resource 2.0: 2019 update**. Nucleic Acids Res. 2020 Jan 8;48(D1):D504-D510. doi: 10.1093/nar/gkz949. PMID: 31665520; PMCID: PMC7145695.

Szklarczyk D, Kirsch R, Koutrouli M, Nastou K, Mehryary F, Hachilif R, Gable AL, Fang T, Doncheva NT, Pyysalo S, Bork P, Jensen LJ, von Mering C. **The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest**. Nucleic Acids Res. 2023 Jan 6;51(D1):D638-D646. doi: 10.1093/nar/gkac1000. PMID: 36370105; PMCID: PMC9825434.


# Core Gene Essentiality

Data supporting core essentiality of a target

## Overview

Core essential target genes are unlikely to tolerate inhibition and are susceptible to causing adverse events if modulated. This is crucial safety information for drug discovery scientists looking to develop inhibition strategies.

To support target prioritisation with an extra focus on target safety, the Open Targets Platform includes key information on target core essentiality in the context of 28 tissues assayed in the [DepMap portal](https://depmap.org/portal/). In this project and its ancillary projects (e.g. [Achilles](https://depmap.org/portal/achilles/)), cell fitness is measured after the inhibition of individual genes across a number of cell lines. As a result, a gene is catalogued as core essential if the majority of cell lines die after inhibition/knockout. Although in cancer, this experiment represents a good proxy of whether loss of function is tolerated across a diverse set of tissues.

**Essentiality assessment:** The [Chronos dependency score](https://depmap.org/portal/gene/KIT?tab=dependency) is based on data from a cell depletion assay. A lower Chronos score indicates a higher likelihood that the gene of interest is essential in a given cell line. A score of 0 indicates a gene is not essential; correspondingly -1 is comparable to the median of all pan-essential genes.&#x20;

Together with a DepMap target annotation widget, essential target genes are annotated with a \`Core essential gene\` chip on the target page.

## **Data source**

### DepMap Portal

The goal of the Dependency Map ([DepMap](https://depmap.org/portal/)) portal is to empower the research community to make discoveries related to cancer vulnerabilities by providing open access to key cancer dependencies analytical and visualisation tools.

## Publications

When using this data please remember to acknowledge the sources:

Pacini C, Dempster JM, Boyle I, Gonçalves E, Najgebauer H, Karakoc E, van der Meer D, Barthorpe A, Lightfoot H, Jaaks P, McFarland JM, Garnett MJ, Tsherniak A, Iorio F. **Integrated cross-study datasets of genetic dependencies in cancer**. Nat Commun. 2021 Mar 12;12(1):1661. doi: 10.1038/s41467-021-21898-7. PMID: 33712601; PMCID: PMC7955067.

[See all relevant DepMap publications](https://depmap.org/portal/home/#/publications).


# Pharmacogenetics

Data supporting pharmacogenetics annotation for a target

## Overview

Pharmacogenetics is the study of how genetic variation may change your response to a specific drug. Through the integration of pharmacogenetics data into the Platform, our objective is to apply clinical annotations available in the Pharmacogenetics Knowledgebase (ClinPGx) to aid target prioritisation for drug discovery.

We also enhance these annotations by adding detailed annotations on variant consequence prediction, drug response categories, and specific drug information and whether the gene is a direct target of the drug. This process involves applying advanced phenotype extraction techniques to offer a refined representation of phenotypes, thereby providing a more precise understanding of the genetic determinants influencing treatment outcomes.

## Data sources

[ClinPGx](https://www.clinpgx.org/) is an NIH-funded comprehensive resource that provides information about how human genetic variation affects response to medications. PharmGKB collects, curates and disseminates knowledge about clinically actionable gene–drug associations and genotype–phenotype relationships, focusing on the impact of genetic variation on drug response for clinicians and researchers.

We have modelled and harmonised the data from the[ ](https://www.pharmgkb.org/clinicalAnnotations)[Clinical Annotation](https://www.pharmgkb.org/clinicalAnnotations) section of ClinPGx, which provides information about variant–drug pairs based upon collating and summarising variant annotations. Variant annotations are curated manually from a scientific publication. Clinical Annotations then provide an overarching summary and curated [level of evidence](https://www.pharmgkb.org/page/clinAnnLevels) for the association between a genetic variant and particular drug responses based on multiple variant annotations. The likely consequence for each genotype of the variant on drug response is represented, in comparison to the other genotypes of that variant. Variant-level annotation (e.g. direction of effect) is also presented, when available.

We have also included annotations for star (\*) alleles, a nomenclature used in the pharmacogenetics field for representing key functional variants involved in drug responses.

***

## Publications

When using this data please remember to acknowledge the sources:

Whirl-Carrillo M, Huddart R, Gong L, Sangkuhl K, Thorn CF, Whaley R, Klein TE. **An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine.** Clin Pharmacol Ther. 2021 Sep;110(3):563-572. doi: 10.1002/cpt.2350. Epub 2021 Jul 22. PMID: 34216021; PMCID: PMC8457105.

Whirl-Carrillo M, McDonagh EM, Hebert JM, Gong L, Sangkuhl K, Thorn CF, Altman RB, Klein TE. **Pharmacogenomics knowledge for personalized medicine**. Clin Pharmacol Ther. 2012 Oct;92(4):414-7. doi: 10.1038/clpt.2012.96. PMID: 22992668; PMCID: PMC3660037.

[See all relevant PharmGKB publications.](https://www.pharmgkb.org/page/citingPharmgkb)


# Disease or Phenotype

## Overview

A disease or phenotype in the Platform is understood as any disease, phenotype, biological process or measurement that might have any type of causality relationship with a human target. The EMBL-EBI [Experimental Factor Ontology](https://www.ebi.ac.uk/efo/) (EFO) is used as scaffold for the disease or phenotype entity.

In order to maximise the alignment of the ontology with a clinical application, a few modifications have been added to EFO. Some high-level terms have been removed (e.g. disease by anatomical region) and others have been rearranged to align them to a less anatomical and more clinical interpretation. For each EFO release, the EFO OTAR slim can be found in parallel with the official [EFO release](https://github.com/EBISPOT/efo/releases).

{% embed url="<https://www.youtube.com/watch?v=jUc_Dpm1hJ0>" %}

## Disease or phenotype annotation data sources

| Annotation data                                                                  | Data source                                                                      |
| -------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
| Description; cross-references; synonyms; location; ontology and classification   | [EFO](https://www.ebi.ac.uk/efo/)                                                |
| [Drugs and Clinical Candidates](/disease-or-phenotype/drugs)                     | [Open Targets](https://github.com/opentargets/clinical_mining)                   |
| [Clinical signs and symptoms](/disease-or-phenotype/clinical-signs-and-symptoms) | [HPO](https://hpo.jax.org/app/) and [MONDO](https://mondo.monarchinitiative.org) |
| [Bibliography](/bibliography)                                                    | [Open Targets](/bibliography)                                                    |


# Drugs and Clinical Candidates

The **Drugs and Clinical Candidates** widget provides a disease-centric view of drugs and clinical candidates with evidence for the disease, based on the **Clinical Indications** dataset.&#x20;

A clinical indication consolidates all clinical reports sharing the same drug and disease into a single record, aggregating evidence across sources and development stages while preserving traceability to each contributing report.

For each drug/indication pair, the widget displays the **maximum clinical stage** reached across all contributing reports. This is determined by ranking harmonised clinical stage values from each clinical report, with one explicit rule: if at least one supporting report has a stage of *Phase IV* or *Withdrawal*, the maximum clinical stage is set to *Approval*, reflecting the assumption that marketing authorisation must have been reached before a post-marketing or withdrawal context can occur.

```mermaid
flowchart LR
    CT["<b>Clinical trial</b> <br>NCT02684006 · Phase III</br>"]
    DL["<b>Drug label</b> <br>FDA label · Approval</br>"]
    CI["<b>Curated indication</b><br>ChEMBL · Phase II</br>"]
    DW["<b>Drug warning</b><br>ChEMBL · Withdrawal</br>"]

    CLIN["<b>Clinical indication</b><br>Drug A · Disease B</br>4 contributing reports</br>Sources: trial, label, ChEMBL"]

    APP["<b>Approval</b>"]

    CT --> CLIN
    DL --> CLIN
    CI --> CLIN
    DW --> CLIN
    CLIN --> APP

    subgraph reports["Clinical Reports"]
        CT
        DL
        CI
        DW
    end

    subgraph indication["Clinical indication"]
        CLIN
    end

    subgraph maxstage["Max stage"]
        APP
    end

    style CT fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style DL fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style CI fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style DW fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style CLIN fill:#dcd8f5,stroke:#7b6fc4,color:#3b2e8a
    style APP fill:#fce8dc,stroke:#c97a50,color:#8b3a10
    style reports fill: transparent, stroke: transparent
    style indication fill: transparent, stroke: transparent
    style maxstage fill: transparent, stroke: transparent
```

{% hint style="info" %}
For a full description of the **clinical stage categories** and their ranking, see [Clinical stage categories](/drug/clinical-report#clinical-stage-categories) in the Clinical Report page.
{% endhint %}


# Clinical signs and symptoms

## **Overview**

The Clinical Signs and Symptoms data in the Platform aims to capture other diseases or phenotypes that occur as a consequence, or in conjunction with a primary disease.

Disease–phenotype relationships are not only useful to better characterise the disease phenotypic space, but also to serve as proxies for additional causal evidence that can help prioritise new or existing targets.

In order to maximise the list of available disease–phenotype links, the Platform ingests data from the [Monarch Merged Disease Ontology](https://mondo.monarchinitiative.org) (MONDO) and the [Human Phenotype Ontology](https://hpo.jax.org) (HPO), the latter in a joint effort with [Orphanet](http://www.orpha.net).

## **Data sources**

### Monarch Merged Disease Ontology (MONDO)

The Monarch Merged Disease Ontology (MONDO) – [https://mondo.monarchinitiative.org/](https://mondo.monarchinitiative.org) – is an open-source semi-automatically constructed ontology that integrates multiple disease resources to build a single, coherent merged ontology. Originally constructed in an entirely automatic way with IDs of source databases and ontologies, manually curated cross-ontology axioms have been added and a native Mondo ID system was developed and implemented to reduce confusion with source databases and ontologies.

### Human Phenotype Ontology (HPO)

The Human Phenotype Ontology (HPO) – <https://hpo.jax.org/app/> – provides a standardised vocabulary of phenotypic abnormalities encountered in human disease and contains over 13,000 terms and over 156,000 annotations to hereditary diseases. HPO has been developed using medical literature, Orphanet, DECIPHER, and OMIM.

## Computational pipelines and datasets

The relationships described in MONDO and HPO are also enriched with annotations such as sex, typical age of onset or frequency of disease patients presenting the phenotype. To improve the interoperability with the rest of the Platform, diseases and phenotypes are mapped to the [Experimental Factor Ontology](https://www.ebi.ac.uk/efo/) (EFO) when possible.

Complete disease–phenotype relationship datasets are available for download on [our data downloads page](https://platform.opentargets.org/downloads).

## Publications

Köhler S, Carmody L, Vasilevsky N, Jacobsen JOB, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D, Cipriani V, Balhoff JP, Conlin T, Blau H, Baynam G, Palmer R, Gratian D, Dawkins H, Segal M, Jansen AC, Muaz A, Chang WH, Bergerson J, Laulederkind SJF, Yüksel Z, Beltran S, Freeman AF, Sergouniotis PI, Durkin D, Storm AL, Hanauer M, Brudno M, Bello SM, Sincan M, Rageth K, Wheeler MT, Oegema R, Lourghi H, Della Rocca MG, Thompson R, Castellanos F, Priest J, Cunningham-Rundles C, Hegde A, Lovering RC, Hajek C, Olry A, Notarangelo L, Similuk M, Zhang XA, Gómez-Andrés D, Lochmüller H, Dollfus H, Rosenzweig S, Marwaha S, Rath A, Sullivan K, Smith C, Milner JD, Leroux D, Boerkoel CF, Klion A, Carter MC, Groza T, Smedley D, Haendel MA, Mungall C, Robinson PN. **Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources**. Nucleic Acids Res. 2019 Jan 8;47(D1):D1018-D1027. doi: 10.1093/nar/gky1105. PMID: [30476213](https://pubmed.ncbi.nlm.nih.gov/30476213/); PMCID: PMC6324074.

Shefchek KA, Harris NL, Gargano M, Matentzoglu N, Unni D, Brush M, Keith D, Conlin T, Vasilevsky N, Zhang XA, Balhoff JP, Babb L, Bello SM, Blau H, Bradford Y, Carbon S, Carmody L, Chan LE, Cipriani V, Cuzick A, Della Rocca M, Dunn N, Essaid S, Fey P, Grove C, Gourdine JP, Hamosh A, Harris M, Helbig I, Hoatlin M, Joachimiak M, Jupp S, Lett KB, Lewis SE, McNamara C, Pendlington ZM, Pilgrim C, Putman T, Ravanmehr V, Reese J, Riggs E, Robb S, Roncaglia P, Seager J, Segerdell E, Similuk M, Storm AL, Thaxon C, Thessen A, Jacobsen JOB, McMurry JA, Groza T, Köhler S, Smedley D, Robinson PN, Mungall CJ, Haendel MA, Munoz-Torres MC, Osumi-Sutherland D. **The Monarch Initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across species**. Nucleic Acids Res. 2020 Jan 8;48(D1):D704-D715. doi: 10.1093/nar/gkz997. PMID: [31701156](https://pubmed.ncbi.nlm.nih.gov/31701156/); PMCID: PMC7056945.


# Variant

Common and rare variation in Open Targets Platform

## Overview

In the Open Targets Platform, a variant refers to any human variation associated with a [disease, trait or phenotype](/disease-or-phenotype) that has been reported in any of our sources. All variation is mapped to GRCh38 build and enriched with functional annotation. The Platform currently captures single nucleotide polymorphisms (SNPs) and insertions/deletions.&#x20;

{% hint style="info" %}
**Variant identifier**

Variant identifiers of SNPs and small indels are created based on genomic location and alleles like: `6_160589086_A_G` where `A` is the reference allele at position `160,589,086` on chromosome `6` and the alternate allele is `G`. Being consistent with gnomAD, we are using a 1-based coordinate system.&#x20;

For longer insertions (200+) and deletions, where keeping the full length of the allele in the variant identifier is impractical, the allele string is hashed to create the identifier, which, when available, might contain the chromosome and position as well. Example: `OTVAR_11_614383_9cc2ae367cc98c283cb510e8ea29c9f0`
{% endhint %}

All variants shown in the Platform are reported in at least one of our [variant-to-phenotype](#variant-to-phenotype) sources.

## Population Allele Frequencies

Alternate allelic frequencies from [gnomAD](https://gnomad.broadinstitute.org/) variation database are reported for all major populations when available.

**Source:** [gnomAD 4.1](https://gnomad.broadinstitute.org/news/2024-04-gnomad-v4-1/)

## Variant effect

Variants are annotated with an integrated view of variant effects from multiple methods. Based on all predictions or annotations, we normalise the variant's likely deleteriousness to a common scale.

To make the predicted variant effects comparable across different methods,  raw predictions from each methods were normalised to a unified scale ranging from likely benign to uncertain to likely deleterious.&#x20;

## Molecular Structure Viewer

For predicted missense variants, we have included a Molecular Structure Viewer on the variant page, with the reference amino acid highlighted within the AlphaFold predicted model for the relevant protein.

The feature includes the option to visualise:

1. Protein pathogenicity - highlighting the AlphaMissense pathogenicity for the substitution corresponding to the variant, and the average AlphaMissense pathogenicity score across all possible amino acid substitutions at other positions
2. Protein domains  (from [Uniprot](https://www.uniprot.org/))
3. Secondary structure (from [AlphaFold](https://alphafold.ebi.ac.uk/))
4. Residues hydrophobicity (from [here](https://www.sigmaaldrich.com/GB/en/technical-documents/technical-article/protein-biology/protein-structural-analysis/amino-acid-reference-chart#hydrophobicity))&#x20;

The widget also includes a linear representation of the protein which updates alongside the structural representation.&#x20;

You can take a look at this example view for variant [7\_44152420\_C\_G](https://platform.opentargets.org/variant/7_44152420_C_G) to discover more about scores and colour coding used to build the viewer.

**Source:** [AlphaPhold DB](https://alphafold.ebi.ac.uk/), [UniProt](https://www.uniprot.org/)&#x20;

## Transcript consequences

Every variant is annotated with the predicted consequence for all canonical transcripts in a +/-500Kb window, allowing us to understand the likely effects in the neighbouring coding or non-coding genes. For all variant-transcript pairs in the region, this information includes:

* Distance from transcription start site (TSS)
* Distance to footprint
* Predicted functional consequence based on Ensembl VEP
* Amino-acid consequence relative to the UniProt reference protein

**Source**: [Ensembl VEP](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0974-4)

## Variant-to-phenotype

The list of variant sources includes:

* [95% GWAS credible sets](/gentropy/fine-mapping)
* [95% Molecular QTL credible sets](/gentropy/fine-mapping)
* [Enhancer-to-gene](/gentropy/enhancer-to-gene-encode-re2g): Epigenetically derived genomic regions that regulate putative genes.
* [ClinVar](/evidence): Submitted variants at all clinical significances
* [Uniprot](/evidence): Literature-based curation of disease-associated variants
* [Pharmacogenetics](/target/pharmacogenetics): Variants corresponding to genotypes associated with drug responses


# Study

## Overview

In the Open Targets Platform, a study is a qualifying genome-wide association study for binary or quantitative traits.

Studies might capture different types of associations, such as curated associations (top hits) from literature or post-processing GWAS summary statistics.

For each study meeting the [inclusion criteria](#study-inclusion-criteria), harmonisation, [fine-mapping](/gentropy/fine-mapping), [colocalisation](/gentropy/colocalisation), and [Locus-to-Gene](/gentropy/locus-to-gene-l2g) analysis are performed, resulting in a list of significant credible sets and their most likely associated genes. The resulting 95% credible sets can be visualised in all study pages. Within the same study, credible sets might result from different fine-mapping pipelines but still be presented in the same study with their respective provenance and confidence.&#x20;

{% hint style="info" %}
**Studies without GWAS-significant associations**

The study will still be presented in the Open Targets Platform if no significant associations are found.&#x20;
{% endhint %}

#### **Key study annotation**

To provide a rich context to downstream interpretation of studies, and their associations, we capture a wide range annotations, whenever possible in a standardised way to enable interoperability.&#x20;

For example:

* The measured trait/phenotype is standardised using the Experimental Factor Ontology ([EFO](https://www.ebi.ac.uk/efo/)), as the background trait when applied
* For molecular QTLs, the affected gene/protein is standardised to Ensembl gene identifiers
* For molecular QTLs, the biological context is standardised to relevant ontology (e.g. tissue — [UBERON](https://www.ebi.ac.uk/ols4/ontologies/uberon) or cell ontology — [CL](https://www.ebi.ac.uk/ols4/ontologies/cl))
* Depending on the granularity of the ingested study metadata, a rich cohort information is captured including sample sizes, ancestry composition and a list of named cohorts
* Indicated if association data of a study was ingested as summary statistics, together with the corresponding quality control measurements performed on the summary statistics
* Besides the Pubmed identifier, rich publication information is also captured including first author, publication year for easier data access

## Study categories

Integrated studies can be split into two major groups that determine how their associations would be utilised in the target identification process.&#x20;

### 1. Genome-wide association studies (GWAS)

A GWAS identifies associations between genetic variants and traits or phenotypes by analysing genome-wide variation data from a population. These studies cover binary or quantitative traits reported in the integrated sources.

{% hint style="info" %}
**Shared trait studies**&#x20;

GWAS studies sharing the same disease or phenotype (EFO) as the study are listed.
{% endhint %}

**Sources**: [GWAS Catalog](https://www.ebi.ac.uk/gwas/), [FinnGen](https://www.finngen.fi/en/access_results)

### 2. molQTL studies

A molQTL study identifies genetic variation that can be significantly associated with changes at the molecular level. As such, associations of these studies provide invaluable help to put GWAS derived loci into context via [colocalisation](/gentropy/colocalisation). The type of molecular trait measured defines the type of QTL:&#x20;

* eQTLs impact gene expression
* pQTLs impact protein abundance
* sQTLs have an effect on gene splicing
* tuQTLs impact transcript usage
* sceQTL impact gene expression at the single-cell level

{% hint style="info" %}
All molQTL studies are annotated with the gene whose expression levels are regulated, referred as the affected gene.

On top of the effected gene, molQTL studies also capture the tissue/cell type in which the change in the molecular trait was measured providing further context for interpretation.
{% endhint %}

Each molQTL study captures the unique combination of the following annotations:

* The publication authoring the study
* The gene product measured in the study&#x20;
* The quantitative method (e.g. aptamer) used to measure the trait, when available
* The cell type or tissue where the trait is measured, when available
* The experimental conditions (e.g. interferon-stimulated macrophages)

This is an example of a [QTL study page](https://platform.opentargets.org/study/sun_2018_aptamer_plasma_tnfrsf1a_2654_19_1_1) and unique ID.

**Sources**: [eQTL Catalogue](https://www.ebi.ac.uk/eqtl/), [UK Biobank Pharma Proteomics Project](https://www.synapse.org/Synapse:syn51364943/wiki/622119) (UKB-PPP)

## Study inclusion criteria

Studies must meet predefined criteria to ensure consistent representation, and increased reliability of associations for downstream analysis.&#x20;

**Each integrated study needs to meet the following criteria**:

* GWAS studies need to have a valid disease or phenotype identifier (EFO)
* QTL studies must have an affected gene and tissue/cell-type valid identifier (biosample)
* Valid study type (e.g `gwas` or a QTL type from a defined set)
* Data licensed for commercial usage
* When summary statistics are available, further [criteria](/gentropy/data-sources#gcss-quality-control) must be met:
  * Mean beta within expected range
  * Genomic control (GC) lambda value within the expected range
  * The PZ value within the expected range
  * Additional [curation](/gentropy/data-sources#gcss-studies-manual-curation) from Open Targets team ensuring study meet standards

## GWAS/molQTL credible sets

The [fine-mapping](/gentropy/fine-mapping) results for all 95% credible sets in the study are displayed with their most relevant metadata, as well as the top [Locus-to-Gene](/gentropy/locus-to-gene-l2g) assignment.

Because the same study might be processed through different fine-mapping pipelines, it's possible to observe credible sets obtained with different methods even within the same study. The [credible set exclusion criteria](/credible-set#credible-set-exclusion-criteria) describes the rules applied to avoid duplicated credible sets resulting from separate fine-mapping pipelines.

Source: [Open Targets](/gentropy/fine-mapping)


# Drug

## Overview

A drug in the Platform is understood as any bioactive molecule with drug-like properties as defined in the EMBL-EBI [ChEMBL](https://www.ebi.ac.uk/chembl/) database.

To further refine how we define a drug in the Platform, we subset ChEMBL's drugbase to capture molecules that meet the following criteria:

* Drugs for which the indication is known
* Drugs for which the target they modulate is identified
* Drugs that are listed in DrugBank
* Drugs that are acknowledged as chemical probes

A drug in the Platform might belong to different modalities, including small molecules, antibodies, or oligonucleotides among others. However, some biologic therapies such as vaccines, blood products, or cell therapies are not represented in our drug set. Moreover, the molecule-centric definition implies multi-ingredient drugs won't be represented and only their individual active moieties might be available on the site.

In the ChEMBL representation of drugs, a clear distinction is made between parent bioactive molecules and their corresponding child molecules. **Parent molecules** encompass the original, unmodified form of the active ingredient, while **child molecules** refer to modified versions, such as salts. In the Platform, both parent and child molecules are included, ensuring comprehensive coverage of the drug landscape.

When it comes to data propagation, parent molecules retain their own distinct information, as well as aggregate the specific details from all their child molecules. On the other hand, child molecules solely capture their own individual information, without incorporating any data from other molecules. This consistent approach extends to various aspects, including indications, mechanism of action, and drug warnings. By adopting this systematic framework, our Platform facilitates accurate representation and analysis of the diverse molecular entities and their associated properties.

{% embed url="<https://www.youtube.com/watch?v=2cHctti3aIQ>" %}

## Drug annotation data sources

| Annotation data                                                                                                                | Data source                                                                                                     |
| ------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------- |
| Molecule information, Mechanisms of action, Drug warnings (black box and withdrawn warnings)                                   | [ChEMBL](https://www.ebi.ac.uk/chembl/)                                                                         |
| [Clinical Report](/drug/clinical-report)                                                                                       | [Open Targets](https://github.com/opentargets/clinical_mining)                                                  |
| [Pharmacovigilance](/drug/pharmacovigilance)                                                                                   | [FAERS](https://www.fda.gov/drugs/surveillance/questions-and-answers-fdas-adverse-event-reporting-system-faers) |
| [Pharmacogenetics ](https://app.gitbook.com/o/-LC3OlEMulAutIN2QOro/s/-MU4dMxOmLaVNWfVNvpC/~/changes/445/drug/pharmacogenetics) | [ClinPGx](https://www.clinpgx.org/)                                                                             |
| [Bibliography](/bibliography)                                                                                                  | [Open Targets](/bibliography)                                                                                   |


# Clinical Report

A **clinical report** is a single, traceable piece of clinical evidence that links a drug to a disease or safety outcome. At a conceptual level, each clinical report captures four core questions:

* Which *drug* is this evidence about?
* Which *disease* or *safety outcome* is it linked to?
* Where does this evidence come from?
* What development or regulatory status does the source describe?

Clinical evidence is heterogeneous in format and origin, so a clinical report is not limited to a single source type. Depending on the source, a report may represent a registered clinical trial record, a regulatory medicine or approval record, a curated indication reference, or a curated warning or withdrawal-related record. Across all source types, the core requirement is the same: the record must support a **clinically meaningful link between a drug and a disease** (or safety outcome) with identifiable provenance.

Clinical reports are later combined into aggregated drug-disease or drug-target views.

{% hint style="info" %}
The clinical reports and all downstream datasets described in this page are generated by the **Clinical Mining Pipeline**. The pipeline is modular and designed to be extended to additional clinical data sources. The code is openly available on [GitHub](https://github.com/opentargets/clinical_mining).
{% endhint %}

### Data sources used to extract clinical reports

We currently integrate six source families. They are complementary: each source contributes a different evidence profile and helps provide a more complete picture of the *journey from discovery to market*.

<table><thead><tr><th width="249">Source</th><th>Description</th><th>Unit of evidence</th><th>Reference</th></tr></thead><tbody><tr><td>ClinicalTrials.gov via AACT</td><td>Structured registry records of clinical studies, including interventions and studied conditions.</td><td>A single trial record (one NCT ID), which may involve multiple drugs and multiple disease conditions</td><td><a href="https://aact.ctti-clinicaltrials.org/">AACT documentation</a></td></tr><tr><td>ChEMBL curated indications</td><td>Curated indication records, linked to external references (for example FDA, EMA, ATC, DailyMed, INN, USAN).</td><td>A single indication record, which may correspond to either a curated drug/indication pair (one drug, one disease) or a DailyMed medicine reference (one label, potentially covering multiple drugs and diseases)</td><td><a href="https://www.ebi.ac.uk/chembl/">ChEMBL database</a></td></tr><tr><td>ChEMBL drug warnings</td><td>Curated warning and withdrawal-oriented records linked to drugs.</td><td>A single warning record associated with a drug (either a black box warning or a withdrawal)</td><td><a href="https://www.ebi.ac.uk/chembl/">ChEMBL database</a></td></tr><tr><td>Therapeutic Target Database (TTD)</td><td>Curated drug and disease information from the Therapeutic Target Database.</td><td>A single drug/indication pair</td><td><a href="https://db.idrblab.net/ttd/">TTD</a></td></tr><tr><td>EMA Human Medicines</td><td>Regulatory evidence for authorised human medicines and therapeutic-use context in Europe.</td><td>A single medicine label, which may cover one or more active ingredients and one or more approved indications</td><td><a href="https://www.ema.europa.eu/en/medicines">EMA medicines</a></td></tr><tr><td>PMDA approvals</td><td>Public approvals information from Japan's Pharmaceuticals and Medical Devices Agency.</td><td>A single drug/indication pair</td><td><a href="https://www.pmda.go.jp/english/review-services/reviews/approved-information/drugs/0002.html">PMDA approved drugs information</a></td></tr></tbody></table>

### Entity extraction for clinical trials

Most clinical reports link a drug to a disease. Depending on the source, this link may already be structured and curated, as with ChEMBL indications, regulatory approvals from the EMA and PMDA, or TTD records. Alternatively, it may need to be extracted from free-text and semi-structured trial metadata, as is the case for clinical trials from AACT.

ClinicalTrials.gov trial records include free-text condition and intervention fields, in the form of MeSH terms that are assigned upon submission. Previously, the pipeline used these MeSH terms directly to identify the diseases and drugs involved in each trial. While this approach works well for straightforward, single-indication trials, MeSH assignment is inconsistent for complex studies where the condition terms may describe the patient population rather than the target indication, or list comorbidities and background conditions alongside the primary disease under investigation.

From release 26.06, drug and disease entities for AACT records are extracted relying on large language models. The model receives the full trial context (title, description, intervention and condition fields, and any linked literature references) and identifies the drugs being investigated and the diseases they are being studied against. The entities identified by the LLM are subsequently mapped to ChEMBL and EFO identifiers using the same normalisation steps applied to all other sources.

{% hint style="info" %}
The extraction method described here applies to **AACT (ClinicalTrials.gov)** reports only. All other source families supply structured drug and disease labels directly. Prompt template and extraction logic are documented in the pipeline repository on [GitHub](https://github.com/opentargets/clinical_mining#2-llm-extraction).
{% endhint %}

### Clinical stage categories

Each source reports development or regulatory status using its own terms. To support cross-source analysis, Open Targets harmonises source-reported values into a **shared clinical-stage framework**.

This harmonisation has two goals:

* make stage labels *comparable* across heterogeneous sources;
* support consistent *ranking* of evidence in downstream clinical precedence views.

The harmonised framework includes the following categories:

| Category          | Description                                                                                                                                                               |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Withdrawal**    | Evidence that a medicine was withdrawn, revoked, lapsed, suspended, or otherwise removed from use                                                                         |
| **Approval**      | Evidence of marketing authorisation or approved status from regulatory or authoritative sources                                                                           |
| **Phase IV**      | Post-marketing interventional evidence from trials conducted after regulatory approval                                                                                    |
| **Preapproval**   | Late-stage regulatory submission evidence before full marketing authorisation, including submitted applications, formal opinions, and equivalent pre-market review stages |
| **Phase III**     |                                                                                                                                                                           |
| **Phase II/III**  | Evidence from trials spanning or bridging mid- to late-stage clinical development                                                                                         |
| **Phase II**      | Mid-stage clinical development evidence (Phase II, including subphases such as IIa/IIb)                                                                                   |
| **Phase I/II**    | Evidence from trials spanning or bridging early to mid-stage clinical development                                                                                         |
| **Phase I**       | Early human clinical development evidence (Phase I, including subphases such as Ib)                                                                                       |
| **Early Phase I** | Exploratory early human studies conducted prior to standard Phase I, typically with a limited number of participants and a primary focus on safety or pharmacokinetics    |
| **IND**           | Investigational New Drug application or equivalent regulatory filing authorising first-in-human studies; the drug has not yet entered clinical trials                     |
| **Preclinical**   | Evidence reported as preclinical or patented/preclinical development                                                                                                      |
| **Unknown**       | Source value is missing, ambiguous, not mappable, or not directly comparable to standard stage labels                                                                     |

### Reason to stop categories

For clinical reports derived from clinical trials, we integrate a machine learning-based classification of the reasons *why a trial ended earlier* than expected. The model was trained on free-text stop reasons from 28,842 stopped trials on ClinicalTrials.gov and classifies them into 17 categories covering negative, neutral, and positive reasons for stoppage. The model is [available on Hugging Face](https://huggingface.co/opentargets/clinical_trial_stop_reasons).

The 17 classes are: *Another Study, Business or Administrative, Negative, Study Design, Invalid Reason, Ethical Reason, Insufficient Data, Insufficient Enrolment, Study Staff Moved, Endpoint Met, Regulatory, Logistics or Resources, Safety and Side Effects, No Context, Success, Interim Analysis, and Covid 19.*&#x20;

This classification is used to [down-weight evidence](/evidence#clinical-precedence) from trials that stopped for negative reasons in downstream scoring.

**Reference:** [*Razuvayevskaya et al., Nature Genetics, 2024*](https://www.nature.com/articles/s41588-024-01854-z)

### Quality filters

Not all clinical reports are used to derive downstream datasets. Before aggregation into clinical indications, clinical targets, or target/disease evidence, each report is evaluated against a set of quality criteria. Reports that fail any of these checks are flagged and excluded from downstream use.

Three exclusion criteria are currently applied:

* Phase IV reports with indications for which we don't have any approval
* Clinical trial reports in which the primary purpose of the study is not to directly measure the effect of the intervention on the target condition or its symptomatology.

Flagged reports remain accessible at the clinical report level but do not contribute to clinical indications, clinical targets, or target/disease evidence.


# Indications

The **Indications** widget provides a drug-centric view of the same Clinical Indications dataset described in the [Drugs and Clinical Candidates](/disease-or-phenotype/drugs) section of the Drug page. In this case, the drug is fixed, and the widget displays all diseases for which it has clinical evidence, aggregated across sources and [development stages](/drug/clinical-report#clinical-stage-categories).

The same aggregation logic applies: clinical reports sharing the same drug and disease are consolidated into a single record, and the maximum clinical stage is assigned using the same ranking rules.

```mermaid
flowchart LR
    CT["<b>Clinical trial</b> <br>NCT02684006 · Phase III</br>"]
    DL["<b>Drug label</b> <br>FDA label · Approval</br>"]
    CI["<b>Curated indication</b><br>ChEMBL · Phase II</br>"]
    DW["<b>Drug warning</b><br>ChEMBL · Withdrawal</br>"]

    CLIN["<b>Clinical indication</b><br>Drug A · Disease B</br>4 contributing reports</br>Sources: trial, label, ChEMBL"]

    APP["<b>Approval</b>"]

    CT --> CLIN
    DL --> CLIN
    CI --> CLIN
    DW --> CLIN
    CLIN --> APP

    subgraph reports["Clinical Reports"]
        CT
        DL
        CI
        DW
    end

    subgraph indication["Clinical indication"]
        CLIN
    end

    subgraph maxstage["Max stage"]
        APP
    end

    style CT fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style DL fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style CI fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style DW fill:#c8f0e8,stroke:#3a9e82,color:#1a5c4a
    style CLIN fill:#dcd8f5,stroke:#7b6fc4,color:#3b2e8a
    style APP fill:#fce8dc,stroke:#c97a50,color:#8b3a10
    style reports fill: transparent, stroke: transparent
    style indication fill: transparent, stroke: transparent
    style maxstage fill: transparent, stroke: transparent
```

{% hint style="info" %}
For a full description of the **clinical stage categories** and their ranking, see [Clinical stage categories](/drug/clinical-report#clinical-stage-categories) in the Clinical Report page.
{% endhint %}


# Pharmacovigilance

Data describing pharmacovigilance annotation for the drug

## Overview

Before approval, new therapeutic drug treatments are extensively tested in clinical trials. However, some of the side effects are only identified when prescribed to larger cohorts of patients, with one or more medical conditions, for a sustained period of time or in combination with other treatments.

For this reason, regulatory agencies—for example, the Food & Drug Administration or the European Medicines Agency—provide pharmacovigilance programs to monitor and survey Adverse Drug Reactions (ADRs).

## Data sources

### FDA Adverse Event Reporting System (FAERS) <a href="#hero-title" id="hero-title"></a>

The FDA Adverse Event Reporting System (FAERS) – [https://open.fda.gov/data/faers](https://open.fda.gov/data/faers/) – is a database that contains millions of public reports with information on adverse event and medication error reports submitted to FDA. The database is designed to support the FDA's post-marketing safety surveillance program for drug and therapeutic biologic products. Adverse events and medication errors are mapped to terms in the Medical Dictionary for Regulatory Activities (MedDRA) terminology.

## Computational pipelines and datasets

While recurrence of a given adverse event is relevant, it's the specificity of the event to the drug what might flag concerns. In order to get a list of significant drug–ADRs associations, we have implemented an analysis similar to the one described by [Maciejewski et al. (2017)](https://europepmc.org/abstract/MED/28786378).

First we apply a set of filters to the reports as described below:

* Only reports submitted by health professionals (*primarysource.qualification in (1,2,3)*)
* Exclude reports that resulted in death (no entries with *seriousnessdeath=1*)
* Only drugs that were considered by the reporter to be the cause of the event *(drugcharacterization=1)*
* Remove [blacklisted](https://github.com/opentargets/platform-etl-openfda-faers/blob/master/src/main/resources/blacklisted_events.txt) events curated manually to exclude uninformative events

Next, we sought to map the drugs in the FAERS reports to the drugs in the Open Targets Platform (ChEMBL IDs). Any of the above listed fields were used when exact matches were available:

| FAERS drugs                   | Open Targets Platform drugs |
| ----------------------------- | --------------------------- |
| `drug.medicinalproduct`       | `drug.medicinalproduct`     |
| `drug.openfda.generic_name`   | `synonyms`                  |
| `drug.openfda.brand_name`     | `pref_name`                 |
| `drug.openfda.substance_name` | `trade_names`               |

The significant drug–ADR pairs were then evaluated using the Likelihood Ratio Test (LRT) as previously described by [Huang et al. (2011)](https://europepmc.org/abstract/MED/23331230). The significance of a given drug–ADR is implicitly corrected by how often a drug is found in a report and how often an event is reported across drugs. This way, we prevent the drug–ADR associations to be biased by overrepresented ADRs (e.g. headache, nausea) or drugs (e.g. paracetamol, ibuprofen). In order to assess significance, an LRT critical value for every drug is calculated using an empirical Monte Carlo simulation, similar to the one implemented by [openFDA](https://openfda.shinyapps.io/LRTest/_w_c5c2d04d/lrtmethod.pdf).

Due to the nature of the surveillance reports, it's relatively common for the indication for which a drug was prescribed to appear in the list of significant ADRs. Given the current structure of the data provided in a FAERS report, we cannot distinguish whether it's a problem with the dosage the drug was prescribed or an excessive phenotypic characterisation of the patient in the report.

All pharmacovigilance data is available for download on [our data downloads page](https://platform.opentargets.org/downloads).

## Publications

Huang L, Zalkikar J, Tiwari RC. **Likelihood ratio test-based method for signal detection in drug classes using FDA's AERS database**. J Biopharm Stat. 2013;23(1):178-200. doi: [10.1080/10543406.2013.736810](https://doi.org/10.1080/10543406.2013.736810). PMID: [23331230](https://pubmed.ncbi.nlm.nih.gov/23331230/).

Maciejewski M, Lounkine E, Whitebread S, Farmer P, DuMouchel W, Shoichet BK, Urban L. **Reverse translation of adverse event reports paves the way for de-risking preclinical off-targets**. Elife. 2017 Aug 8;6:e25818. doi: [10.7554/eLife.25818](https://doi.org/10.7554/elife.25818). PMID: [28786378](https://pubmed.ncbi.nlm.nih.gov/28786378/); PMCID: [PMC5548487](https://europepmc.org/article/PMC/PMC5548487).


# Pharmacogenetics

Data supporting pharmacogenetics annotation for a drug

## Overview

Pharmacogenetics is the study of how genetic variation may change your response to a specific drug. Through the integration of pharmacogenetics data into the Platform, our objective is to apply clinical annotations available in the Pharmacogenetics Knowledgebase (ClinPGx) to aid target prioritisation for drug discovery.

We also enhance these annotations by adding detailed annotations on variant consequence prediction, drug response categories, and specific drug information and whether the gene is a direct target of the drug. This process involves applying advanced phenotype extraction techniques to offer a refined representation of phenotypes, thereby providing a more precise understanding of the genetic determinants influencing treatment outcomes.

## Data sources

[ClinPGx](https://www.clinpgx.org/) is an NIH-funded comprehensive resource that provides information about how human genetic variation affects response to medications. ClinPGx collects, curates and disseminates knowledge about clinically actionable gene–drug associations and genotype-phenotype relationships, focusing on the impact of genetic variation on drug response for clinicians and researchers.

We have modelled and harmonised the data from the[ ](https://www.pharmgkb.org/clinicalAnnotations)[Clinical Annotation](https://www.pharmgkb.org/clinicalAnnotations) section of ClinPGx which provides information about variant–drug pairs based upon collating and summarising variant annotations. Variant annotations are curated manually from a scientific publication. Clinical Annotations then provide an overarching summary and curated [level of evidence](https://www.pharmgkb.org/page/clinAnnLevels) for the association between a genetic variant and particular drug responses based on multiple variant annotations. The likely consequence for each genotype of the variant on drug response is represented, in comparison to the other genotypes of that variant. Variant-level annotation (e.g. direction of effect) is also presented, when available.

We have also included annotations for star (\*) alleles, a nomenclature used in the pharmacogenetics field for representing key functional variants involved in drug responses.

## Publications

When using this data please remember to acknowledge the sources:

Whirl-Carrillo M, Huddart R, Gong L, Sangkuhl K, Thorn CF, Whaley R, Klein TE. **An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine.** Clin Pharmacol Ther. 2021 Sep;110(3):563-572. doi: 10.1002/cpt.2350. Epub 2021 Jul 22. PMID: 34216021; PMCID: PMC8457105.

Whirl-Carrillo M, McDonagh EM, Hebert JM, Gong L, Sangkuhl K, Thorn CF, Altman RB, Klein TE. **Pharmacogenomics knowledge for personalized medicine**. Clin Pharmacol Ther. 2012 Oct;92(4):414-7. doi: 10.1038/clpt.2012.96. PMID: 22992668; PMCID: PMC3660037.

[See all relevant PharmGKB publications.](https://www.pharmgkb.org/page/citingPharmgkb)


# Credible Set

## Overview

A **credible set** is a set of genetic variants near a genetic association signal that is predicted, with a specific probability, to include the causal variant for that signal. The results of the fine-mapping analysis determine this, assigning each variant in the region a posterior probability of being causal when considering the observed statistics and the population structure. The variants covering the top 95% likelihood of containing the causal variant define the credible sets in the Platform.

A credible set results from statistical analysis on a specific locus in a study. As a consequence, all credible sets are defined as:

* **Study** in which the association is reported
* **Lead variant -** Variant with the highest posterior probability in the credible set
* **Fine-mapping** method and statistics

The Platform contains every credible set resulting from [fine-mapping](/gentropy/fine-mapping) all our [sources](/gentropy/data-sources) after applying certain [exclusion criteria](#credible-set-exclusion-criteria).

{% hint style="info" %}
**Credible set identifier**

Different from other entities in the Open Targets Platform, credible sets are identified with an alphanumeric list of characters that have no semantic meaning. The identifier is derived from a combination of fields that define a credible set's uniqueness, and it will remain the same as long as the credible set metadata hasn't changed.&#x20;
{% endhint %}

## Credible set exclusion criteria

Credible sets fulfilling any of the next rules are excluded from the Platform:

1. The lead variant is within the MHC region (chr6:25726063-33400556)
2. The credible set is not in a [valid study](/study#inclusion-criteria)
3. There is another fine-mapped SuSiE credible set from the same region and study
4. Being a GWAS Catalog fine-mapped top-hit, there is a GWAS Catalog fine-mapped credible set from summary statistics for the same region and study
5. The lead variant is reported in an invalid chromosome (1:22, X, Y, XY, MT)
6. The sum of PIPs in the credible is not within the \[0.95,1] range

## Credible set confidence

Credible sets are categorised based on the fine-mapping confidence, derived from the available association data, [fine-mapping](/gentropy/fine-mapping) framework and availability of linkage disequilibrium (LD) information for the specific study/population . The categories in descending order of confidence are:&#x20;

* SuSiE or SuSiE-inf fine-mapped credible set with in-sample LD
* SuSiE or SuSiE-inf fine-mapped credible set with out-of-sample LD
* PICS fine-mapped credible set extracted from summary statistics
* PICS fine-mapped credible set based on reported curated association
* Unknown confidence

The confidence values are symbolised by 0 to 4 stars on the UI. It is important to note that this confidence does not reflect the strength of the association or the effect size.

## Credible set variants

All variants in the credible set are annotated with:

* **P-value**. Unconditioned p-value from the study when available.
* **Beta**. Corresponds to SuSiE *mu* (SuSiE fine-mapped GWAS Catalog, UKBB-PPP and FinnGen) or beta (PICS fine-mapped GWAS Catalog and SuSiE fine-mapped eQTL Catalog)&#x20;
* **Standard error**. Only available for PICS-fine-mapped credible sets.
* **LD (r^2)**. Linkage-disequilibrium information. Only available for PICS-fine-mapped credible sets.
* **Posterior probability.** Posterior inclusion probability (PIP) of variant being causal after fine-mapping.
* **log(BF)**. The logarithm of the Bayes factor. Only available for SuSiE credible sets.
* **Predicted consequence**. The most severe consequence is across all overlapping canonical transcripts, as reported by the Ensembl VEP.

## Locus-to-Gene (L2G)

Machine Learning prioritisation of likely causal genes based on available features. L2G integrates multiple features to predict what's the most likely causal gene in the neighbourhood of the observed association. All predictions for protein-coding genes with a score above 0.05 are displayed. See [Locus-to-Gene](/gentropy/locus-to-gene-l2g) section for a description of the methodology.

### Explaining L2G predictions

Using the [SHAP](https://shap.readthedocs.io/en/latest/example_notebooks/overviews/An%20introduction%20to%20explainable%20AI%20with%20Shapley%20values.html) (SHapley Additive exPlanations) library, we have extracted feature importance values for all L2G predictions. Shapley values provide a principled approach based on game theory to explain the contribution of individual features or groups of features, revealing how each group influences the final L2G score. These contributions are approximated to be additive, meaning the sum of the Shapley values for all feature groups equals the total L2G score or reasonably close to it.

<figure><img src="/files/FNtIoxGreUOZgkz0slbL" alt=""><figcaption></figcaption></figure>

<sub>L2G prioritisation for a</sub> [<sub>credible set linking to psoriasis</sub>](https://platform.opentargets.org/credible-set/b8a67437f19eb1607f9219ea17adebe7)

We have aggregated the Shapley values into the main feature groups to understand their relative importance.

> The **base value** represents the baseline before any feature-specific information is considered, and is therefore equivalent for all genes and credible sets.

This approach helps to identify which types of evidence (e.g. distance, colocalisation or functional impact) are most influential for a given locus-gene association. In addition, we can visualise the individual contribution of each feature within the group. As the features within the group are highly correlated, the individual values are less interpretable than the group contribution.

{% hint style="warning" %}
**Additivity of SHAP values**

Because the SHAP analysis is an approximation, in some occasions, the sum of all shapley values might not result in the L2G score.
{% endhint %}

A full description of all features is available here.

**Source**: [Open Targets](/gentropy/locus-to-gene-l2g)

## Enhancer-to-gene predictions

Enhancer-to-gene widget shows the genes relevant to the credible set as determined by overlap with epigenetically derived datasources.  This widget is equivalent to the [enhancer-to-gene widget](/gentropy/enhancer-to-gene-encode-re2g) found on the variant page of the lead variant in the credible set.&#x20;

## Colocalisation

Credible sets are compared against other credible sets to find overlapping signals. Two overlapping credible sets are those that share at least one variant in the set. The Platform contains all the overlaps between all GWAS vs all GWAS and all GWAS vs all molQTL studies.

For the overlapping pairs of credible sets, estimates for two colocalisation methods are computed:

* [COLOC-PIP](/gentropy/colocalisation#coloc-colocalisation)
* [eCAVIAR](/gentropy/colocalisation#ecaviar-colocalisation)&#x20;

Source: [Open Targets](/gentropy/colocalisation)

### Colocalisation directionality&#x20;

A directionality assessment is included in the colocalisation analysis to help the user interpret the relationship between the two overlapping credible sets.

For every overlapping variant in a pair of overlapping credible sets, the sign of the ratio between both [beta](#credible-set-variants) estimates are calculated (+1 or -1) indicating the individual variant directionality. The average of these signs across all overlapping variants in both credible sets is used to estimate the sign estimates. When the average approximates to +1, the two credible sets are assessed to share the **same** directionality. If the average approximates to -1, the two credible sets are interpreted to have **opposite** directionality. In any other case, the assessment is declared inconclusive (N/A).

Source: [Open Targets](/gentropy/colocalisation)


# Target–disease evidence

## Target–disease evidence

Every event or set of events pinpointing a target as a potential causal gene or protein for a disease represents the unit of information, most often referred to as **evidence**. Within the Open Targets Platform, a series of pipelines ensure information is retrieved from its sources and standardised in a way that can be immediately applied to answer drug development queries.

All evidence is mapped to the reference target entity identifier (Ensembl gene) and disease or phenotype identifier (experimental factor ontology, EFO), as well as other reference controlled vocabularies and ontologies when appropriate. Evidence is also reviewed to minimise the presence of duplicates within the same data source.

Data sources are also grouped into bigger categories abstracting the type of evidence they predominantly capture. In the platform, these categories are usually referred to as data types, as opposed to the individual resource data referred to as **data sources**.

The Open Targets Platform provides a scoring framework for each data source to contextualise the relative importance of each piece of evidence. This score will be more relevant when understanding the association scoring in later sections.

## Evidence data sources

### Clinical Precedence

Clinical evidence in this data source represents any target/disease relationship that can be explained by a drug targeting the gene product and indicated for the disease, whether approved or in clinical development. Each piece of evidence corresponds to a single clinical report: a record linking a drug to a disease with an identifiable source and development stage. The target/disease link is inferred by joining the drug/disease relationship captured in the clinical report with drug mechanism of action data.

Clinical reports are drawn from multiple source families, including clinical trial registries, regulatory drug labels, curated indication references, and drug warning records.

{% hint style="info" %}
For a full description of the **source types** or **clinical stage categories**, see the [Clinical Report](/drug/clinical-report) page.
{% endhint %}

**Data type**: Clinical

**Evidence scoring:** Clinical precedence evidence is scored in a 2-step process. In Step 1, a score is assigned to every piece of evidence based on the clinical stage:

| Clinical Precedence                      | Evidence score |
| ---------------------------------------- | -------------- |
| Unknown                                  | 0.01           |
| Preclinical                              | 0.01           |
| IND                                      | 0.05           |
| Early Phase I                            | 0.05           |
| Phase I                                  | 0.1            |
| Phase I/II                               | 0.15           |
| Phase II                                 | 0.2            |
| Phase II/III                             | 0.5            |
| Phase III                                | 0.7            |
| Preapproval                              | 0.8            |
| Approval                                 | 1.0            |
| Phase IV (only for approved indications) | 1.0            |
| Withdrawal                               | 1.0            |

In Step 2, for those clinical trials that have stopped early, the score is down-weighted based on the [classification of the reason to stop](/drug/clinical-report#reason-to-stop-categories). In this way, less importance is attributed to evidence of studies that have been stopped due to negative outcomes or safety concerns:

| Reason to stop class   | Score weight |
| ---------------------- | ------------ |
| Negative               | 0.5          |
| Safety or side effects | 0.5          |

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th align="center">Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th align="center">Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td align="center"><p>Activators = GoF</p><p>Inhibitors = LoF</p></td><td align="center">Assumption of Protective</td></tr></tbody></table>

**Source**: [Open Targets Clinical Mining](https://github.com/opentargets/clinical_mining)

### GWAS associations

The GWAS associations data source aggregates target-disease relationships supported by significant genome-wide associations (GWAS) in the context of other functional genomics data.

The evidence in this data source results from a comprehensive statistical genetics analysis described in [GWAS and functional genomics](/gentropy) section. The aim of this analysis is to identify GWAS-significant signals across an [extensive set](/gentropy/data-sources) of GWAS studies covering binary and quantitative traits. To address linkage disequilibrium, all significant signals are [fine-mapped](/gentropy/fine-mapping) and the resulting credible sets [colocalised](/credible-set#colocalisation) against molQTL studies. All GWAS and functional genomic features are leveraged by the [Locus-to-Gene](/gentropy/locus-to-gene-l2g) machine-learning method aimed to prioritise likely causal genes in the region.

The GWAS association evidence is defined as any credible set in a GWAS trait associated with a gene with a Locus2Gene (L2G) > 0.05. The feature contributions for the L2G predictions are also [explained](/credible-set#explaining-l2g-predictions) by SHAP analysis helping with the interpretation of the observed features. All credible sets can also be futher interrogated in their own [credible set](/credible-set) page, including an interpretation of the directionality in the context of colocalising molQTL studies.

**Data type**: Genetic association

**Evidence scoring**: [Locus-to-Gene score](/gentropy/locus-to-gene-l2g), filtered to use scores above 0.05

### Gene Burden

Gene burden data comprises gene–phenotype relationships observed in gene-level association tests using rare variant collapsing analyses. The Platform integrates burden tests carried out by several sources:

* **REGENERON** (Backman et al., 2021), a whole-exome sequencing analysis of individuals from the UK Biobank.
* **AstraZeneca PheWAS Portal** (Wang et al., 2021), a whole-exome sequencing analysis of individuals from the UK Biobank.
* **Genebass** (Karczewski et al., 2022): Gene-based Association Summary Statistics (Genebass),  a whole-exome sequencing analysis of individuals from the UK Biobank.
* The results of whole-exome and whole-genome sequencing analysis based on the **SPARK cohort** bring evidence of novel targets implicated in **autism spectrum disorder** (Zhou et al., 2022).
* **The SCHEMA consortium** (Singh et al., 2022), a whole-exome sequencing analysis of individuals with **schizophrenia**.
* **The Epi25 collaborative** (Epi25 Collaborative, 2019), a whole-exome sequencing analysis of individuals with **epilepsy**.
* **The Autism Sequencing Consortium** (Satterstrom et al., 2020), a whole-exome sequencing analysis of individuals with **autism spectrum disorder.**
* The results of an **Open Targets project** (Bomba et al., 2022), a whole-exome sequencing analysis of individuals from the INTERVAL cohort testing for associations between rare coding variants and **blood metabolites**.
* The results of a pan-ancestry whole-exome sequencing analysis identify relevant genes associated with **fat distribution** (Akbari et al., 2022).
* The results of whole-exome and whole-genome sequencing analysis on **Parkinson disease** and promoted by the **AMP-PD initiative**, and other collaborators (Makarious et al., 2022).
* The results of gene-based analyses of rare variants and circulating metabolic biomarkers relevant to **cardiovascular disease** (Riveros-McKay et al., 2020).
* The results of rare coding variant analyses from whole exome sequencing of Black South African men to identify genes significantly associated with **prostate cancer** (Soh et al., 2023)
* The **FinnGen (R12)** gene-based burden test results from collapsing loss of function variants, based on **genotyping** data from the Finnish population. [Find out more in their documentation](https://finngen.gitbook.io/documentation/methods/lof-variant-burden).
* The **Broad CVDI Human Disease Portal** is a pan-ancestry whole-exome and whole-genome sequencing analysis combining data from three major biobanks: the **UK Biobank**, **All of Us**, and the **Mass General Brigham Biobank**. This comprehensive study analysed rare variant associations across nearly 750,000 individuals, identifying 363 significant gene-disease associations for 123 genes across 165 diseases. The analysis includes both cohort-specific burden tests and meta-analyses across ancestries (Jurgens et al., 2024).

These associations are a result of collapsing rare variants in a gene into a single burden statistic and regress the phenotype on the burden statistic to test for the combined effects of all rare variants in that gene. The different collapsing methods inform about the filters used to select the set of qualifying variants, mostly based on their pathogenicity and frequency in the population.

**Data type**: Genetic association

**Evidence scoring:** Scaled p-value from 0.25 (p = 1e-7) to 1 (p < 1e-17).

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th width="331" align="center">Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th align="center">Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td align="center">Assumption of all variants LoF</td><td align="center"><p>               Beta values &#x3C; 0 = Protective </p><p>   Beta values >0 = Risk</p></td></tr><tr><td align="center"></td><td align="center"><p>           Odds ratios &#x3C;1=Protective </p><p>Odds ratios>1 =Risk</p></td></tr></tbody></table>

**Source:** [AstraZeneca PheWAS Portal](https://azphewas.com), [Genebass](https://app.genebass.org)

**References:** [Wang, Q. et al, 2021](https://doi.org/10.1038/s41586-021-03855-y); [Backman, J.D. et al, 2021](https://doi.org/10.1038/s41586-021-04103-z); [K.K., Karczewski et al., 2022](https://www.medrxiv.org/content/10.1101/2021.06.19.21259117v4); [Zhou X. et al, 2022](https://www.nature.com/articles/s41588-022-01148-2), [Singh et al., 2022](https://rdcu.be/cPZP3); [Epi25 Collaborative, 2019](https://www.cell.com/ajhg/fulltext/S0002-9297\(19\)30207-1); [Satterstrom et al., 2020](https://www.sciencedirect.com/science/article/pii/S0092867419313984); [Bomba et al., 2022](https://www.cell.com/ajhg/fulltext/S0002-9297\(22\)00157-4); [Akbari, P., 2022](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9399235/); [Makarious et al., 2022](https://www.medrxiv.org/content/10.1101/2022.11.08.22280168v1.full.pdf); [Riveros-McKay et al., 2020](https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1008605); [Soh et al., 2023](https://doi.org/10.1038/s41467-023-43726-w); [Jurgens et al., 2024](https://doi.org/10.1038/s41588-024-01894-5)

### ClinVar

ClinVar is an NIH public archive of reports of the relationships among human variations and phenotypes, with supporting evidence. The ClinVar data source in the Open Targets Platform captures the subset of ClinVar that refers to germline variants (as opposed to somatic variants). Each evidence in the platform aims to capture an individual RCV record in ClinVar.

Information on variants is covered extensively for both single point and structural variants. When available, genomic coordinates are reported with RS numbers, or by following the CHROM\_POS\_REF\_ALT and HGVS notations.

**Data type**: Genetic association

**Evidence scoring**: ClinVar evidence is scored in a 2-step process. In Step 1, a score is assigned to every piece of evidence based on the clinical significance:

| Clinical significance                        | Evidence score |
| -------------------------------------------- | -------------- |
| association not found                        | 0              |
| benign                                       | 0              |
| not provided                                 | 0              |
| likely benign                                | 0              |
| conflicting data from submitters             | 0.3            |
| conflicting interpretations of pathogenicity | 0.3            |
| low penetrance                               | 0.3            |
| other                                        | 0.3            |
| uncertain risk allele                        | 0.3            |
| uncertain significance                       | 0.3            |
| established risk allele                      | 0.5            |
| risk factor                                  | 0.5            |
| affects                                      | 0.5            |
| likely pathogenic                            | 0.7            |
| association                                  | 0.9            |
| confers sensitivity                          | 0.9            |
| drug response                                | 0.9            |
| protective                                   | 0.9            |
| pathogenic                                   | 0.9            |

In Step 2, the score is modulated based on the ClinVar review status:

| Confidence                                           | Evidence score modifier |
| ---------------------------------------------------- | ----------------------- |
| no assertion provided                                | +0                      |
| no assertion criteria provided                       | +0                      |
| no assertion for the individual variant              | +0                      |
| criteria provided, single submitter                  | +0.02                   |
| criteria provided, conflicting interpretations       | +0.02                   |
| criteria provided, multiple submitters, no conflicts | +0.05                   |
| reviewed by expert panel                             | +0.07                   |
| practice guideline                                   | +0.1                    |

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th align="center">Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th align="center">Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td align="center">LoF variants</td><td align="center"><p>Pathogenic/Likely pathogenic = Risk</p><p>Protective = Protective</p></td></tr></tbody></table>

**Source**: [ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/) (via [European Variation Archive](https://www.ebi.ac.uk/eva/))

**References**: [Cezard T. et al, 2021](https://doi.org/10.1093/nar/gkab960); [Shen A. et al, 2024](https://doi.org/10.1093/bioadv/vbae018); [Landrum, M. et al, 2014](https://doi.org/10.1093/nar/gkt1113); [Landrum, M. et al, 2020](https://doi.org/10.1093/nar/gkz972)

### Genomics England (GEL) PanelApp

The Genomics England PanelApp is a knowledge base that combines crowdsourced expertise with curation to provide gene–disease relationships. Virtual gene panels related to human disorders are reviewed by experts within the clinical and scientific community to support the interpretation of genomes within the 100,000 Genomes Project. Within a panel, genes are rated based on the level of evidence supporting the association with the phenotypes identified by the panel. Genes are then classified according to a traffic light system with red/stop, amber/pause, and green/go classifications. To receive a green rating (diagnostic-grade) on a version 1+ panel, the gene requires "evidence from 3 or more unrelated families or from 2-3 unrelated families where there is strong additional functional data" and "genes that do not meet these criteria are rated as Amber (borderline) or Red (low level of evidence)."

The Open Targets Platform includes "green" and "amber" genes from version 1+ panels along with their phenotypes, providing the latter can be mapped to a disease or phenotype ontology. As we standardise our evidence to EFO, some of the phenotypes cannot be mapped and included in the Platform; please visit the [Genomics England PanelApp website](https://panelapp.genomicsengland.co.uk) for the full set.

**Data type**: Genetic association

**Evidence scoring**: Based on Genomics England gene rating:

| Gene Rating in GEL Panel | Evidence score |
| ------------------------ | -------------- |
| Amber                    | 0.5            |
| Green                    | 1              |

**Source**: [Genomics England PanelApp](https://panelapp.genomicsengland.co.uk)

**References**: [Martin, A. et al, 2019](https://doi.org/10.1038/s41588-019-0528-2)

### Gene2Phenotype

The data in Gene2Phenotype (G2P) is produced and curated from the literature by different sets of panels formed by consultant clinical geneticists. The G2P data is designed to facilitate the development, validation, curation, and distribution of large-scale, evidence-based datasets for use in diagnostic variant filtering. Each G2P entry associates an allelic requirement and a mutational consequence at a defined locus with a disease entity. A confidence level and evidence link are assigned to each entry. This confidence level follows the terminology described by [GenCC](https://thegencc.org/about.html) for describing gene–disease validity.

G2P evidence in the Platform is the result of any target-disease curation by any of the expert panels.

**Data type:** Genetic association

**Evidence scoring:**

| Gene2Phenotype confidence | Evidence score |
| ------------------------- | -------------- |
| Limited                   | 0.01           |
| Moderate                  | 0.5            |
| Strong                    | 1              |
| Both RD and IF            | 1              |
| Definitive                | 1              |

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th>Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th>Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td>LoF and GoF variants</td><td>Assumption of Risk</td></tr></tbody></table>

**Source**: [Gene2Phenotype](https://www.ebi.ac.uk/gene2phenotype)

**References**: [Thormann, A. et al, 2019](https://doi.org/10.1038/s41467-019-10016-3)

### UniProt literature

The Universal Protein Resource (UniProt) provides a large compendium of sequence and functional information at the protein level. As part of their functional annotation effort, UniProt curators also annotate proteins with publications supporting their involvement on pathogenic processes.

All publications supporting a given target disease relationship are aggregated into one single Platform evidence.

**Data type**: Genetic association

**Evidence scoring**:&#x20;

| **Uniprot confidence** | Evidence score |
| ---------------------- | -------------- |
| Medium                 | 0.5            |
| High                   | 1              |

**Source**: [UniProt](https://www.uniprot.org)

**References**: [The UniProt Consortium, 2021](https://academic.oup.com/nar/article/49/D1/D480/6006196)

### UniProt curated variants

The Universal Protein Resource (UniProt) also curate variants supported by publications that are known to alter protein function on disease. Curated mutations are predominantly protein coding or in regulatory regions clearly associated with the causal protein.

All publications supporting a given variant in connection with a disease constitute individual evidence. All supporting publications are aggregated within the same evidence.

**Data type**: Genetic association

**Evidence scoring**:&#x20;

| **UniProt confidence** | Evidence score |
| ---------------------- | -------------- |
| Medium                 | 0.5            |
| High                   | 1              |

**Source**: [UniProt](https://www.uniprot.org)

**References**: [The UniProt Consortium, 2021](https://academic.oup.com/nar/article/49/D1/D480/6006196)

### Orphanet

Orphanet is an international network that offers a range of resources to improve the understanding of rare disorders of genetic origin. These resources include an inventory of rare disease and gene associations, classification of the gene–disease relationship, information on the kind of mutation, and supporting publication references.

**Data type**: Genetic association

**Evidence scoring:**

| Orphanet Disorder Gene Association Status | Evidence score |
| ----------------------------------------- | -------------- |
| Not yet assessed                          | 0.5            |
| Assessed                                  | 1              |

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th>Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th>Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td>LoF and GoF variants</td><td>Assumption of Risk</td></tr></tbody></table>

**Source**: [Orphanet Genes Associated with Rare Diseases](https://www.orpha.net/consor/cgi-bin/Disease_Genes.php?lng=EN)

**References**: [Orphanet](https://www.orpha.net); [Orphadata](http://www.orphadata.org/cgi-bin/index.php)

### ClinGen

The Clinical Genome Resource (ClinGen) Gene–Disease Validity Curation aims to evaluate the strength of evidence supporting or refuting a claim that variation in a particular gene causes a particular disease. ClinGen provides a framework of guidelines to assess clinical validity in a semi-quantitative manner allowing curators to classify the validity of given gene–disease pair.

All gene–disease pairs mapped to EFO constitute individual evidence in the Platform.

**Data type**: Genetic association

**Evidence scoring:**

| **ClinGen classification** | Evidence score |
| -------------------------- | -------------- |
| No reported evidence       | 0.01           |
| Refuted                    | 0.01           |
| Disputed                   | 0.01           |
| Limited                    | 0.01           |
| Moderate                   | 0.5            |
| Strong                     | 1              |
| Definitive                 | 1              |

**Source**: [ClinGen Gene-Disease Validity](https://clinicalgenome.org/curation-activities/gene-disease-validity/)

**References**: [Strande, N. et al., 2017](https://doi.org/10.1016/j.ajhg.2017.04.015)

### **Cancer Gene Census**

Cancer Gene Census (CGC) is part of the Wellcome Sanger Institute Catalogue of Somatic Mutations in Cancer ([COSMIC](http://cancer.sanger.ac.uk/cosmic)). CGC is an effort to catalogue genes which contain mutations that have been causally implicated in cancer. The exhaustive curation of the CGC covers individual studies as well as pan-cancer sequencing efforts, including The Cancer Genome Atlas (TCGA) and the International Cancer Genome Consortium (ICGC) among others.

In the Platform, CGC evidence is aggregated at the target–disease level to provide a summary of all curated evidence supporting the involvement of a target with a particular cancer type.

**Data type**: Somatic mutations

**Evidence scoring**: Scoring is based on the [Cancer Gene Census tier system](https://cancer.sanger.ac.uk/census)

| Modulator | Condition                                                                                         |
| --------- | ------------------------------------------------------------------------------------------------- |
| -0.25     | Only 1 mutated sample                                                                             |
| +0.25     | Gene mutated more frequently in particular disease compared to other diseases                     |
| +0.25     | Mutations in gene occur more frequently than in other genes of similar length in the same disease |

**Source**: [Cancer Gene Census](https://cancer.sanger.ac.uk/census)

**References**: [Sondka, Z. et al, 2018](https://doi.org/10.1038/s41568-018-0060-1)

### **IntOGen**

IntOGen provides a framework to identify potential cancer driver genes using large-scale mutational data from sequenced tumour samples. By harmonising tumour sequencing data from the ICGC/TCGA Pan-Cancer Analysis of Whole Genomes ([PCAWG](https://dcc.icgc.org/pcawg)) and other comprehensive efforts, IntOGen aims to provide a consensus assessment of cancer driver genes. Several state-of-the-art driver methodologies aiming to cover different approaches (e.g. dN/dS, Hotspots, etc.) are included to finally produce a consensus q-value for each driver gene in every tumour.

In the Platform, independent target–disease evidence are defined as any significant driver gene detected in any individual cohort. Information regarding the individual driver methods is also provided within each evidence.

**Data type**: Somatic mutations

**Evidence scoring**: Scaled [combined q-values](https://intogen.readthedocs.io/en/latest/drivers_combination.html) from 0.25 (q = 0.1) to 1 (q < 1e-10)

**Source**: [intOGen](http://www.intogen.org/search)

**References**: [Martínez-Jiménez, F. et al, 2020](https://doi.org/10.1038/s41568-020-0290-x)

### **ClinVar (somatic)**

ClinVar is an NIH public archive of reports of the relationships among human variations and phenotypes, with supporting evidence. The ClinVar (somatic) data source in the Open Targets Platform captures the subset of ClinVar that refers to somatic variants (as opposed to germline variants).

Information on variants is covered extensively for both single point and structural variants. When available, genomic coordinates are reported with RS numbers, or by following the CHROM\_POS\_REF\_ALT and HGVS notations. &#x20;

Each evidence in the Platform aims to capture an individual RCV record in ClinVar.

**Data type**: Somatic mutations

**Evidence scoring**: ClinVar evidence is scored in a 2-step process. In Step 1, a score is assigned to every piece of evidence based on the clinical significance:

| Clinical significance                        | Evidence score |
| -------------------------------------------- | -------------- |
| association not found                        | 0              |
| benign                                       | 0              |
| not provided                                 | 0              |
| likely benign                                | 0              |
| conflicting interpretations of pathogenicity | 0.3            |
| other                                        | 0.3            |
| uncertain significance                       | 0.3            |
| risk factor                                  | 0.5            |
| affects                                      | 0.5            |
| likely pathogenic                            | 0.7            |
| association                                  | 0.9            |
| drug response                                | 0.9            |
| protective                                   | 0.9            |
| pathogenic                                   | 0.9            |

In Step 2, scored is modulated based on the ClinVar review status:

| Confidence                                           | Evidence score modifier |
| ---------------------------------------------------- | ----------------------- |
| no assertion provided                                | +0                      |
| no assertion criteria provided                       | +0                      |
| no assertion for the individual variant              | +0                      |
| criteria provided, single submitter                  | +0.02                   |
| criteria provided, conflicting interpretations       | +0.02                   |
| criteria provided, multiple submitters, no conflicts | +0.05                   |
| reviewed by expert panel                             | +0.07                   |
| practice guideline                                   | +0.1                    |

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th align="center">Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th align="center">Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td align="center">LoF variants</td><td align="center"><p>Pathogenic/Likely pathogenic = Risk</p><p>Protective = Protective</p></td></tr></tbody></table>

**Source**: [ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/) (via [European Variation Archive](https://www.ebi.ac.uk/eva/))

**References**: [Cezard T. et al, 2021](https://doi.org/10.1093/nar/gkab960); [Shen A. et al, 2024](https://doi.org/10.1093/bioadv/vbae018); [Landrum, M. et al, 2014](https://doi.org/10.1093/nar/gkt1113); [Landrum, M. et al, 2020](https://doi.org/10.1093/nar/gkz972)

### Cancer Biomarkers

One of the aims of the Cancer Genome Interpreter is to identify how variations in the tumour genome may influence its response to anti-cancer therapies. The Cancer Biomarkers database features biomarkers of drug sensitivity, resistance, and toxicity for drugs targeting specific targets in cancer, curated by clinical and scientific experts in precision oncology, and classified by cancer type.

**Data type:** Affected Pathway

**Evidence scoring:** All manually curated evidence in Cancer Biomarkers has a score of 0.5

**Source:** [Cancer Genome Interpreter](https://www.cancergenomeinterpreter.org/biomarkers)

**References:** [Tamborero, D. et al, 2018](https://genomemedicine.biomedcentral.com/articles/10.1186/s13073-018-0531-8)

### CRISPR screens

One of the most powerful approaches to uncover gene function is the experimental perturbation of genes followed by the observation of related phenotypes. The perturbation of gene function in human cells has been greatly facilitated by developments in CRISPR technology.

CRISPRbrain is a database for functional genomics screens in differentiated human brain cell types. We have prioritised genome-wide [CRISPRi/a/KO screens](https://crisprbrain.org/background/) (healthy vs KO) for integration in the Platform to generate target–disease evidence.

We have linked cell types to diseases, meaning these diseases are often characterised with abnormal phenotypes in these cell types — hence the association.\
If knocking out a gene causes significant perturbation in the cell type, it might indicate a potential targeting strategy in the disease.

**Data type**: Affected pathway

**Evidence Scoring**: The Platform uses the linearised CRISPRbrain's assessment of statistical significance to assign a score, including hits from both the upper and lower end of the distribution

**Source**: [CRISPRbrain](https://crisprbrain.org/)&#x20;

**Reference**: [Tian, R et al, 2021](https://www.nature.com/articles/s41593-021-00862-0)

### Project Score

Project Score is a Wellcome Sanger Institute resource that aims to identify dependencies in cancer cell lines to guide precision medicine. The project combines gene fitness effects derived from whole-genome CRISPR-Cas9 synthetic lethality screenings with tractability data, genomic biomarkers and various target annotation enabling a systematic prioritisation of potential targets. The resulting inferences are then mapped from the cancer cell lines in which the experiment is performed to their corresponding tumours.&#x20;

In the Platform, any Project Score prioritised target with priority score reaching 36.0 is included as independent evidence; however, pan-cancer dependencies are excluded from the integration.

**Data type**: Affected pathway

**Evidence scoring**: Project Score priority score divided by 100

**Source**: CRISPR (via[ Project Score](https://score.depmap.sanger.ac.uk))

**References**: [Pacini et al, 2024](https://pubmed.ncbi.nlm.nih.gov/38215750/)

### Reactome

The Reactome database manually curates and identifies reaction pathways that are affected by a disease. Reactome annotation includes information regarding the causal target–disease link either being a protein coding mutation or an altered expression.

In the Platform, any mutation or altered expression event affecting a different reaction is captured in a different target–disease evidence.

**Data type**: Affected pathway

**Evidence scoring**: All manually curated evidence in Reactome has a score of 1.

**Source**: [Reactome](https://reactome.org)

**References**: [Jassal, B. et al, 2020](http://doi.org/10.1093/nar/gkz1031)

### **Europe PMC**

The EMBL-EBI's Europe PMC enables access to a worldwide collection of life science publications and preprints from trusted sources. The Europe PMC data source aims to identify target–disease co-occurrences in the literature and provide an assessment on the confidence of the relationship. This pipeline uses deep-learning based Named Entity Recognition (NER) to identify gene/proteins and diseases when mentioned in the text, to later normalise them to the target or disease/phenotype entities in the Platform. All co-occurrences of both types of entities in the same sentence are considered evidence.

In the Platform, a piece of Europe PMC evidence is the result of aggregating all co-occurrences of the same target and disease within the same publication.

**Data type**: Literature

**Evidence scoring**: Score based on weighted document sections, sentence locations, and title for full text articles and abstracts as described in [Kafkas et al., 2017](https://doi.org/10.1186/s13326-017-0131-3). The aggregated scores of each gene/disease co-occurrence in the publication are further normalised between 0 and 1.

**Source**: [Europe PMC](http://europepmc.org)

**References**: [The Europe PMC Consortium, 2015](https://doi.org/10.1093/nar/gku1061); [Kafkas et al., 2017](https://doi.org/10.1186/s13326-017-0131-3)

### Expression Atlas

The EMBL-EBI Expression Atlas provides a differential expression pipeline aiming to identify genes that are differentially expressed in disease vs control samples. Only contrasts from studies with enough replicates and minimum quality criteria are included in the processing.

In a given contrast, to consider a gene significantly regulated in a contrast, all the following rules are required:

* Absolute log2 fold change > 1
* Adjusted p-value <= 0.05
* Maximum significant genes probes per contrast = 1000

In the Platform, each contrast from independent studies capturing differentially regulated genes constitutes independent evidence.

**Data type**: RNA expression

**Evidence scoring**: ExpressionAtlas scoring is the result of the product of:

* Scaled p-value from 0 (p = 1) to 1 (p<1e-10)
* Absolute log2 fold change divided by 10
* Percentile rank divided by 100

**Source**: [Expression Atlas](https://www.ebi.ac.uk/gxa/home)

**References**: [Papatheodorou, I. et al, 2020](https://doi.org/10.1093/nar/gkz947)

### **IMPC**

The genotype–phenotype associations made available by the International Mouse Phenotypes Consortium (IMPC) are used to identify models of human disease based on phenotypic similarity scores.

The Wellcome Sanger Institute PhenoDigm is an algorithm aimed at capturing the similarity between a knockout mouse and the clinical manifestations (phenotype) of a human disease. The premise is that if a gene knock-out causes an equivalent phenotype in mouse, the human counterpart is likely to be related with the cause of the disease.

It uses a semantic approach to map between clinical features observed in humans and mouse phenotype annotations. The phenotypic effects in mice are then mapped to phenotypes associated with human diseases. The matches are identified and a similarity score between a mouse model and a human disease is computed.

**Data type**: Animal model

**Evidence scoring**: The evidence score indicates the degree of concordance between the mouse and disease phenotypes, as described by [Smedley et al 2013](https://europepmc.org/abstract/MED/23660285).

**Direction of Effect assessment:**

<table data-full-width="false"><thead><tr><th>Direction on Target (Gain of Function (GoF) / Loss of Function (LoF))</th><th>Direction on Trait (Risk/Protective)</th></tr></thead><tbody><tr><td>Assumption of all variants LoF</td><td>Assumption of Risk</td></tr></tbody></table>

**Source**: [IMPC](https://www.mousephenotype.org)

**References**: [Smedley, D. et al, 2013](https://doi.org/10.1093/database/bat025)


# Target–disease associations

Each unique target–disease pair in the Open Targets Platform is defined as an **association**. For example, while there might be several pieces of evidence referring to CFTR and Cystic fibrosis from multiple sources, one single association contextualises all this information within the Platform. Also, since multiple pieces of evidence might refer to the same or similar associations, the Platform undertakes a series of steps to quantify their relative strength for a given association.

## Ontological selection of evidence

The Platform associations aim to aggregate all evidence referring to the target and disease, but the complex phenotypic representation of the disease might sometimes cause different pieces of evidence to be annotated against slightly different levels of granularity of the disease.&#x20;

For example, the association between inflammatory bowel disease and NOD2 is annotated with multiple pieces of evidence referring to these two terms specifically. The Platform refers to associations described by aggregated evidence between two specific terms in our data sources as **direct associations**. By default, direct associations are displayed in the web application when listing associated diseases or phenotypes with a target of interest.

However, evidence can sometimes be informative to discriminate targets in similar diseases or phenotypes. For example, when evaluating the inflammatory bowel disease and NOD2 association, other pieces of evidence describing the relationship between Crohn's disease and NOD2 might also be informative.&#x20;

To approach this problem systematically, the Platform makes use of the properties of the disease ontology (EFO), to select all evidence referring to NOD2 in the context of inflammatory bowel disease or any of its ontology descendants (including Crohn's disease). This type of association is referred to in the platform as an **indirect association**. Indirect associations are displayed in the web application when displaying all the evidence for a target-disease association or listing associated targets with a disease or phenotype of interest.

{% hint style="info" %}
**Summary**

An association page for diseases associated with a target (e.g. [NOD2 associations page](https://platform.opentargets.org/target/ENSG00000167207/associations)) includes **direct evidence only**.

An association page for targets associated with a disease (e.g. [Inflammatory Bowel Disease associations page](https://platform.opentargets.org/disease/EFO_0003767/associations)) includes **both direct and indirect evidence**.&#x20;

An evidence page (e.g. [NOD2 and inflammatory bowel disease](https://platform.opentargets.org/evidence/ENSG00000167207/EFO_0003767)) displays **both direct and indirect evidence**. &#x20;
{% endhint %}

\
\
The same calculations are applied to calculate the association scores, the only difference is the evidence included in the calculation.

Both direct and indirect associations can be queried using the GraphQL API or [our data downloads page](https://platform.opentargets.org/downloads).

{% hint style="info" %}
**Note**

RNA expression data type evidence is not propagated in the ontology. We made this decision to prevent parent terms from having long lists of associated targets with weak RNA expression association scores.
{% endhint %}

## Association scores

Deciding what constitutes a strong association is open to interpretation. While some data sources can be more deterministic about the underlying causal evidence, others can occasionally point to targets indirectly linked to the disease. The Platform relies on a series of heuristics to maximise the transparency of the target and disease rankings.

### By data source

The Platform's scoring by data source aims to take into account variations in how the data sources organise their evidence.

For example, some data sources are more stringent when it comes to defining what constitutes a single piece of evidence, whereas others rely on their internal evidence score to stratify the strength of the evidence they present.&#x20;

Additionally, some data sources will capture the meaningful association in one single piece of evidence. In other data sources, the repetition of evidence increases our confidence in the association.&#x20;

The scoring by data source aims to balance all these differences and provide a consensus view on the strength of the underlying evidence for a particular source.

For all cases, the Platform defines a **data source association score** by calculating a harmonic sum using the full vector of [evidence scores](/evidence#evidence-data-sources) as defined for each data source using the following the next steps:

1. The pieces of evidence are sorted in descending order and assigned an incremental value that indicates their position in the sorted list (the top-scoring item has a positional id of 1, the second has a positional id of 2, and so on).
2. The harmonic sum for each data source is then calculated by summing the result of dividing each evidence score by (positional id^2).

![](/files/-MYFH7AR3l7rAPLVV-Wr)

3. To ensure the result is between 0 and 1, the harmonic sum is normalised by dividing the result by the maximum theoretical harmonic sum, which is the one calculated using an infinite vector of ones. The platform derives this calculation (which approximates to 1.644) by using a vector of 1,000 ones.

{% hint style="info" %}
**Example**

To calculate the **data source association score** for a vector of evidence scored 1, 0.9 and 0.8 the Platform will follow the next logic

*Step 1: Sorting/Indexing*

```
evidence with score=1.0 -> positional id=1
evidence with score=0.9 -> positional id=2
evidence with score=0.8 -> positional id=3
```

*Step 2: Harmonic sum calculation*

```
harmonic sum score = 1.0/1^2 + 0.9/2^2 + 0.8/3^2
```

*Step 3: Scaling*

```
max. theoretical harmonic sum score = 1.0/1^2 + 1.0/2^2 + 1.0/3^2 + 1.0/4^2 + ... ~ 1.644
normalised harmonic sum score = harmonic sum score / max. theoretical harmonic sum score
```

{% endhint %}

### By data type

The **data type association score** aims to capture the strength of the supporting evidence at the data type level (e.g. Genetic Associations). A second harmonic sum is calculated by using the vector of [data source association scores](#data-source-weights) weighted by the [data source weights](#data-source-weights).

{% hint style="warning" %}
**Association score by data type scaling**

While the harmonic sum calculation remains mostly the same for **data sources, data types** and **overall**, the scaling factor is slightly modified in the **association by** **data type** calculation.&#x20;

So that data types only featuring one data source (e.g. text mining) are not penalised, the maximum theoretical harmonic sum score is calculated based on a vector of as many ones as data sources are in the respective datatype. In this way, the scaling factor of a data type with one data source will be `1.0/1^2 = 1`, whereas a datatype with 3 data sources will be scaled by the result of calculating: `1.0/1^2 + 1/2^2 + 1/3^2 = 1.36`.
{% endhint %}

### Overall

The **overall association score** aims to summarise all the aggregated evidence for a given target-disease association. The score is derived by calculating the harmonic sum of the association score by data source weighted by the [data source weights](#data-source-weights), regardless of their data type categorisation. The algorithm to compute the scores is the same as the association by data source, resulting in a score between 0 and 1.

### Data source weights

To calculate both **data type** and **overall** association scores, evidence is weighted using a factor that aims to calibrate the relevance of each data source relative to others. The default weights used in the web application can be modified by the user to adjust to different prioritisation strategies, both in the API and in the user interface through the "Advanced Option" tab from the new [Associations on the Fly](https://platform-docs.opentargets.org/web-interface/associations-on-the-fly-beta) page.

| Data source                          | Weight factor |
| ------------------------------------ | ------------- |
| Europe PMC                           | 0.2           |
| Expression Atlas                     | 0.2           |
| IMPC                                 | 0.2           |
| OTAR Projects (partner preview only) | 0.5           |
| Cancer Biomarkers                    | 0.5           |
| Others                               | 1             |

### Interpreting association scores

There are a few important considerations regarding association scores. As described above, association scores are a heuristic based on the availability of data. While scores are useful to rank lists of targets or diseases, **they should not be interpreted as a confidence score for the target-disease association**.&#x20;

For example, under-studied diseases are unlikely to produce high-scoring targets due to the lack of available evidence. In such diseases, a relatively low-scoring target might still be the top-ranked target and potentially a very interesting lead from a therapeutic standpoint.

Similarly, not all associations with available target–disease evidence should be considered legitimate target–disease associations. Some of our data sources rely on predictions to assess the relationship between a target and a disease. Thus, they should be considered with caution and always take their relative support into consideration.

<figure><img src="/files/DF0CiZD8B9V0otFW8tuc" alt=""><figcaption></figcaption></figure>

<sub>Schematic of how data source, data type, and overall associations scores are calculated in the Open Targets Platform.</sub>


# GWAS & functional genomics

Open Targets post-GWAS analysis pipelines

The Platform is built upon a significant effort to analyse genome-wide associations studies and interpret them in the context of functional genomics studies. This effort to inform target identification and prioritisation datasets leverages [gentropy](/gentropy/gentropy) a highly-scalable python framework for post-GWAS analysis.

A more detailed view on the data and methodology is available in:

* [Data sources](/gentropy/data-sources) used to ingest variant annotation, GWAS and functional genomics information.&#x20;
* [Fine-mapping](/gentropy/fine-mapping) pipelines to identify likely causal variants in GWAS-significant signals in different sources.
* [Enhancer-to-gene](/gentropy/enhancer-to-gene-encode-re2g) dataset to connect variants to genes using rE2G model scores.
* [Colocalisation](/gentropy/colocalisation) performed on overlapping GWAS and molQTL credible sets.
* [Locus-to-Gene](/gentropy/locus-to-gene-l2g) predictions to assess likely causal genes near identified credible sets.


# Data sources

Overview of the GWAS association and functional genomics data sources

The Platform ingests GWAS, functional genomics and additional datasets to aid the interpretation of the observed signals. These include:

## NHGRI-EBI GWAS Catalog

🌍 [Website](https://www.ebi.ac.uk/gwas/) — 📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/gwas_catalog/_gwas_catalog/)

The [NHGRI-EBI GWAS Catalog](https://www.ebi.ac.uk/gwas/) is a data source that provides detailed, structured genome-wide association study data in standardised format of summary statistics and curated associations (top-hits) to [EFO](https://www.ebi.ac.uk/efo/) traits. We rely on the GWAS Catalog for a rich source of genetic associations, utilising the data for analysis and interpretation.

The GWAS Catalog information feeds the Open Targets Platform by providing study metadata and GWAS associations from two sources:&#x20;

### **GWAS Catalog literature-curated associations (GCCA)**

The Platform ingests the GWAS Catalog curated associations (top-hits) from the table “All associations - with added ontology annotations, GWAS Catalog study accession numbers and genotyping technology” (see [GWAS Catalog download page](https://www.ebi.ac.uk/gwas/docs/file-downloads)).  The GWAS Catalog curation process may result in multiple GWAS being grouped under one study accession, in which case one study accession will need to be split into multiple studies in the Platform. An example of such a study is: [GCST001762](https://www.ebi.ac.uk/gwas/studies/GCST001762), where the reported trait is 'Obesity-related traits', but the underlying association data informs which top-hit comes from which specific trait, for example 'BMI z-score change'. In this case, we create as many new studies as there are association-level trait annotations. The new study identifiers are generated from the original study accession plus a suffix. The disease annotations of the newly created studies are inherited from the association data. A similar action is required when results from different ancestries are pooled under one study accession.&#x20;

The harmonisation process includes a number of steps that flag any associations with quality concerns. When harmonising GCCA, the following cases were flagged:

* Variant interaction associations
* Variant location is not available in an association source
* Variant location data is inconsistent (chromosome and position values don’t match)
* Lead variant is duplicated for the same study
* The provided variant location doesn’t match any GnomAD variant
* The risk allele couldn't be mapped to the GnomAD reference
* The lead variant had palindromic alleles
* The lead variant p-value was above the genome-wide significant threshold (p-value>1e-8).

Curated associations were converted into the Gentropy `StudyLocus` format and uniformly flagged with `Study locus from curated top hit` quality control flag. Corresponding `StudyIndex` are generated to captures the study metadata.

### **GWAS Catalog Summary Statistics  (GCSS)**

The Platform ingests GWAS Catalog studies from full summary statistics when they have been harmonised following the [protocol](https://www.ebi.ac.uk/gwas/docs/methods/summary-statistics). Each harmonised study is converted into a Gentropy `SummaryStatistics` format as described in the [Gentropy documentation](https://opentargets.github.io/gentropy/python_api/datasets/summary_statistics/) including additional quality controls:

1. filtering out SNPs with unavailable beta, standard error, or p-value
2. filtering out SNPs with zero or Inf values for beta and standard error
3. filtering out SNPs with negative p-values or standard error
4. filtering out SNPs with p-values equal to 1

In some cases, this results in empty `SummaryStatistics`datasets and these studies are excluded from further analysis.&#x20;

#### **GCSS studies manual curation**

The Open Targets team performs additional manual curation for the GCSS studies from “All studies - With study accession numbers, ontology annotations, genotyping technology, cohort identifiers and full summary statistics availability”. The following information is extracted from the corresponding publication:

1. **Study type** — The GCSS studies identified as pQTL or microbiome GWAS are flagged and excluded from further analysis.
2. **Analysis type** — studies are flag if performed with any of the following analysis: multivariate analysis, ExWAS, non-additive model, metabolite, GxG, GxE, case-case study. All analysis flags but “metabolite” are excluded from SuSiE fine-mapping and fine-mapped using PICS.

A `StudyIndex`dataset was created as part of this pipeline containing all the available metadata for the included GWAS Catalog studies. If available, the ancestries from GWAS Catalog were mapped to gnomAD ancestry suffixes using [the dictionary here](https://github.com/opentargets/gentropy/blob/dev/src/gentropy/assets/data/gwas_population_2_LD_panel_map.json). The ancestry wasn't assigned if the ancestry was unavailable or wasn't presented in the dictionary.

#### **GCSS** quality control

GWAS summary statistics quality control (QC) is performed for all GWAS Catalog studies with available summary statistics, following the methods described in [Winkler et al. (2014)](https://www.nature.com/articles/nprot.2014.071):

1. **The P-Z test**. This check estimates the mean and standard deviation of the difference between the log p-values reported in the study and the reported betas and standard errors. If at least one value for the study was greater than 0.05, the study fails QC and is flagged with the label `The PZ QC check values are not within the expected range`.
2. **The mean beta check**. This check estimates the mean value of the beta across all SNPs in the study. If the absolute mean value is more than 0.05, the study fails QC and is flagged with the label `The mean beta QC check value is not within the expected range`.
3. **The genomic control (GC) lambda check**. This check estimates the additive GC lambda of the study (see [Tsepilov et al. (2013)](https://pubmed.ncbi.nlm.nih.gov/24358113/)). If the GC lambda value is outside the \[0.7,2.5] range the study fails QC and is flagged with the label `The GC lambda value is not within the expected range`.
4. **Number of variants.** All summary statistics with fewer than 2,000,000 variants do not fail QC but are flagged with the label `The number of SNPs in the study is below the expected threshold`. These studies are not eligible for SuSiE fine-mapping.

#### **GCSS** heritability estimate

GWAS summary statistics which pass certain quality criteria — including a minimum effective sample size of 10,000 and a compatible ancestry and study design — will have SNP heritability estimated using LD Score Regression ([LDSC](https://www.nature.com/articles/ng.3211)). This is implemented using precomputed LD scores derived from the HapMap3 SNP panel with gnomAD v2.1.1 reference genotypes, across five ancestry groups (`AFR`, `AMR`, `EAS`, `FIN`, `NFE`). Studies are matched to the most appropriate reference panel based on their reported population structure. Heritability estimates and associated statistics (`h2`, `h2_se`, `intercept`, `intercept_se`, `mean_chisq`, `lambda_gc`) are stored in the `sumstatQCValues` array of the study index. The `lambda_gc` flag is obtained from the interception between the summary stattistics and the variants derived from the LD score panel.\
\
Estimates from runs where the regression is unlikely to be reliable are flagged with `runStatus = "degenerate"` and excluded from downstream use.

## FinnGen

🌍 [Website](https://www.finngen.fi/en) — 📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/finngen/_finngen/)

[FinnGen](https://www.finngen.fi/en) is an academia-industry partnership that aims to produce genome variant data for 500,000 Finns. The genomic data is then combined with phenotype data collected by national health registries, including extensive longitudinal registry data available on all Finns. More details on the protocols are available [in the FinnGen documentation](https://finngen.gitbook.io/documentation/releases). Because the Finnish population has been genetically isolated, fine-mapping was performed with a suitable reference panel of LD. The Platform includes the fine-mapping results from the FinnGen team using SuSiE and FINEMAP, based on a reference panel of whole genome sequencing data from Finns.

The Platform includes the 95% credible sets as [described in the documentation](https://finngen.gitbook.io/documentation/data-description) after converting them Gentropy `StudyIndex` and `StudyLocus` objects. Credible sets with a lead SNP p-value>=1e-5 are marked with a `Subsignificant p-value` flag.

## eQTL Catalogue

🌍 [Website](https://www.ebi.ac.uk/eqtl/) — 📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/eqtl_catalogue/_eqtl_catalogue/)

The [eQTL Catalogue](https://www.ebi.ac.uk/eqtl/) aims to provide unified gene expression, protein-level and splicing QTLs from publicly available human public studies. Using a standardised pipeline ([QTLmap](https://github.com/eQTL-Catalogue/qtlmap)), the eQTL catalogue is integrated in the Platform as a rich source of molecular QTLs (molQTLs). The pipeline uses SuSiE as the method of fine-mapping and results in 95% CSs.

All credible sets and study information are reformatted to Gentropy's `StudyIndex` and `StudyLocus` datasets. Credible sets with a lead SNP p-value>=1e-3 are marked with a `Subsignificant p-value` flag.&#x20;

## The Pharma Proteomics UK Biobank Project (PPP-UKBB)&#x20;

&#x20;📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/ukb_ppp_eur/_ukb_ppp_eur/)

[The Pharma Proteomics Project](https://www.synapse.org/Synapse:syn51364943/wiki/622119) is a pre-competitive biopharmaceutical consortium characterising the plasma proteomic profiles of 54,219 participants in the UK Biobank.&#x20;

The Platform includes GWAS full summary statistics for 2,954 proteins of European ancestry [from the Synapse platform](https://www.synapse.org/Synapse:syn51364943/wiki/622119). Data is converted to `SummaryStatistics` and `StudyIndex` format applying the following modifications:

* SNPs with MAF<1e-4 and INFO<0.8 are filtered out
* Align the order of effective and reference alleles with the gnomAD annotation: if the alleles are reversed, the sign of the effect size is changed; if the allele combination do not match the reference, the SNP is filtered out
* Flag QTLs as *cis* or *trans* based on a 5Mb window from the affected gene.

## gnomAD

🌍[Website](https://gnomad.broadinstitute.org/) — 📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/gnomad/_gnomad/)

[gnomAD](https://gnomad.broadinstitute.org/) (Genome Aggregation Database) is a comprehensive resource that provides aggregated genomic data from large-scale sequencing projects. It encompasses variants from diverse populations and is widely used for variant annotation and population genetics studies.

### Variant annotation

GnomAD (v4) variant annotation is used to provide extra annotation for variants included in the Platform and assist some of our harmonisation pipelines. Among the most relevant annotations included are: population allelic frequencies, in silico predictors and cross-references. All variants are available in GRCh38 coordinates.

### LD matrices

gnomAD v2.1.1 [LD matrices](https://gnomad.broadinstitute.org/news/2018-10-gnomad-v2-1/) are used to create a multi-ancestry LD reference for the [PICS fine-mapping method](/gentropy/fine-mapping#pics-fine-mapping). All variants are liftovered to build GRCh38.

## Pan-UKBB LD matrices

🌍 [Website](https://pan.ukbb.broadinstitute.org/docs/summary/index.html)

The Platform uses [Pan-UK Biobank project](https://pan.ukbb.broadinstitute.org/docs/summary) LD matrices for three ancestries: Non-Finnish European (NFE), Central/South Asian (CSA) and African (AFR). See the descriptive summary [from Pan-UK Biobank](https://pan.ukbb.broadinstitute.org/docs/technical-overview). The LD was computed for each chromosome in 10 Mb radius. Only SNPs with INFO > 0.8, MAC > 20 in each population are used to calculate LD. Matrices are stored in hail BlockMatrix format. We use these LD matrices for SuSiE fine-mapping.

## EMBL-EBI Ensembl

🌍 [Website](https://www.ensembl.org/index.html) — 📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/ensembl/_ensembl/)

The Platform uses Ensembl as a source of gene, transcript and variant annotations. Affected genes in QTL studies are validated using fixed version of Ensembl and all variants are annotated using Ensembl VEP.

## Biosample

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/biosample_ontologies/_cell_ontology/)

The Experimental Factor Ontology, UBERON and Cell Ontology are composed as a meta-ontology to capture every tissue or cell type that could be described in a QTL study.

## Enhancer-to-Gene

🌍 [Website](https://www.encodeproject.org/) - 📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasources/intervals/_intervals/)

The **ENCODE Consortium** data (ENCODE rE2G **thresholded element gene links**) contains scores calculated by a **supervised classifier** that integrates molecular features of chromatin state and 3D physical interactions to predict **enhancer-gene regulatory interactions** in a given cell type.&#x20;

The rE2G dataset is used as a resource to explain the interaction of credible set variants and regulatory (**enhancer** and **promoter**) regions. The interaction score is used to add the regulatory context to L2G features.


# Fine-mapping

Fine-mapping is a statistical analysis technique used to pinpoint the specific genetic variant(s) most likely responsible for a trait association identified in a GWAS. The main result is a credible set (CS): the minimal list of variants with assigned posterior probabilities (PIPs) to be causal that form a predefined probability. The Platform uses 95% CS, meaning the CS has a 0.95 probability of containing the causal variant. The expected sum of PIPs for all variants in CSs have to be within the range of \[0.95,1].

**Open Targets Platform fine-mapping strategies (see more in** [**Fine-mapping pipelines**](#fine-mapping-pipelines)**):**

<table><thead><tr><th>Source</th><th>Clumping</th><th>Fine-mapping</th><th>LD</th><th data-type="rating" data-max="4">Confidence</th></tr></thead><tbody><tr><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#gwas-catalog-literature-curated-associations-gcca">GCCA</a></td><td><a href="#distance-based-clumping">Distance-based</a> + <a href="#ld-based-clumping">LD-based</a></td><td><a href="#pics-fine-mapping">PICS</a></td><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#ld-matrixes">gnomAD</a></td><td>1</td></tr><tr><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#gwas-catalog-summary-statistics-gcss">GCSS</a></td><td><a href="#distance-based-clumping">Distance-based </a>+ <a href="#ld-based-clumping">LD-based</a></td><td><a href="#pics-fine-mapping">PICS</a></td><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#ld-matrixes">gnomAD</a></td><td>2</td></tr><tr><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#gwas-catalog-summary-statistics-gcss">GCSS</a></td><td><a href="#locus-breaker">Locus breaker</a></td><td><a href="#susie-inf-fine-mapping">SuSiE-inf</a></td><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#pan-ukbb-ld-matrices">pan-UKBB</a></td><td>3</td></tr><tr><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#finngen">FinnGen</a></td><td>Study-specific</td><td>SuSiE</td><td>In-sample</td><td>4</td></tr><tr><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#eqtl-catalogue">eQTL catalogue</a></td><td>Study-specific</td><td>SuSiE</td><td>In-sample</td><td>4</td></tr><tr><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#the-pharma-proteomics-uk-biobank-project-ppp-ukbb">UKBB-PPP</a></td><td><a href="#locus-breaker">Locus breaker</a></td><td><a href="#susie-inf-fine-mapping">SuSiE-inf</a></td><td><a href="/pages/J9HJWrdUhl7XVNFjARQu#pan-ukbb-ld-matrices">pan-UKBB</a></td><td>3</td></tr></tbody></table>

## Clumping

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/methods/clumping/)

Clumping is a technique for selecting the most significant variants within a region or set of variants in LD, essentially collapsing signals into loci with an increased chance of containing a single causal variant for the phenotype. Three different methods are used for clumping:

### **Distance-based clumping**

The method is based on an iterative procedure in which the variant with the strongest p-value is selected and all other variants within a predefined distance (radius) from this variant are clumped together. The procedure is repeated as long as there is at least one significant variant. The output of the method is a list of lead variants. There are two parameters: the distance to clump and the p-value significance threshold.

### **LD-based clumping**

The method is applied to the results of the distance-based clumping (list of lead variants). It is again based on an iterative procedure in which the variant with the strongest p-value is selected and all other variants with high LD (r2>=0.5) are clumped together.&#x20;

### **Locus breaker**

The method consists of three steps:&#x20;

1. In the first step, we perform the regular distance-based clumping.
2. In the second step, we filter the input summary statistics by the baseline p-value (default is 1e-5). We then clump variants that are closer to each other than the cutoff distance (default is 250,000 bp). Next, we filter clumps by having at least one variant with a p-value above the genome-wide significance threshold. At this stage, we define a list of loci consisting of information about the lead variant and the locus boundaries (the leftmost and rightmost variants in the clump). To each of the locus boundaries we subtract/add (to the left/right boundaries respectively) the flanking distance (the default is 100,000 bp) to avoid situations where the locus size consisting of only one lead variant is 0.&#x20;
3. In the third step we select the loci with the size greater than the specified large locus size threshold (1,500,000 bp by default) and 'break' each of the large loci using the lead variants from the distance-based clump that lie within the boundaries of the large locus. For each of the distance-based lead variants, we assigned the boundaries as +/-half the size of the large locus size threshold. Thus, each large locus is divided by several overlapping loci of the size of the large locus size threshold. The small loci are unaffected by the splitting.&#x20;

The general procedure results in a list that contains information about the lead variant and the locus boundaries (the most left and right variants in the cluster), and the largest locus size doesn't exceed the large locus size threshold. The boundaries of the locus are then used to define the region for fine-mapping and LD matrix ingestion. The reason for using this procedure is because locus breaker results in smaller locus sizes on average compared to those defined by window-based clumping.

**Acknowledgements:** We are grateful to our colleagues at Human Technopole, Sodbo Sharapov and Nicola Pirastu, for advice on this method.

## PICS fine-mapping

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/methods/pics/)

The PICS algorithm was originally implemented in [Farh et al. (2015)](https://www.nature.com/articles/nature13835) investigating the fine-mapping of causal autoimmune disease variants. It is a method to fine-map the most likely causal variants associated with a trait or disease within a haplotype. The algorithm is based on the calculation of the Posterior Inclusion Probabilities (PIP) of tag variants linked to the lead variant by LD within the target population. Only five populations are used for PICS fine-mapping: African-American (AFR), American Admixed/Latino (AMR), East Asian (EAS), Finnish (FIN), and Non-Finnish European (NFE). The calculation is performed by using all the proxy variants where r2 ≥ 0.5 and the default parameter k=6.4 as reported in the original paper.

## SuSiE-inf fine-mapping

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/methods/susie_inf/)

If both summary statistics and high-precision LD are available for the locus, we use SuSiE-inf  fine-mapping (see [Cui et al. (2024)](https://www.nature.com/articles/s41588-023-01597-3) for more details). This method is a generalisation of the original SuSiE, allowing modelling of infinitesimal effects alongside fewer larger causal effects.

SuSiE-inf has two approaches for updating estimates of the variance components: Method of Moments and Maximum Likelihood Estimator ('MoM' / 'MLE'). The function takes an array of Z-scores and a numpy array matrix of variant LD to perform fine-mapping. There is a boolean option “est\_tausq” that enables the estimation of infinitesimal effects. If it is disabled, it performs fine-mapping similar to the original SuSiE method. We use SuSiE-inf only with the combination of pan-UK Biobank LD matrices for three ancestries: NFE, CSA and AFR.

## Fine-mapping pipelines

Different pipeline strategies have been defined for different sources based on a combination of the above methodologies:

### **Clumping and fine-mapping of GCSS and GCCA PICS**

Distance and LD-based clumping is applied to all available GCSS and GCCA. P-value significance is defined at <=1e-8 threshold and distance-based clumping is performed using a 500,000 bp radius. LD-clumping is performed on the same window using the LD dataset from gnomAD v2.1.1. [PICS](#pics-fine-mapping) fine-mapping is then performed using the same LD information and credible sets are defined based on 95% inclusion probability. In the case of multiple ancestry studies, we used Non-Finnish Europeans (NFE) if it was in the list of ancestries, or major ancestry otherwise. In the case the study's ancestry is not in the list available in LD index ancestries, no fine-mapping results are produced. If the lead variant is not present in the LD-index, the credible set contains only that variant and it's flagged with the label `Variant not found in LD reference`. All resulting credible sets objects were additionally flagged with the label `Study locus fine-mapped without in-sample LD reference`.

### **Clumping and fine-mapping of GCSS SuSiE**

GWAS-significant study-locus are identified using a p-value threshold <=1e-8 and the [locus-breaker](#locus-breaker) method. [SuSiE-inf fine-mapping](#susie-inf-fine-mapping) was applied on all study-loci following the next criteria:

1. Major ancestry for the study is NFE, AFR, CSA or EAS. If the major ancestry was EAS, we used CSA instead
2. The study type is "gwas"
3. The study has no analysis flags except "metabolite"
4. Study has no quality control flags
5. The locus didn't overlap with the MHC region. We didn’t exclude X or Y chromosomes if they were presented in GWAS
6. The number of variants in the locus after overlapping with the LD matrix was in the range \[100, 15,000]
7. Study is not a meta analysis ( The major LD ancestry relative sample size ≥ 90%)

For all eligible study loci, the SuSiE-inf method was applied without estimation of infinitesimal effects (equivalent to the classical SuSiE method) using [pan-UKBB LD matrices](/gentropy/data-sources#pan-ukbb-ld-matrices). Filtering of the resulting 95% CSs is performed based on CS log(BF)<=2, minimum R2 purity <=0.25 and lead variant p-value>=1e-5. Additionally, the pairwise r^2 between lead variants within the locus is calculated, removing the less significant of the CSs if the r2>=0.8. As the locus breaker procedure can create overlapping loci, we can obtain the redundant credible sets within the same study. Thus, if two CSs from different loci within a study had the same lead variant, we removed one of them, leaving the CS with the largest CS log(BF). All resulting credible sets from this pipeline are flagged with the label `Study locus fine-mapped without in-sample LD reference`.

### UKBB-PPP clumping and fine-mapping

The [UKBB-PPP](/gentropy/data-sources#the-pharma-proteomics-uk-biobank-project-ppp-ukbb) summary statistics are clustered into locus using the [locus breaker](#locus-breaker) method using a p-value significance threshold <=1.7e-11. The resulting study-loci are fine-mapped using the [SuSiE-inf](#susie-inf-fine-mapping) method with [Pan-UKBB LD matrices](/gentropy/data-sources#pan-ukbb-ld-matrices) for the EUR population. The resulting 95% CSs are filtered similar to GCSS SuSiE pipeline described above. All resulting StudyLocus objects were additionally flagged with the label `Study locus fine-mapped without in-sample LD reference`.


# Enhancer-to-Gene (ENCODE rE2G)

Enhancer-to-gene scores are directly ingested from the ENCODE consortium, which utilises a mixture of epigenetic datasets to connect genomic regions (defined as chromosome, start, end) to putative genes. These regions are likely regulatory elements and affect the transcriptional activity of their annotated genes.

The input datasets are based on publicly available resources from the [ENCODE project](https://www.encodeproject.org/), these include&#x20;

* histone modification ChIP-seq,&#x20;
* open chromatin DNase-seq&#x20;
* ATAC-seq
* 3D chromatin conformation structure determined by Hi-C.&#x20;

These inputs are normalised and used as features in a machine learning approach described in the [original rE2G publication](https://www.biorxiv.org/content/10.1101/2023.11.09.563812v1).  The final output connects potential regulatory regions to their target genes, along with a score that ranges between 0 to 1 that indicates the confidence of a given assignment.

### rE2G in the platform

Enhancer-to-gene scores from rE2G are ingested into the Open Targets ecosystem through the orchestration and can be&#x20;

* browsed through the variant page, where the regulatory regions overlapping a given variant of interest are displayed in the enhancer-to-gene widget. &#x20;
* browsed through the credible set page, which shows the overlapping rE2G scores for the lead variant of the credible set.
* viewed in the form of `e2gMean` and `e2gNeighbourhoodMean` features in the L2G predictions.

### Applied transformations

The dataset is downloaded in the latest version using the ENCODE API and transformed to the [Interval](https://opentargets.github.io/gentropy/python_api/datasets/intervals/) gentropy format. The mapping between original biosamples provided by the source was based on [manual curation](https://github.com/opentargets/curation/blob/25.12.2/genetics/E2G_biosample_mapping.csv). We apply post-transformation quality controls and flagging system that includes:

* Validation of gene identifiers from input against latest Ensembl version&#x20;
* Validation of biosample identifiers against the latest Biosample dataset
* Score stringent filtering

We have implemented a stringent filter of 0.6 on the rE2G dataset to reduce computational costs and redundancy.  The selection of the filter was based on an analysis performed on rE2G overlaps with eQTL credible sets.

<figure><img src="/files/6PnXyk0PzUtdLt0I7JQX" alt=""><figcaption><p><strong>Effect of rE2G-score filtering on gene prioritisation for eQTL credible sets.</strong><br>True positives are defined as the eQTL target gene; <strong>Sensitivity</strong> (orange) is TP recall, and <strong>FDR</strong> (blue) = 1 − precision among retained cs–gene pairs. Points are thresholds labelled by percentiles (Px) of the <strong>rE2G score</strong> distribution. <strong>Moderate filtering</strong> removes significant amounts of raw rE2G entries while <strong>retaining most TP assignments</strong>. <strong>FDR changes little until very aggressive cutoffs</strong>, indicating many non-TP—but potentially interesting—gene links persist.</p></figcaption></figure>


# Colocalisation

Pairs of credible sets overlapping at least one variant are subject to additional analysis to infer the likelihood of them sharing the same causal variant. The Platform compares all GWAS vs all GWAS credible sets and all GWAS vs all molQTL credible sets. For overlapping credible sets, colocalisation metrics are estimated using the following methods:

## COLOC-PIP

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/methods/coloc_pip/)

COLOC-PIP is an adaptation of the COLOC method ([Giambartolomei et al. , 2014](https://www.ncbi.nlm.nih.gov/pubmed/24830394)), where instead of using log Bayes factors the H3 and H4 probabilities are calculated directly from the variant-level posterior inclusion probabilities. This means the method can be applied to credible sets with posterior probabilities derived from fine-mapping methods other than SuSiE. The method assumes the probabilities of H0, H1, and H2 are zero, as any overlap is already the result of two fine-mapped associations.&#x20;

* **H0**: No association with either trait
* **H1**: Association with trait 1, not with trait 2
* **H2**: Association with trait 2, not with trait 1
* **H3**: Association with trait 1 and trait 2, two independent SNPs
* **H4**: Association with trait 1 and trait 2, one shared SNP

## eCAVIAR

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/methods/ecaviar/)

eCAVIAR is an a heuristic algorithm that uses SNP’s PIPs (see [Hormozdiari et al. (2016)](https://pubmed.ncbi.nlm.nih.gov/27866706/)) to estimate the colocalisation posterior probability (CLPP). CLPP is computed as the sum over the product of the variant fine-mapping probabilities between the two overlapping credible sets.

{% hint style="info" %}
**Credible set-based colocalisation**

All co-localisations presented in the Platform are based on credible sets. No estimates are based on full locus colocalisation
{% endhint %}


# Locus-to-Gene (L2G)

Overview of the Open Targets Locus-to-Gene algorithm

Based on genetic and functional genomics traits, the Locus-to-Gene (L2G) machine-learning algorithm ranks the most likely causative genes at each GWAS credible set. The likelihood that a gene is causal for a particular GWAS locus is measured by the L2G score, ranging from 0 to 1.

{% hint style="warning" %}
The L2G model included in the Platform is based on the original method published by Mountjoy *et al.* Nature Genetics (2021), but contains several enhancements that can result in different performance.
{% endhint %}

## Features

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasets/l2g_features/_l2g_feature/)

The predictive features used by the L2G algorithm are designed to capture various genetic and genomic contexts that influence the likelihood of a gene being causal at a given GWAS credible set.&#x20;

Features are divided into five categories:&#x20;

{% tabs %}
{% tab title="Distance" %}

<table><thead><tr><th width="362.13671875">Feature Name</th><th width="355.55859375">Description</th><th width="183.2421875">Range</th></tr></thead><tbody><tr><td><code>distanceTssMean</code></td><td>Average distance between all variants in the credible set and the TSS of a gene's canonical transcript. The distance of each variant is weighted by its posterior probability</td><td>(0, 1)</td></tr><tr><td><code>distanceTssMeanNeighbourhood</code></td><td>Ratio between <code>distanceTssMean</code> for a gene and the maximum <code>distanceTssMean</code> for any given gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>distanceSentinelTss</code></td><td>Distance between the sentinel variant and the TSS of a gene's canonical transcript</td><td>(0, 1)</td></tr><tr><td><code>distanceSentinelTssNeighbourhood</code></td><td>Ratio between <code>distanceSentinelTss</code> for a gene and the maximum <code>distanceSentinelTss</code> for any given gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>distanceSentinelFootprint</code></td><td>Distance between sentinel variant and a gene's footprint</td><td>(0, 1)</td></tr><tr><td><code>distanceSentinelFootprintNeighbourhood</code></td><td>Ratio between <code>distanceSentinelFootprint</code> for a gene and the maximum <code>distanceSentinelFootprint</code> for any given gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>distanceFootprintMean</code></td><td>Average distance between all variants in the credible set and a gene's footprint. The distance of each variant is weighted by its posterior probability</td><td>(0, 1)</td></tr><tr><td><code>distanceFootprintMeanNeighbourhood</code></td><td>Ratio between <code>distanceFootprintMean</code> for a gene and the maximum <code>distanceFootprintMean</code> for any given gene in the vicinity</td><td>(0, 1)</td></tr></tbody></table>
{% endtab %}

{% tab title="Colocalisation" %}

<table><thead><tr><th width="335.04296875">Feature Name</th><th width="287.00390625">Description</th><th>Range</th></tr></thead><tbody><tr><td><code>eQtlColocClppMaximum</code></td><td>Maximum CCLP across all eQTL studies for a gene</td><td>(0, 1)</td></tr><tr><td><code>pQtlColocClppMaximum</code></td><td>Maximum CCLP across all pQTL studies for a gene</td><td>(0, 1)</td></tr><tr><td><code>sQtlColocClppMaximum</code></td><td>Maximum CCLP across all sQTL and tuQTL studies for a gene</td><td>(0, 1)</td></tr><tr><td><code>eQtlColocH4Maximum</code></td><td>Maximum H4 across all eQTL studies for a gene</td><td>(0, 1)</td></tr><tr><td><code>pQtlColocH4Maximum</code></td><td>Maximum H4 across all pQTL studies for a gene</td><td>(0, 1)</td></tr><tr><td><code>sQtlColocH4Maximum</code></td><td>Maximum H4 across all sQTL and tuQTL studies for a gene</td><td>(0, 1)</td></tr><tr><td><code>eQtlColocClppMaximumNeighbourhood</code></td><td>Ratio between <code>eQtlColocClppMaximum</code> for a gene and the maximum <code>eQtlColocClppMaximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>pQtlColocClppMaximumNeighbourhood</code></td><td>Ratio between <code>pQtlColocClppMaximum</code> for a gene and the maximum <code>pQtlColocClppMaximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>sQtlColocClppMaximumNeighbourhood</code></td><td>Ratio between <code>sQtlColocClppMaximum</code> for a gene and the maximum <code>sQtlColocClppMaximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>eQtlColocH4MaximumNeighbourhood</code></td><td>Ratio between <code>eQtlColocH4Maximum</code> for a gene and the maximum <code>eQtlColocH4Maximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>pQtlColocH4MaximumNeighbourhood</code></td><td>Ratio between <code>pQtlColocH4Maximum</code> for a gene and the maximum <code>pQtlColocH4Maximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>sQtlColocH4MaximumNeighbourhood</code></td><td>Ratio between <code>sQtlColocH4Maximum</code> for a gene and the maximum <code>sQtlColocH4Maximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>transPQtlColocH4Maximum</code></td><td>Maximum H4 colocalisation score between the GWAS locus and any trans-pQTL locus whose measured gene physically interacts with the candidate gene (via STRING, IntAct, or curated databases)</td><td>(0, 1)</td></tr><tr><td><code>transPQtlColocH4MaximumNeighbourhood</code></td><td>Ratio between <code>transPQtlColocH4Maximum</code> for a gene and the maximum <code>transPQtlColocH4Maximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr></tbody></table>

{% endtab %}

{% tab title="Variant Effect" %}

<table><thead><tr><th width="233.84765625">Feature Name</th><th width="384.11328125">Description</th><th>Range</th></tr></thead><tbody><tr><td><code>vepMaximum</code></td><td>Maximum VEP score across all variants in the credible set</td><td>(0, 1)</td></tr><tr><td><code>vepMaximumNeighbourhood</code></td><td>Ratio between <code>vepMaximum</code> for a gene and the maximum <code>vepMaximum</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr><tr><td><code>vepMean</code></td><td>Average VEP score between all variants in the credible set and a gene's footprint. The score of each variant is weighted by its posterior probability</td><td>(0, 1)</td></tr><tr><td><code>vepMeanNeighbourhood</code></td><td>Ratio between <code>vepMean</code> for a gene and the maximum <code>vepMean</code> for any protein-coding gene in the vicinity</td><td>(0, 1)</td></tr></tbody></table>

{% endtab %}

{% tab title="Other" %}

<table><thead><tr><th>Feature Name</th><th width="491.48046875">Description</th><th>Range</th></tr></thead><tbody><tr><td><code>geneCount500kb</code></td><td>Number of genes 250kb up- and down-stream the sentinel variant of a credible set</td><td>(0, 60000)</td></tr><tr><td><code>proteinGeneCount500kb</code></td><td>Number of protein-coding genes 250kb up- and down-stream the sentinel variant of a credible set</td><td>(0, 60000)</td></tr><tr><td><code>credibleSetConfidence</code></td><td>Degree of confidence we assign to the credible set definition based on the fine mapping methodology: 1 when fine-mapped with SuSIE using in-sample LD: 0.75 when fine-mapped with SuSIE using out-of-sample LD: 0.5, when fine-mapped with PICS and the locus is based on the analysis of summary statistics; 0.25, when fine-mapped with PICS and the locus was reported as a top hit according to the GWAS Catalog</td><td>(0, 1)</td></tr></tbody></table>

{% endtab %}

{% tab title="EnhancerToGene" %}

| Feature Name         | Description                                                                                                                                                                     | Range |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- |
| e2gMean              | Sum of the highest rE2G interaction score between regulatory region-gene and credible set overlap weighted by variant posterior probability and aggregated over gene and locus. | (0,1) |
| e2gNeighbourhoodMean | Ratio between `e2gMean` for a gene and the maximum `e2gMean` for any gene in the vicinity.                                                                                      | (0,1) |
| {% endtab %}         |                                                                                                                                                                                 |       |
| {% endtabs %}        |                                                                                                                                                                                 |       |

{% hint style="info" %}
**Neighbourhood features**

While some features are computed independently for each gene, others reflect the comparative relative context of a given gene compared to the other genes in the neighbourhood (+/- 500,000bp).&#x20;
{% endhint %}

{% hint style="info" %}
For a more detailed description of how each feature is computed, see [the L2G Feature documentation](https://opentargets.github.io/gentropy/python_api/datasets/l2g_features/_l2g_feature/).
{% endhint %}

## Training set

The L2G model is trained based on prior knowledge of gene-trait associations collated from different sources, to later bring representative credible sets supporting that association. The next steps describe the methodology to compose the training set:

#### **Effector gene list**

The **effector gene list (EGL)** represents the set of biologically validated gene-trait associations. This list is derived from several sources:

1. Manually curated [gold standards](https://github.com/opentargets/genetics-gold-standards) (“medium” and “high” confidence) from OTG 22.10.
2. Gene-indication pairs for the pharmacological target in all phase III or IV clinical trials according to the latest ChEMBL release.
3. Gene-disease or phenotype mappings with evidence score ≥ 0.95 from [ClinVar](https://platform-docs.opentargets.org/evidence#clinvar), [UniProt](https://platform-docs.opentargets.org/evidence#uniprot-variants), [Gene2Phenotype](https://platform-docs.opentargets.org/evidence#gene2phenotype), [Genomics England PanelApp](https://platform-docs.opentargets.org/evidence#genomics-england-panelapp) and [ClinGen](https://platform-docs.opentargets.org/evidence#clingen) from the latest Open Targets Platform release.

To ensure the uniqueness of gene-EFO pairs, the combined list was de-duplicated.

#### **Positives**

For each gene-trait pair from the effector gene list, positive gene-credible set pairs were extracted using the following criteria:

1. Only protein-coding genes.
2. Removed any pair with a `distanceSentinelTSS` feature less than 0.1.
3. Only include gene-trait pairs supported by at least two credible sets in different studies sharing the same lead variant. We added this criterion to reduce the chance of false positives among the credible sets.
4. Among all positive credible set-gene pairs, we removed duplications based on functional genomics features.
5. Removed credible sets involved in more than two positives.

#### **Negatives**

All other protein-coding genes in the window are classified as negatives as long as they don't have a strong functional interaction (STRINGdb score > 0.8) with any positive genes for the trait.

## The L2G model

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/methods/l2g/model/)

The Locus-to-Gene model is trained for every data release using the training set and feature matrix following. Similarly to Mountjoy *et al.*, the model is trained based on a gradient-boosting algorithm with the `scikit-learn` library, and using nested cross-validation and hyperparameter tuning.

{% embed url="<https://huggingface.co/opentargets/models>" %}

{% hint style="info" %}
The L2G model will be updated in each release to include new features and/or fixes. You should expect some variation in the prediction scores with each release.
{% endhint %}

## L2G predictions

📖 [Gentropy](https://opentargets.github.io/gentropy/python_api/datasets/l2g_prediction/)

The trained L2G model is applied to every GWAS credible set in the Open Targets Platform. All predictions below 0.05 are filtered out and feature contributions are added using SHAP analysis as described [here](/credible-set#explaining-l2g-predictions). These results feature as Target-Disease evidence, as well as L2G annotation for every credible in the Platform.&#x20;

{% hint style="info" %}
&#x20;The **GWAS Associations** evidence set is constructed based on GWAS credible sets linking to a protein-coding gene with an L2G score higher than 0.05.
{% endhint %}

### Reference

Mountjoy, E., Schmidt, E.M., Carmona, M. et al. [An open approach to systematically prioritize causal variants and genes at all published human GWAS trait-associated loci.](https://www.nature.com/articles/s41588-021-00945-5) Nat Genet 53, 1527–1533 (2021)


# Gentropy

Gentropy is an open-source Python package to facilitate the interpretation and analysis of GWAS and functional genomic studies for target identification. The Platform leverages Gentropy to perform post-GWAS analysis and derive the evidence and datasets for web portal visualisation.

Gentropy provides a set of data models, data ingestion methods and statistical analysis organised in discrete steps to maximise re-usability. Gentropy is designed with scalability in mind, so it's suitable for small dedicated analysis but also for the high-performing orchestrated tasks required to generate the data-lake that populates the Open Targets Platform.

<figure><img src="/files/ZrrpM4M8NdjZekvvfcZN" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
See more in the Gentropy 📖 [documentation](https://opentargets.github.io/gentropy/).
{% endhint %}


# Bibliography

Computational pipeline generating lists of publications based on similar entities

## Overview

The Platform Bibliography tool aims to provide context on the scientific background available in the literature regarding the target, disease or phenotype or drug entities of interest. The Platform aims to provide not only what publications have referenced the entities of interest, but also what other entities frequently occur in the literature in conjunction with the entities of interest. Open Targets has developed in collaboration with Europe PMC a pipeline that tries to maximise the literature information extraction by using a combination of Natural Entity Recognition, ontological normalisation and data analysis.

On our entity annotation pages, users also have the option of filtering the bibliography data to a specific timeframe.

## Data sources

#### Europe PMC

Europe PMC — [https://europepmc.org](https://europepmc.org/About) — is an open science database that facilitates access to a comprehensive collection of life science publications, preprints, and patents. Europe PMC provides researchers with freely available data available through their website, APIs, and bulk downloads.

In conjunction with the Europe PMC corpus, all the information required to build the Platform entities is required in order to establish a dictionary of terms and synonyms associated with each entity.

## Computational pipeline and datasets

#### Natural entity recognition (NER)

A machine learning model trained using BioBERT and a combination of curated datasets is applied on the Europe PMC corpus formed by abstracts and full-text articles. The model tags every gene-protein (GP), disease (DS) or drug (DR) entity in every publication. The resulting set provides some metadata on the publications, the sections where the matches are found and the sentences, where more than one type of entity co-occur.

#### Entity normalisation

[OnToma](https://github.com/opentargets/OnToma), our Python package for ontology mapping, is used to ground the tagged text to the Platform entities. This involves applying a series of standard Natural Language Processing tools (e.g. stemming, stop-words, etc.) followed by entity mapping using a dictionary-based approach. When entity mappings are ambiguous, disambiguation is performed to ensure only correct entity mappings remain. As a result, a fraction of the tagged sentences are grounded to the Platform entities and annotated with their respective identifiers.

The result of the normalisation is used in parallel to build the Platform Literature evidence covered [in a different section](/evidence#europe-pmc).

#### Literature-based similar entities

In order to define a strategy to find similar entities based on literature, the resulting matched entities are later used to train a Word2Vec model; the hyper-parameter settings used to train the model are taken following the suggestion from the benchmark conducted by Benjamin et al. (2020).

By using this model, the user can query what are other entities similar to the one selected, based on all the corpus of literature. Similarly, several entities can be selected and the product of their vectors can propose additional entities similar to all the selected entities.

The Bibliography section uses this algorithm to navigate the universe of publications. As a user keeps selecting entities, the universe is narrowed to show the intersection of all the publications that mention all of the selected entities.

## OpenAI literature summarisation tool

For data features that link to publications, the Platform now provides the option for users to ask for a natural language summary of the target-disease evidence presented in the publication (when the full-text article is available and free to re-use).

Using LangChain, we ask OpenAI’s GPT-4o mini model to summarise relevant portions of the text which we then ask it to summarise with the following prompt: “Can you provide a concise summary about the relationship between \[target] and \[disease] according to this study?“. The resulting text is presented to the user (see screenshots with example).

<figure><img src="/files/Lvx4GBEZamyittiwq6z0" alt=""><figcaption><p>Arrows annotate source of publication evidence and query prompt button for the feature</p></figcaption></figure>

<figure><img src="/files/KnIKze4qC3C9CO5GDg9o" alt=""><figcaption><p>Arrow shows prompted OpenAi summary</p></figcaption></figure>

We hope this feature will help the user better understand the available bibliography evidence, and it may actually highlight cases in which the publication does not in fact provide evidence for the target-disease relationship in question.

You can find details on the OpenAI [terms of use](https://openai.com/policies/terms-of-use) here.

## Publications

Rosonovski S, Levchenko M, Bhatnagar R, et al. Europe PMC in 2023. Nucleic Acids Research. 2024 Jan;52(D1):D1668-D1676. DOI: 10.1093/nar/gkad1085. PMID: [37994696](https://pubmed.ncbi.nlm.nih.gov/37994696/); PMCID: [PMC10767826](https://europepmc.org/article/MED/37994696).

Benjamin P. Chamberlain, Emanuele Rossi, Dan Shiebler, Suvash Sedhain, and Michael M. Bronstein. 2020. Tuning Word2vec for Large Scale Recommendation Systems. In Fourteenth ACM Conference on Recommender Systems (RecSys '20). Association for Computing Machinery, New York, NY, USA, 732–737. DOI:<https://doi.org/10.1145/3383313.3418486>


# Web interface

The web interface available at [https://platform.opentargets.org](https://platform.opentargets.org/) constitutes the first point of access for most of the Platform users. The site provides a unified search box connected to a series of tools allowing users to query different therapeutic hypotheses.

When possible, users can download the displayed information, but we invite users to visit the [Data access](/data-access) section when more complex queries are under consideration.


# Associations on the Fly

Learn about the updated associations view

“[Associations on the Fly](https://platform.opentargets.org/disease/EFO_0005774/associations)” is a revamp the Open Targets Platform association page with new facets and additional built-in functionalities. This view replaced the classic associations page.

In the Associations on the Fly page, a target or disease/phenotype is fixed and the prioritised list of alternative entities is displayed. A more detailed explanation on associations is available in the [Target-Disease associations](/associations) section.

### Key features

* Rapid comparison of evidence for different associations
* User control over the weighting of contributing evidence from each data source
* Ability to include and filter by specific data sources (Note: this is an OR filter)&#x20;
* Searching and applying filters by various target, disease or phenotype categories
* Ability to 'pin' a list of targets to create a customised list
* Upload a list of interested entities and export results
* Ability to propt a 'target interactors' subview

### Data source score weights

The data source weights are presented as an advanced option in the new user interface.\
They have been designed to allow users to dynamically modify the relative importance of the different data sources from the defaults set by Open Targets. The view automatically recomputes the association scores based on the new user-defined weights, giving this view its name: Associations on the Fly.

The new feature can be adjusted to modify the preset [Open Targets data sources scores](https://platform-docs.opentargets.org/associations#data-source-weights) and their effect on the global association score “on the fly”. This interaction ultimately enables a more tailored therapeutic hypothesis formulation.

### Evidence

Where evidence is available for a data source, clicking on the button will reveal the **detail widget** for that data source. Evidence displayed in the widget includes **indirect evidence**, so the user can interrogate evidence annotated with descendants of the disease or phenotype of interest.

### Filtering functionality

The Platform comes with a re-designed functionality which allows filtering on the Associations on the Fly and Target Prioritisation pages:

* **On a disease or phenotype** **page**: Users can search and apply target specific filters which filters the association and prioritisation pages by a particular target or a target category (categories details in the table below). The default is All Categories, however you can also select and view filter suggestions for a specific category from the drop down menu.&#x20;

<table><thead><tr><th width="136">Filter category</th><th width="426">Explanation</th><th>Example</th></tr></thead><tbody><tr><td>Names</td><td>Name of the target (<code>approvedName</code>)</td><td>interleukin 13, tyrosine kinase 2</td></tr><tr><td>Symbol</td><td>Target symbol (<code>approvedSymbol</code>)</td><td>IL13, TYK2 </td></tr><tr><td>ChEMBL Target Class</td><td>Class of drug target from the ChEMBL database</td><td>Enzyme, Kinase, Surface antigen</td></tr><tr><td>GO:BP</td><td>Gene Ontology: Biological Process<br>The larger processes, or ‘biological programs’ accomplished by multiple molecular activities</td><td>DNA repair, Intracellular signal transduction</td></tr><tr><td>GO:CC</td><td>Gene Ontology: Cellular Component<br>A location, relative to cellular compartments and structures, occupied by a macromolecular machine</td><td>Cytoskeleton, Clathrin complex</td></tr><tr><td>GO:MF</td><td>Gene Ontology: Molecular Function<br>Molecular-level activities performed by gene products</td><td>Oxidoreductase activity, Transporter activator activity</td></tr><tr><td>Reactome</td><td>Pathways from the Reactome database</td><td>Circadian clock, Interleukin-6 signaling</td></tr><tr><td>Subcellular Location</td><td>Subcellular location from UniProt and HPA</td><td>Cell membrane, Cytoplasm</td></tr><tr><td>Target ID (ENSG)</td><td>Ensembl gene IDs of the target, beginning with ENSG (<code>id</code>)</td><td>ENSG00000169194, ENSG00000105397</td></tr><tr><td>Tractability Antibody</td><td><a href="https://platform-docs.opentargets.org/target/tractability#antibody">Tractability assessments for a target</a> with data on an accessible epitope for antibody based therapy</td><td>UniProt loc high conf, Human Protein Atlas loc</td></tr><tr><td>Tractability Other Modalities</td><td><a href="https://platform-docs.opentargets.org/target/tractability#assessments">Tractability assessments for a target</a> with data on compound in clinical trials with a modality other than small molecule or antibody</td><td>Approved Drug</td></tr><tr><td>Tractability PROTAC</td><td><a href="https://platform-docs.opentargets.org/target/tractability#protac">Tractability assessments for a target</a> with data on using Proteolysis Targeting Chimeras (PROTACs)</td><td>UniProt Ubiquitination, Half-life Data</td></tr><tr><td>Tractability Small Molecule</td><td><a href="https://platform-docs.opentargets.org/target/tractability#small-molecule">Tractability assessments for a target</a> with data on binding site suitable for small molecule binding</td><td>Structure with Ligand, High-Quality Pocket</td></tr><tr><td>All Categories</td><td>Search and apply all of the above filters</td><td>IL13, ENSG00000105397</td></tr></tbody></table>

* **On a target page**: Users can search and apply disease or phenotype specific filters allowing them to filter the association page by a particular disease/phenotype or a disease/phenotype category \[Disease, Therapeutic Area (eg Infectious disease, Endocrine system disease)].

{% hint style="info" %}

* Whenever you select a specific category, a few suggestions from the selected category are shown by default.
* When multiple filters from different categories are selected, they are applied using an **AND** operator. When multiple filters within the same category are selected, they are applied with an **OR** operator.
  {% endhint %}

Check out this tutorial video to know more about this feature:

{% embed url="<https://youtu.be/-OU3ha7XPaU?si=9unxwQgGVReDUSd3>" %}

### Upload functionality

Users can now upload a custom list of targets or diseases or phenotype of interest to obtain a tailored association/target prioritisation view.

The feature is accessible from the "upload" icon on the Associations on the Fly page. This gives users the option of uploading a file containing a custom list of targets or diseases or phenotype. *The file should have one entity per row*. There are multiple allowed file formats for the uploaded file (.txt)\* , (.csv/.tsv/.xlsx)\*\* , (.json). An example format has also been provided for each file format in the feature.

The Platform then suggests potential matches between the entities in uploaded list and the ones in the Platform; the matches are provided through their platform entity ids. The users also have the option to select specific results that they wish to be displayed on the final view. Clicking the 'Pin hits' tab prompts the "Associations on the Fly" page to build up a custom view with the entities from the uploaded list.&#x20;

Watch our video describing the feature:

{% embed url="<https://www.youtube.com/watch?v=Mqvr2mwA7DM>" %}

{% hint style="info" %}
*\* For .txt file format, please create your input list using a text editor.*

*\*\* For .csv/.tsv/.xlsx file formats, please ensure that the file has a header called* `id`*.*
{% endhint %}

### Export functionality

We have designed and developed an export functionality for the Associations on the Fly/Target Prioritisation pages, allowing users to download:

* Entire dataset view (default status)
* Customised dataset view including custom controls changes, subset of data types (aggregations) and/or data from pinned targets only
* TSV and JSON formats

### Target Interactors View

This interactive view enables access to target interactions information within the main "Associations on the Fly" disease page.

The interactors subview is prompted by one of the context menu options and will contain target molecular interactions from the current Platform data feeds:

* IntAct for binary physical interactions
* Reactome for pathway-based interactions
* Signor for directional, causal interactions
* String for functional interactions

Users can select their favourite molecular interactions data source from a drop down list located on the interactions menu bar. The paginated interactors list is then sorted by their individual target-disease associations score, while their interaction score is shown by the target name.&#x20;

Additionally, a slider on the menu bar allows changing the default interaction score cutoff for the target interactors (please note - the default cutoff interaction score is 0.42 for IntAct and 0.75 for String. No scores reported by Reactome or Signor).

### Walkthrough of the new Associations on the Fly view

{% embed url="<https://www.youtube.com/watch?v=2A9bksboAag>" %}

## Reference

Cruz-Castillo et al. (2025). [Associations on the Fly, a new feature aiming to facilitate exploration of the Open Targets Platform evidence](https://academic.oup.com/bioinformatics/article/41/4/btaf070/8010255). *Bioinformatics.*


# Target Prioritisation

Learn about the target prioritisation view

The new [target prioritisation page](https://platform.opentargets.org/disease/EFO_0005774/associations?table=prioritisations) can be accessed by clicking onto the **Target prioritisation factors** tab from the Associations on the Fly page when searching for targets associated with a disease or phenotype.

The view focuses on displaying target-specific properties in a disease agnostic way, which have been aggregated into four main sections—**Precedence, Tractability, Doability, Safety—**&#x61;nd individually scored by the Open Targets team.

A "traffic light" system has been designed to visually inform on target prioritisation, with the aim to facilitate target recommendations. Using a colour scale, green indicates potentially positive attributes and red indicates potentially negative attributes, providing information to help users assess the targets for further prioritisation or deprioritisation, respectively.

## Precedence section

### Target in clinic

**Definition:** Gene is targeted by available drugs in any clinical stage for any indication.

**Source of Data:** Drugs and Clinical Precedence for a Target ([Open Targets](/target/drugs))

**Scoring:** Maximum harmonised clinical stage score across all drugs acting on the target, independently of the indication. Scores follow the same scale as the [Clinical Precedence Evidence](/evidence#clinical-precedence) dataset, ranging from 0.01 (Preclinical or Unknown) to 1.0 (Approval, Phase IV, or Withdrawal).

{% hint style="info" %}
For a full description of the **clinical stage categories** and their ranking, see [Clinical stage categories](/drug/clinical-report#clinical-stage-categories) in the Clinical Report page.
{% endhint %}

## Tractability section

### Membrane Protein

**Definition:** Target is annotated to be located in the cell membrane.

**Source of Data:** Platform Subcellular location widget \[HPA ([Human Protein Atlas](https://www.proteinatlas.org/)) and [UniProt](https://www.uniprot.org/)]

**Scoring:**

* 1 = Protein target is located (at least) in the cell or plasma membrane.&#x20;
* 0 = Protein target is not located in the cell membrane but some location information is accessible.&#x20;
* NA = No location information available.

### Secreted protein

**Definition:** Target is secreted or predicted to be secreted.

**Source of Data:** Platform Subcellular location widget \[HPA ([Human Protein Atlas](https://www.proteinatlas.org/)) and [UniProt](https://www.uniprot.org/)]

**Scoring:**

* 1 = Protein target is (at least) secreted or predicted to be secreted.&#x20;
* 0 = Not secreted but with location information.&#x20;
* NA = No location information available.

{% hint style="info" %}
Note: When contradictions between HPA (Human Protein Atlas) and UniProt exist (i.e. target is secreted according to HPA but in membrane according to UniProt), the information from HPA is taken.
{% endhint %}

### Ligand binder

**Definition:** Target binds at least one High-Quality Ligand according to ChEMBL tractability bucket.

**Source of Data:** Platform tractability widget ([Open Targets tractability](https://platform-docs.opentargets.org/target/tractability))

**Scoring:**

* 1 = Target has a high-quality ligand reported.&#x20;
* 0 = Target does not have high-quality ligand reported.
* NA = No information available.

### Small molecule binder

**Definition:** Target has been co-crystallised with a small molecule, reported in the Protein Data Bank.

**Source of Data:** Platform tractability widget ([Open Targets tractability](https://platform-docs.opentargets.org/target/tractability))

**Scoring:**

* 1 = Target has a small molecule reported.&#x20;
* 0 = Target does not have a small molecule reported.
* NA = No information available.

### Predicted pockets

**Definition:** Target has a [DrugEBIlity](https://chembl.blogspot.com/2010/10/drugebility-structure-based-component.html) score equal or above 0.7, which is predictive of harbouring a high-quality pocket.

**Source of Data:** Platform tractability widget ([Open Targets tractability](https://platform-docs.opentargets.org/target/tractability))

**Scoring:**

* 1 = Target contains a high-quality predicted pocket.&#x20;
* 0 = Target does not have a high-quality predicted pocket.
* NA= No information available.

## Doability section

> Models, tools and/or reagents that allow target assessment in preclinical settings to enable exploration of a given target

### Mouse ortholog identity

**Definition:** Mouse orthologs maximum identity percentage. A mouse harbouring an ortholog for the target of interest could be useful for *in vivo* assaying.

**Source of Data:** Platform comparative genomics widget ([Ensembl Compara](https://www.ensembl.org/info/docs/api/compara/index.html))

**Scoring:**

From 0 to 1 are linearly scored those targets with at least one ortholog in mice harbouring at least 80% with the target.

* 1 = There is at least one gene in mice that contains a sequence with a 100% of identity with the target.
* 0 = There are no genes in mice containing a sequence with at least 80% of identity with the target.
* NA = No ortholog information.

{% hint style="info" %}
Note: Here we consider mouse orthologs and display the "query percentage identity" (percentage of the human target sequence that matches to the mouse gene) when there is an 80% identity or more. In the cases of targets with more than one ortholog, we take the one with the maximum query % ID.
{% endhint %}

### Chemical probes

**Definition:** Target has high quality chemical probes.&#x20;

> [Chemical probes](https://platform-docs.opentargets.org/target/chemical-probes-and-teps#chemical-probes) are small molecules acting as chemical modulators, binding reversibly to the target.

**Source of Data:** Platform Chemical probes widget ([Probes & Drugs](https://www.probes-drugs.org/home/))

**Scoring:**

* 1 = Target has high-quality chemical probes.
* 0 = Target does not have high-quality chemical probes.
* NA = No information available.

## Safety section

### Genetic constraint

**Definition:** Genes that are important for human physiology are seen to be depleted of deleterious variants. The Genome Aggregation Database ([gnomAD](https://gnomad.broadinstitute.org/)) has developed a continuous measurement of intolerance to loss of function (LoF) variants per gene, based on observed/expected LoF variant analysis. As recommended by gnomAD and implemented in the Open Targets platform, the rank of genes regarding their loss-of-function observed/expected upper bound fraction (LOEUF) metric is used ([LOEUF score](https://gnomad.broadinstitute.org/help/constraint#loeuf)).

**Source of Data:** Platform genetic constraint widget ([gnomAD](https://gnomad.broadinstitute.org/))

**Scoring:**

A score from -1 to 1 is given to genes depending on their LOEUF metric rank, being -1 the least tolerant to LoF variation and 1 the most tolerant.

### Mouse models

**Definition:** The international database Mouse Genome Informatics contains information about reported phenotypes when a gene is knocked-out in this animal model. These phenotypes are categorised in multiple phenotype classes, using an organ/system classification. We retrieve this information (available in our platform) and score the phenotypes classes regarding their severity (from 0 to -1). After aggregating all phenotypes with their scores according to the phenotype class they belong to, we use the harmonic sum to build a continuous score, which is normalised from 0 to -1.

**Source of Data:** Platform mouse phenotypes widget (Mouse Phenotypes, feeded from [MGI](https://www.informatics.jax.org/), a reference database for mice knockouts)

**Scoring:**

* Below 0 to -1 = When the target has been knocked-out in mice there were multiple and severe phenotypes reported, with a score higher than the first quartile.&#x20;
* 0 = Either the target has non-severe phenotypes reported or is in the first quartile of the normalised score.&#x20;
* NA = No information available.

{% hint style="info" %}
Note: Below you can find how we scored the mouse phenotype classes (-1 being the "most severe" and 0 "non relevant" phenotypes
{% endhint %}

<table><thead><tr><th>id</th><th width="395">Phenotype Class</th><th>score</th></tr></thead><tbody><tr><td>MP:0005370</td><td>liver/biliary system phenotype</td><td>-1</td></tr><tr><td>MP:0005385</td><td>cardiovascular system phenotype</td><td>-1</td></tr><tr><td>MP:0010768</td><td>mortality/aging</td><td>-1</td></tr><tr><td>MP:0003631</td><td>nervous system phenotype</td><td>-0.75</td></tr><tr><td>MP:0005388</td><td>respiratory system phenotype</td><td>-0.75</td></tr><tr><td>MP:0005367</td><td>renal/urinary system phenotype</td><td>-0.75</td></tr><tr><td>MP:0005376</td><td>homeostasis/metabolism phenotype</td><td>-0.75</td></tr><tr><td>MP:0005386</td><td>behavior/neurological phenotype</td><td>-0.75</td></tr><tr><td>MP:0005381</td><td>digestive/alimentary phenotype</td><td>-0.5</td></tr><tr><td>MP:0005379</td><td>endocrine/exocrine gland phenotype</td><td>-0.5</td></tr><tr><td>MP:0005382</td><td>craniofacial phenotype</td><td>-0.5</td></tr><tr><td>MP:0005377</td><td>hearing/vestibular/ear phenotype</td><td>-0.5</td></tr><tr><td>MP:0005384</td><td>cellular phenotype</td><td>-0.5</td></tr><tr><td>MP:0005380</td><td>embryo phenotype</td><td>-0.5</td></tr><tr><td>MP:0005394</td><td>taste/olfaction phenotype</td><td>-0.5</td></tr><tr><td>MP:0002006</td><td>neoplasm</td><td>-0.5</td></tr><tr><td>MP:0005375</td><td>adipose tissue phenotype</td><td>-0.5</td></tr><tr><td>MP:0005389</td><td>reproductive system phenotype</td><td>-0.5</td></tr><tr><td>MP:0005397</td><td>hematopoietic system phenotype</td><td>-0.5</td></tr><tr><td>MP:0005387</td><td>immune system phenotype</td><td>-0.5</td></tr><tr><td>MP:0005391</td><td>vision/eye phenotype</td><td>-0.5</td></tr><tr><td>MP:0005390</td><td>skeleton phenotype</td><td>-0.5</td></tr><tr><td>MP:0005369</td><td>muscle phenotype</td><td>-0.25</td></tr><tr><td>MP:0001186</td><td>pigmentation phenotype</td><td>-0.25</td></tr><tr><td>MP:0005378</td><td>growth/size/body region phenotype</td><td>-0.25</td></tr><tr><td>MP:0005371</td><td>limbs/digits/tail phenotype</td><td>-0.25</td></tr><tr><td>MP:0010771</td><td>integument phenotype</td><td>-0.25</td></tr><tr><td>MP:0002873</td><td>normal phenotype</td><td>0</td></tr></tbody></table>

### Gene essentiality

**Definition:** This target prioritisation feature flags common essential genes suggested by the aggregate analysis of a large number of cell lines across numerous cancer types. [DepMap](https://depmap.org/portal/) assigns common essential status to a gene by gene effect rank in the 90th percentile least dependent line. The emphasis is on the consistency rather than strength of the dependency.

**Source of Data:** Gene essentiality (`CRISPRInferredCommonEssentials.csv`) is sourced from the [DepMap Portal](https://depmap.org/portal/).&#x20;

**Scoring:**

* -1 = Target reported as essential
* 0 = Target not reported as essential
* NA = No information available

### **Known safety events**

**Definition:** Target is associated with curated adverse events.

**Source of Data:** Safety liability data from Platform safety widget ([Open Targets Safety](https://platform-docs.opentargets.org/target/safety)) and Open Targets downstream analysis of  toxicity datasets from [ClinPGx](https://www.clinpgx.org/).

**Scoring:**

* -1 = The target has at least one adverse event.
* NA = No information available.

### **Cancer driver gene**

**Definition:** Target is classified as an oncogene and/or tumour suppressor gene.

**Source of Data:** Platform Cancer Hallmarks widget ([COSMIC](https://cancer.sanger.ac.uk/census))

**Scoring:**

We use the attribute information from the cancer hallmarks, in the target profile. Here, targets considered as "cancer driver genes" are flagged as tumour suppressor, oncogene, or both

* -1 = Target is catalogued as driver gene (tumour suppressor, oncogene or both).
* NA = No information available.

### **Paralogues**

**Definition:** Paralogue maximum identity percentage.

**Source of Data:** Platform comparative genomics widget ([Ensembl Compara](https://www.ensembl.org/info/docs/api/compara/index.html))

**Scoring:**

* Below 0 to -1 are linearly scored those targets with at least one paralogue in human harbouring at least 60% of identity with the target.
* 0 = Those targets with paralogues harbouring less than 60% of identity.
* NA = No information available about paralogues for that target.

### Tissue/Cell type specificity <a href="#tissue-specificity" id="tissue-specificity"></a>

**Definition:** CELLEX calculation of tissue-specific target expression.

**Source of Data:**&#x50;latform baseline expression widget ([GTEx](https://www.gtexportal.org/home/), [Tabula Sapiens](https://tabula-sapiens.sf.czbiohub.org/)). We used the assessment for every target from the RNA expression data that covered the whole body.

**Summary:** Tissue/Cell type specificity is scored using the CELLEX ESmu which is based on the average of four specificity metrics, detailed in [Timshel et al. 2020](https://elifesciences.org/articles/55851). We calculated ESmu on each of the aggregated \[tissues/cell types]. In the target prioritisation view, we display the maximum score across all \[GTEx and Tabula Sapiens/Tabula Sapiens] tissue/cell type ESmu values.

**Scoring:**

<table><thead><tr><th width="579">Tissue specificity HPA assessment</th><th>Score</th></tr></thead><tbody><tr><td>This target is the most specifically expressed gene in a [tissue/cell type]</td><td>1</td></tr><tr><td>This target is in the top 25% of specifically expressed genes in a [tissue/cell type]</td><td>0.5</td></tr><tr><td>This target is in the top 50% of specifically expressed genes in a [tissue/cell type]</td><td>0</td></tr><tr><td>This target is in the top 75% of specifically expressed genes in a [tissue/cell type]</td><td>-0.5</td></tr><tr><td>Not specifically expressed in any of the tissues/cell types</td><td>-1</td></tr><tr><td>Not calculated</td><td>NA</td></tr></tbody></table>

### Tissue/Cell type **distribution** <a href="#tissue-distribution" id="tissue-distribution"></a>

**Definition:**  Distribution of any detectable baseline expression for the target across tissue/celltype(s)

**Source of Data:** Platform baseline expression widget ([GTEx](https://www.gtexportal.org/home/), [Tabula Sapiens](https://tabula-sapiens.sf.czbiohub.org/)). We used the assessment for every target from the RNA expression data that covered the whole body.

**Summary:** We compute a score representing the breadth of expression within a dataset, defined as 1 minus the proportion of tissue/celltype(s) in which the target is expressed above a threshold of 0.5 (CPM/TPM) median expression. The threshold is chosen to minimise the impact of noise (e.g. from ambient RNA). For display in the target prioritisation view, a score of 0 is mapped to -1, indicating ubiquitous expression.<br>

**Scoring:**

<table><thead><tr><th width="580">Tissue distribution HPA assessment</th><th>Score</th></tr></thead><tbody><tr><td>This target is expressed in every single [tissue/cell type]</td><td>1</td></tr><tr><td>This target is expressed in 75% of [tissue/celltype]s</td><td>0.75</td></tr><tr><td>This target is expressed in 50% of [tissue/celltype]s</td><td>0.5</td></tr><tr><td>This target is expressed in 25% of [tissue/celltype]s</td><td>0.25</td></tr><tr><td>Not expressed in any [tissue/cell type]</td><td>-1</td></tr><tr><td>Not calculated</td><td>NA</td></tr></tbody></table>


# Evidence pages

Navigate the Associations view in the Open Targets Platform, from which you can access the Evidence pages

Evidence pages aim to summarise all the available evidence for a given target-disease pair. Evidence displayed in the page includes **indirect evidence**, so the user can interrogate evidence annotated with descendants of the disease or phenotype of interest.

The structure of the page follows the same structure as the profile pages including descriptions, summary widgets, and detail widgets.

To navigate to an evidence page, from the associations page, use the kebab menu icon, next to the name of a target or disease / phenotype.

<figure><img src="/files/x8pKpCOq9VVMoTiuq5m2" alt="" width="563"><figcaption><p>Navigating a target - disease evidence page</p></figcaption></figure>


# Entity profile pages

Profile pages aim to describe the entity and provide relevant annotation that might become informative at different stages of the drug development process.

At the top of each profile page, we provide a description for the entity, a series of cross-references to other databases, as well as synonyms provided by some of the upstream data sources.

Next, a list of **summary widgets** display the availability of certain category; summary widgets in grey report the lack of information for the given category.

Further down the page, **detail widgets** provide the full extent of information available in the Platform about a particular summary widget. A description explains the nature of the displayed information, as well as the source of the data.

Open Targets Platform currently has the following entity pages:&#x20;

1. [Target profile page](/target)
2. [Disease or phenotype profile page](/disease-or-phenotype)
3. [Drug profile page](/drug)
4. [Variant profile page](/variant)
5. [Study profile page](/study)
6. [Credible set profile page](https://www.ebi.ac.uk/ols4/ontologies/uberon)

<figure><img src="/files/oSUnN6XebABRUVWtMPwa" alt=""><figcaption><p>BRAF profile page</p></figcaption></figure>


# Data and code access

To support systematic identification and prioritisation of drug targets, we are committed to making our data open and available in a variety of formats to support academic research purposes and commercial activities. For more information, see our [Licence documentation](/licence).

For queries about a **single entity or target-disease association**, we recommend that you use:

* Our intuitive [web interface](/web-interface) with data tables and data visualisations that can be downloaded/exported in multiple formats, including TSV and JSON
* Our robust [GraphQL API](/data-access/graphql-api) with endpoints that can be accessed with the programming language of your choice and an [interactive playground](http://api.platform.opentargets.org/api/v4/graphql/browser) where you can try out sample queries

For more **complex, systematic queries**, we recommend that you use:

* Our comprehensive set of [data downloads](https://platform.opentargets.org/downloads) available via [FTP](https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/), [Google Cloud](https://console.cloud.google.com/marketplace/product/bigquery-public-data/open-targets-platform?project=open-targets-genetics\&inv=1\&invt=Abwwaw), [Microsoft Azure Open Datasets](https://learn.microsoft.com/en-us/azure/open-datasets/dataset-open-targets) and [AWS Open Registry Dataset](https://platform-docs.opentargets.org/data-access/release-outputs-on-aws) .
* Our [Google BigQuery](/data-access/google-bigquery) instance that supports SQL-like queries and allows you to export data to your own Google Cloud Storage bucket. This data is available as a [BigQuery Public Dataset](https://cloud.google.com/bigquery/public-data).

If you use our data in your research or commercial product, please [cite our latest publication](/citation).

### New export functionality

We have designed and developed a new export functionality for the Associations On the Fly/Target Prioritisation pages, allowing users to download:

* Entire dataset view (default status)
* Customised dataset view including custom controls changes, subset of data types (aggregations) and/or data from pinned targets only
* TSV and JSON formats


# Download datasets

To support more complex and systematic queries, we provide all datasets as data downloads.

A list of all datasets is available in the [Platform Data Downloads](https://platform.opentargets.org/downloads) page.

All Platform **datasets** are available as a distributed collection of data. This implies that for each dataset, there will be a directory with a list of partitioned files. Currently, we produce our datasets in Parquet. This formats allow us to expose nested information in a machine-readable way.&#x20;

Archive datasets, as well as input files and other secondary products, are also made available in the [FTP server](http://ftp.ebi.ac.uk/pub/databases/opentargets/platform/) and [Google Cloud Platform](https://console.cloud.google.com/storage/browser/open-targets-data-releases).

Below, we describe how to download, access and query this information in a step-by-step guide.

## Download

Below is a walkthrough on how to download the `disease` dataset from the `25.03` release in Parquet format using different approaches.

We recommend using **lftp** with a command line client, and when using tools like *wget*, *curl*, etc., use *https\://* rather than *ftp\://*

#### Using rsync

`rsync` is a command line tool for efficiently transferring and synchronising files between a computer and an external hard drive.

```bash
rsync -rpltvz --delete rsync.ebi.ac.uk::pub/databases/opentargets/platform/25.03/output/disease .
```

#### Using wget

`wget` is a command line tool that retrieves content from web servers and widely available in Unix systems.

```bash
wget --recursive --no-parent --no-host-directories --cut-dirs 8 \
https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/25.03/output/disease
```

#### Using Google Cloud Platform (paywalled after 1TB)

Users with Google Cloud Platform account can download the datasets through the Google Cloud Console or using `gsutil` command-line tool.

```bash
gsutil -m cp -r gs://open-targets-data-releases/25.03/output/disease
```

#### Other ways to access data

If you are using a non-Linux or non-Unix machine (e.g. Windows), you can access our FTP service using an FTP client like FileZilla or the Windows `ftp` command. For more information, including tips and workarounds, see the [Community Windows `ftp` thread](https://community.opentargets.org/t/how-to-query-api-by-drug-trade-name/316/3?u=ahercules).

## Accessing and querying datasets

To read the information available in the partitioned datasets, there is no need to manipulate or concatenate files. Datasets can be read directly using the dataset path.

The next scripts provide a proof-of-concept example using the ClinVar evidence provided by the European Variation Archive. The next scripts show how to:

* Read a dataset
* Explore the schema of the dataset
* Select a subset of information (columns)
* Display the information

First of all the dataset needs to be downloaded as described in the previous section. For simplicity, only EVA evidence is downloaded, but all evidence can be downloaded at once using the same approach.

```bash
gsutil -m cp -r gs://open-targets-data-releases/25.03/output/evidence/sourceId=eva
```

The next scripts make use of Apache Spark ([PySpark](https://spark.apache.org/docs/latest/api/python/index.html) or [Sparklyr](https://spark.rstudio.com)) to read and query the dataset using modern functional programming approaches. These packages need to be installed in their respective environments.

The next query only displays 6 fields of the ClinVar evidence but there are other non-null values available. The **schema** is the best way to explore what's available and query the most relevant information. All Platform evidence share the same schema, so there will be a long list of fields that might not be informative for ClinVar but will be relevant if trying to query other data sources.

Dealing with nested information can sometimes be tedious. The Platform aims to minimise the nestiness of the data, however some level of structure is sometimes required. Spark provides a series of functions to deal with complex nested information. The scripts provide an example on how the `clinicalSignificances` array is flattened using the `explode` function.

Once loaded into Python or R, the user can decide to continue using Spark, write the output to a file or use alternative libraries to process the information (e.g. `pandas`, `tidyverse`, etc.).

{% tabs %}
{% tab title="Python" %}

```python
from pyspark import SparkConf
from pyspark.sql import SparkSession
import pyspark.sql.functions as F

# path to ClinVar (EVA) evidence dataset 
# directory stored on your local machine
evidencePath = "local directory path - e.g. /User/downloads/sourceId=eva"

# establish spark connection
spark = (
    SparkSession.builder
    .master('local[*]')
    .getOrCreate()
)

# read evidence dataset
evd = spark.read.parquet(evidencePath)

# Browse the evidence schema
evd.printSchema()

# select fields of interest
evdSelect = (evd
 .select("targetId",
         "diseaseId",
         "variantRsId",
         "studyId",
         F.explode("clinicalSignificances").alias("cs"),
         "confidence")
 )
 evdSelect.show()

# +---------------+--------------+-----------+------------+--------------------+--------------------+
# |       targetId|     diseaseId|variantRsId|     studyId|                  cs|          confidence|
# +---------------+--------------+-----------+------------+--------------------+--------------------+
# |ENSG00000153201|Orphanet_88619|rs773278648|RCV001042548|uncertain signifi...|criteria provided...|
# |ENSG00000115718|  Orphanet_745|       null|RCV001134697|uncertain signifi...|criteria provided...|
# |ENSG00000107147|    HP_0001250|rs539139475|RCV000720408|       likely benign|criteria provided...|
# |ENSG00000175426|Orphanet_71528|rs142567487|RCV000292648|uncertain signifi...|criteria provided...|
# |ENSG00000169174|   EFO_0004911|rs563024336|RCV000375546|uncertain signifi...|criteria provided...|
# |ENSG00000140521|  Orphanet_298|rs376306906|RCV000763992|uncertain signifi...|criteria provided...|
# |ENSG00000134982|   EFO_0005842| rs74627407|RCV000073743|               other|no assertion crit...|
# |ENSG00000187498| MONDO_0008289|rs146288748|RCV001111533|uncertain signifi...|criteria provided...|
# |ENSG00000116688|Orphanet_64749|rs119103265|RCV000857104|uncertain signifi...|no assertion crit...|
# |ENSG00000133812|Orphanet_99956|rs562275980|RCV000367609|uncertain signifi...|criteria provided...|
# +---------------+--------------+-----------+------------+--------------------+--------------------+
# only showing top 10 rows

# Convert to a Pandas Dataframe
evdSelect.toPandas()
```

{% endtab %}

{% tab title="R" %}

```r
library(dplyr)
library(sparklyr)
library(sparklyr.nested)

## path to ClinVar (EVA) evidence dataset 
## directory stored on your local machine
evidencePath <- "local directory path - e.g. /User/downloads/sourceId=eva"

## establish connection
sc <- spark_connect(master = "local")

## read evidence dataset
evd <- spark_read_parquet(sc,
                          path = evidencePath)
## Browse the evidence schema
columns <- evd %>%
  sdf_schema() %>%
  lapply(function(x) do.call(tibble, x)) %>%
  bind_rows()

## select fields of interest
evdSelect <- evd %>%
  select(targetId,
         diseaseId,
         variantRsId,
         studyId,
         clinicalSignificances,
         confidence) %>%
  sdf_explode(clinicalSignificances)

##  # Source: spark<?> [?? x 6]
##    targetId   diseaseId   variantRsId studyId  clinicalSignific… confidence     
##    <chr>      <chr>       <chr>       <chr>    <chr>             <chr>          
##  1 ENSG00000… Orphanet_8… rs773278648 RCV0010… uncertain signif… criteria provi…
##  2 ENSG00000… Orphanet_7… NA          RCV0011… uncertain signif… criteria provi…
##  3 ENSG00000… HP_0001250  rs539139475 RCV0007… likely benign     criteria provi…
##  4 ENSG00000… Orphanet_7… rs142567487 RCV0002… uncertain signif… criteria provi…
##  5 ENSG00000… EFO_0004911 rs563024336 RCV0003… uncertain signif… criteria provi…
##  6 ENSG00000… Orphanet_2… rs376306906 RCV0007… uncertain signif… criteria provi…
##  7 ENSG00000… EFO_0005842 rs74627407  RCV0000… other             no assertion c…
##  8 ENSG00000… MONDO_0008… rs146288748 RCV0011… uncertain signif… criteria provi…
##  9 ENSG00000… Orphanet_6… rs119103265 RCV0008… uncertain signif… no assertion c…
## 10 ENSG00000… Orphanet_9… rs562275980 RCV0003… uncertain signif… criteria provi…
## # … with more rows

# Convert to a dplyr tibble
evdSelect %>%
  collect()
```

{% endtab %}
{% endtabs %}

## File formats&#x20;

The Open Targets data generation pipeline produces outputs only in **Parquet** file format. The pipeline no longer produces outputs in JSON file format. This is due to Parquet file format having favourable features like:&#x20;

* built-in schema and data typing
* size-efficiency when compressed
* efficient reading
* the wide availability of interfaces with most dataframe libraries

If you are new to Parquet and switching over from JSON, the change should be simple and your pipeline should be faster at reading the data. There are various examples of Parquet file readers from popular data frame libraries in [R](https://www.rdocumentation.org/packages/arrow/versions/0.14.1/topics/read_parquet), [Spark](https://spark.apache.org/docs/3.5.1/sql-data-sources-parquet.html), [Polar](https://docs.pola.rs/user-guide/io/parquet/), [Pandas](https://pandas.pydata.org/docs/reference/api/pandas.read_parquet.html). Typically the reader is built on the Apache Arrow [library](https://arrow.apache.org/), which itself has APIs in many languages should you need them.

If you don’t wish to read data into dataframes and instead want to read the data as JSON (newline delimited), Open Targets has its [in-house tool `p2j` (python)](https://github.com/opentargets/p2j) for converting parquet to newline delimited JSON.

{% hint style="info" %}
Post 25.03, the data downloads paths have changed as now only parquet file format is available. Also there are minor changes to the name of the dataset (snake\_case & singular). More details can be found [here](https://community.opentargets.org/t/issues-with-gcp-big-query/1704/5).

* Previous releases (till 24.09):
  * `https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/24.09/output/etl/parquet/associationByOverallDirect/`
* 25.03 release (and therafter):&#x20;
  * `https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/25.03/output/association_by_datasource_direct/`
    {% endhint %}

## Tutorials and how-to guides

For more information on how to access and work with [our data downloads](https://platform.opentargets.org/downloads/data) and example scripts based on actual use cases and research questions, check out the [Open Targets Community](http://community.opentargets.org).


# Platform datasets on AWS

### Accessing Open Targets Platform data hosted on AWS

From 2026, Open Targets Platform datasets are also made available through Amazon's [Registry of Open Data](https://registry.opendata.aws/opentargets/) initiative.&#x20;

**Dataset location:**

`s3://open-targets-public-data-releases/platform/`

This [notebook](https://colab.research.google.com/github/opentargets/notebooks/blob/main/notebooks/reading_data_from_aws.ipynb#scrollTo=79ce23e9-8572-4d7c-b2dd-a7a022920ebb) provides guided examples demonstrating how to access these datasets from the dedicated AWS S3 bucket using [pandas](https://pandas.pydata.org/), [polars](https://pola.rs/) and [pyspark](https://spark.apache.org/docs/latest/api/python/index.html).


# Google BigQuery

To support more complex queries and advanced informatics workflows that use Google Cloud services, the Open Targets Platform data is also available as a Google Cloud public dataset via our Google BigQuery instance — [open-targets-prod](https://console.cloud.google.com/bigquery?p=open-targets-prod\&d=platform_21_06).

## What is Google BigQuery?

Google BigQuery is a data warehouse that enables researchers to run super-fast, asynchronous SQL queries using Google's cloud infrastructure. After running your query, you can either export into various formats or copy into a Google Cloud bucket for further downstream analyses.

Open Targets Platform data is publicly accessible as a [Google Cloud public dataset](https://console.cloud.google.com/marketplace/product/bigquery-public-data/open-targets-platform?project=open-targets-genetics). Users only pay for the queries they perform on the data, and through this program, the first 1 TB per month is free.

## BigQuery access points

Open Targets has uploaded all of our data to Google BigQuery. You can run queries via:

* [Cloud Console](https://cloud.google.com/bigquery/docs/quickstarts/quickstart-web-ui)
* [Command line `bq` tool](https://cloud.google.com/bigquery/docs/quickstarts/quickstart-command-line)
* [Client libraries, including Python](https://cloud.google.com/bigquery/docs/quickstarts/quickstart-client-libraries)

For more information on BiqQuery, please review the [BigQuery documentation](https://cloud.google.com/bigquery/docs).

## Example BigQuery SQL queries

Below is a sample query that uses our `association_overall_direct` dataset to return a list of targets associated with psoriasis (EFO\_0000676) and the overall association score.

```sql
SELECT
  associations.targetId AS target_id,
  targets.approvedSymbol AS target_approved_symbol,
  associations.diseaseId AS disease_id,
  diseases.name AS disease_name,
  associations.score AS overall_association_score
FROM
  `open-targets-prod.platform.association_overall_direct` AS associations
JOIN
  `open-targets-prod.platform.disease` AS diseases
ON
  associations.diseaseId = diseases.id
JOIN
  `open-targets-prod.platform.target` AS targets
ON
  associations.targetId = targets.id
WHERE
  associations.diseaseId='EFO_0000676'
ORDER BY
  associations.score DESC
```

Similarly, you can use our `drug_molecule` dataset and pass a list of drug trade names to find relevant information:

```sql
DECLARE
  my_drug_list ARRAY<STRING>;
SET
  my_drug_list = [ 'Premarin',
  'Calcium disodium versenate',
  'Keytruda',
  'Vioxx',
  'Humira' ];
SELECT
  id AS drug_id,
  name AS drug_chembl_name,
  tradeNameList.element AS drug_trade_name,
  drugType AS drug_type,
  isApproved AS drug_is_approved,
  blackBoxWarning AS drug_blackbox_warning,
  hasBeenWithdrawn AS drug_withdrawn,
FROM
  `open-targets-prod.platform.drug_molecule`,
  UNNEST (tradeNames.list) AS tradeNameList
WHERE
  (tradeNameList.element) IN UNNEST(my_drug_list)
```

## Tutorials and how-to guides

For more information on how to use BigQuery to access Platform data and example queries based on actual use cases and research questions, check out the [Open Targets Community](https://community.opentargets.org) and [our Google Cloud dataset homepage](https://console.cloud.google.com/marketplace/product/bigquery-public-data/open-targets-platform?project=open-targets-prod).&#x20;


# GraphQL API

The Open Targets Platform GraphQL — available at [http://api.platform.opentargets.org](http://api.platform.opentargets.org/api/v4/graphql/schema) — is our new API that allows for language-agnostic access to our data, along with other key benefits:

1. You can construct a query that returns only the fields that you need
2. You can build graphical queries that traverse a data graph through resolvable entities and this reduces the need for multiple queries
3. You can access the [GraphQL API playground](http://api.platform.opentargets.org/api/v4/graphql/browser) with built-in documentation and schema showing required and optional parameters
4. You can view the [schema](http://api.platform.opentargets.org/api/v4/graphql/schema) that shows the available fields for each object along with a description and data type attribute
5. You only have to use `POST` requests with a simple query string and variables object

Our GraphQL API supports queries for a single target, disease/phenotype, drug, or target-disease association. For more systematic queries (e.g. for multiple targets), please use [our data downloads](/data-access/datasets) or [our Google BigQuery instance](/data-access/google-bigquery).

## Available endpoints

The base URL endpoint for our new GraphQL API is:

```
https://api.platform.opentargets.org/api/v4/graphql
```

You can then access relevant data from the following endpoints:

**/target:** contains annotation information for targets including tractability assessments, mouse phenotype models, and baseline expression; also contains data on diseases and phenotypes association with the given target

**/disease:** contains annotation information for diseases and phenotypes including ontology, known drugs, and clinical signs and symptoms; also contains data on targets associated with the given disease or phenotype

**/drug:** contains annotation information for compounds and drugs including mechanisms of action, indications, and pharmacovigilance data

**/variant:** contains annotation information for variants including population allele frequencies, variant effect, transcript consequences, and credible sets associated with complex traits containing the variant.

**/studies:** contains annotation information for studies including trait or phenotype, publication, cohort information and list of credible sets associated with the study.

**/credibleSet:** contains annotation information for credible sets including the complete sets of variants in the credible set, gene assignment based on our L2G predictions and colocalisation metrics.

**/search**: contains index of all entities contained within the Platform

## Example GraphQL query

Below is an example GraphQL query for AR (ENSG00000169083) that will return Genetic Constraint and Tractability data.

```graphql
query targetInfo {
  target(ensemblId: "ENSG00000169083") {
    id
    approvedSymbol
    biotype
    geneticConstraint{
      constraintType
      exp
      obs
      score
      oe
      oeLower
      oeUpper
    }
    tractability{
      label
      modality
      value
    }
  }
}
```

[Run this query in our GraphQL API playground](https://api.platform.opentargets.org/api/v4/graphql/browser?query=query%20targetInfo%20%7B%0A%20%20target%28ensemblId%3A%20%22ENSG00000169083%22%29%20%7B%0A%20%20%20%20id%0A%20%20%20%20approvedSymbol%0A%20%20%20%20biotype%0A%20%20%20%20geneticConstraint%20%7B%0A%20%20%20%20%20%20constraintType%0A%20%20%20%20%20%20exp%0A%20%20%20%20%20%20obs%0A%20%20%20%20%20%20score%0A%20%20%20%20%20%20oe%0A%20%20%20%20%20%20oeLower%0A%20%20%20%20%20%20oeUpper%0A%20%20%20%20%7D%0A%20%20%20%20tractability%20%7B%0A%20%20%20%20%20%20label%0A%20%20%20%20%20%20modality%0A%20%20%20%20%20%20value%0A%20%20%20%20%7D%0A%20%20%7D%0A%7D%0A)

Using GraphQL's [query strings](https://graphql.org/learn/queries/) and [variables object](https://graphql.org/graphql-js/passing-arguments/) constructs, you can also access the data using a programming language that supports HTTP `POST` requests. While this is a valid approach, we discourage users from repeatedly querying the GraphQL API one entity at a time. Instead, our comprehensive [datasets available for download ](/data-access/datasets)provide a simpler and more performant strategy to achieve the same result.

### Sample scripts

Below is an example script using the same AR query above, but written for Python and R:

{% tabs %}
{% tab title="Python" %}

```python
#!/usr/bin/env python3

# Import relevant libraries to make HTTP requests and parse JSON response
import requests
import json

# Set gene_id variable for AR (androgen receptor)
gene_id = "ENSG00000169083"

# Build query string to get general information about AR and genetic constraint and tractability assessments 
query_string = """
  query target($ensemblId: String!){
    target(ensemblId: $ensemblId){
      id
      approvedSymbol
      biotype
      geneticConstraint {
        constraintType
        exp
        obs
        score
        oe
        oeLower
        oeUpper
      }
      tractability {
        label
        modality
        value
      }
    }
  }
"""

# Set variables object of arguments to be passed to endpoint
variables = {"ensemblId": gene_id}

# Set base URL of GraphQL API endpoint
base_url = "https://api.platform.opentargets.org/api/v4/graphql"

# Perform POST request and check status code of response
r = requests.post(base_url, json={"query": query_string, "variables": variables})
print(r.status_code)

# Transform API response from JSON into Python dictionary and print in console
api_response = json.loads(r.text)
print(api_response)
```

{% endtab %}

{% tab title="R" %}

```r
# Install relevant library for HTTP requests
library(httr)

# Set gene_id variable for AR (androgen receptor)
gene_id <- "ENSG00000169083"

# Build query string to get general information about AR and genetic constraint and tractability assessments 
query_string = "
  query target($ensemblId: String!){
    target(ensemblId: $ensemblId){
      id
      approvedSymbol
      biotype
      geneticConstraint {
        constraintType
        exp
        obs
        score
        oe
        oeLower
        oeUpper
      }
      tractability {
        label
        modality
        value
      }
    }
  }
"

# Set base URL of GraphQL API endpoint
base_url <- "https://api.platform.opentargets.org/api/v4/graphql"

# Set variables object of arguments to be passed to endpoint
variables <- list("ensemblId" = gene_id)

# Construct POST request body object with query string and variables
post_body <- list(query = query_string, variables = variables)

# Perform POST request
r <- POST(url=base_url, body=post_body, encode='json')

# Print data to RStudio console
print(content(r)$data)
```

{% endtab %}
{% endtabs %}

## Tutorials and how-to guides

{% embed url="<https://youtu.be/_sZR0VxpwqE>" %}

For more information on how to use the GraphQL API and example queries based on actual use cases and research questions, check out the [Open Targets Community](https://community.opentargets.org).


# Model Context Protocol

{% hint style="warning" %}
The MCP is under active development and experimentation. All feedback is welcome in <https://community.opentargets.org>
{% endhint %}

The Open Targets Platform Model Context Protocol (MCP) server provides a purpose-built interface and instructions to access and interpret the data and analyses in the Open Targets Platform using AI.

MCP Servers are part of a standardised interface for AI agents to interact with data sources using the [Model Context Protocol](https://modelcontextprotocol.io/docs/getting-started/intro).

### Ways to connect

* **Remote Server**: Connect directly to Open Targets hosted endpoint available in `https://mcp.platform.opentargets.org/mcp`
* **Local Server:** Run a local MCP server using uvx or Docker

More detailed information can be found in the [MCP server github repository](https://github.com/opentargets/open-targets-platform-mcp).

Claude provides [a detailed tutorial](https://claude.com/resources/tutorials/using-the-open-targets-connector-in-claude) on how to add the MCP as a connector.

<figure><img src="/files/XQGahytQRfDPMDyoqwkL" alt=""><figcaption><p>Remote Server + Sonnet 4.5 + Open Targets Platform 25.12 data</p></figcaption></figure>

{% hint style="info" %}
The Platform documentation you are reading here is also available as an MCP in `https://platform-docs.opentargets.org/~gitbook/mcp`
{% endhint %}


# Platform infrastructure

Overview of the technical infrastructure that supports the Open Targets Platform

## Introduction

The Open Targets Platform infrastructure stack is composed of three layers, data, backend and frontend.

1. Data layer
   1. Opensearch — Contains the bulk of the data
   2. Clickhouse — Contains data related to associations view
2. Backend layer
   1. API — The main backend application, providing the GraphQL engine
   2. OpenAI-API — The summarizing engine for literature section
3. Frontend layer
   1. Web — A React SPA web application

<figure><img src="/files/6GsuaJijfcTFesGelc9f" alt="" width="563"><figcaption></figcaption></figure>

Currently this infrastructure is hosted on Google Cloud, using a load-balanced, globally distributed and highly scalable deployment based on Terraform.

## GitHub repositories

### Backend

* [platform-api](https://github.com/opentargets/platform-api) — GraphQL API
* [ot-ai-api](https://github.com/opentargets/ot-ai-api) — OpenAI API router
* [open-targets-platform-mcp](https://github.com/opentargets/open-targets-platform-mcp) — MCP

### Frontend

* [ot-ui-apps](https://github.com/opentargets/ot-ui-apps) — Open Targets web applications

### Deployment

* [terraform-google-opentargets](https://github.com/opentargets/terraform-google-opentargets-platform) — Open Targets Infrastructure definition, scripts in charge of setting up the data layer with the relevant disk images and spinning up the rest of the infrastructure.

## Open source contributions

As a consortium committed to developing open-source, freely available tools that support systematic drug target identification and prioritisation, we actively encourage and accept open source contributions to our various repositories.

Please review [our contribution guidelines](https://github.com/opentargets/platform/blob/master/Contribution%20Guidelines.md) and check out the [Open Targets Community](https://community.opentargets.org) to get started.

If you have further questions, please get in touch with us on the [Open Targets Community](https://community.opentargets.org/).


# Data pipeline

The Open Targets data pipeline is a complex process orchestrated in Apache Airflow, and it is divideded into data acquisition, transformation and data output.

## Introduction

The data pipeline is composed of multiple elements:

1. Data and evidence generation processes
2. Input stage
3. Transformation stage and ETL processes
4. Output stage
5. Gentropy-specific processes
6. Orchestration

## GitHub repositories

### Data and evidence

* [curation](https://github.com/opentargets/curation) — Open Targets curation repository
* [json\_schema](https://github.com/opentargets/json_schema) — evidence object schema used for evidence and association scoring
* [OnToma](https://github.com/opentargets/OnToma) — Python package for mapping ontology terms to the disease, target, and drug indices

### Gentropy

* [gentropy](https://github.com/opentargets/gentropy) — Open Targets' genomics toolkit

See [here](https://app.gitbook.com/o/-LC3OlEMulAutIN2QOro/s/-MU4dMxOmLaVNWfVNvpC/~/changes/506/gentropy) for more info on the Gentropy pipelines.&#x20;

### Orchestration&#x20;

* [orchestration](https://github.com/opentargets/orchestration) — Open Targets data pipelines orchestrator

See detailed orchestration documentation [here](https://github.com/opentargets/orchestration/tree/dev/docs).

The Platform ETL (“extract, transform, and load”) and the [Genetics ETL](https://app.gitbook.com/o/-LC3OlEMulAutIN2QOro/s/-MU4dMxOmLaVNWfVNvpC/~/changes/506/gentropy#genetics-etl) were separate processes before, but they are now merged into one single pipeline. This means that the data produced for both Genetics ETL and the Platform are released at the same time. Herein, we refer to this joint pipeline as the "unified pipeline".

The orchestration occurs on Google Airflow using Google Cloud as the cloud resource provider. The logic of the orchestration is based on the steps. The combination of steps forms directed acyclic graphs (DAGs).&#x20;

<figure><img src="/files/oA77Q4b7piU1qIzmdc6u" alt=""><figcaption><p>Schematic overview of Open Targets pipelines</p></figcaption></figure>

The unified pipeline uses many [static assets](https://app.gitbook.com/o/-LC3OlEMulAutIN2QOro/s/-MU4dMxOmLaVNWfVNvpC/~/changes/506/gentropy#static-assets) (link), like Open Targets related data and data needed to run Genetics ETL.

### Unified pipeline&#x20;

* [otter](https://github.com/opentargets/otter) — **O**pen **T**argets' **T**ask **E**xecuto**R** i.e. scripts that process and prepare data for our ETL pipelines
* [pts](https://github.com/opentargets/pts) — **P**ipeline **T**ransformation **S**tage i.e. scripts that convert files into formats and structures used by the Open Targets data pipeline
* [platform-output-support](https://github.com/opentargets/platform-output-support): scripts for infrastructure tasks and generating a Platform release

If you have further questions, please get in touch with us on the [Open Targets Community](https://community.opentargets.org/).


# FAQs

**Do you have questions for us? Ask on** [**Open Targets Community**](https://community.opentargets.org/c/frequently-asked-questions/6)**!**

* **Why can't I find a variant that I search for?**&#x20;

In the Open Targets Platform, a variant refers to any human variation that is associated with a [disease, trait or phenotype](/disease-or-phenotype) that has been reported in at least one of our [variant-to-phenotype](/variant#variant-to-phenotype) sources. This accounts for only 1% of the total variants reported in [gnomAD](https://gnomad.broadinstitute.org/) (6.5M vs 700M). See [here](/variant) for more information.

* **Why can't I find a study I search for?**

Each [study](https://app.gitbook.com/o/-LC3OlEMulAutIN2QOro/s/-MU4dMxOmLaVNWfVNvpC/~/changes/506/study) undergoes quality control and validation procedures before ingestion into the Platform.&#x20;

We remove studies that had unsupported study types, invalid target ID or invalid biosample ID (for molQTLs), and invalid ontology trait mapping. See [here ](https://app.gitbook.com/o/-LC3OlEMulAutIN2QOro/s/-MU4dMxOmLaVNWfVNvpC/~/changes/506/study#study-inclusion-criteria)for study inclusion criteria.

For GWAS Catalog studies with summary statistic, we perform [summary statistic quality control](/gentropy/data-sources#gcss-quality-control). If this fails, we exclude it.

&#x20;For GWAS Catalog curated associations (top hits), we tightened the p-value significance threshold compared to OTG 22.10 (1e-8 instead of 5e-8). This may also lead to differences in the study list.

* **What is a Credible Set?**

A credible set is the set of variants near a genetic association signal that have a 95% probability of containing the true causal variant(s) for that signal.

* **What are the quality control criteria for fine mapping?**

Credible sets (CSs) are filtered out based on the following criteria:

1. The lead variant maps within the MHC region (`chr6:25726063-33400556`)
2. The lead variant has invalid chromosome coding (`1:22, X, Y, XY, mt`)
3. The CS is not on the list of valid studies
4. It is the GWAS Catalog Summary Statistics PICS CSs and it has valid SuSiE CS from the same region and study
5. It is the GWAS Catalog Curated Association PICS CS and it has valid GWAS Catalog Summary Statistics PICS CS from the same region and study
6. The sum of PIPs in the CS is not within the \[0.95, 1] range
7. The CS didn't pass the study-specific p-value threshold (e.g. GWAS Catalog Curated Association p-value<=1e-8)

* **Why is the V2G score not available anymore?**

The variant-to-gene (V2G) score was developed as an aggregated score for the assignment of the variant to the gene in the region using different sources, such as association with molQTLs and VEP. &#x20;

Although intuitive, V2G does not reflect the complexity of modern data.&#x20;

We have therefore introduced a new score - [Locus-to-Gene score (L2G)](/gentropy/locus-to-gene-l2g), based on a machine learning algorithm that assesses the assignment score of the credible set (not only variant) to gene and does not use the old V2G pipeline anymore. It is a more accurate and sophisticated way of assigning genes to associated GWAS loci. In some cases, the L2G could be interpreted as V2G if the credible set is the size of only one variant. In the new Platform datasets, the L2G score has completely superseded V2G.

Full information about variant assignment to genes based on distance or VEP is available on all [variant ](/variant)pages in the corresponding widgets. There is also information about whether the variant is part of the molQTL credible set, which can also be used to assign this variant to the gene of interest (available in the molQTL widget).&#x20;

* **Why are some L2G associations I found in OT Genetics 22.10 not found in the L2G predictions in OT Platform 25.03?** &#x20;

There could be several reasons for this:

1. We don't have the matching study anymore (due to quality control or validation).
2. We don't have the corresponding credible set anymore due to a different fine-mapping approach or p-value filtering thresholds (e.g. for GWAS Catalog curated associations we now use p-value < 1e-8 instead of p-value < 5e-8 in OTG 22.10).
3. The L2G model is different and can lead to different L2G estimates even when the same study and variant are presented. We don't display L2G assignments when L2G < 0.05.

* **Where can I find studies or credible sets excluded from the Platform?**

To ensure high-quality outputs, the data processing pipelines perform a number of validation steps across different datasets. For data provenance considerations the excluded part of the datasets are also available both on FTP (`ftp://ftp.ebi.ac.uk/pub/databases/opentargets/platform/{release}/excluded`) and Google Cloud (`gs://open-targets-data-releases/{release}/excluded/`). In this folder you'll find the list of excluded credible sets (`credible_set`), evidence (`evidence`), interactions (`interaction`), studies (`study`) and target validation (`target_validation`) dataset. These datasets also provide context about the reason for exclusion.


# Release notes

Summary of release highlights for the Open Targets Platform

## 26.06

#### **Release date**

24 June 2026

### Highlights

#### Data updates

Data updates in this release include:

* New ChEMBL 37
* New PanelApp NHSE Genomic Medicine Service panel
* Latest IMPC data
* Latest GWAS Catalog studies

The GWAS Catalog update adds over 14,000 studies from 111 publications, including 27,689 credible sets and more than 105,000 variants.

* Alignment to EFO 3.88

EFO 3.88 replaces many disease identifiers with Mondo identifiers. The Platform has been updated to reflect these ontology changes.

#### New product features

**New baseline expression**&#x20;

We have completely revamped and expanded our baseline expression dataset with:

* Single-cell RNA sequencing data from [Tabula Sapiens](https://tabula-sapiens.sf.czbiohub.org/)
* Bulk RNA sequencing data from [Genotype-Tissue Expression (GTEx)](https://gtexportal.org/home/) and the [Database of Immune Cells (DICE)](https://dice-database.org/)
* Mass spectrometry proteomics datasets from the [PRoteomics IDEntifications Database (PRIDE)](https://www.ebi.ac.uk/pride/), which was set up as part of an Open Targets project.
* New visualisations

Users can now explore the new baseline expression data through a [redesigned widget](https://staging.platform.opentargets.org/target/ENSG00000133703), at both tissue and cell type level.

**Target prioritisation using baseline expression data**

* The target prioritisation view has also been updated with new expression distribution and specificity scores for tissues and cell types. Please refer to the dedicated [documentation page](https://platform-docs.opentargets.org/web-interface/target-prioritisation) for more info on the target prioritisation assessment method.

**Subcellular location widget**

* The subcellular location widget now displays isoform-specific localisation information from UniProt sources.

**New Drug molecule representations**

* Drug pages now use updated molecular structure images from ChEMBL.

#### Pipeline updates

**Clinical Mining**

* We have improved extraction of drug-disease relationships from AACT clinical trial data using LLM-based methods, increasing accuracy and reducing false-positive association.
* This update has added 4,742 additional clinical reports from AACT, 32,518 drug indication pairs, and 292,325 clinical precedence evidence.

**Locus-to-Gene**

* Two new trans-pQTL features (*transPQtlColocH4Maximum* and  *transPQtlColocH4MaximumNeighbourhood*) have been incorporated into the L2G pipeline, improving causal gene prioritisation by leveraging molecular interaction data.

**Literature**

* Updates to the literature pipeline, including improved ontology mapping, entity disambiguation, and co-occurrence generation, have increased evidence coverage while reducing false positive associations.&#x20;
* Overall, these changes add approximately 1.8 million evidence records and remove around 600,000 direct associations.

#### Technical improvements

**New search now supporting:**

* Clinical trial NCT identifiers
* Additional variant query formats
* Disease/ontology identifiers using either colon or underscore notation

**Associations page URL synchronisation**

* URLs now preserve page state, including filters, scoring settings, pinned targets and open widgets

**Infrastructure changes**&#x20;

* Started creation of a public Helm Chart for deploying a whole Platform in Kubernetes
* New revision mechanics to support data versioning and release management
* Migration of POS to Otter 26
* Created a knowledge base with technical docs on operations, release process
* New POS workflow step for loading data into AWS

### More info and overall metrics

* 78,691 targets
* 47,080 diseases and phenotypes
* 22,407 drugs and compounds
* 42,394,639 evidence strings
* 17,199,165 target-disease associations
* 7,538,243 variants

Check out the 26.06 release [blog post](https://blog.opentargets.org/open-targets-platform-26-06-has-been-released/) for more information on the new features and datasets introduced in this release.

Visit the Open Targets Community [26.06 release post](https://community.opentargets.org/t/26-06-release-now-live/2045) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 26.03

#### Release date

23 March 2026

### Highlights

#### Data updates

* The latest GWAS Catalog data adds 710 new studies from 97 publications, resulting in over 5,000 new credible sets
* This update includes 780 credible sets from the biggest study for hypothyroidism ever ingested in the Platform:  ([Rand SA et al. Nat Genet, 2025](https://www.nature.com/articles/s41588-025-02410-z))
* New data from ClinVar and ClinPGx through the European Variation Archive, as well as updates from String DB and IntAct

#### New product features

* Designed and implemented a new [clinical mining pipeline](https://github.com/opentargets/clinical_mining), which expands the data sources for clinical information,  integrating and annotating data from:
* [ClinicalTrials.gov](http://clinicaltrials.gov) via AACT,
* ChEMBL curated indications,&#x20;
* ChEMBL drug warnings,
* Therapeutic Target Database (TTD),
* European Medical Agency Human Medicines (EMA),&#x20;
* Japan’s Pharmaceuticals and Medical Devices Agency Approvals (PMDA)

The outputs of the new pipeline are presented to users through our re-designed clinical-centric widgets: Indications (Drug pages), Drugs and Clinical Candidates (Target and Disease and Target Prioritisation pages) and Clinical Precedence (new evidence feed in the Associations page)

See the new [clinical reports section](https://platform-docs.opentargets.org/drug/clinical-report) in Platform documentation for more details.

* Added enhancer-gene regulatory predictions from the [ENCODE Project rE2G model](https://www.biorxiv.org/content/10.1101/2023.11.09.563812v1) to the Locus-to-Gene (L2G) framework, introducing new features (`e2gMean` and `e2gNeighbourhoodMean`) that incorporate probabilistic regulatory interactions to improve gene prioritisation
* Enhanced association data with \~98% date coverage to improve novelty estimation, and integrated novelty/time-series calculations directly into the association pipeline by introducing a unified schema with an embedded `timeseries` field (capturing yearly scores, evidence and novelty) to support upcoming novelty features. The association dataset schema has been revised to better reflect this integration - please take a look at our [download page](https://platform.opentargets.org/downloads) for details&#x20;
* Resolved disease–phenotype duplications in the ontology by merging evidence for overlapping terms, removing \~300 duplicates and improving data consistency across the Platform
* Updated LD annotation by correcting the liftover process, fixing coordinate errors that affected finemapping in <1% of credible sets

#### Data access

* **AWS:** Open Targets Platform data is now available on Amazon Web Services (AWS) through the Open Data Program. You can find more information on how to access the Open Targets AWS buckets [here](https://platform-docs.opentargets.org/data-access/platform-datasets-on-aws) or through the 'access data' tabs from our [download page](https://platform.opentargets.org/downloads)
* **Open Targets MCP:** We have released an update to the MCP, with improved biological domain awareness, reducing token usage and enhancing query efficiency

#### Technical  enhancements&#x20;

* Implemented first version of End-to-End testing to improve the stability of the Platform. See [here](https://github.com/opentargets/ot-ui-apps/tree/main/packages/platform-test#readme) for more details
* Migrated all non-search datasets from OpenSearch to ClickHouse, resulting in a \~2× improvement in API query performance
* Resolved minor scoring inconsistencies between the API and web interface by aligning default settings
* Variant page viewer component refactored for reusability
* Fixed a number of FE bugs

Check out the [26.03 release **blog post**](https://blog.opentargets.org/open-targets-platform-26-03-has-been-released) for more information on the new features and datasets introduced in this release.

#### Overall data metrics

* 78,691 targets
* 47,030 diseases and phenotypes
* 22,230 drugs and compounds
* 34,086,838 evidence strings
* 12,466,856 target-disease associations
* 7,432,549 variants

Visit the [**Open Targets Community 26.03** **release thread**](https://community.opentargets.org/t/26-03-release-now-live/1987) for more data metrics for this release, including a per datasource breakdown of evidence strings

## 25.12

### Release date

10th December 2025

### Highlights

#### Data updates

* The latest GWAS Catalog data adds an impressive 78% additional credible sets, most of which are from:

1. UK Biobank Whole-Genome Sequencing Consortium’s [study of 490,640 UK Biobank participants](https://www.nature.com/articles/s41586-025-09272-9)&#x20;
2. Karczewski, Gupta, and Kanai’s [Pan-UK Biobank genome-wide association analyses](https://www.nature.com/articles/s41588-025-02335-7)
3. Zoodsma M, et al. [UK Biobank human metabolite meta analysis](https://www.nature.com/articles/s41588-025-02355-3)

* New update from [CHEMBL 36](https://chembl.blogspot.com/2025/09/chembl-36-is-out.html), including new  molecules, indications and drug warnings
* New data from [Ensembl 115](https://www.ensembl.info/2025/09/02/ensembl-115-has-been-released/), Probes\&Drugs, Reactome, and EVA (through ClinVar)

#### New product features

* New COLOC-PIP colocalisation methodology in Gentropy for the colocalisation of overlapping GWAS-GWAS and GWAS-molQTL credible sets. This method was adapted from [Giambartolomei et al. , 2014](https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1004383)
* Removed all SuSiE fine-mapping credible sets for multi-ancestry GWAS
* Our credible set pages now have a new widget showcasing Enhancer-to-Gene (E2G) predictions from the ENCODE-rE2G model ([Gschwind\*, Mualim\*, Karbalayghareh\*, Sheth\*, Dey\*, Jagoda\*, Nurtdinov\*, and Xi\* et al., bioRxiv](https://www.biorxiv.org/content/10.1101/2023.11.09.563812v1)). Please note that we have also renamed the widget “Enhancer-to-Gene” rather than “Intervals”
* Measurement traits are now filtered out by default from the target association view, with an option to include them back if needed. This new functionality aims to highlight direct links between target and diseases
* Three sources of target-disease evidence were removed from our associations view -  PROGENy, SLAPenrich, and Gene Signatures (SysBio) - as their data has been superseded by the information contained in other sources
* [Download files](https://platform.opentargets.org/downloads) are now split by data sources, with the aim to facilitate investigation of individual evidence
* Our GraphQL API schema documentation was expanded&#x20;

#### Technical  enhancements&#x20;

* Large-scale refactoring of the [orchestration](https://github.com/opentargets/orchestration) of Open Targets Platform pipelines
* Following up from a rewrite in [OnToma](https://github.com/opentargets/OnToma), our Python package for ontology mapping, we have enhanced mapping of disease phenotypes for several evidence sources such as ClinGen, Gene2Phenotype, Orphanet, Genomics England PanelApp, IMPC, Gene Burden, and Pharmacogenetics

Check out the [25.12 release blog post](https://blog.opentargets.org/open-targets-platform-25-12-has-been-released/) for more information on the new features and datasets introduced in this release.

#### Key metrics

| Metric                    | Count      |
| ------------------------- | ---------- |
| Targets                   | 78,725     |
| Diseases/phenotypes       | 46,960     |
| Drugs/clinical candidates | 18,475     |
| Evidence                  | 32,515,132 |
| Associations              | 12,010,760 |

Visit the [Open Targets Community 25.12 release thread](https://community.opentargets.org/t/25-12-release-now-live/1951) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 25.09

### Release date

17 September 2025

### Highlights

* New intervals datasets on variant page: includes over 13 million enhancer-gene regulatory interactions across 352 cell types and tissues from [Gschwind et al.’s](https://www.biorxiv.org/content/10.1101/2023.11.09.563812v1) ENCODE-rE2G model
* Added full list of 95% molQTL credible sets on target page&#x20;
* New version of variant page structural viewer (from new FE component) - now including additional options for users to navigate structure confidence, pathogenicity, domains, secondary structure, residue hydrophobicity&#x20;

### Data updates

* New GWAS Catalog update (+30,000 new credible sets)
* Expression Atlas:  new differential expression dataset, adding around 6,000 new unique target-disease associations to the Platform
* Gene-level synonymous, missense, and loss-of-function constraint scores have been updated from gnomAD v2 to v4
* Updated DepMap data (25Q2) used for the essentiality widget, introducing survival results for 17876 genes for 5 new cell lines
* Added 67 new chemical probes following up from the latest Probes\&Drugs release

### Product enhancements and bug fixes

* 99% of evidence now has date (evidenceDate), with 100% coverage for GWAS credible set derived evidence. This is part of a [recently preprinted](https://www.researchsquare.com/article/rs-5669559/v1) Open Targets project to help assess novelty of disease target associations
* Deployment on new Kubernetes-based cluster infrastructure
* Streamlined L2G pipeline&#x20;
* PharmGKB, a source of pharmacogenetics data, has been rebranded to ClinPGx. This change is now reflected in the Platform
* Bug fixes and improvements spanning across data, back-end and front-end

Check out the [25.09 release blog](https://blog.opentargets.org/open-targets-platform-25-09-release/) post for more information on the new features and datasets introduced in this release

### Overall data metrics

* 78,726 targets
* 39,530 diseases and phenotypes
* 18,119 drugs and compounds
* 30,396,274 evidence strings
* 10,989,518 target-disease associations

Visit the [Open Targets Community 25.09 release thread](https://community.opentargets.org/t/25-09-platform-release-now-live/1929) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 25.06

### Release date

18 June 2025

### Highlights

**New features**

* A new Target Interactors view, which allows you to view association evidence target interactors directly in the disease associations page. Users can choose one of four sources of molecular interactors and view the association evidence for the top scoring interactors for that database.
* A revamped data downloads page, making Open Targets Platform data more FAIR. The data follows the [Croissant](https://mlcommons.org/working-groups/data/croissant/) metadata standard format, based on JSON-LD, developed by ML-Commons.

**Data updates**

* The latest GWAS Catalog data adds 36% more credible sets, most of which are from the [VA Million Veteran Program](https://www.research.va.gov/mvp/) study.
* Burden evidence from the [Broad CVDI Human Disease Portal](https://hugeamp.org:8000/research.html?pageid=600_traits_app_home).
* Experimental Factor Ontology (EFO) replaced measurement terms with terms from the Ontology of Biological Attributes (OBA), and is now reflected in the Platform.
* New data from Reactome, Europe PMC, COSMIC and EVA (through ClinVar).

**Product features**

* Updated version of the Molecular Structure viewer on target profile pages. We also added a new version of the viewer on missense variant pages, which indicates the location of the variant in the AlphaFold model and view AlphaMissense pathogenicity scores.
* Pharmacogenetics widgets on the variant, target, and drug profile pages now have an additional Directionality column.

**Product enhancements and bug fixes**

* Strengthened our search functionality with performance optimisations and additional filtering capabilities.
* Case-case studies have been removed from our GWAS data as these are difficult to map to the correct disease.
* Method descriptions to variant effect widget and bug fixes on variant page.

Check out the [25.06 release blog post](https://blog.opentargets.org/open-targets-platform-25-06-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 78,726 targets
* 38,959 diseases and phenotypes
* 18,081 drugs and compounds
* 29,602,753 evidence strings
* 10,563,905 target-disease associations

Visit the [Open Targets Community 25.06 release thread](https://community.opentargets.org/t/25-06-platform-release-now-live/1828/1) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 25.03

### Release date

19 March 2025

### Highlights

**New features**

* Variant, study, and credible set information is now available in the Open Targets Platform. This unites the Open Targets Platform and Open Targets Genetics into a single interface for human genetic and target discovery information.
* Interpret gene-disease evidence from both common and rare variation in one resource, and in multiple ancestries.
* The Platform now has three additional entities:
  * [Variant](/variant): functional context for 6.5M rare and common variants
    * Please note: the Platform only integrates variants associated with a disease, trait, or phenotype
  * [Study](/study): GWAS and molQTL studies
  * [Credible Set](/credible-set): 2.6 million credible sets derived from various sources
    * Colocalisation is now based on credible set overlaps
* A new [Locus-to-Gene (L2G)](/gentropy/locus-to-gene-l2g) machine learning model which prioritises likely causal genes at each GWAS locus by using functional genomics features. The Platform also uses [Shapley values](/credible-set#explaining-l2g-predictions) as part of the L2G predictions to illustrate the relative contribution of each feature.

**Data updates**

* A substantial increase in the literature evidences due to improvements in resolving disambiguation of entities.
* Updated gene burden data through FinnGen R12.
* New NHS Genomic Medicine Service panels from GEL PanelApp and a new hearing loss panel to the Gene2Phenotype evidence set.
* Rewritten Uniprot pipeline with new associations from Uniprot variants
* Updated data from Probes\&Drugs and DepMap.
* New data from Reactome, ChEMBL, Europe PMC, COSMIC and EVA (through ClinVar).

**Product features**

* A new Scalable and reproducible genetic analyses pipeline available as a Python package for post-GWAS analysis: [Gentropy](https://opentargets.github.io/gentropy/).
* An updated data downloads page which has a more detailed description of each file. (Temporary removal of schema which will be brought back in the subsequent release).
* [otter](https://github.com/opentargets/otter) - **O**pen **T**argets' **T**ask **E**xecuto**R** i.e. scripts that process and prepare data for our ETL pipelines.
* Various improvements to our web interface:
  * Users can search the UI using variants and study id.
  * Improved searching, filtering and sorting of entities on our associations pages. In particular, there are now separate sections for uploaded entity lists and pinned entities, and you can remove individual filters from the view.
  * The platform and the Target Prioritisation view has a more accessible colour scheme.\
    Option to select desired columns in the UI.
  * A graphical Comparative Genomics view in Target Prioritisation.
  * Preview on hover: Ability to view details of an entity without navigating to it.

**Product enhancements and bug fixes**

* Filtered out Phase IV clinical trial evidence that lacks regulatory approval for the specific indication from our target-disease association data.
* The BE infrastructure has been upgraded to Scala 3
* Improved and restructured documentation.

{% hint style="info" %}
Post 25.03, the data downloads paths have changed as now only parquet file format is available. Also there are minor changes to the name of the dataset (snake\_case & singular). More details can be found [here](https://community.opentargets.org/t/issues-with-gcp-big-query/1704/5).

* Previous releases (till 24.09):
  * `https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/24.09/output/etl/parquet/associationByOverallDirect/`
* 25.03 release (and therafter):&#x20;
  * `https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/25.03/output/association_by_datasource_direct/`
    {% endhint %}

Check out the [25.03 release blog post](https://blog.opentargets.org/open-targets-platform-25-03-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 78,766 targets
* 28,513 diseases and phenotypes
* 18,081 drugs and compounds
* 28,168,992 evidence strings
* 10,162,821 target-disease associations
* 6,493,882 variants

Visit the [Open Targets Community 25.03 release thread](https://community.opentargets.org/t/25-03-platform-release-now-live-open-targets-genetics-data-update/1708) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 24.09

### Release date

18 September 2024

### Highlights

**New features**

* Users now have the ability to apply target-specific or disease/phenotype-specific filters to the target-disease association and target prioritisation pages. Read more details about the feature in our [documentation](https://platform-docs.opentargets.org/web-interface/associations-on-the-fly#filtering-functionality).

**Data updates**

* New safety liabilities associated with several targets which are routinely used by pharmaceutical companies have been added, from Brennan et al. (2024).
* New Gene Burden evidence from the LoF burden analyses in FinnGen’s latest public release (R11).
* New data from Reactome, Europe PMC, COSMIC and EVA (through ClinVar).

**Product features**

* Improved visualisation for Gene essentiality data from Cancer DepMap.
* Updated frontend table design and functionality leading to better searching/filtering and sorting functionality for most tables in the UI.
* The classic associations view has been deprecated from the Platform..

**Product enhancements and bug fixes**

* OpenAI model in the literature summarisation tool was updated to GPT-4o-mini.
* Changes to `variant` field in the cancer biomarker evidence.
* Aggregated the granularity of the description of phenotypes inside `cohortPhenotypes` to resolve duplication in ChEMBL evidence.
* Bug fixes: pharmacogenetics schema, molecule dataset.

Check out the [24.09 release blog post](https://blog.opentargets.org/open-targets-platform-24-09-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 63,121 targets
* 28,327 diseases and phenotypes
* 18,041 drugs and compounds
* 17,853,184 evidence strings
* 8,155,988 target-disease associations

Visit the [Open Targets Community 24.09 release thread](https://community.opentargets.org/t/24-09-platform-release-now-live/1556) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 24.06

### Release date

19 June 2024

### Highlights

**New features**

* Uploading a custom list of targets or diseases to obtain a tailored associations view allowing users to view selected target-disease evidence for the specific entities.

**Data updates**

* Integration of ChEMBL 34 which now contains data from the European Medicines Agency (EMA) increasing drug-indication coverage.
* Updates to the clinical trials and tractability data.
* Updating the gene burden results from AstraZeneca’s PheWAS portal [version 5](https://azphewas.com/about).
* New gene burden evidence for schizophrenia from the SCHEMA consortium and ancestry-specific evidence for prostate cancer.
* A new Gene2Phenotype panel with musculo-skeletal implications.
* New data from Reactome, COSMIC and EVA (through ClinVar), Probes and Drugs and GEL PanelApp increasing our coverage of data.

**Product features**

* Ability to handle pharmacogenetic evidence involving drug combinations.

**Product enhancements and bug fixes**

* Exclusion of splice QTLs from the assessment for the direction of effect and selecting the beta from the evidence with the lowest p-value instead of the largest effect size.
* Improvements in the AotF GQL Playground Component.
* Dropping `isHumanApplicable` field from target safety.
* Bug fixes to the ClinVar (somatic) widget loading state.
* Resolving an error in phenotype mapping in GEL PanelApp, resolving incorrect tractability precedence and removing categorical burden tests from Genebass data based on community feedback.

Check out the [24.06 release blog post](https://blog.opentargets.org/open-targets-platform-24-06-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 63,226 targets
* 28,198 diseases and phenotypes
* 18,041 drugs and compounds
* 17,703,456 evidence strings
* 8,079,215 target-disease associations

Visit the [Open Targets Community 24.06 release thread](https://community.opentargets.org/t/24-06-platform-release-now-live/1455) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 24.03

### Release date

20 March 2024

### Highlights

**New features**

* Implementation of direction of effect assessment for eight different sources of target-disease association evidence.
* Filtering bibliography data based on a publication date.

**Data updates**

* Integration of the latest dataset from Project Score.
* Integration of the 23Q4 version of DepMap ([depmap.org](https://depmap.org)).
* New data from Reactome, COSMIC and EVA (through ClinVar) increasing our coverage of data.

**Product features**

* Inclusion of star alleles from PharmGKB and a new *Direct Drug Target* column to the pharmacogenetics widget.
* Use of pharmacogenetics data to inform adverse drug response as an additional source of information on target safety.
* New dedicated GraphQL API query playground for Associations-on-the-Fly and target prioritisation view.

**Product enhancements and bug fixes**

* Redesigned context menu with a new navigation and pinning behaviour.
* Updated protvista-uniprot viewer library to v2.11.1.
* Bug fixes to download file schema and search issues.

Check out the [24.03 release blog post](https://blog.opentargets.org/open-targets-platform-24-03-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 63,226 targets
* 25,817 diseases and phenotypes
* 17,111 drugs and compounds
* 17,317,290 evidence strings
* 7,802,260 target-disease associations

Visit the [Open Targets Community 24.03 release thread](https://community.opentargets.org/t/24-03-platform-release-now-live/1374) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 23.12

### Release date

30 November 2023

### Highlights

**New features**

* 'Target Prioritisation' view: A new view for assessment of the target features considered when prioritising (or deprioritising) targets for drug discovery. Watch a detailed video [here](https://www.youtube.com/watch?v=WQwQn6I4jkw).
* A new widget in the Platform adding pharmacogenetics data from PharmGKB.

**Data updates**

* New data available on Baseline RNA and protein expression data for targets via the API and FTP.
* New data from Reactome and EVA (through ClinVar) increasing our coverage of data.

**Product features**

* Users can easily export the entire associations table and the target prioritisation table in json or tsv format.
* Updated search bar design with search suggestions.

**Product enhancements and bug fixes**

* Transition to OpenSearch from Elasticsearch.
* Bug fixes in the widgets in the Association On The Fly view and styling issues.

Check out the [23.12 release blog post](https://blog.opentargets.org/open-targets-platform-23-12-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 62,733 targets
* 25,246 diseases and phenotypes
* 17,095 drugs and compounds
* 16,710,896 evidence strings
* 7,994,180 target-disease associations

Visit the [Open Targets Community 23.12 release thread](https://community.opentargets.org/t/23-12-platform-release-now-live/1294) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 23.09

### Release date

21 September 2023

### Highlights

**New features**

* ‘Associations on the Fly’ - revamp of the current Open Targets Platform association page with new facets and additional built-in functionalities like view data directly in the associations table, control weights of contributing evidence, filter by datasource and data type (OR filters) and pin rows. Watch a detailed video [here](https://www.youtube.com/watch?v=2A9bksboAag).
* OpenAI Literature Summarisation tool - For data features that link to publications, users can ask for a natural language summary of the target-disease evidence presented in the publication using LangChain and OpenAI’s GPT3.5 Turbo model.

**Data updates**

* Updated Molecular Interactions data source [STRING Database](https://string-db.org/) to version 12.0
* Increase in Europe PMC literature evidence by 9.9% to 10,355,423
* New data from ChEMBL, COSMIC and EVA (through ClinVar) increasing our coverage of data

**Product features**

* Easy access to the schema of the files available for download in the Open Targets Platform

**Product enhancements and bug fixes**

* Expanded the definition of a drug to include all probes as reported by [Probes & Drugs Portal (P\&D)](https://www.probes-drugs.org/home/) as chemical probes are useful from a target's doability perspective
* Open Targets Platform user interface migration to Material UI v5
* Refactoring of the sections in the frontend codebase - Components and sections moved into the packages/sections and packages/ui

Check out the [23.09 release blog post](https://blog.opentargets.org/open-targets-platform-23-09-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 62,733 targets
* 25,209 diseases and phenotypes
* 17,096 drugs and compounds
* 16,232,046 evidence strings
* 7,922,844 target-disease associations

Visit the [Open Targets Community 23.09 release thread](https://community.opentargets.org/t/the-latest-release-22-09-is-now-live/1212) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 23.06

### Release date

26 June 2023

### Highlights

**New features**

* Addition of a CRISPR Screens widget featuring data from [CRISPRBrain](https://crisprbrain.org/)
* Introduction of a Cancer DepMap widget showcasing gene essentiality data from the [Cancer DepMap Portal](https://depmap.org/portal/)

**Data updates**

* Updated data from ChEMBL, including adverse event drug warning data and more granular information on clinical phases
* New data from IntoGEN, Europe PMC, and EVA (through ClinVar) increasing our coverage of data

**Product features**

* Missense variants in the OT Genetics, UniProt variants and ClinVar widgets now link to [ProtVar](https://www.ebi.ac.uk/ProtVar/), a new tool to interpret the functional consequences of human missense variants

**Product enhancements and bug fixes**

* Fixes - homology widget, fixes to the data
* More meaningful 404 error message
* Fixed bugs in the API Playground

Check out the [23.06 release blog post](https://blog.opentargets.org/open-targets-platform-23-06-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 62,685 targets
* 24,713 diseases and phenotypes
* 13,210 drugs and compounds
* 15,117,741 evidence strings
* 7,835,247 target-disease associations

Visit the [Open Targets Community 23.06 release thread](https://community.opentargets.org/t/23-06-platform-release-now-live/1125) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 23.02

### Release date

22 February 2023

### Highlights

In addition to regular updates from our data providers, we have a number of new features in this release:

**New evidence for target-disease associations**

* Additional data for metabolic biomarkers added to our Gene Burden widget
* QTL-based direction of effect included in evidence from Open Targets Genetics

**Improved target annotation data**

* Integration of Target safety evidence from AOPWiki
* New data from Probes and Drugs’ 04.2022 release

**Literature updates**

* Preprints and patents now included in our bibliography

**Development updates**

* Redesigned search
* Provenance metadata

Check out the [23.02 release blog post](https://blog.opentargets.org/open-targets-platform-23-02-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 62,678 targets
* 24,713 diseases and phenotypes
* 12,854 drugs and compounds
* 10,446,771 evidence strings
* 6,656,559 target-disease associations

Visit the [Open Targets Community 23.02 release thread](https://community.opentargets.org/t/23-02-platform-release-now-live/962) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 22.11

### Release date

24 November 2022

### Highlights

In addition to continuous updates from our data providers, we have introduced the following new features:&#x20;

* Gene burden data for Parkinson’s disease
* Updated classifications for clinical trial stop reasons
* Variant functional consequences, available to browse in the Gene2Phenotype and Orphanet widgets
* Other improvements and bug fixes

Check out the [22.11 release blog post](https://blog.opentargets.org/open-targets-platform-22-11-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 62,678 targets
* 22,274 diseases and phenotypes
* 12,854 drugs and compounds
* 14,611,717 evidence strings
* 6,960,486 target-disease associations

Visit the [Open Targets Community 22.11 release thread](https://community.opentargets.org/t/22-11-platform-release-now-live/870) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 22.09

### Release date

29 September 2022

### Highlights

New data, in particular:

* Open Targets Genetics
* Genomics England PanelApp
* Gene burden
* Probes and drugs
* New data integrity file, in line with FAIR principles

Check out the[ 22.09 release blog post](https://blog.opentargets.org/open-targets-platform-22-09-release/) for more information on the new features and datasets introduced in this release.<br>

### Overall data metrics

* 61,888 targets
* 20,931 diseases and phenotypes
* 12,854 drugs and compounds
* 14,229,684 evidence strings
* 7,003,171 target-disease associations

Visit the [Open Targets Community 22.09 release thread](https://community.opentargets.org/t/22-09-platform-release-now-live/783) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 22.06

### Release date

24 June 2022

### Highlights

* New data: five additional gene burden analyses from Genebass
* New feature: new visualisation of subcellular locations of targets now available to users
* New ontology term: “medical procedure”&#x20;

Check out the [22.06 release blog post](https://blog.opentargets.org/open-targets-platform-22-06-release/) for more information on the new features and datasets introduced in this release.&#x20;

### Overall data metrics

* 61,524 targets
* 23,074 diseases and phenotypes
* 12,854 drugs and compounds
* 14,455,104 evidence strings
* 7,247,865 target-disease associations

Visit the [Open Targets Community 22.06 release thread](https://community.opentargets.org/t/the-latest-release-22-06-is-now-live/675) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 22.04

### Release date

28 April 2022

### Highlights

* New datasource: gene burden analyses from Regeneron and the AstraZeneca
* Integration of structural variants from ClinVar
* Additional information from DailyMed drug label text-mining
* NLP classification of why clinical trials stopped
* New data: Gene2phenotype cardiac panel

Check out the[ 22.04 release blog post](https://blog.opentargets.org/open-targets-platform-22-04-release/) for more information on the new features and datasets introduced in this release.&#x20;

### Overall data metrics

* 61,524 targets
* 18,520 diseases and phenotypes
* 12,854 drugs and compounds
* 13,829,174 evidence strings
* 7,541,360 target-disease associations

Visit the [Open Targets Community 22.04 release thread](https://community.opentargets.org/t/open-targets-platform-22-04-is-out-now/555) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 22.02

### Release date

28 February 2022

### Highlights

* Gene2Phenotype terminology updated in line with the Gene Curation Coalition (GenCC)
* Data updates from a range of providers including Open Targets Genetics and ChEMBL

Check out the [22.02 release blog post](https://blog.opentargets.org/open-targets-platform-22-02-release/) for more information on the new features and datasets introduced in this release.&#x20;

### Overall data metrics

* 61,524 targets
* 18,468 diseases and phenotypes
* 12,594 drugs and compounds
* 10,880,832 evidence strings
* 7,980,448 target-disease associations

Visit the [Open Targets Community 22.02 release thread](https://community.opentargets.org/t/open-targets-platform-22-02-has-been-released/476) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 21.11

### Release date

29 November 2021

### Highlights

* New Cancer Biomarkers evidence data from the Cancer Genome Interpreter
* Updated genetic association evidence from Open Targets Genetics
* Embedded GraphQL API playground for each data table and query

Check out the [21.11 release blog post](https://blog.opentargets.org/open-targets-platform-21-11-release/) for more information on the new features and datasets introduced in this release.&#x20;

### Overall data metrics

* 60,636 targets
* 18,706 diseases and phenotypes
* 12,594 drugs and compounds
* 10,481,189 evidence strings
* 7,787,231 target-disease associations

Visit the[ Open Targets Community 21.11 release thread](https://community.opentargets.org/t/open-targets-platform-21-11-has-been-released/413) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 21.09

### Release date

30 September 2021

### Highlights

* Integration of new PROTAC tractability data from [Schneider et al. (2021)](https://doi.org/10.1038/s41573-021-00245-x)
* Integration of Genetic Constraint data from gnomAD and new Chemical Probes data from Probes & Drugs database
* Data updates from EFO, ChEMBL, and Mouse Genome Informatics
* Other improvements and bug fixes

Check out our [21.09 release blog post](https://blog.opentargets.org/open-targets-platform-21-09-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 60,636 targets
* 18,663 diseases and phenotypes
* 12,594 drugs and compounds
* 11,071,233 evidence strings
* 7,927,820 target-disease associations

Visit the [Open Targets Community 21.09 release thread](https://community.opentargets.org/t/21-09-platform-release-now-live/364) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## 21.06

### Release date

30 June 2021

### Highlights

* Updated Open Targets Genetics Portal evidence, which included the integration of FinnGen biobank data (R5) and new GWAS Catalog studies
* Integration of gene-disease data from Orphanet
* Improvements to the user interface (e.g. datatype chips on evidence page)
* Bug fixes (e.g. users can download Known Drugs table, association scores in datasets match values returned by API)

Check out our [21.06 release blog post](https://blog.opentargets.org/open-targets-platform-21-06-release/) for more information on the new features and datasets introduced in this release.

### Overall data metrics

* 60,606 targets
* 18,507 diseases and phenotypes
* 13,185 drugs and compounds
* 13,267,236 evidence strings
* 9,216,710 target-disease associations

Visit the [Open Targets Community](https://community.opentargets.org/t/21-06-platform-release-now-live/244) for more data metrics for this release, including a per datasource breakdown of evidence strings.

## Archive

For release notes for previous releases, check out the [Open Targets Community News & Announcement section](https://community.opentargets.org/c/news-and-announcements/5).


# Citation

## Latest publication

If you use the Open Targets Platform in your work, please cite our latest publication:

Buniello, A. et al. (2025). [Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery](https://academic.oup.com/nar/article/53/D1/D1467/7917960). *Nucleic Acids Research.*

## Previous publications

We have additional publications about the Open Targets Platform:

* Ochoa, D. et al. (2023). [The next-generation Open Targets Platform: reimagined, redesigned, rebuilt.](https://academic.oup.com/nar/advance-article/doi/10.1093/nar/gkac1046/6833237) *Nucleic Acids Research.*
* Ochoa, D. et al. (2021). [Open Targets Platform: supporting systematic drug–target identification and prioritisation](https://doi.org/10.1093/nar/gkaa1027). *Nucleic Acids Research.*
* Carvalho-Silva, D. et al. (2019). [Open Targets Platform: new developments and updates two years on](https://doi.org/10.1093/nar/gky1133). *Nucleic Acids Research*.
* Koscielny, G. et al. (2017). [Open Targets: a platform for therapeutic target identification and validation](https://doi.org/10.1093/nar/gkw1055). *Nucleic Acids Research.*

You can also refer to our publications about Open Targets Genetics and our Locus-to-Gene (L2G) pipeline:

* Ghoussaini, M., et al. (2021) [Open Targets Genetics: systematic identification of trait-associated genes using large-scale genetics and functional genomics](https://doi.org/10.1093/nar/gkaa840). *Nucleic Acids Research*.
* Mountjoy, E., et al. (2021) [An open approach to systematically prioritize causal variants and genes at all published human GWAS trait-associated loci](https://doi.org/10.1038/s41588-021-00945-5). *Nature Genetics.*


# Licence

As one of Open Targets flagship informatics products, the team that maintains the Open Targets Platform is committed to building open source tools and supporting open access research.

If you use our code and/or our data, please cite [our latest publication](/citation#latest-publication).

## **Data**

Open Targets Platform is marked with [CC0 1.0](http://creativecommons.org/publicdomain/zero/1.0?ref=chooser-v1). This dedicates the data to the public domain, allowing downstream users to consume the data without restriction.

The Platform also conforms to the EBI long-term data preservation [policies](https://www.ebi.ac.uk/long-term-data-preservation).

Please contact us on [the Open Targets Community](https://community.opentargets.org/) if you have questions about our licence or using our data and codebases.

## **Codebases**

The codebases that power the Platform - including our pipelines, GraphQL API, and React UI - are all open source and licensed under the [APACHE LICENSE, VERSION 2.0](https://www.apache.org/licenses/LICENSE-2.0).

You can find all of our code repositories on GitHub at <https://github.com/opentargets>.

## Data sources licensing

See list below if you are interested in knowing more about our data sources and their licensing status.

Please note that all data sources listed in the table, including the ones labeled with 'commercial use for Open Targets', have agreed for their data to be used without restriction by all Open Targets users.

| Data source                                                                                                     | Licence                                                                                                   |
| --------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
| [AlphaMissense](https://github.com/google-deepmind/alphamissense)                                               | CC-BY 4.0                                                                                                 |
| [Cancer Gene Census](https://cancer.sanger.ac.uk/census)                                                        | Commercial use for Open Targets                                                                           |
| [Cancer Genome Interpreter](https://www.cancergenomeinterpreter.org/biomarkers)                                 | CC-BY 4.0                                                                                                 |
| [Cell Ontology](https://obofoundry.org/ontology/cl.html)                                                        | CC BY 4.0                                                                                                 |
| [ChEMBL](https://www.ebi.ac.uk/chembl/)                                                                         | CC BY-SA 3.0                                                                                              |
| [ClinGen](https://clinicalgenome.org/)                                                                          | CC0 1.0                                                                                                   |
| [ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/) (via [EVA](https://www.ebi.ac.uk/eva/))                        | [EMBL-EBI terms of use ](https://www.ebi.ac.uk/about/terms-of-use/)                                       |
| [EFO](https://www.ebi.ac.uk/efo/)                                                                               | Apache 2.0                                                                                                |
| [Ensembl](https://www.ensembl.org/index.html)                                                                   | CC-BY 4.0                                                                                                 |
| [Europe PMC](http://europepmc.org/)                                                                             | CC-BY, CC-BY-NC or CC0 license or open access (depending on publication licenses)                         |
| [eQTL Catalogue](https://www.ebi.ac.uk/eqtl/)                                                                   | CC-BY 4.0                                                                                                 |
| [Expression Atlas](https://www.ebi.ac.uk/gxa/home)                                                              | CC-BY 4.0                                                                                                 |
| [FAERS](https://www.fda.gov/drugs/surveillance/questions-and-answers-fdas-adverse-event-reporting-system-faers) | CC0 1.0                                                                                                   |
| [FinnGen R12](https://www.finngen.fi/en)                                                                        | Summary statistics can be requested [here](https://elomake.helsinki.fi/lomakkeet/124935/lomake.html).     |
| [Gene Ontology](http://geneontology.org/)                                                                       | CC-BY 4.0                                                                                                 |
| [Gene Signature](https://www.hsls.pitt.edu/obrc/index.php?page=URL1268854187)                                   | Apache 2.0                                                                                                |
| [Gene2Phenotype](https://www.ebi.ac.uk/gene2phenotype)                                                          | [EMBL-EBI terms of use](https://www.ebi.ac.uk/about/terms-of-use/)                                        |
| [Genomics England PanelApp](https://panelapp.genomicsengland.co.uk/)                                            | Commercial use for Open Targets                                                                           |
| [gnomAD](https://gnomad.broadinstitute.org/)                                                                    | CC0 1.0                                                                                                   |
| [GTEx](https://www.gtexportal.org)                                                                              | CC-BY 4.0                                                                                                 |
| [GWAS Catalog](https://www.ebi.ac.uk/gwas/home)                                                                 | [EMBL-EBI terms of use ](https://www.ebi.ac.uk/about/terms-of-use/)(Summary statistics are under CC0 1.0) |
| [HeCaTos](https://cordis.europa.eu/project/id/602156/results)                                                   | CC-BY 4.0                                                                                                 |
| [HPO](https://hpo.jax.org/app/)                                                                                 | CC0 1.0                                                                                                   |
| [Human Protein Atlas](http://www.proteinatlas.org/)                                                             | CC BY-SA 3.0                                                                                              |
| [Intact](https://www.ebi.ac.uk/intact/home)                                                                     | [EMBL-EBI terms of use](https://www.ebi.ac.uk/about/terms-of-use/)                                        |
| [IntOGen](http://www.intogen.org/search)                                                                        | CC0 1.0                                                                                                   |
| [MGI](http://www.informatics.jax.org/phenotypes.shtml)                                                          | CC-BY 4.0                                                                                                 |
| [MONDO](https://mondo.monarchinitiative.org)                                                                    | CC-BY 4.0                                                                                                 |
| [Orphanet](https://www.orpha.net/consor/cgi-bin/index.php)                                                      | CC-BY 4.0                                                                                                 |
| [panUKB](https://pan.ukbb.broadinstitute.org/)                                                                  | CC BY 4.0                                                                                                 |
| [PhenoDigm](https://www.sanger.ac.uk/tool/phenodigm/)                                                           | CC-BY 4.0                                                                                                 |
| [Probes and Drugs](https://www.probes-drugs.org/home/)                                                          | CC-BY 4.0                                                                                                 |
| [PROGENy](https://saezlab.github.io/progeny/)                                                                   | Apache 2.0                                                                                                |
| [Project Score](https://score.depmap.sanger.ac.uk/)                                                             | Commercial use for Open Targets                                                                           |
| [Reactome](https://reactome.org/)                                                                               | CC-BY 4.0                                                                                                 |
| [Sequence Ontology](http://www.sequenceontology.org/)                                                           | CC-BY 4.0                                                                                                 |
| [Signor](https://signor.uniroma2.it/) (via [Intact](https://www.ebi.ac.uk/intact/home))                         | [EMBL-EBI terms of use](https://www.ebi.ac.uk/about/terms-of-use/)                                        |
| [SLAPenrich](https://saezlab.github.io/SLAPenrich/)                                                             | MIT terms of use                                                                                          |
| [String DB](https://string-db.org/)                                                                             | CC-BY 4.0                                                                                                 |
| [TEP](https://www.thesgc.org/tep)                                                                               | CC-BY 4.0                                                                                                 |
| [Tox21](https://tox21.gov/overview/about-tox21/)                                                                | CC0 1.0                                                                                                   |
| [Uberon](https://obophenotype.github.io/uberon/)                                                                | CC BY 3.0                                                                                                 |
| [UKB-PPP](https://www.synapse.org/Synapse:syn51364943/wiki/622119)                                              | Apache 2.0                                                                                                |
| [UniProt](https://www.uniprot.org/)                                                                             | CC-BY 4.0                                                                                                 |


# Terms of use

Terms of use for the Open Targets Platform - updated June 2021

## **General**

1. These Terms of Use reflect Open Targets' objective to develop, implement and rapidly disseminate to the wider scientific community, new informatics tools, experimental methods, platforms and associated data related to target validation. They impose no additional constraints on the use of the contributed data than those provided by the data owner.
2. Open Targets expects attribution (e.g. in publications, services or products) for any of its online services, databases or software in accordance with good scientific practice. The expected attribution will be indicated on the appropriate web page.
3. Any feedback provided to Open Targets on the Open Targets Platform will be treated as non-confidential unless the individual or organisation providing the feedback states otherwise.
4. Open Targets is not liable to you or third parties claiming through you, for any loss or damage.
5. Personal data will only be released in exceptional circumstances when required by law or judicial or regulatory order.
6. We reserve the right to update these Terms of Use at any time. When alterations are inevitable, we will attempt to give reasonable notice of any changes by placing a notice on our website, but you may wish to check each time you use the website. The date of the most recent revision will appear on this page. If you do not agree to these changes, please do not continue to use our online services. We will also make available an archived copy of the previous Terms of Use for comparison.
7. Any questions or comments concerning these Terms of Use can be addressed to: The Operations Director, Open Targets, Wellcome Genome Campus, Hinxton CB10 1SD
8. Users of the Open Targets Platform agree not to attempt to use any Open Targets computers, files or networks apart from through the service interfaces provided.
9. Open Targets will make all reasonable effort to maintain continuity of the Open Targets Platform and provide adequate warning of any changes or discontinuities. However, Open Targets accepts no responsibility for the consequences of any temporary or permanent discontinuity in service.
10. Any attempt to use the Open Targets Platform to a level that prevents, or looks likely to prevent, Open Targets providing services to others, will result in the use being blocked.
11. Software that can be run from the Open Targets webpages may be used by any individual for any purpose unless specific exceptions are stated on the web page.
12. Open Targets does not accept responsibility for the consequences of any breach of the confidentiality of Open Targets site by third parties.

## **Data Services**

1. The online data services and databases of Open Targets are generated in part from data contributed by the community who remain the data owners.
2. Open Targets itself places no additional restrictions on the use or redistribution of the data available via its online services other than those provided by the original data owners.
3. Open Targets does not guarantee the accuracy of any provided data, generated database, software or online service nor the suitability of databases, software and online services for any purpose.
4. The original data may be subject to rights claimed by third parties, including but not limited to, patent, copyright, other intellectual property rights, biodiversity-related access and benefit-sharing rights. It is the responsibility of users of the Open Targets Platform to ensure that their exploitation of the data does not infringe any of the rights of such third parties.

## **Privacy**

Please review our updated [Privacy Notice](https://www.ebi.ac.uk/data-protection/privacy-notice/embl-ebi-public-website/).

{% file src="/files/izkeJiMGFOZS5bLgDGcV" %}


# Partner Preview Platform

The [Open Targets Partner Preview Platform](https://partner-platform.opentargets.org/) (PPP) is provided exclusively to Open Targets consortium members.&#x20;

All data and results of queries must remain confidential and must not be shared publicly. Please note that data from OTAR projects is pre-publication, being actively worked on by projects teams and therefore subject to change through further analysis — our release notes contain details of any known issues with data sets.&#x20;

All pre-publication data incorporated into the PPP will be publicly released by the project teams at the time of publication.

For more information about the PPP, please read the [PPP specific documentation.](http://home.opentargets.org/ppp-documentation)


