It’s no secret that data scientists and researchers spend 80% of their time on the less glamorous tasks of chasing down data, cleaning it up, and making sure it’s not full of nonsense. Researchers in the biomedical domain are facing these challenges as only the first step of a long process. One of the most prominent examples of such work is the drug discovery and development process. It typically spans 10-15 years, including 4-7 years dedicated to identifying and validating a target.
During the target identification phase of drug development, several challenges related to data can impede progress. For example, incomplete, inconsistent, and erroneous data, along with heterogeneous sources and semantic inconsistencies complicate the integration and interpretation of biological datasets. Also, access barriers to proprietary data or data siloes further hinder collaborative efforts. The sheer volume and complexity of biological data, coupled with biases, noise, and annotation gaps pose additional hurdles.
In the rapidly evolving field of biomedical research, having immediate access to comprehensive and integrated data is a game-changer. Imagine a single framework where researchers, data scientists, and Healthcare professionals can find all the information they need, interconnected and ready for analysis. This vision is becoming a reality.
Ontotext’s LinkedLifeData Inventory
Ontotext’s LinkedLifeData (LLD) Inventory is a FAIR data-centric solution that seamlessly integrates a growing collection of interconnected datasets and ontologies in the biomedical domain. These include UMLS, ChEMBL, UniProt, SemMedDB, ClinicalTrials.gov, NCBI Gene, DrugCentral as well as niche ones like OpenTargets, EBI Expression Atlas, MarkerDB, and Human Protein Atlas.
LLD Inventory is a semantic data fabric integrating data from disparate sources into knowledge graphs. These enable researchers, engineers, and data scientists to explore links between different biomedical concepts, discover new insights, and build innovative applications in areas such as drug discovery, personalized medicine, and biomedical research.
Our approach to semantic integration
Building on the foundation of LLD Inventory, Ontotext employs a sophisticated strategy to integrate and manage these diverse datasets. Here’s how this approach is executed in practice.
An interdisciplinary team of biologists, bioinformaticians, and data engineers develops a semantic model and the automation components necessary for producing RDF serializations.
Creating a semantic model involves identifying key concepts and relationships within the dataset and defining them in a structured manner. This process requires deep domain expertise to accurately capture the complexities of biological and medical information.
The automation includes several critical steps. First, data is retrieved from different sources in various formats. Then, a pipeline is designed to systematically processing the original data, converting it into an RDF format, and ensuring that the resulting data is accurate, consistent, and ready for use. This automated pipeline enables continuous updates and integration of new data, making sure that the semantic model remains current and comprehensive.
Where not directly encoded, semantic mappings are added to connect concepts across datasets. This involves aligning different terminologies, classifications, and data structures so that information from diverse sources can be precisely and meaningfully integrated.
For example, consider mapping the gene identifiers in the NCBI Gene database with protein identifiers in UniProt. If NCBI Gene uses a specific gene symbol for a human gene, while UniProt uses a different identifier for the protein encoded by that gene, a semantic mapping would link these identifiers. This allows researchers to connect genetic information from NCBI Gene with protein data from UniProt, facilitating a more holistic understanding of gene-protein interactions.
Building the knowledge graph
The LLD Inventory team follows rigorous standards to generate metadata, which describes the data’s content, context, and structure. Metadata is crucial for data discovery, understanding, and management. The generated metadata is then published in a data catalogue – a centralized repository that provides detailed information about the available datasets.
Data and metadata are synchronized in Ontotext GraphDB using dedicated tooling. With just a click, we can subscribe to the dataset we want and start querying it right away. Updating a dataset when a new version becomes available is also very easy.
LLD Inventory helps create a comprehensive knowledge graph used by Ontotext’s Target Discovery solution. Knowledge graphs enhance research and decision-making in Life Sciences and Healthcare by integrating diverse data sources into a unified format. They map relationships between entities such as genes, diseases, and drugs, revealing hidden connections and patterns that support hypothesis generation and advanced querying.
By enriching data with contextual information, knowledge graphs facilitate deeper insights and inference of new relations. They also aid in data quality and standardization, provide valuable resources for analytical processing and machine learning, and promote interdisciplinary collaboration. All this ultimately leads to more informed decision-making and innovative discoveries.
Empowering Target Discovery with LLD Inventory
To illustrate the utility of LLD Inventory for research, we picked the case of drug repurposing. This type of drug development process aims to identify new therapeutic uses for already existing drugs. In this way it aims to reduce the time and cost of finding a new treatment as the compound has already been proven safe.
One way to approach drug repurposing is to investigate possible applications of a selected drug and formulate a hypothesis for a therapeutic effect downstream or upstream of a known target. We will demonstrate how this can be quickly and easily achieved with the help of interconnected data provided by LLD Inventory. This data will be consumed through Target Discovery, facilitating the streamlined discovery process and analytics.
We present an example of the drug Alectinib – an anticancer medication approved for the treatment of non-small-cell lung cancer (NSCLC). Our goal is to find potential new applications for the drug and validate our approach by also showing that we can identify its primary indication (NSCLC in this case).
Accessing well-integrated data with LLD Inventory
Investigating such a use case requires gathering information from various databases to get a 360° view of the drug, its mechanism of action (MoA), target interactions, pathological relevance and importantly, and up-to-date data from clinical trials and publications. In Target Discovery, all of this data can be integrated through LLD Inventory, enabling researchers to explore well-known facts. They can also uncover novel relationships and utilize advanced analytics on top of any combination of datasets in an intuitive way.
Let’s put ourselves in the shoes of the scientist. As the start, we will find general information about the drug Alectinib in the DrugCentral database that holds data such as its trade name, synonyms, and adverse events. In another data source, ChEMBL, we can find the known targets and the molecular MoA of the drug. LLD Inventory allows us to get this information at once due to the mappings between DrugCentral and ChEMBL that it provides.
Getting an overview of the drug and its action
We can easily find all of this information in Target Discovery on a single dashboard. This is possible because LLD Inventory normalizes and interconnects disparate data sources by establishing links between the various identifiers of the same entity. As a result, researchers can search in Target Discovery for Alectinib, Alecensa, or any other available synonym of the given drug and find all available information about it in a single place.
The Target Discovery dashboard presents the MoA and shows that Alectinib is an inhibitor of the RET and ALK kinases and the ALK receptor.

Exploring the target interactants to broaden the view
We select one of the known targets of Alectinib, for example the ALK receptor, and continue our journey in other data sources exploring the molecules that this protein interacts with. Researchers can simply open a Target Discovery dashboard to view the protein level data of the ALK kinase and identify relevant interactions coming from diverse sources, such as Uniprot, NCBI, StringDB, WikiPathways, and AI-derived data from text documents.
Here, LLD Inventory identifies the interaction partner in these databases by connecting the identifiers of a Uniprot protein, NCBI gene, WikiPathways interactant, and StringDB protein. For the AI-based analytics, LLD Inventory recognizes the correct mention of the protein of interest with the help of custom pipelines. These scan the scientific publications or clinical trials for mentions of ALK and match this string to the corresponding gene identifier in NCBI or protein identifier in Uniprot.
LLD Inventory enables interconnecting all of this knowledge in a comprehensive knowledge graph. This allows researchers not only to gain a richer picture of the available interactions, but also to easily determine the confidence of a single piece of information through the redundancy of the evidence availability.

Focusing on the target and the pathways it is part of
For the goals of drug repurposing, researchers would try to find new pathological mechanisms that a drug might address. Since the target is already known to be associated with or even causal for the primary indication of the drug, researchers can find new potential indications by exploring the target’s interaction network and the gene-disease relations downstream or upstream of a particular interaction. This means that by inhibiting the target and disrupting the direct interaction, researchers would be searching for novel therapeutic applications in the space of the interactant instead of the target.
To showcase this, let’s select an interactant with a good evidence score, EMAP-4, and explore its disease relevance in terms of expression profile. LLD Inventory facilitates this by providing multimodal expression data – the expression data is available both on protein and on gene level thanks to the extensive mappings between the data sources.
On protein level, the expression data is provided as protein staining profiles in different cancer types. The protein staining value is the number of patients (maximum 12 patients) that show high, medium, or low expression of the protein. We have selected to showcase three types of cancer with relatively high numbers of high/medium expression. One is the original application of Alectinib – Lung cancer. There are also two potential indication candidates based on the inhibited interaction – Pancreatic cancer and Stomach cancer.

On gene level we can explore the expression using the mapping of the protein (EMAP-4) to the gene that encodes it (EMAP like 4). LLD Inventory allows us to extract the gene expression in tissues of different cancer types, as categorized in Clinical Trials data. We have visualized examples of the relevant cancer types: NSCLC (non-small-cell lung cancer), PAAD (pancreatic adenocarcinoma), and STAD (stomach adenocarcinoma).

Exploring seamlessly with Target Data
The goal of this demonstration was to find out whether the interaction partner of the target protein has a high expression in another disease, for which the original target may not have indication. We observed that the selected interactant has high expression in lung cancer (the original application of Alectinib) and relatively high expression in Pancreatic cancer. In comparison, stomach cancer is lower expression.
The whole process can be applied in bulk for many targets simultaneously using a streamlined application such as Target Discovery. Consistent evidence of expression of interactants in other diseases will help in formulating a sound hypothesis.
The video below shows all the steps of using Target Discovery for our example of drug repurposing research we’ve described so far.
Validating the direction of the research
To validate our observations, we can check out if this has already been researched. Searching for clinical trials that include Alectinib and the disease of interest, provides the following results:
- There are many clinical trials for non-small-cell lung cancer as this is the primary application of the drug
- One clinical trial was found for pancreatic neoplasms
- No clinical trials are available for stomach cancer and Alectinib
Our investigation showed that Alectinib could potentially be used in treatment of pancreatic cancer and this was already explored in a phase 2 clinical study. This suggests that the indications based on the expression of the interactant may lead us to potential repurposing.
Main Takeaways
By providing a rich collection of interconnected biomedical datasets, LLD inventory speeds up analytical and research processes significantly. Its datasets and mappings can be consumed directly through querying. However, having a specialized application for a use-case of interest can make research much easier. Here, we have demonstrated how by utilizing LLD in Target Discovery researchers can make early drug R&D discoveries.
LLD Inventory forms the semantic foundation for Ontotext’s Target Discovery solution. It offers a robust platform for leveraging biomedical data to drive research and development forward. Target Discovery provides customizable dashboards and analytics allowing researchers to interact with the data seamlessly.
By centralizing and organizing the data, a knowledge graph not only supports individual research projects but also fosters collaboration across the scientific community. Researchers can share knowledge and resources, collectively tackling complex health challenges and achieving advancements that benefit society as a whole.



