AI to acceleratetranscriptomic research.
Geneversity: millions of RNA profiles to identify cells, estimate their proportions and explore new biomarkers.
single-cell datasets
single-cell profiles
bulk RNA datasets
microarray and RNA-seq samples
Millions of profiles.New biological questions.
Transcriptomic experiments generate a vast but heterogeneous body of data. This project collects and standardises microarray, bulk RNA and single-cell data to explore relationships between gene expression, cell types, tissues and diseases.
- Single-cell RNA-seq
- Bulk RNA
- Multitask learning
Inside the platform.
Three views connect cell selection, bulk composition and gene-expression comparisons. Interface design reconstructed from the project prototype; chart profiles are illustrative.
From an individual cell to its context.
The selector follows the prototype’s cell populations and connects them to the single-cell foundation: 28 million profiles, including 5 million labelled profiles. The side panel shows the binary test on T lymphocytes and epithelial cells, separately from the coverage of 32 cell types.
Explore the populations within a tissue.
A bulk sample combines signals from multiple cell populations. The project explores how deep learning can estimate their proportions. The view compares samples and components: the bars are illustrative; the 317 datasets and 26,000 samples describe the project’s data foundation.
Compare groups. Form new hypotheses.
Select the populations of interest and compare gene expression across groups while keeping the context visible. The view retains the prototype’s two example groups of 120 and 35 individuals. Gene X and the bars illustrate the comparison; they do not represent validated biomarkers.
What we designed.
The aim is to learn from multiple data modalities and address several tasks within a shared framework: cell annotation, proportion estimation and biomarker exploration. An organised foundation helps researchers move from data preparation to biological questions.
Organise biological knowledge
Collect, standardise and annotate data by cell type and tissue: the information foundation comes before the model.
Learn across modalities
A multitask approach to single-cell and bulk RNA data for cell classification, cell proportion estimation and biomarker research.
Make hypotheses explorable
The prototype lets researchers select cell populations and compare gene expression across groups. Sample context accompanies the differences and helps formulate hypotheses about a disease or biological state.
Microarray, bulk RNA and single-cell
Standardisation + multitask deep learning
Cell annotation and biomarker hypotheses
One research foundation.Two scales of observation.
Single-cell and bulk RNA data brought together in a structured foundation for transcriptomic analysis.
From cell profile to tissue.
Single-cell
381 datasets and 28 million profiles, including 5 million labelled by cell type and tissue.
- 32 cell types
- 18 tissues
- T lymphocytes and epithelial cells
Bulk RNA
317 datasets with 26,000 microarray and RNA-seq samples labelled by cell type and tissue.
- 15 tissues
- Over 20 tissue subtypes and cell types
- 16 diseases, including rheumatoid arthritis and psoriasis
From RNA data.To biological questions.
A journey connecting data preparation, learning and exploration.
Make different sourcescomparable.
Microarray, bulk RNA-seq and single-cell data are collected and standardised. Cell type and tissue organise the available labels.
- Microarray
- Bulk RNA-seq
- Single-cell RNA
A structured data foundation for learning.
Multiple modalities.Multiple biological tasks.
The deep learning framework connects multimodal data and different objectives through a multitask approach. Cell annotation is one of the studied applications.
- Cell types
- Proportions
- Biomarkers
Shared representations to explore across different tasks.
A biological question.A comparison to investigate.
The researcher selects populations and compares groups and expression profiles. Observed differences guide hypotheses and subsequent experimental investigation.
- Populations
- Groups
- Gene expression
Contextualised research hypotheses to investigate.
A focused test.Measured results.
The project includes a binary classifier for T lymphocytes and epithelial cells: 99% precision, 98% recall, 99% specificity and 99% accuracy on the binary comparison test set.
The test covers only binary classification between T lymphocytes and epithelial cells. These results do not measure every framework application and are not clinical validation.
The questions driving the research.
- Which cell types are present in the sample?
- How does gene expression differ between two groups?
- Which signals warrant further investigation as potential biomarkers?
What it makespossible.
- A single-cell and bulk RNA foundation collected, standardised and labelled by cell type and tissue.
- A binary cell-annotation test with 99% precision, 98% recall, 99% specificity and 99% accuracy.
- An exploration environment to compare populations and formulate new biomarker hypotheses.
Turn your datainto new research questions.
Turn your data into new research questions.
Tell us about your projectAutomated data extraction from property appraisals.
Upload appraisal reports. AI extracts cadastral data, occupancy, discrepancies and financial values for structured property profiles.