molecular data science for precision health

Understanding individual health through molecular measurements, statistical inference, and artificial intelligence

Research group of Dr. Kosmas Kepesidis

We develop statistical and machine-learning methods to extract health information from infrared measurements of human blood. Our work combines Fourier-transform infrared (FTIR) spectroscopy and electric-field molecular fingerprinting with clinical data, probabilistic modelling and artificial intelligence.

Building on studies of cancer detection, disease severity and other health-related traits, our central priority is personalised health monitoring. We investigate how repeated molecular measurements can define an individual's baseline and reveal changes that may signal a developing health problem. We aim to establish the scientific basis for earlier, more informative and affordable health assessment.

 

personalised health monitoring and early disease detection

A molecular change can be unusual for one person even when it falls within the normal range of a population. Using longitudinal cohorts, including Health for Hungary (H4H), we study the stability of individual blood fingerprints and distinguish ordinary within-person fluctuations from differences between people and measurement variability.

We develop methods for representing personal reference states and recognising persistent departures from expected health trajectories. Bayesian inference, information geometry and sequential analysis provide a framework for updating these states, quantifying uncertainty and accumulating evidence across visits. A key objective is to determine when a molecular deviation warrants further investigation and how reliably it can anticipate a clinically confirmed change in health.

 

multimodal health profiles and adaptive screening

Different measurements capture different aspects of physiology. We combine infrared fingerprints with routine clinical laboratory data to learn compact representations of an individual's molecular state. We investigate which information is shared across assays, which is complementary, and what remains uncertain when some measurements are unavailable.

This work supports research into adaptive screening: selecting additional tests according to their expected value for a defined clinical decision, while accounting for cost, patient burden and delay. Our broader programme aims to extend this approach to proteomics, metabolomics and lipidomics, linking scalable monitoring with more detailed molecular investigation when it is informative.

 

generative models and synthetic cohorts

We develop statistical and deep generative models of blood-based infrared spectra to study molecular variation under controlled computational conditions. Conditional models generate profiles for specified characteristics, such as age, sex, and body mass index, supporting virtual cohorts, exploratory ageing trajectories, and targeted augmentation of underrepresented groups.

Synthetic data provide a resource for benchmarking analysis methods and investigating the effects of cohort composition. We evaluate their fidelity and usefulness against real measurements. Our research also explores causal generative models and the assumptions and evidence needed to move from conditional simulation towards meaningful counterfactual questions about individual health.

 

informative and interpretable molecular representations

We investigate how to represent complex spectroscopic signals while retaining the information relevant to health. Our approaches include unsupervised deep learning and physically motivated features, such as the timing of zero crossings in electric-field molecular fingerprints.

Alongside representation learning, we develop information-theoretic approaches to compare molecular assays and assess the stable, person-specific information they capture. The goal is to connect measurable features with biological variation, make models easier to interpret, and identify combinations of measurements that contribute useful additional information.

 

reliable machine learning and clinical evidence

Reliable biomedical inference requires careful treatment of both biological and technical variation. In collaborative clinical studies, we investigate molecular signatures associated with cancer, disease severity, prognosis and multiple health-related traits. We combine medical statistics, study design and model evaluation to assess the influence of demographic differences, clinical covariates and other potential confounders.

We also develop domain-adaptation and harmonisation methods to improve transfer between instruments and measurement settings. Reproducible data processing, quality control and validation support all of these activities. For longitudinal screening, the next step is prospective evaluation of clinically useful lead time, false-alert burden and the value of the follow-up decisions triggered by a molecular change.

 

A mobile version for attoworld.de is under construction.