Lentzen, Manuel: Leveraging AI for Real-World Health Data : From Clinical Notes to Wearable Sensors. - Bonn, 2026. - Dissertation, Rheinische Friedrich-Wilhelms-Universität Bonn.
Online-Ausgabe in bonndoc: https://nbn-resolving.org/urn:nbn:de:hbz:5-92098
@phdthesis{handle:20.500.11811/14419,
urn: https://nbn-resolving.org/urn:nbn:de:hbz:5-92098,
doi: https://doi.org/10.48565/bonndoc-952,
author = {{Manuel Lentzen}},
title = {Leveraging AI for Real-World Health Data : From Clinical Notes to Wearable Sensors},
school = {Rheinische Friedrich-Wilhelms-Universität Bonn},
year = 2026,
month = aug,

note = {The digitalization of healthcare has generated unprecedented volumes of Real-World Data (RWD) from both clinical and nonclinical sources. Clinical data streams encompass structured Electronic Health Records (EHRs) containing diagnoses, medications, and laboratory values, alongside unstructured clinical notes, while nonclinical sources include Remote Monitoring Technologies (RMTs) such as wearable sensors and smartphone applications. Although this abundance of data holds promise for more precise diagnostics and personalized treatments, translating heterogeneous RWD into actionable clinical insights remains challenging due to irregular sampling, coding artifacts, informative missingness, and source-specific limitations. This dissertation investigates how computational models, from classical Machine Learning (ML) to advanced transformer architectures, can be systematically leveraged to address the distinct characteristics of heterogeneous real-world health data.
Three complementary studies demonstrate this principle of method-data fit across clinical and nonclinical data regimes. First, to address unstructured clinical text, we developed BioGottBERT, a domain-adapted transformer for German clinical notes that exemplifies non-English biomedical Natural Language Processing (NLP). Systematic evaluation across five biomedical corpora and multiple Named Entity Recognition (NER) and document classification tasks demonstrated that continued pretraining outperforms training from scratch in low-resource settings. Second, for structured clinical data, we introduced ExMed-BERT, a transformer architecture that processes longitudinal EHR sequences from a U.S. claims database, integrating diagnoses, medications, and demographics while employing harmonized vocabularies (Phecodes, Anatomical Therapeutic Chemical (ATC) codes) to reduce dimensionality and improve transferability. Trained on this database, the model outperformed tree- and LSTM-based baselines for Coronavirus Disease 2019 (COVID-19) disease progression prediction. Third, for nonclinical data streams, we evaluated multiple RMTs for Alzheimer's Disease (AD) staging in a cross-sectional study, applying classical ML methods appropriate for small samples while maintaining interpretability. The best-performing technology could distinguish prodromal from healthy states, though early preclinical detection remained challenging.
Collectively, these studies illustrate that method-data fit is essential when applying computational methods to heterogeneous health data: model selection should be driven by data structure, sample size, and clinical context rather than by architectural novelty alone. Beyond this principle, the work contributes practical resources, including the publicly available BioGottBERT and ExMed-BERT models, as well as systematic evaluations that can inform future applications of ML in real-world healthcare settings.},

url = {https://hdl.handle.net/20.500.11811/14419}
}

Die folgenden Nutzungsbestimmungen sind mit dieser Ressource verbunden:

Namensnennung 4.0 International