Univ. Prof. Dr.
Data Integration, Interoperability & Standards
Tools & Services
Publikationen
Meet NUM-ENRICH: A Collaborative National Effort to Extend and Harmonize Research Infrastructures Within the German Network University Medicine.
During the Corona Pandemic, the German Network University Medicine has established robust research data infrastructures for data and biospecimen-driven research. However, due to the fast set-up, realizing the full potential has been hindered by challenges, such as siloed and unharmonized data, varying data access procedures, and infrastructure-specific governance frameworks. These issues impede cross-infrastructure data sharing and interoperability and easy access. A collaborative, national effort involving all German University clinics and stakeholders in the health landscape will therefore work towards the harmonization of infrastructures, data pools, data handling procedures and governance frameworks. The ENRICH consortium is committed to maximise the value of NUM data and biosamples for the research community as a whole.
Catnip for MedCAT: Optimizing the Input for Automated SNOMED CT Mapping of Clinical Variables
Introduction: Mapping local medical data assets to international data standards such as medical ontology SNOMED CT fosters data harmonization and, thereby, global progress in medical research. Since its intense resource requirements often hinder manual SNOMED CT mapping, automated mapping tools such as MedCAT have been developed. We investigated how the formulation of study variable names (VNs) influences the efficacy and accuracy of the SNOMED CT concepts identified by MedCAT.
Methods: We extracted 763 VNs from the GEPESTIM database hosted locally in REDCap and created three VNs using different REDCap metadata items for MedCAT-based SNOMED CT mapping. A fourth VN version was created manually. The mapping was evaluated based on the number and quality of identified SNOMED CT concepts, using manual scoring to assess concept accuracy while ensuring a blind evaluation process.
Results: Increasing the expressiveness of VNs by adding more metadata items led to more SNOMED CT concepts being mapped, but also introduced mismatches, particularly when additionally included metadata contained misleading terms. The best overall mapping performance was achieved on the manually specified VNs while a basic VN version with minimal extra information from the metadata resulted in similarly good results.
Conclusion: Our study identified key challenges in using MedCAT for automatically mapping study variables to SNOMED CT concepts. To improve accuracy, we recommend refining VNs reducing misleading terms and iteratively improving VN phrasing for optimal mapping outcome. Furthermore, it appears reasonable to always conduct a final manual review of the mapping outcome especially for critical variables and for those VNs containing negations or abbreviations.
Keywords: Automation; Computerized Medical Record System; Controlled Vocabulary; Natural Language Processing; SNOMED CT.
What prevents us from reusing medical real-world data in research
Medical real-world data stored in clinical systems represents a valuable knowledge source for medical research, but its usage is still challenged by various technical and cultural aspects. Analyzing these challenges and suggesting measures for future improvement are crucial to improve the situation. This comment paper represents such an analysis from the perspective of research.
Kurse
RDM4Researchers: Designing Reproducible Life Science Research Across the Data Lifecycle
Dieser Kurs richtet sich an Lebenswissenschaftler, die möchten, dass ihre Daten lange nach Abschluss eines Experiments verständlich, nutzbar und zuverlässig bleiben. Anstatt Management von Forschungsdaten (RDM) als Bürokratie oder Compliance zu behandeln, nähert sich der Kurs diesem als praktischen Bestandteil guter Forschung. Anhand von Beispielen aus den Bereichen Omics, Bildgebung, Mikroskopie und Molekularbiologie erkunden die Teilnehmer, wie alltägliche Entscheidungen – wie Daten benannt, dokumentiert, analysiert und geteilt werden – die Reproduzierbarkeit, Wiederverwendbarkeit und den langfristigen Wert im gesamten Forschungsdatenlebenszyklus beeinflussen.
In KLIPS anzeigenOpen Science Essentials: Tools, Principles, and Practices
Lecture introduces the principles and practice of Open Science and the FAIR (Findable, Accessible, Interoperable, Reusable) principles in a clear and accessible way for researchers across disciplines. It explores why transparency, collaboration, and responsible data sharing are becoming central to research, particularly in light of funder requirements under programmes such as Horizon Europe. Participants will gain an overview of key concepts, policy context, and practical steps to integrate Open Science and FAIR principles into their everyday research workflows.
In KLIPS anzeigen