Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Abstract

BACKGROUND: Generating synthetic patient data is crucial for medical research, but common approaches build up on black-box models which do not allow for expert verification or intervention. We propose a highly available method which enables synthetic data generation from real patient records in a privacy preserving and compliant fashion, is interpretable and allows for expert intervention.

METHODS: Our approach ties together two established tools in medical informatics, namely OMOP as a data standard for electronic health records and Synthea as a data synthetization method. For this study, data pipelines were built which extract data from OMOP, convert them into time series format, learn temporal rules by 2 statistical algorithms (Markov chain, TARM) and 3 algorithms of causal discovery (DYNOTEARS, J-PCMCI+, LiNGAM) and map the outputs into Synthea graphs. The graphs are evaluated quantitatively by their individual and relative complexity and qualitatively by medical experts.

RESULTS: The algorithms were found to learn qualitatively and quantitatively different graph representations. Whereas the Markov chain results in extremely large graphs, TARM, DYNOTEARS, and J-PCMCI+ were found to reduce the data dimension during learning. The MultiGroupDirect LiNGAM algorithm was found to not be applicable to the problem statement at hand.

CONCLUSION: Only TARM and DYNOTEARS are practical algorithms for real-world data in this use case. As causal discovery is a method to debias purely statistical relationships, the gradient-based causal discovery algorithm DYNOTEARS was found to be most suitable.

OriginalspracheEnglisch
Aufsatznummer136
ZeitschriftBMC Medical Research Methodology
Jahrgang24
Ausgabenummer1
Seiten (von - bis)136
DOIs
PublikationsstatusVeröffentlicht - 22.06.2024

Fördermittel

Some experimental parts of this work were conducted by N.A.S. for his graduate project. During the preparation of this manuscript, the authors used (generative) AI-powered tools like DeepL, Grammarly, and ChatGPT to improve the language and style of some parts of the paper. These tools were not used to generate any content of the paper itself. AI-CARE Working Group Institution Members 2Hamburg Cancer Registry, Ministry of Science, Research, Equality and Districts, Free and Hanseatic City of Hamburg, S\u00FCderstra\u00DFe 30, 20097 Hamburg, Germany Alice Nennecke2, Henrik Kusche2, Ole Johanns2, Vera Heinrichs2 6Bremen Cancer Registry, Leibniz Institute for Prevention Research and Epidemiology - BIPS, Achterstra\u00DFe 30, 28359 Bremen Andrea Eberle6, Sabine Luttmann6 7Hessian Cancer Registry, Hessian Office of Health and Care, Lurgiallee 10, 60439 Frankfurt Khalid Abnaof7, Soo-Zin Kim-Wanner7 4Saarland Cancer Registry, State Ministry of Labour, Social Affairs, Women and Health, Neugel\u00E4ndstra\u00DFe 9, 66117 Saarbr\u00FCcken, Germany Bernd Holleczek4, Katharina Rausch4, Natalie Rath4 8German Research Center for Artificial Intelligence (DFKI), Ratzeburger Allee 160, 23562 L\u00FCbeck, Germany Heinz Handels8, Sebastian Germer8 9Baden-Wuerttemberg Cancer Registry, Klinische Landesregisterstelle Baden-W\u00FCrttemberg GmbH, Birkenwaldstra\u00DFe 149, 70191 Stuttgart, Germany Marco Halber9, Martin Richter9 10Johann Wolfgang Goethe-Universit\u00E4t Frankfurt, Universit\u00E4tsklinikum Frankfurt, Institut f\u00FCr Medizininformatik, Theodor-Stern-Kai 7, 60590 Frankfurt am Main Martin Pinnau10, David Reinert10, Jannik Schaaf10, Holger Storf10 11Clinical Cancer Registry Lower Saxony, Sutelstra\u00DFe 2, 30659 Hannover, Germany Tobias Hartz11, Nils Goeken11, Janina B\u00F6sche11 12Institute for Community Medicine, Section Epidemiology of Health Care and Community Health, University Medicine Greifswald, Ellernholzstra\u00DFe 1-2, 17475 Greifswald, Germany Alexandra Stein12, Kerstin Weitmann12, Wolfgang Hoffmann12 13Institut f\u00FCr Sozialmedizin und Epidemiologie, Universit\u00E4t zu L\u00FCbeck Louisa Labohm13, Alexander Katalinic5,13 5Institut f\u00FCr Krebsepidemiologie an der Universit\u00E4t zu L\u00FCbeck, Registerstelle des Krebsregisters Schleswig-Holstein Christiane Rudolph5, Alexander Katalinic5,13 1Universit\u00E4tsklinikum Hamburg-Eppendorf, Institut f\u00FCr Angewandte Medizininformatik, Martinistra\u00DFe 52, 20246 Hamburg Christopher Gundler1, Frank \u00DCckert1 Open Access funding enabled and organized by Projekt DEAL. The work is partly funded through the third-party funded project \u201CAI-CARE\u201D by the German Ministry of Health. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. For the open-access publication itself, we acknowledge support from the Open Access Publication Fund of UKE - Universit\u00E4tsklinikum Hamburg-Eppendorf and DFG - German Research Foundation.

TrägerTrägernummer
Ministry of Science, ICT and Future Planning
Hessian Office of Health and Care
British Institute of Persian Studies
Universitätsklinikum Frankfurt
Leibniz-Institut für Präventionsforschung und Epidemiologie
Deutsche Forschungsgemeinschaft
Universitätsklinikum Hamburg-Eppendorf
Deutsches Forschungszentrum für Künstliche Intelligenz
German Ministry of Health
Institut für Medizininformatik17475

    UN SDGs

    Dieser Output leistet einen Beitrag zu folgendem(n) Ziel(en) für nachhaltige Entwicklung

    1. SDG 3 – Gesundheit und Wohlergehen
      SDG 3 – Gesundheit und Wohlergehen

    Strategische Forschungsbereiche und Zentren

    • Profilbereich: Zentrum für Bevölkerungsmedizin und Versorgungsforschung (ZBV)

    DFG-Fachsystematik

    • 2.22-02 Public Health, gesundheitsbezogene Versorgungsforschung, Sozial- und Arbeitsmedizin

    Fingerprint

    Untersuchen Sie die Forschungsthemen von „Learning debiased graph representations from the OMOP common data model for synthetic data generation“. Zusammen bilden sie einen einzigartigen Fingerprint.

    Zitieren