Resources


Database Credentialed Access

PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions

Daeun Kyung, Hyunseung Chung, Seongsu Bae, Jiho Kim, Jae Ho Sohn, Taerim Kim, Soo Kim, Edward Choi

PatientSim is a patient simulator that simulates realistic and diverse personas for clinical scenarios, enabling robust training and evaluation of doctor-patient interactions in multi-turn dialogues.

electronic health records multi-turn dialogue llm simulation doctor-patient consultation

Published: Oct. 18, 2025. Version: 1.0.0


Database Contributor Review

Salzburg Intensive Care database (SICdb), a freely accessible intensive care database

Niklas Rodemund, Andreas Kokoefer, Bernhard Wernly, Crispiana Cozowicz

The SICdb dataset, version 1.0.8 contains 27350 admissions to an ICU in an Austrian tertiary care institution.

clinical intensive care critical care open data machine learning

Published: Sept. 10, 2024. Version: 1.0.8


Database Open Access

SHDB-AF: a Japanese Holter ECG database of atrial fibrillation

Kenta Tsutsui, Shany Biton Brimer, Joachim Behar

Holter ECG database from Japan, containing data from 100 unique patients with paroxysmal AF including expert annotations of Supraventricular arrhythmias at the beat level.

atrial fibrillation ecg holters

Published: April 16, 2025. Version: 1.0.1


Database Open Access

Auditory evoked potential EEG-Biometric dataset

Nibras Abo Alzahab, Angelo Di Iorio, Luca Apollonio, Muaaz Alshalak, Alessandro Gravina, Luca Antognoli, Marco Baldi, Lorenzo Scalise, Bilal Alchalabi

Recording of electroencephalogram (EEG) signals with the aim to develop an EEG-based Biometric. The Data includes resting-state and auditory stimuli experiments.

eeg biometric electroencephalogram auditory stimuli resting-state

Published: Dec. 1, 2021. Version: 1.0.0

Visualize waveforms

Database Restricted Access

Visual Question Answering evaluation dataset for MIMIC CXR

Timo Kohlberger, Charles Lau, Tom Pollard, Andrew Sellergren, Atilla Kiraly, Fayaz Jamil

This dataset provides 224 VQAs for 40 test set cases, and 111 VQAs for 23 validation set cases of the MIMIC CXR dataset.

Published: Jan. 28, 2025. Version: 1.0.0


Database Credentialed Access

MIMIC-III and eICU-CRD: Feature Representation by FIDDLE Preprocessing

Shengpu Tang, Parmida Davarmanesh, Yanmeng Song, Danai Koutra, Michael Sjoding, Jenna Wiens

Features and labels from MIMIC-III and eICU-CRD produced by FIDDLE, an EHR preprocessing pipeline.

preprocessing electronic health record machine learning

Published: April 28, 2021. Version: 1.0.0


Database Open Access

Simultaneous physiological measurements with five devices at different cognitive and physical loads

Marcus Vollmer, Dominic Bläsing, Julian Elias Reiser, Maria Nisser, Anja Buder

Dataset to support comparison of usability and accuracy from simultaneous measurements collected from 13 subjects including five devices: NeXus-10 MKII, eMotion Faros 360°, Hexoskin Hx1, SOMNOTouch NIBP, Polar RS800 Multi.

holter multiparameter photoplethysmogram noise accelerometer heart rate movement temperature hrv respiration ecg

Published: Jan. 18, 2023. Version: 1.0.2

Visualize waveforms

Database Contributor Review

COVID Data for Shared Learning (CDSL): A comprehensive, multimodal COVID-19 dataset from HM Hospitales

Álvaro Ritoré, Andreea M Oprescu, Alberto Estirado Bronchalo, Miguel Ángel Armengol de la Hoz

COVID Data for Shared Learning (CDSL) is a multimodal database comprising de-identified structured health data and radiological images from 4,479 patients with COVID-19, as a comprehensive toolkit for developing predictive models.

covid-19 multimodal database radiological images open data healthcare data machine learning and ai

Published: Oct. 25, 2024. Version: 1.0.0


Database Restricted Access

MIMIC-IV-Ext-Apixaban-Trial-Criteria-Questions

Elizabeth Woo, Michael Craig Burkhart, Emily Alsentzer, Brett Beaulieu-Jones

We created 23 questions resembling eligibility criteria from the apixaban clinical trial and evaluated them on a random sample of 100 patient notes from MIMIC-IV. We release the 2300 total question-answer pairs as a dataset here.

clinical q and a evaluation set clinical trial eligibility

Published: April 30, 2025. Version: 1.0.0


Database Contributor Review

COVID Data for Shared Learning (CDSL): A comprehensive, multimodal COVID-19 dataset from HM Hospitales

Álvaro Ritoré, Andreea M Oprescu, Alberto Estirado Bronchalo, Miguel Ángel Armengol de la Hoz

COVID Data for Shared Learning (CDSL) is a multimodal database comprising de-identified structured health data and radiological images from 4,479 patients with COVID-19, as a comprehensive toolkit for developing predictive models.

covid-19 multimodal database radiological images open data healthcare data machine learning and ai

Published: Oct. 25, 2024. Version: 1.0.0