Resources


Database Credentialed Access

FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark

Mingjie Li, Wenjia Cai, Rui Liu, et al.

Benchmark dataset for report generation based on fundus fluorescein angiography images and reports.

fundus fluorescein angiography medical report generation vision and language explainable and reliable evaluation

Published: Jan. 21, 2025. Version: 1.1.0


Database Credentialed Access

FFA-IR: Towards an Explainable and Reliable Medical Report Generation Benchmark

Mingjie Li, Wenjia Cai, Rui Liu, et al.

Benchmark dataset for report generation based on fundus fluorescein angiography images and reports.

fundus fluorescein angiography medical report generation vision and language explainable and reliable evaluation

Published: Jan. 21, 2025. Version: 1.1.0


Database Open Access

Brno University of Technology ECG Quality Database (BUT QDB)

Andrea Nemcova, Radovan Smisek, Kamila Opravilová, et al.

The database is intended for the development and objective comparison of algorithms designed to assess the quality of ECG records. It also enables objective comparison of results between authors.

Published: July 22, 2020. Version: 1.0.0

Visualize waveforms

Database Restricted Access

TAME Pain: Trustworthy AssessMEnt of Pain from Speech and Audio for the Empowerment of Patients

Tu-Quyen Dao, Eike Schneiders, Jennifer Williams, et al.

TAME Pain is a dataset that captures acoustic signals of pain and is augmented by annotating every sentence the participant speaks with details such as background and foreground noise, speech errors, and non-speech vocal features.

audio speech pain cold pressor task

Published: Jan. 21, 2025. Version: 1.0.0


Database Open Access

Image-derived cardiomegaly biomarker values for 96K chest X-rays in MIMIC-CXR/MIMIC-CXR-JPG

Benjamin Duvieusart, Felix Krones, Guy Parsons, et al.

Automatically extracted cardiomegaly biomarkers - cardiothoracic ratio (CTR) and cardiopulmonary area ratio (CPAR) - for all posterior-anterior chest x-ray scans in MIMIC-CXR/MIMIC-CXR-JPG.

biomarkers mimic-cxr cpar ctr cardiomegaly

Published: Aug. 23, 2024. Version: 1.0.0


Database Open Access

Hillel Yaffe Glaucoma Dataset (HYGD): A Gold-Standard Annotated Fundus Dataset for Glaucoma Detection

Or Abramovich, Hadas Pizem, Jonathan Fhima, et al.

HYGD is a rigorously annotated fundus image dataset with gold-standard clinical labels designed to improve and benchmark deep learning models for accurate glaucoma detection.

ophthalmology retina glaucoma dfi gon fundus gold-standard

Published: March 16, 2026. Version: 1.1.0


Database Open Access

VitalDB Arrhythmia Database: An Anesthesiologist-Validated Large-Scale Intraoperative Arrhythmia Dataset with Beat and Rhythm Labels

Dain Eun, Kayoung Shim, Hyunsoo Lee, et al.

We present a comprehensive intraoperative arrhythmia dataset with 734,528 seconds of ECG recordings from 482 patients, featuring over 660,000 beats annotated and validated by five anesthesiologists.

ppg vitaldb ecg arterial waveform intraoperative dataset

Published: Feb. 26, 2026. Version: 1.0.0


Database Open Access

PSG-IPA: A PolySomnoGraphic Inter-scorer Performance Assessment database

Diego Alvarez-Estevez

The HMC-IPA dataset comprises 20 PSG recordings, each with manual and computer-assisted scorings by 12 sleep technologists, for studying inter-scorer variability and evaluating automated sleep analysis algorithms

Published: Jan. 8, 2026. Version: 1.0.0

Visualize waveforms

Database Restricted Access

TN-Mammo: A Multi-view Mammography Dataset for Breast Density Classification

Binh Nguyen, Cat Le, Loc Vu, et al.

We release the first version of TN-Mammo (June 2024), a mammogram dataset of 676 cases with breast density labels, providing high-quality data to support machine learning and early breast cancer detection.

Published: Oct. 4, 2025. Version: 1.0.0


Database Credentialed Access

MIMIC-IV-Ext-22MCTS: A 22 Millions-Event Temporal Clinical Time-Series Dataset with Relative Timestamp

Jing Wang, Xing Niu, Tong Zhang, et al.

It is a time series clinical events dataset with concrete temporal information. The dataset consists of 22,588,586 clinical events and related timestamps from 267,284 discharge summaries of the MIMIC-IV-Note.

mimic clinical event annotation time series temporal annotation

Published: Sept. 29, 2025. Version: 1.0.0