Database Credentialed Access

MIMIC-Ext-SynNotes: Large-Scale Synthetic Clinical Notes via Rephrasing MIMIC-III/IV with LLMs

Jinghui Liu ,  Sarvesh Soni ,  Anthony Nguyen

Published Oct. 5, 2026 · Version 1.0
When using this resource, please cite:

Liu, J., Soni, S., & Nguyen, A. (2026). MIMIC-Ext-SynNotes: Large-Scale Synthetic Clinical Notes via Rephrasing MIMIC-III/IV with LLMs (version 1.0). PhysioNet. RRID:SCR_007345. https://doi.org/10.13026/f6se-e243

Please include the standard citation for PhysioNet:

Pollard, T., Moody, B. E., Lehman, L., Gow, B., Fernandes, C., Xie, C., Johnson, A., Mark, R. G., & Heldt, T. (2026). PhysioNet as a global platform for biomedical research. Nature Health. https://doi.org/10.1038/s44360-026-00096-z. Available from: https://rdcu.be/faatM

Abstract

MIMIC-III and MIMIC-IV are foundational resources for clinical NLP, but their scale is modest relative to the data volumes needed to train language models. Building on recent advances in synthetic data generation through LLM-based rephrasing, we construct MIMIC-Ext-SynNotes by rephrasing all clinical notes from MIMIC-III/IV with five open-weight LLMs. The release comprises eight synthetic corpora totalling 13.7B words, roughly a tenfold expansion of the free text in MIMIC. An accompanying peer-reviewed study evaluates the corpus through intrinsic similarity analysis, downstream clinical NLP tasks, fairness assessment, LLM-based fact-checking, and human error analysis. It finds that the synthetic notes preserve core clinical information and predictive utility for coarse-grained tasks but lose fine-grained detail — a loss largely mitigated by rephrasing in chunks rather than whole notes. We release the resource to support research in clinical language modeling, synthetic data generation, and domain adaptation for healthcare LLMs.


Background

Synthetic data generation using large language models (LLMs) offers adaptability and cost-effectiveness. LLMs have recently been used to generate pretraining corpora for other LLMs [1,2], motivated by the saturation and inconsistent quality of web data, with filtering [3,4] and rephrasing [5,6] proposed to improve synthetic data quality at scale.

This need is acute in healthcare, where domain-specific adaptation is essential [7–9] but access to patient data is limited. Most pretraining corpora for clinical language models rely on MIMIC [10,11], which at approximately 1.4 billion words of free text remains modest compared to general-domain datasets. This has motivated recent work [12,13] using LLMs to generate synthetic clinical notes, demonstrating that LLMs can produce text closely resembling real clinical documentation.

Source material. All synthetic text was generated by rephrasing notes drawn exclusively from MIMIC-III Clinical Database v1.4 and MIMIC-IV-Note v2.2. No external clinical text sources were used in any rephrasing workflow — no other hospital corpora, publicly available note collections, web-scraped medical text, or clinician-authored text. Every released row corresponds one-to-one to a source note or chunk and records its source database and original note identifier. No documents beyond the source text were injected into the generation context.

The construction, evaluation, and limitations of this resource are reported in an accompanying peer-reviewed publication [14], which users should consult before using the data.


Methods

Notes were sourced from MIMIC-III v1.4 and MIMIC-IV-Note v2.2 only. After removing exact duplicates, 4.64M unique notes (~1.4B words) remained. The sole preprocessing step was replacing MIMIC-III deidentification placeholders with underscores ("___") to match MIMIC-IV; no other normalization, sectioning, or content removal was applied.

Models. Three open-weight instruction-tuned LLMs of 7–9B parameters were used for primary generation: Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and Qwen-2.5-7B-Instruct. Two additional models, Gemma-2-9B-it and Mixtral-8x7B-Instruct-v0.1, were used for chunk-level generation only (see below). Rephrasing grounds generation in real notes rather than relying on the LLM as a knowledge bank, an approach shown effective on web data [5,6]. All context lengths (32K–128K tokens) accommodate MIMIC note lengths. All models were used exactly as published, with no fine-tuning or adaptation of any kind. No weights were updated and no MIMIC data was used to train or adapt any model. All models were served in FP8, except Qwen-2.5-7B, which used the official GPTQ Int8 release. Inference ran on NVIDIA H100 GPUs with vLLM [15].

Prompting. Each model received a system prompt matching its own chat template ("You are a medical artificial intelligence assistant. The assistant gives truthful, detailed, and professional answers to the requests.") followed by a paraphrasing instruction found effective in prior work [5,13] ("For the following paragraph give me a diverse paraphrase of the same in high-quality English language as in clinical notes written by medical professionals:"). Prompts were identical across models, notes, and granularities. No few-shot demonstrations, task labels, or retrieved context were used.

Decoding parameters. Sampling used temperature 0.75 and top-p 0.9 throughout. For note-level rephrasing, notes were binned by word count — (0,500], (500,1000], (1000,2000], (2000,4000], >4000 — with maximum new tokens of 1,000 / 2,000 / 4,000 / 8,000 / 10,000 respectively; chunk-level generation used 512. No constraints or guardrails were applied beyond these parameters. Row-level filtering applied after generation is described under Data Description; it did not alter the text of any retained output.

Chunk segmentation. Because LLMs omit detail as input length grows [5,6], we also rephrase by chunk. Segmentation ran on the original note: sentences were detected, then greedily accumulated until adding the next would exceed a ~150-word target, at which point a new chunk began.

  • Boundaries are strictly non-overlapping. No sentence, token, or character appears in more than one chunk; no sliding window or context carry-over was used. A note's chunks form an ordered partition.
  • Chunks are cut only at sentence boundaries. A sentence longer than the ~150-word target is emitted as its own oversized chunk rather than split.
  • Each chunk was rephrased independently, with no cross-chunk context. This drives the trade-off documented in [14]: chunking improves retention of clinical detail but raises hallucination, since grounding context may sit in a neighbouring chunk.

Segmentation yields an average of 2.3 chunks per note.


Data Description

The release comprises eight synthetic corpora: three LLMs (Llama, Mistral, Qwen) applied at both note and chunk level, plus Gemma and Mixtral at chunk level only — 13.7B words in total. The peer-reviewed evaluation [14] analyses the six Llama/Mistral/Qwen corpora (9.1B words) in its main results and reports complementary findings for Gemma and Mixtral, which did not differ substantially.

File Structure and Formats

All corpora are UTF-8 CSV files in data/. Sizes are uncompressed and row counts are logical CSV records, excluding the header.

Filename Size Rows A row is Words Avg words/row
data/synthetic-notes-by-Llama.csv 8.6 GB ~4.64M one note 1.21B 260
data/synthetic-notes-by-Mistral.csv 9.3 GB ~4.64M one note 1.34B 289
data/synthetic-notes-by-Qwen.csv 8.0 GB ~4.29M one note 1.13B 263
data/synthetic-chunks-by-Llama.csv 16.3 GB ~10.88M one chunk 2.24B 206
data/synthetic-chunks-by-Mistral.csv 12.6 GB ~10.88M one chunk 1.74B 160
data/synthetic-chunks-by-Qwen.csv 10.2 GB ~9.67M one chunk 1.41B 146
data/synthetic-chunks-by-Gemma.csv 17.8 GB ~10.87M one chunk 2.40B 221
data/synthetic-chunks-by-Mixtral.csv 11.4 GB ~10.87M one chunk 1.62B 149

Note-level files contain four columns; chunk-level files contain the same four plus chunk_id, inserted before text:

  • source: database of origin, either mimic_iii_1.4 or mimic_iv_2.2.
  • note_id: identifier of the original note — ROW_ID in MIMIC-III, note_id in MIMIC-IV. Repeats across a note's chunks.
  • category: MIMIC note category, such as Discharge summary, Nursing, or Radiology.
  • chunk_id (chunk-level files only): zero-based position of the chunk within the original note.
  • text: the LLM-generated synthetic text.

(source, note_id) is unique in note-level files and (source, note_id, chunk_id) in chunk-level files. deduplicated_note_ids.csv lists the ~93K exact-duplicate notes removed before synthesis, with columns source and note_id. Each file's header row matches this documentation exactly.

Quality Filtering

Two row-level filters were applied before release; neither alters the text of a retained generation. Empty or failed generations were removed, and non-English outputs were removed using a fastText language-identification score [16]. These account for the differing row counts across corpora, most visibly Qwen. Because filtering is per row, a note may appear in one corpus and not another, so users combining corpora should join on (source, note_id) rather than assume aligned row order. No other filtering was applied: nothing was removed on the basis of length, repetition, or similarity to the source.

Synthetic Artifacts

MIMIC is deidentified and all protected health information appears as the placeholder ___ in the source, but the LLMs frequently rewrite these placeholders into plausible surface forms. The released text therefore may sometimes contain real-looking dates, times, ages, phone-number-like strings, names, or residual placeholder tokens. Every such string is an LLM-generated artifact: invented during generation, not present in or derived from MIMIC, and corresponding to no real patient, clinician, institution, date, or identifier. Because the models received only deidentified input, no identifying information was ever available for them to reproduce, so these artifacts cannot be used to re-identify any patient or clinician and any resemblance to a real name, date, or number is coincidental. They should be treated as noise and masked or discarded rather than used as signal.


Usage Notes

Intended Uses

This resource is intended as large-scale, task-agnostic text for language model development: continued pretraining and domain adaptation of clinical language models, research on synthetic data generation and rephrasing, corpus-level analysis of LLM-generated clinical text, and data augmentation for supervised clinical NLP. The accompanying evaluation [14] found that augmenting real training data with these notes improved ICD coding performance, with the largest gains on rare codes, despite the corpus being generated without task-specific conditioning. Because each synthetic note is paired with its human-written source, the corpus also supports paired study of how LLMs preserve and distort clinical information at scale.

Known Limitations

This is not a privacy-preserving substitute for real clinical notes: the text is derived from real MIMIC notes and remains subject to credentialed access and the data use agreement. It is not a source of clinical truth and should not be used, without validation, for clinical decision support, for deriving facts about patients, or in any patient-facing system.

The synthetic notes differ measurably from their human-written sources in length, sentence structure, readability, and lexical and semantic distribution. Moreover, LLM rephrasing introduces recurring errors identified through human error analysis [14]:

  • Omission of content — the dominant failure mode, worsening as note length increases and substantially mitigated by rephrasing in chunks.
  • Misinterpretation of clinical context — a fact restated with altered certainty, negation, or subject.
  • Temporal confusion and measurement errors — likely arising from clinical shorthand and failure to interpret numerical values in context.
  • Hallucinated clinical detail — conversely increased by chunking, where grounding context may lie in an adjacent chunk.
  • Fabricated identifiers in place of ___ placeholders.

Chunking therefore tends to trade retention of clinical detail against factual precision. The corpora are also not interchangeable, as each LLM leaves a distinct stylistic and semantic signature. More quantitative results and the evaluation methodology are reported in [14].

Users should assume that any specific detail may be unsupported by the source note, and that any name-, date-, time-, or phone-like token is fabricated. Where facts are needed, the original MIMIC note is linkable through source and note_id in each table.

Additional quality checks and filtering are likely appropriate depending on the application, for which users may consider near-duplicate detection, length filters, perplexity or quality-classifier filtering as used in web-corpus curation [3,4], or entailment-based factuality filtering. Permissive filtering is usually adequate for pretraining, but any use requiring factual content warrants stricter checks. As mentioned in Data Description, we only removed failed generations and non-English generations for the current release, and excluded ~93K exact-duplicate notes before synthesis, which are listed in deduplicated_note_ids.csv for mapping purposes. Future releases of the dataset will consider additional filters as well as more advanced generation methodologies.


Release Notes

MIMIC-Ext-SynNotes v1.0.0

Version 1.0.0: Initial public release of the dataset. This release contains eight synthetic corpora totaling 13.7B words, generated by rephrasing all clinical notes from MIMIC-III v1.4 and MIMIC-IV-Note v2.2 with five open-weight LLMs at both note and chunk level.


Ethics

The collection of patient information and creation of the source databases were reviewed by the Institutional Review Board at the Beth Israel Deaconess Medical Center, which granted a waiver of informed consent and approved the data-sharing initiative [10,11]. This dataset is derived exclusively from those two deidentified databases — MIMIC-III Clinical Database v1.4 and MIMIC-IV-Note v2.2 — and no other patient data or clinical text was used at any stage.

Any name-like, date-like, or phone-like token appearing in the synthetic text is an LLM-generated artifact and does not correspond to a real identifier or individual; see Data Description for details.

All LLM inference and data-generation steps were performed in a secure environment consistent with the MIMIC Data Use Agreement, using local open-weight models on institutionally controlled infrastructure, and note text was never transmitted to any third-party or commercial LLM API. Where the accompanying study [14] used a closed-source model for fact-checking on a small sample, this followed PhysioNet's guidance on the responsible use of LLMs on MIMIC data.


Conflicts of Interest

The authors declare no conflicts of interest.


References

  1. Gunasekar S, Zhang Y, Aneja J, Mendes CCT, Del Giorno A, Gopi S, et al. Textbooks Are All You Need. arXiv [cs.CL]. 2023. Available from: http://arxiv.org/abs/2306.11644
  2. Kang F, Ardalani N, Kuchnik M, Emad Y, Elhoushi M, Sengupta S, et al. Demystifying synthetic data in LLM pre-training: A systematic study of scaling laws, benefits, and pitfalls. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. p. 10750–69. Available from: http://dx.doi.org/10.18653/v1/2025.emnlp-main.544
  3. Penedo G, Kydlíček H, Allal LB, Lozhkov A, Mitchell M, Raffel C, et al. The FineWeb datasets: decanting the web for the finest text data at scale. arXiv [cs.CL]. 2024. Available from: http://arxiv.org/abs/2406.17557
  4. Li J, Fang A, Smyrnis G, Ivgi M, Jordan M, Gadre SY, et al. DataComp-LM: In search of the next generation of training sets for language models. In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track [Internet]. 2024. Available from: https://openreview.net/forum?id=CNWdWn47IE
  5. Maini P, Seto S, Bai R, Grangier D, Zhang Y, Jaitly N. Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024. pp. 14044–14072. https://aclanthology.org/2024.acl-long.757
  6. Su D, Kong K, Lin Y, Jennings J, Norick B, Kliegl M, et al. Nemotron-CC: Transforming common crawl into a refined long-horizon pretraining dataset. In: Che W, Nabende J, Shutova E, Pilehvar MT, editors. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics; 2025. p. 2459–75. Available from: http://dx.doi.org/10.18653/v1/2025.acl-long.123
  7. Lehman E, Hernandez E, Mahajan D, Wulff J, Smith MJ, Ziegler Z, et al. Do We Still Need Clinical Language Models? In: Proceedings of the Conference on Health, Inference, and Learning. PMLR; 22 Jun--24 Jun 2023. p. 578–97. (Proceedings of Machine Learning Research; vol. 209). Available from: https://proceedings.mlr.press/v209/eric23a.html
  8. Christophe C, Raha T, Maslenkova S, Salman MU, Kanithi P, Pimentel MAF, et al. Beyond fine-tuning: Unleashing the potential of continuous pretraining for clinical LLMs. In: Findings of the Association for Computational Linguistics: EMNLP 2024. p. 10549–61. Available from: http://dx.doi.org/10.18653/v1/2024.findings-emnlp.618
  9. Tian Y, Gan R, Song Y, Zhang J, Zhang Y. ChiMed-GPT: A Chinese medical large language model with full training regime and better alignment to human preferences. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. p. 7156–73. Available from: http://dx.doi.org/10.18653/v1/2024.acl-long.386
  10. Johnson AEW, Pollard TJ, Shen L, Lehman L-WH, Feng M, Ghassemi M, et al. MIMIC-III, a freely accessible critical care database. Scientific data. 2016 May;3:160035. Available from: http://dx.doi.org/10.1038/sdata.2016.35
  11. Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, et al. MIMIC-IV, a freely accessible electronic health record dataset. Scientific data. 2023 Jan;10(1):1. Available from: http://dx.doi.org/10.1038/s41597-022-01899-x
  12. Kweon S, Kim J, Kim J, Im S, Cho E, Bae S, et al. Publicly shareable clinical large language model built on synthetic clinical notes. In: Findings of the Association for Computational Linguistics: ACL 2024. p. 5148–68. Available from: https://aclanthology.org/2024.findings-acl.305
  13. Liu J, Nguyen A. Rephrasing electronic health records for pretraining clinical language models. In: Proceedings of the 22nd Annual Workshop of the Australasian Language Technology Association. 2024. p. 164–72. Available from: https://aclanthology.org/2024.alta-1.13
  14. Liu J, Soni S, Nguyen A. Systematic evaluation of the quality of synthetic clinical notes rephrased by LLMs at million-note scale. In: Proceedings of the 25th Workshop on Biomedical Language Processing (BioNLP) 2026. p. 353–71. Available from: https://aclanthology.org/2026.bionlp-1.28
  15. Kwon W, Li Z, Zhuang S, Sheng Y, Zheng L, Yu CH, et al. Efficient memory management for large language model serving with PagedAttention. In: Proceedings of the 29th Symposium on Operating Systems Principles. p. 611–26. Available from: http://dx.doi.org/10.1145/3600006.3613165
  16. Joulin A, Grave E, Bojanowski P, Mikolov T. Bag of tricks for efficient text classification. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics. 2017. p. 427–31. Available from: https://aclanthology.org/E17-2068

Files