Dataset Overview
Definition:MRI Cardiac Imaging Datasetis provided byChangsha Langhui Information Technology Co., Ltd.A large-scale MRI-series imaging database built。Based on 100,000 Cases。Final settlement is based on the de-duplicated count of accepted CMR examinations / Studies。Qualified MRI (magnetic resonance imaging) examinations as the main body, in DICOM standard raw data format, retaining complete sequence information, spatial geometric parameters and metadata。Covering more than 12 core disease categories, from multi-center sources, with strict quality control; suitable for medical imaging AI model training and assisted-diagnosis capability building。
| Dataset Name | MRI Cardiac Imaging Dataset |
| Total Data Volume | 100,000 Cases。Final settlement is based on the de-duplicated count of accepted CMR examinations / Studies。(CMR Examination / Study Level) |
| Imaging Modality | MRI (Magnetic Resonance Imaging) |
| Scan Protocol | Multi-Plane Cine SSFP; T2/STIR or T2 Mapping (if available); Native/Post-Contrast T1 Mapping (if available); first embedding (if available); Late Gadlinium Employment / LGE (if available); Case-Contrast Flow (on demand). The protocol name must match the actual document... |
| Raw Data Format | De-identified raw DICOM with full hierarchy and metadata preserved |
| Core Conditions (P0) | Myocarditis, dilated cardiomyopathy, hypertrophic cardiomyopathy, ischemic cardiomyopathy / myocardial infarction, valvular disease, etc. |
| Extended Conditions (P1) | Cardiac amyloidosis, cardiac sarcoidosis / other infiltrative diseases, arrhythmogenic cardiomyopathy, congenital heart disease, etc. |
| Source Institution | It is recommended to cover 10–20 institutions, with a minimum of 10 in principle and no upper limit;The specific number of partner Grade-A tertiary hospitals is adjusted flexibly according to the delivered data volume |
| Quality Control Standards | Kappa ≥ 0.75, five-tier L1–L5 quality grading |
| Use Cases | AI model training, assisted diagnosis, radiomics research, algorithm validation |
Disease Distribution Overview
Delivery Specifications and Field Requirements
The dataset strictly follows unified delivery index requirements to ensure data quality and traceability。
Delivery & Counting:One row corresponds to one CMR examination / Study;Multiple sequences, phases, reconstructions, views or repeated exports within the same exam may not be counted separately。
Composition of Positive Cases:In disease-oriented subsets, clearly positive/abnormal cases are in principle no less than 90%;Normal heart / volunteer controls are grouped separately; contraindications to contrast or absence of LGE must be explicitly flagged。Findings such as "suspected", "considered", "cannot be excluded" and "post-treatment changes" must be retained separately from pathologically confirmed diagnoses and must not be forcibly merged into a definite diagnosis。
Coverage:Full coverage of the heart and the heart bag; functional parameters derived from the recording of the standard length axis, room movement, tissue characteristics and blood flow sequences by clinical problem must be traced back to the original sequence.
Raw Data Requirements:De-identified raw DICOM with key metadata preserved, including Study/Series/Instance hierarchy, spatial geometry, magnetic field strength, coil, sequence name, TR/TE/TI, slice thickness, b-value, contrast agent and dynamic phases。Screenshots, film photographs, report PDFs or key-frame collages may not replace the agreed-upon original images/videos/slides。
Field Group Coverage:Batch source and location, anonymized patient information, exam primary key, modality and body part, protocol and device, specialized fields, file hierarchy, report text, diagnostic labels, clinical context, related fields, quality and disposition
Structured Text:At minimum, retain the exam name, exam findings, diagnostic conclusion/impression, primary diagnosis, fine-grained labels, and negative and uncertain semantics;When original reports are available, they must be linked case by case。
Data Volume Growth Trend
2026 Frontier AI Research Progress
Domain Review:In 2025–2026, MRI-series AI entered the foundation model era, with large-scale pretrained models achieving breakthrough progress in multi-task generalization and few-shot learning。Top journals such as Nature, Nature Medicine and The Lancet have published multiple clinical-grade validation studies, driving MRI-series AI from the laboratory to clinical deployment。
Med-PaLM M Multimodal Medical Foundation Model
Nature Biomedical Engineering, 2025The general-purpose medical AI model from Google DeepMind achieves state-of-the-art performance across 14 modalities and 114 tasks, including MRI image interpretation, report generation and clinical Q&A。
MIMIC-CXR Foundation Model Progress
Nature Medicine, 2025The foundation model built on millions of chest X-rays demonstrates strong generalization in chest MRI/CT transfer learning through self-supervised pretraining, improving AUC by 8–12%。
nnU-Net v2: A New Standard for Medical Image Segmentation
Nature Methods, 2025The second-generation nnU-Net framework refreshed records across MRI multi-organ segmentation benchmarks and supports a 3D fully convolutional Transformer hybrid architecture。
Self-Supervised MRI Representation Learning
IEEE TMI, 2026The masked autoencoder-based MRI pretraining strategy improves over the supervised baseline by more than 15% in few-shot brain/spine/joint segmentation tasks。
Annotation Workflow & Quality Control
Image Acquisition & De-identification
Standardized acquisition workflow: complete data is exported directly from the devices, and patient identifiers (PHI) are removed before storage to ensure data compliance。
Initial Annotation (Specialist Physician)
Attending physicians annotate case by case against the standard, including lesion localization, morphological description, disease diagnostic labels and specialized indicators。
Review (Associate Chief Physician or Above)
Experts with the title of Associate Chief Physician or above review each initial annotation result item by item, correcting erroneous annotations and supplementing missing dimensions to ensure annotation accuracy。
Consistency Assessment
10% of samples are randomly selected and independently annotated by 3 physicians to calculate Fleiss' Kappa; below 0.Dimensions scored 75 are flagged for rework。
Quality Acceptance Checklist
License and Usage Agreement
Academic Research License
For universities and research institutions, supporting academic research related to medical AI。After signing the agreement, a de-identified data subset is provided; the source must be acknowledged:Langhui Technology DataAssetsAPI。
Commercial License
For medical device companies and AI diagnostics companies, supporting medical AI product R&D and medical device registration。Provides full datasets + custom annotation + incremental update services。
Data Compliance Statement
All images come from legally authorized sources; PHI fields are fully de-identified and contain no information that can directly identify an individual, in compliance with the Personal Information Protection Law and the Data Security Law。
Customized services
Supports extended requirements such as adding specific disease types, multimodal annotation expansion, and custom training of AI-assisted diagnostic models。
Quick Facts
AI Frontier Research
In 2025–2026, MRI-series AI entered the foundation model stage; multimodal self-supervised learning performed excellently on multiple benchmark tasks, and several top-journal papers advanced clinical-grade AI diagnosis。