Langhui
  1. Home Page
  2. /
  3. Data assets
  4. /
  5. Longitudinal Multimodal Medical Record Dataset
NEW Added Data Set Healthcare Longitudinal Multimodal Multimodal

Longitudinal Multimodal Medical Record Dataset

5000+ longitudinal total mosaic data · 3 years time span ⁇ 8 + 7 medical image mosaics ICD-10-CN coded map

High-quality longitudinal multi-modular clinical data sets for training in time series and reasoning for large medical models. Each case integrates structured medical records, clinical documents, medical images, testing, drug records, pathological reports, and quantitative scales, scoring seven large data patterns, with complete clinical tracks in chronological order at each point of time for the same patient.

5000+
Total number of eligible cases
≥3年
Time-Span Requirements
7Category
Medical image modelling
8Section
Priority Clinical Section

Dataset Overview

Dataset NameLongitudinal Multimodal Medical Record Dataset(Longitudinal Multimodal Electronic Health Records Dataset)
Total ScaleNot less than 5,000 例Qualified cases, covering 8 major focus units + supplementary specialist strains
Time Span3-year vertical check-up record per case, at least 2 visits/tests/treatment/follow-up at different points
Data time frame2016-2025 (priority for patients with complete medical records and follow-up data)
Department CoverageOncology + ICU/psychopathology/Ears, nose and throat/gynaecology and gynaecology/ rheumatology
数据模态Structured medical records — medical images — medical tests — medical records — pathology reports — quantitative scores
Imaging ModalityCT MIS PT-CT Ultrasity ENDRIGHT X-L/DR Pathology WSI
Encoding StandardICD-10-CN Diagnostic Code Map, with original ICD code, code version and map rules
Quality standardsKey field filling rate 95% % % total scrutinisation 10% video specialisation % double dissensitisation
Compliance requirementsThe Personal Information Protection Act, the Data Security Act, complies with the full amount of dissensitisation and data not to leave the country
Use CasesThe medical megamodel time series of the training • Vertical efficacy prediction • clinical trajectory modelling • polymodular integration studies • disease progress prediction
Supply FormatHard Drive · Cloud Drive · API · Data Infrastructure
ProviderChangsha Langhui Information Technology Co., Ltd.

Long-term sequence structure: vertical treatment track restored

The core difference value of this data set is thatLong-time sequence relevance– Not only does each eligible case meet the three-year time horizon, but more importantly, there is a consistent internal correlation between diagnosis, consultation, testing, examination, medication, paperwork, images, pathology, stats, and end-of-life information within the same dissensitized patient, which can be restored in chronological order to the main clinical trajectory.Time series reasoning(temporal training)

01

Vertical Time Range

Chronic or progressive diseases such as COPD, IBD, cirrhosis, CKD, type 2 diabetes, Parkinson's disease, Alzheimer's disease, etc. are prioritized to provide longer time series data.

02

More Time-Cast

The video reports refer to the previous data of the "comparable pre-film" and "relation" over the previous one, and provides, to the extent possible, corresponding checks and reports, guaranteeing continuous traceability of the main clinical tracks.

03

Internal traceability chain

The same patient can be traced back to the same patient’s attendance.

Long-term sequence track (single patient)

T0 * Baseline

2019.03 • First diagnosis · Hospital admission + Diagnosis (ICD-10:C34.1) + CT chest + pathological biopsy + baseline

T1 Treatment

2019.05 • Post-operative chemotherapy cycle 1 • drug record + test (blood routine/pastal function) + CT review + pathology record

T2 • Follow-up visits

2020.03 • One year follow-up • CT precomparison + oncology marker + physical scoring (ECOG) + follow-up document

T3 Review

2021.06 · 2-year review · MRI + PET-CT + pathology + drug-adjusted record + quality of life Quality of life table

T4 . End of the story

2022.09 · Last follow-up · End of story + Image assessment + Last drug + Total survival record

Multimodular Data Dimension System

7 data mosaics are integrated for each eligible case, with the core acceptance standard beingInternal relevance of multiple types of data under the same patient, same visit, same examination or same treatment. All the modes are verifiable links of connection by the number of the dissensitive patient, the number of the consultation, the number of the examination, the number of the report, the index of the document.

📋

StructuredEHR / Records

The diagnosis record (ICD-10-CN), the clinical/hospital record, the test results, the examination record, the drug log, the surgical/operational record.

📝

Original clinical document

The full text is guaranteed without serious interruption.

🔬

Medical Images

CT/MRI/PET-CT (Dinox DIC), ultrasound/endoscope (original format), X-ray/DR (DICOM), pathology WSI (SVS/NDPI/TIFF/KFB).

💊

Medication Records

The drug’s name, dose, frequency, route of delivery, starting time, and drug-based adjustment record.

🧬

Pathology Report

Pathological diagnostic reports, immuno-group results, molecular/genetic tests, and TNM phases.

📊

Measurement score

Specialized assessment tables for NIHSS, mRS, NYHA, LVEF, ECOG, APACHE II, SOFA, DAS28, SLEDAI, etc., support vertical efficacy evaluation.

📈

Inspection

You can also check the results of tests on blood, biochemicals, coagulation, tumor markers, microorganisms, genetics, etc., and functional tests.

🏥

Follow-up outcome

Follow-up, survival, relapse/transfer, and complications. Support training in disease progress modelling and prognosis models.

数据模态 Delivery Format Requirements Association Keys
Structured Data Diagnosis (ICD-10-CN), consultation, testing, examination, medication, surgical records CSV / Excel / Database patient_id + encounter_id
Clinical documents Hospital admissions, discharges, medical records, transfer records, full surgical records Text / PDF patient_id + encounter_id
CTImaging Thin + General Layer Thick Sequences, preferred 512 x 512 Matrix, Retain Phase Information De-identified DICOM study_id + series_id
MRIImaging T1WI/T2WI/FLAIR/DWI/ADC/Strengthened sequences, retention of sequence description and scanning parameters De-identified DICOM study_id + series_id
PET-CT PET corresponds to the C.T. sequence, retaining integration and SUV metabolic information De-identified DICOM study_id + series_id
Ultrasound/oversight Static images + dynamic video, association inspection reports and key measurements DICOM/ Original Export study_id + report_id
PathologyWSI Full slice digital image, retention multiplier/scale/chromosomal type/slice number SVS / NDPI / TIFF / KFB study_id + report_id
Medication Records Name, dose, frequency, route of delivery, starting time CSV/ Database Table patient_id + encounter_id

Distribution of sections and disease spectrum

The number of units and diseases is distributed in a targeted manner, and the final delivery structure is based on the mutually confirmed delivery list.

Section/category Target ratio Number of recommendations Focused disease spectrum Baseline information requirements
Oncology About 35-38 per cent 1,750-1,900 Lung cancer, breast cancer, colon cancer, liver cancer, stomach cancer, edible cancer, carcinoma of the neck, pancreatic cancer, ovarian cancer, etc. Image reports, periodic reports, pathological reports; molecular/immunological indicators available on a realistic basis
Cardiovascular About 20-22% 1,000-1,100 Coronary heart disease/acute heart infarction, post-PCI, heart failure, room tremors, etc. Coronary artery, heart ultrasound, electrocardiogram, heart myase spectrum, LVEF, NYHA
Neurology About 13-15% 650-750 Illustrative, haemorrhagic, Parkinson's, cognitive disorders/Atzheimer's, etc. Head CT/MRI, description of the disease, NIHSS, mRS, cognitive/motorized table
Respiratory Section Approximately 7-8 per cent 350-400 COPD, bronchial asthma, community access to pneumonia, bronchial expansion, etc. Lung function, chest image, grade or symptoms control assessment
Indigestion Section 约6%-7% 300-350 Hepatic cirrhosis, IBD, digestive ulcer, GERD, etc. Stomach/intestinal/image reports, pathological reports, liver function ratings, disease activity ratings
Nephrology 约6%-7% 300-350 CKD, Diabetes Nephrosis, End-of-life kidney disease, dialysis, post-transplant follow-up and complications, etc. eGFR, UACR, urine tests, kidney function, kidney pathology, dialysis records
Endocrinology About 5-6 per cent 250-300 Type 2 diabetes mellitus and complications, thyroid glands/functional abnormalities/tumours, osteoporosis, etc. HbA1c, blood sugar records, diagnosis of complications, thyroid function/ultrasound, bone density
Blood. About 3-4% 150-200 Leukemia, lymphoma, multiple osteoporosis, etc. Osteomymystalgia/live, flow, FISH, stratification/scoring, seroprotein electron swim
Other specialized supplements ≤3% ≤150 ICU (suspensive/ARDS), psychiatric, oral, ophthalmic, ear, nose and throat, gynaecology and obstetrics, rheumatism immunisation Implementation according to the confirmation list by the parties
Total 100% ≥5,000 8 key sections + 7 additional specialist

ICD coding and diagnostic distribution

Encoding retention requirements

If you need to map to ICD-10-CN, provide map rules, map sources and map pre- and post-magnification fields.

Statistical delivery requirements

The results can be summarized by section, disease spectrum, ICD code, diagnostic name, number of cases, time span, and data model coverage.

Medical Image Technology Code

Image data are used for AI training, giving priority to the delivery of thin, raw resolution and sequence complete data. The project covers multi-pathological, multi-dimensional, multi-historic data, without using a single layer of thick thresholds as a condition for core access. DIOCOM-type images retain key metadata and spatial positioning information necessary for AI training.

ImagingType Delivery Format Key technical requirements
CT De-identified DICOM Priority thin layer sequences and original layer thickness/pixel spacing; same examination thin layer + conventional layer thickness delivered simultaneously to the extent possible; matrix thallium 512 x 512; enhanced examination of retention period phase information
MRI De-identified DICOM Retain major diagnostic sequences such as T1WI/T2WI/FLAIR/DWI/ADC/enhanced; serial name, scanning location, layer thickness, pixel spacing identifiable
PET-CT De-identified DICOM PET corresponds to the CT sequence, retaining integration; report/metadata reflects metabolic information such as SUV
X-line/DR De-identified DICOM Retain the place of delivery, check the part, pixel size and report association; may not be replaced by low-resolution preview
Ultrasound DICOM/Video/Preliminary Export Static images + dynamic video; linkage inspection reports, key measurements, parts and conclusions retained
内镜 Original Image/Video/system Export Maintain inspection sites, time, reports/records, key images/videos; establish pathological correspondence
PathologyWSI SVS/NDPI/TIFF/KFB Retain multiples, scales, scan levels, dye types, slice numbers; priority H&E and diagnostic-related slices

DICOM key field retention requirement

__KEP_core_mark

StudyInstanceUID, SeriesInstanceUID, SOPInstanceUID, AccessionNumber

Space positioning

ImagePositionPatient, ImageOrientationPatient, SliceLocation, FrameOfReferenceUID

Pixels

Rows, Columns, PixelSpacing, SliceThickness, SpacingBetweenSlices

扫描参数

KVP, Exposure/mAs, ConvolutionKernel, Pitch, Manufacturer, ModelName

⁇ Desensitization does not remove key fields that affect the use of 3D reconstruction, serial recognition, spatial positioning and training.

Data Sample

Example 1 . Long-time indexing structure

{
  "patient_id": "LH_PT_2024_003172",
  "demographics": {
    "gender": "M",
    "age_at_baseline": 58,
    "age_unit": "year"
  },
  "primary_diagnosis": {
    "icd_code_original": "C34.1",
    "icd_version": "ICD-10",
    "icd_code_mapped": "C34.1",
    "diagnosis_name": "肺上叶恶性肿瘤",
    "department": "肿瘤科",
    "disease_spectrum": "实体瘤-肺癌"
  },
  "temporal_trajectory": {
    "time_span_years": 3.5,
    "encounter_count": 8,
    "first_record_date": "2019-03-15",
    "last_record_date": "2022-09-20",
    "encounters": [
      {
        "encounter_id": "ENC_001",
        "date": "2019-03-15",
        "type": "inpatient",
        "department": "肿瘤科",
        "phase": "baseline_diagnosis",
        "linked_data": {
          "clinical_notes": ["入院记录", "首次病程", "出院小结"],
          "imaging": [
            {"study_id": "IMG_CT_001", "modality": "CT", "body_part": "CHEST", "report_id": "RPT_001"}
          ],
          "pathology": [
            {"report_id": "PATH_001", "type": "biopsy", "finding": "非小细胞肺癌,腺癌"}
          ],
          "lab_results": ["血常规", "生化", "肿瘤标志物"],
          "medications": [],
          "scales": [{"name": "ECOG", "score": 1}]
        }
      },
      {
        "encounter_id": "ENC_004",
        "date": "2020-03-10",
        "type": "outpatient",
        "phase": "follow_up_1y",
        "linked_data": {
          "imaging": [
            {"study_id": "IMG_CT_004", "modality": "CT", "body_part": "CHEST",
             "comparison": "对比前片IMG_CT_003", "report_id": "RPT_004"}
          ],
          "lab_results": ["肿瘤标志物"],
          "scales": [{"name": "ECOG", "score": 0}]
        }
      }
    ]
  }
}

Example 2 . Multimodular Image Link Index

Other Organiser
"Patient id": "LH PT 2024 00317."
"Encounter id": "ENC 006",
"Study id."
"Study datetime": "2021-06-18T09:30:00",
"modality": "MRI,"
"body part": "Brain,"
"Study description": "MRI Sweeping of Head + Enhancement,"
"series":
Other Organiser
"series id": "SER 006 01",
"series description": "T1WI,"
"slice thickness": 5.0,
"pixel spacing": [5, 0.5],
"rows": 512,
"Columns": 512,
"file count": 24,
"file path": "patient 00317/enc 006/mri/ser 01/"
{\cHFFFFFF}{\cH00FFFF}
Other Organiser
"series id": "SER 006 02",
"series description": "T2WI FLAIR,"
"slice thickness": 5.0,
"pixel spacing": [5, 0.5],
"rows": 512,
"Columns": 512,
"file count": 24,
"file path": "patient 00317/enc 006/mri/ser 02/"
{\cHFFFFFF}{\cH00FFFF}
Other Organiser
"series id": "SER 006 05",
"series description": "T1WI+C"
"slice thickness": 5.0,
"pixel spacing": [5, 0.5],
"rows": 512,
"Columns": 512,
"file count": 24,
"file path": "patient 00317/enc 006/mri/ser 05/",
"contrast case": "post contrast"
♪ I'm sorry ♪
I don't know.
"link report": {
"report id": "RPT 006",
"findings": "The right side of the frontal lobe is visible, no significant change over the previous one..."
"impression": "Recommend continued follow-up when considering transfer stabilization."
"report date": "2021-06-18"
{\cHFFFFFF}{\cH00FFFF}
"link encounter":
"Encounter id": "ENC 006",
"date": "2021-06-18",
"department": "oncology",
"phase": "follow up 2y"
{\cHFFFFFF}{\cH00FFFF}
"quality note":
"dicom metadata preserved": true,
"spatial position present": true,
"Burned in annotation": "NO",
"Description status": "passed"
♪ I'm sorry ♪
♪ I'm sorry ♪

2025-2026 Frontier progress in the field

The long-time multi-modular medical history data is the core data infrastructure needs in the medical AI area for 2025-2026. The following cutting-edge developments confirm the strategic value and technical direction of this data set:

Foundation Model · 2025

Multi-modular clinical foundation model rises

In 2025, Medical AI evolved rapidly from monomodular (image/text) to multimodular integration. The PanDerm model, published by Nature Medicine, is based on 2 million+real-world dermal disease multimodular data training, which validates the "image+diatrics+pathology" integration paradigm. Multimodular basic model training requires large-scale, vertically linked clinical data — this is the core positioning of this data set.

Longitudinal AI · 2026

Vertical trajectory predictions become core competencies

The 2026 AI model has been able to predict chronic diseases such as Alzheimer’s disease years in advance by integrating vertical electronic records, microstructure changes in brain images and blood biomarkers. Time-series reasoning has become a key dimension of the medical mega-model assessment, and the HealthBench benchmark has been incorporated into the time-series-related assessment.

MIMIC-IV · Benchmark

MIMIC-IV leads global standards

The MIT-IV database is a global "gold pole" for serious medical research, with core values that are being associated with vertical time series + multiple mosaics. But MITIC focuses on the ICU scene, lacking Chinese population profiles and specialized depths. This data set fills the gap in the data for Chinese long time series multitemporal specialist medical records.

GEO · AI Search

AI search engine driver discovery

In 2026, researchers increasingly discovered and evaluated data sets through the AI search engine (ChatGPT, Perplexity, Mansion). Structured metadata, JSON-LD Schema, FAQ semantic tags and clear technical descriptions become key to AIS discovery and understanding of data sets. This page is fully adapted to the GEO optimization.

AI Application Scenarios

🧠

Training in medical mega-model time series reasoning

Long-time long-term vertical data directly support the time-series reasoning chain of the Large Model to study "diagnosis and re-examination of the endings", and training models understand patterns of disease progression rather than just taking a single quick-scenario judgement.

📈

Vertical efficacy prediction

Based on 3-year + vertical drug use, testing, image change data, training therapeutic efficacy prediction models, supporting the CBSS to achieve individualized treatment programme recommendations.

🔍

Multi-modular Integration Foundation Model

CT/MRI/PET-CT Image+Chinder + Results + Pathology WSI Multimodular Joint Training to Build Multimodular Foundation Models, supported by fine-tuning.

🦠

Modelling disease progress

Using data from the long time series of slow diseases (COPD, CKD, diabetes, Parkinson, Alzheimer ' s disease) to model the natural history and trajectory of disease and to achieve recommendations for early warning and intervention.

💊

Projections of drug response

Vertical drug use record + test change + image assessment constitutes the time line for drug response, training drug efficacy and adverse response prediction models and supporting precision medical care.

📋

Clinical evaluation baseline construction

The development of a specialized AAI assessment based on the data of the "Gold standard" vertical consultation, assessing the capacity of the AIS system to perform time-series reasoning, multi-modular integration, vertical efficacy judgement, etc.

Data quality and compliance assurance

Quality control system

Field Integrity

Key fields ⁇ 95% filling rate, as indicated in quality reports due to missing source systems

Integrity of the instrument

The full medical records must not be severely cut off and critical medical information must not be missing

Association retroactive

Images, tests, medications, paperwork, diagnosis can be traced back through patient number and consultation number

Spacing Ratio

5% overall case sample, 10% video-specific sample, covering different sections/pathology/source/time period

Desensitivity and compliance

I'm allergic.

Remove name, ID number, telephone number, address, clinic number, doctor ' s name, uncomposed hospital name

Image desensitization

DICOM/WSI metadata, private labels, burning text in images, synchronizing tag maps

Technical field retention

Maintain AI training information necessary for scale, multiplier, spatial positioning, sequence recognition

LawRegCompliance

Data are prohibited from leaving the country in compliance with the Personal Information Protection Act, the Data Security Act

Complete list of deliverables

✓ Structured data files/database tables (CSV/Excel/Database backup)
✓ Documents/text of medical records (full text, relevant)
✓ Medical and image files (de-sensitive original files)
✓ Data dictionary (field name/mean/type/source)
✓ Document index list (pathways/patients/diagnostics/test numbers)
✓ List of diseases and cases (number/pathology/section/time span)
✓ ICD Code and Diagnostic Distribution Statistics (original + map)
✓ Batch statistical tables (section/pathology/modular/temporal/image distribution)
✓ Quality review reports (missing/duplicating/unusual/association/image quality)
✓ De-sensitization statement (scope/rule/reserved field)
✓ Hospital sources and authorization instructions (separately examined)

Common problems

How does the long-temporal multi-modular patient data set differ from the general electronic medical data set?

This data set requires that each case be of three years duration and that it contain a record of visits at least two or more different points of time, and that the information on diagnosis, testing, examination, medication, paperwork, images and end-effects within the same patient be restored in chronological order to the main clinical trajectory. The normal electronic medical history data set is usually a single-diagnostic snapshot, lacking vertical time-series linkages and not being able to support training in time-series reasoning models.

What medical image models do data sets support?

Supports seven types of medical image modelling for CT, MRI, PET-CT, ultrasound, endoscopy, X-ray/DR, pathological whole-slice digital images (WSI). DICOM-type images retain the necessary metadata for training in _KEEP_core_labels and spatial positioning (ImagePositint/ImageOrganizationPatient), sequence parameters (story thickness/pixel spacing/scan parameters).

How do data sets guarantee privacy compliance?

The full amount of data is automated and manually de-sensitized, removing identifiable information such as the patient’s name, ID number, telephone number, etc., while retaining original values of age to guarantee analytical value. DIOCOM and WSI data are processed in a synchronized manner, private labels, and image burning. Data are prohibited from leaving the country in compliance with the requirements of the Personal Information Protection Act, the Data Security Act.

Are the video data containing the disease-proof signs?

The project does not require the marking of results such as a frame, profile, manual classification label, etc., unless the parties agree otherwise in writing. The focus of video data delivery is on originality and internal relevance - image documentation, image reports, examination records, patient numbers, and patient numbers, and should correspond to each other and be included in the long-term treatment chain of the same patient.

Which AI model training scenarios are suitable for the data set?

The core scenario includes: 1 training in the time sequence of a large medical model (learning the diagnostic treatment review of the causal chain of the endings and consequences); 2 basic training in multimodular integration models (image+ clerical+ testing+ pathology combination); 3 modelling of vertical efficacy predictions and disease progression; 4 drug response predictions; 5-pharmaceutical AI benchmarking. ICD-10-CN coding maps and standardized data dictionarys of the data set support direct access to the mainstream training framework.

How to obtain a complete catalogue of data sets and quotations?

Please contact the data expert team at Chang Longway Information Technology. 137-5502-0164, or visits www.langhuiai.com/Contact The programme is flexible and tailored by section, disease, pattern, time span.

A complete catalogue of long-time sequenced multi-modular patient data sets

5000+ longitudinal total simulator data 8 Big Section 7 Type image simulation ICD-10-CN code 10 Compliance commercial

• Medical AI Data Hole Service

Compliant Data · Customized Delivery · Professional Technical Support