Langhui
LH
Langhui AI

Million-Scale Disease Management Dataset

One-million-degree, full-scale management data covering hypertension, diabetes, coronary heart disease, slow lung resistance, five major chronic diseases in the brain, with electronic medical records, follow-up records, drug dependence and multi-dimensional dimensions of lifestyle

Chronic disease management Electronic Medical Records Follow-up Million-degree records

1,000,000+
Patient management records
5
Core chronic diseases
200+
Structured Fields
3 Years+
Follow-up cycle coverage

Disease coverage

High blood pressure.

Blood pressure trend Damage to target organs Drug use programme Complications

2

Blood sugar surveillance HbA1c Insulin programme Complication screening

Coronary heart disease

EKG Coronary CTA After the rack. Blood resin management

Chronic obstructive pulmonary disease

Lung function GoLD Level Acute intensification of history Inhalation treatment

In the head.

NIHSS Rating Rehabilitation assessment Secondary prevention Image Followup

Data field specification

Basic informationIID, Age Group, Gender, BMI, Years of Medical Experience, Year of First Clinic (reserved only)
ICD-10 code, date of diagnosis (year only), disease chronology/classification, list of complications
Medication RecordsDrugs generic name, dosage, frequency, start date, cut-off date, subject-to-use rating
Key indicators such as blood pressure/smalt/slippera/hepatic kidney function/urea routine and test date
Follow-up recordsDate of follow-up visit, change of symptoms, complications, adjustment of treatment programme, self-reported end of patient
LifestyleSmoking status, alcohol consumption, frequency of exercise, diet (semi-quantitative)

Data quality and standardization

Standardization of terminologyDiagnosis using ICD-10 code, drugs using ATC classification, and inspection projects using LOINC mapping
Data integrityCore field (diagnostic + drug) completeness rate ≥ 95%
Time SpanLongitudinal follow-up data covering no less than three years to support disease progress modelling
De-sensitization criteriaHIPAA Safeports Code, removing all 18 categories of protected health information identifiers

AI Training Application Scenarios

Disease risk prediction

Training in risk prediction models for chronic disease complications based on vertical follow-up data, identifying high-risk populations in advance and triggering interventions.

Drug dependence analysis

Analysis of patient drug behaviour patterns, prediction of the risk of de-addiction, and support for the construction of individualized drug management programmes.

Modelling disease progress

Training in the trajectory of chronic diseases using time-series data to support clinical decision-making and treatment path optimization.

Personalized health management

Integration of multi-dimensional data to train health risk assessment and intervention referral systems, enabling intelligent slow disease management platforms.

Data Security and De-identification

All data has been thoroughly de-identified, with all information that could directly or indirectly identify an individual removed; only parameters and labels relevant to medical research are retained。

De-identified FieldsAll personal identifications, including name, ID number, telephone number, address, hospital number, medical insurance number, etc.
Reserved FieldsDiagnosis codes, drug records, inspection indicators, follow-up data, lifestyle information