One-million-degree, full-scale management data covering hypertension, diabetes, coronary heart disease, slow lung resistance, five major chronic diseases in the brain, with electronic medical records, follow-up records, drug dependence and multi-dimensional dimensions of lifestyle
| Basic information | IID, Age Group, Gender, BMI, Years of Medical Experience, Year of First Clinic (reserved only) |
|---|---|
| ICD-10 code, date of diagnosis (year only), disease chronology/classification, list of complications | |
| Medication Records | Drugs generic name, dosage, frequency, start date, cut-off date, subject-to-use rating |
| Key indicators such as blood pressure/smalt/slippera/hepatic kidney function/urea routine and test date | |
| Follow-up records | Date of follow-up visit, change of symptoms, complications, adjustment of treatment programme, self-reported end of patient |
| Lifestyle | Smoking status, alcohol consumption, frequency of exercise, diet (semi-quantitative) |
| Standardization of terminology | Diagnosis using ICD-10 code, drugs using ATC classification, and inspection projects using LOINC mapping |
|---|---|
| Data integrity | Core field (diagnostic + drug) completeness rate ≥ 95% |
| Time Span | Longitudinal follow-up data covering no less than three years to support disease progress modelling |
| De-sensitization criteria | HIPAA Safeports Code, removing all 18 categories of protected health information identifiers |
Training in risk prediction models for chronic disease complications based on vertical follow-up data, identifying high-risk populations in advance and triggering interventions.
Analysis of patient drug behaviour patterns, prediction of the risk of de-addiction, and support for the construction of individualized drug management programmes.
Training in the trajectory of chronic diseases using time-series data to support clinical decision-making and treatment path optimization.
Integration of multi-dimensional data to train health risk assessment and intervention referral systems, enabling intelligent slow disease management platforms.
All data has been thoroughly de-identified, with all information that could directly or indirectly identify an individual removed; only parameters and labels relevant to medical research are retained。
| De-identified Fields | All personal identifications, including name, ID number, telephone number, address, hospital number, medical insurance number, etc. |
|---|---|
| Reserved Fields | Diagnosis codes, drug records, inspection indicators, follow-up data, lifestyle information |