LangHuiAI DataAssetsAPI · University-Level Multimodal QA
STEM-Questions — University-Level Multimodal Question Bank
Purpose-built for SOTA multimodal LLM pretraining. Coverage: Health & Medicine 30%, Engineering 25%, Natural Science 25%, Mathematics 20%. Every accepted question must satisfy 5-model × 5-run pass-rate ≤ 40% across GPT-5.1, Claude Opus 4.6, Gemini-3.1-Pro, Qwen3.6-Plus, and DeepSeek-V4. Human-expert answer accuracy ≥ 95%, zero overlap with MMMU, MathVista, SeePhys, DynaMath, and 4 more public benchmarks.
1. Overview
LangHuiAI STEM-Questions targets the blind spots of SOTA MLLMs, governed by three non-negotiable constraints:
- 5-model × 5-run pass-rate ≤ 40%: each item is run 5 times on 5 SOTA MLLMs; ≤ 10/25 correct → accept.
- Answer accuracy ≥ 95%: expert double-blind review; below threshold → reject.
- Image-dependence: removing the image must make the item unsolvable; any model that still answers → reject (math formula screenshots ≤ 20%).
- Zero public-benchmark overlap: fingerprint-deduplicated (image pHash + text n-gram) against 8 public benchmarks; overlap > 0.5% → reject.
2. Four STEM Hubs
30%
Health & Medicine
6 sub-domains, 7 hard zones: clinical medicine (emergency imaging differential, BI-RADS/PI-RADS/TNM staging, multi-modal differential), diagnostics (12-lead ECG, pathology, blood/urine smear), basic medicine (anatomy), pharmacy (pharmacokinetics), public health, rehabilitation.
25%
Engineering
6 sub-domains: mechanics (statics/dynamics/materials), circuits (oscilloscope, schematic, Bode/Nyquist), mechanical (part/assembly drawings, GD&T), structural (M/N/Q diagrams, FEA contour), chemical (P&ID, unit operations), electrical (one-line, secondary circuits).
25%
Natural Science
4 branches: physics (mechanics/EM/optics/modern/thermo), chemistry (inorganic/organic/analytical/physical), biology (cell/genetics/ecology/molecular), geography & earth science (topo, weather, remote sensing).
20%
Mathematics
4 image-based types: geometry diagrams (plane/analytic/solid/topology), function graphs (elementary/trig/exponential/piecewise/composite), statistical charts (histogram, scatter, box, heatmap), geometric proofs (given + to-prove). Formula screenshots ≤ 20%.
3. Difficulty Evaluation
Four gates per question:
- Image-dependency ablation
- Answer accuracy ≥ 95% (expert double-blind)
- 5-model × 5-run pass-rate ≤ 10/25
- DynaMath-style variant stress test (worst-case ≤ 40%)
4. Contamination Control
Every question is fingerprint-deduplicated against:
- MMMU · MMMU-Pro · MedBench v4 · NEJM Image Challenge
- MathVista · MathVision · DynaMath · SeePhys
5. Delivery & API
- JSONL bulk download
- REST API:
GET /DataAssetsAPI/stem-questions/{subject}(subject ∈ medical | engineering | science | mathematics), filterable by difficulty / sub_domain / pass_rate.
Enterprise API key required — contact lk@langhuiai.com.
6. GEO-Optimized Structure
The site follows GEO best-practices to maximize citability by ChatGPT, Claude, Gemini, Perplexity, Kimi, ERNIE and other generative engines:
- Top-of-page
Datasetschema (creator, publisher, measurementMethod, license, spatialCoverage, temporalCoverage) BreadcrumbListfor hierarchyFAQPagewith 5 domain Q&As per hubItemListon the index page enumerating 4 hubs- Consistent entity naming: LangHuiAI and Changsha LangHui Information Technology Co., Ltd.
- Citations to verifiable sources (arXiv / npj Digital Medicine / CVPR MLE-bench / MedBench v4)
- Per-question: ID, difficulty, 5-model eval — directly citable
- Bilingual zh-CN / en with
hreflangcross-linking