ANUJ
DATA×AI×ENGINEERING
INITIALIZING SYSTEM
ANUJ MUNDU
PROJECT 08•DATA SCIENCE•DATA SCIENCE

EndoGuard CDSS™ — Clinical Diabetes Risk Suite

FHIR-native Clinical Decision Support System with Tabular Deep MLP, Stacked Ensembles & HL7 LOINC

PythonTabular Deep MLPCatBoostXGBoostHL7 FHIR R4SHAPScikit-LearnHealthcare AI
CV Accuracy
95.28%
10-Fold Stratified Cross-Validation
ROC-AUC Score
0.9810
Tabular Deep Neural Network
Interoperability
HL7 FHIR R4
LOINC clinical codes integration
Clinical Cohort
2,500 Records
Multi-center harmonized patient records
// INTERACTIVE SYSTEM TELEMETRY & DIAGNOSTIC LAB
INTERACTIVE DECISION THRESHOLD CALIBRATOR (τ)

Slide the classification cut-off threshold to evaluate precision vs recall trade-offs and net ROI.

OPTIMAL τ:0.42 (Brier Calibrated)
DECISION BOUNDARY (THRESHOLD τ):0.42
0.15 (High Sensitivity / Catch All)0.80 (High Specificity / Conservative)
Precision
80.8%
Targeting accuracy
Recall
81.2%
Churners captured
F1 Score
81%
Harmonic mean
Net Monthly Value
+$84,723
Revenue preserved
At τ = 0.42: Caught 349 of 430 churners; 147 false alarm outreaches.
Model: XGBoost + Isotonic CalibratedCV

01 // SYSTEM OVERVIEW

EndoGuard CDSS™ is an enterprise-grade, FHIR-native Clinical Decision Support System designed for early-stage diabetes detection, risk stratification, and clinician-in-the-loop (HITL) triaging. It harmonizes 2,500 clinical patient records and implements Tabular Deep Neural Networks (MLP 128-64) and Stacking ensembles with full HL7 FHIR R4 standard compliance.

02 // THE PROBLEM & ENGINEERING SIGNIFICANCE

The Core Challenge

Undiagnosed Type 2 diabetes leads to severe macrovascular and microvascular complications. Traditional risk scoring is fragmented and detached from Electronic Health Record (EHR) systems.

Why This Matters

Early intervention through lifestyle and pharmacotherapy can reverse prediabetes, but clinicians need automated, explainable alerts integrated directly into EHR workflows.

Key Constraints:
  • •Harmonizing heterogeneous patient records with differing laboratory biomarker standards.
  • •Achieving ultra-high sensitivity while maintaining specificity to prevent clinical alert fatigue.
  • •Complying with healthcare data standards (HL7 FHIR R4, LOINC terminology).

03 // DATA PIPELINE & PREPROCESSING

Input Format: HL7 FHIR R4 Patient and Observation JSON Bundles / Tabular Clinical Laboratory RecordsSample Volume: 2,500 harmonized multi-center clinical cohort records
Transformation Steps:
  • 25 engineered clinical biomarkers including HOMA-IR Proxy, Metabolic Syndrome Index, and Age-Glucose interactions
  • Robust outlier clipping based on physiological feasibility thresholds
  • LOINC code mapping (`1558-6` Fasting Glucose, `8462-4` Diastolic BP, `39156-5` BMI, `20448-7` Insulin)
Cleaning Strategy: Missing laboratory records imputed via iterative multivariate chained equations (MICE).

04 // SYSTEM ARCHITECTURE & DATA FLOW

HL7 FHIR Ingest -> Biomarker Engineering -> Stacked Super Learner (Deep MLP + CatBoost + XGBoost) -> Youden Calibration -> SHAP Clinician Report -> EHR Export.

STEP 01HL7 FHIR · Python
FHIR R4 Ingestion Engine

Parses standard clinical observation bundles and extracts LOINC laboratory values.

STEP 02NumPy · Pandas
Biomarker Feature Pipeline

Computes physiological interaction indices and metabolic syndrome composites.

STEP 03PyTorch · Scikit-Learn
Deep Tabular MLP & Ensembles

128-64 hidden layer architecture trained with adaptive Adam optimizer.

STEP 04SHAP · Matplotlib
Explainability & Reporting

Generates SHAP clinical attribution waterfalls for clinician verification.

05 // MODEL ENGINEERING & HYPERPARAMETERS

Base Architecture: Tabular Deep Neural Network (MLP 128-64) + Stacked Super Learner (CatBoost, XGBoost, LightGBM)

10-Fold Stratified Cross-Validation with Youden's J index threshold calibration (0.650).

Hyperparameters & Training Dynamics:
  • • Hidden Layers: [128, 64]
  • • Dropout: 0.3
  • • Learning Rate: 0.001
  • • CatBoost Depth: 6
Loss Function: Binary Cross-Entropy with class weight calibration
Trade-off Rationale: Prioritized Deep MLP with CatBoost stacking to achieve an industry-leading 0.9810 ROC-AUC.

06 // FAILURE ANALYSIS & ZERO-TRUST SAFEGUARDS

OBSERVED FAILURE MODES UNDER STRESS
  • • Patients with atypical steroid-induced hyperglycemia not captured by standard metabolic profiles.
  • • Missing insulin lab records in outpatient clinic settings.
Mitigation & Fallback: Dual-model pathway: full biomarker model when insulin is present, fallback surrogate model when only fasting glucose and BMI are available.

07 // PRODUCTION DEPLOYMENT SPECS

Serving Framework
FastAPI with FHIR R4 Bundle Endpoints
Containerization
Docker container with strict HIPAA-compliant configuration templates
P95 SLA
22ms per patient bundle evaluation
Throughput
80 evaluations/sec

08 // ARCHITECTURAL DECISIONS & TRADE-OFFS

Engineered HL7 FHIR R4 interoperability parser directly into the platform.
Why: Ensures the CDSS can be integrated directly into Epic, Cerner, or other hospital EHR systems without proprietary adapter layers.
Alternative Discarded: Proprietary bespoke CSV/JSON formats.
Used CatBoost and Deep Tabular MLP stacking.
Why: CatBoost's symmetric oblivious trees are highly resistant to tabular noise and clinical outliers.
Alternative Discarded: Standard Random Forest alone.

09 // PLANNED IMPROVEMENTS & NEXT REVISIONS

  • →Incorporate continuous glucose monitor (CGM) real-time streaming telemetry.
  • →Conduct prospective multi-site clinical pilot studies to evaluate clinician alert adoption rates.