ANUJ
DATA×AI×ENGINEERING
INITIALIZING SYSTEM
ANUJ MUNDU
PROJECT 07•COMPUTER VISION•AI / ML•FLAGSHIP CAPSTONE

YOLOv5-CASP Clinical CADx — Lung Nodule Suite

Deep Learning CADx suite for pulmonary nodule detection using YOLOv5-CASP with CBAM, ASPP & CoT3

PyTorch 2.5.1CUDA 12.1OpenCV 4.9.0YOLOv5-CASPCBAMASPPCoT3DICOMPACSStreamlit
Sensitivity
94.2%
Small pulmonary nodules (< 6mm)
mAP@0.5
91.8%
Validation on chest CT & CXR cohorts
Triaging Protocol
Lung-RADS™
Automated categorical risk scoring
PACS Workstation
Interactive
Streamlit clinical viewer with Grad-CAM overlays
// INTERACTIVE SYSTEM TELEMETRY & DIAGNOSTIC LAB
LIVE RTSP INFERENCE [640×640 INT8]
40.3 FPS
P95 LATENCY: 24.8ms
NMS IoU: 0.45
GRAD-CAM LOCALIZATION HEATMAP
FOCAL LOSS // RESNET-50
MICRO-CRACK: (x: 218, y: 92)
CONFIDENCE: 98.2%
HEATMAP BLEND: 65%

01 // SYSTEM OVERVIEW

Lung cancer causes nearly 1.8 million deaths annually worldwide. Early detection via low-dose CT and CXR dramatically improves 5-year survival. This research thesis and clinical suite develops YOLOv5-CASP: an enhanced object detector incorporating CBAM attention, ASPP multi-scale context, and CoT3 contextual transformers to overcome high false-positive rates on ambiguous sub-centimeter lung nodules.

02 // THE PROBLEM & ENGINEERING SIGNIFICANCE

The Core Challenge

Small pulmonary nodules (< 6mm) blend into vascular structures, ribs, and soft tissue, leading to high false-negative rates in standard clinical screenings.

Why This Matters

Early stage I detection increases survival rates to over 60%, but radiologists face high cognitive fatigue reviewing hundreds of axial slices per patient.

Key Constraints:
  • •Resolving low contrast between benign pulmonary parenchyma and malignant micro-nodules.
  • •Handling extreme scale variance: nodules range from 3mm punctate lesions to 30mm masses.
  • •Providing interpretable spatial attention maps so radiologists can verify algorithmic reasoning.

03 // DATA PIPELINE & PREPROCESSING

Input Format: DICOM, High-Resolution 16-bit CT Volumes, and Chest Radiographs (CXR)Sample Volume: Multi-center clinical cohorts with comprehensive expert ground-truth annotations
Transformation Steps:
  • Hounsfield Unit (HU) lung-window clipping (-1000 to +400 HU) for tissue contrast normalization
  • Histogram equalization, CLAHE enhancement, and multi-slice axial slice projection
  • Ablation-specific augmentation: mosaic, mixup, random affine, and perspective warping
Cleaning Strategy: Exclusion of motion-corrupted scans; radiologist consensus labeling for ambiguous border margins.

04 // SYSTEM ARCHITECTURE & DATA FLOW

Custom YOLOv5-CASP backbone and neck: Feature extraction through CSP-Darknet -> CBAM (Channel & Spatial Attention) -> ASPP (Dilated context) -> CoT3 (Contextual Transformer) -> Multi-Scale Detection Heads.

STEP 01PyTorch Module
CBAM Attention Block

Channel and spatial attention gates suppress non-relevant vascular background noise.

STEP 02Dilated Convolutions
ASPP Multi-Scale Receptive Field

Atrous convolutions capture both micro-nodule details and surrounding lobar morphology.

STEP 03Transformer Neck
CoT3 Contextual Transformer

Contextual self-attention models spatial correlations across neighboring anatomical regions.

STEP 04Streamlit · FastAPI
Lung-RADS PACS & API

Streamlit radiologist workstation with interactive thresholding and REST inference API.

ENGINEERING ITERATION & REFACTORING CHRONOLOGYv0 Prototype → v2 Cloud Production
v0 // PROOF OF CONCEPT

Standard YOLOv5s baseline on 2D PNG exports; suffered 2.87 false alarms per scan and failed on nodules < 5mm.

v1 // MODULARIZATION & REFACTOR

Custom architectural injection of CBAM channel/spatial attention and ASPP dilated convolutions; reduced false positives by 42%.

v2 // PRODUCTION CLOUD SYSTEM

Production clinical PACS CADx suite with CoT3 contextual transformer neck, automated Lung-RADS triaging, Grad-CAM interpretability, and live Streamlit workstation.

05 // MODEL ENGINEERING & HYPERPARAMETERS

Base Architecture: YOLOv5-CASP (Custom PyTorch 2.5.1 + CUDA 12.1)

Trained with SGD optimizer, cosine learning rate scheduler, warm-up epochs, and automatic mixed precision (AMP).

Hyperparameters & Training Dynamics:
  • • Batch Size: 16
  • • Image Size: 640x640
  • • Epochs: 150
  • • Initial LR: 0.01
  • • Weight Decay: 0.0005
Loss Function: CIoU Bounding Box Loss + Focal Classification Loss + Objectness BCE
Trade-off Rationale: Added 4.2% computational overhead with CoT3 transformer neck to achieve a 6.8% boost in small-nodule mAP.

06 // FAILURE ANALYSIS & ZERO-TRUST SAFEGUARDS

OBSERVED FAILURE MODES UNDER STRESS
  • • Pleural-based subpleural nodules abutting the chest wall.
  • • Motion artifacts from severe patient breathing during scan acquisition.
Mitigation & Fallback: Integrated multi-slice spatial context and automated quality check flags that prompt radiologist review.
ENGINEERING POST-MORTEM & CONSTRAINT MITIGATION
Root-Cause Analysis & Production Telemetry
Operational Constraint:

Clinical safety mandate: Missing a malignant nodule is catastrophic, but false alarms cause radiologist alarm fatigue.

Bottleneck Encountered:

Standard cross-entropy loss biased model predictions toward background parenchyma due to 99.8% negative pixel imbalance.

Architectural Solution:

Engineered Focal Loss with gamma=2.0 and alpha=0.25 paired with CIoU geometric bounding box penalty.

Empirical Outcome:

Boosted small-nodule sensitivity to 94.2% while reducing false-positive detections per scan to 1.12.

07 // PRODUCTION DEPLOYMENT SPECS

Serving Framework
FastAPI REST Server + Streamlit PACS Interface
Containerization
Dockerized container with NVIDIA CUDA runtime support
P95 SLA
32.6ms per 640x640 CT slice on GPU / 180ms on CPU
Throughput
30 slices/sec

08 // ARCHITECTURAL DECISIONS & TRADE-OFFS

Integrated ASPP (Atrous Spatial Pyramid Pooling) into the neck architecture.
Why: Allows multi-scale contextual aggregation without expanding parameter size or losing spatial resolution.
Alternative Discarded: Standard Feature Pyramid Network without dilated convolutions.
Implemented automated Lung-RADS™ categorical risk scoring.
Why: Directly bridges raw algorithmic coordinates with standard clinical guidelines used by radiologists.
Alternative Discarded: Raw probability percentage outputs only.

09 // PLANNED IMPROVEMENTS & NEXT REVISIONS

  • →Extend pipeline to native 3D volumetric convolution (V-Net / 3D Swin) across entire contiguous CT volumes.
  • →Integrate automated longitudinal scan comparison to track nodule volume doubling time (VDT).
HUMAN-IN-THE-LOOP UI/UX ERGONOMICS:

Designed strictly for dark-environment radiology reading rooms: High-contrast monochrome DICOM viewer with non-blinding cyan/amber lesion bounding boxes and 1-click Lung-RADS score export.

OPEN REPRODUCIBILITY & TEST SUITE COMMAND:python -m pytest tests/test_casp_architecture.py -v

Validated CIoU loss, CBAM attention gate tensors, and Lung-RADS threshold categorization.

TECHNICAL INTERVIEW DISCUSSION PROMPTS
  • Q1:"Why did you choose YOLOv5-CASP over 3D U-Net or Mask R-CNN for this clinical screening workflow?"
  • Q2:"How did Hounsfield Unit (HU) windowing specifically impact model convergence during DICOM preprocessing?"
  • Q3:"How does the CoT3 contextual transformer neck assist in differentiating subpleural nodules from chest wall structures?"