System Note 02 · Research Platforms

Multi-Omics Data Pipeline

Johns Hopkins University

PythonRSparkMachine LearningHPCSVM-RFE
Private build
What it is

A pipeline that pulls together 750+ TB of cancer data from several public sources, cleans and quality-checks it, and runs machine-learning models to surface candidate biomarkers.

750+ TB · TCGA / PCAWG / ENCODE / PRIDE · HPC clusters

Why I built it

Biomarker discovery was bottlenecked on wrangling enormous, messy, mismatched datasets before any modeling could start.

Outcome
750+TB data
8novel biomarkers
40%faster validation

Evidence before confidence.

← Back to Build Notes