Multi-Omics Data Pipeline
Johns Hopkins University
A pipeline that pulls together 750+ TB of cancer data from several public sources, cleans and quality-checks it, and runs machine-learning models to surface candidate biomarkers.
750+ TB · TCGA / PCAWG / ENCODE / PRIDE · HPC clusters
Biomarker discovery was bottlenecked on wrangling enormous, messy, mismatched datasets before any modeling could start.
750+TB data
8novel biomarkers
40%faster validation
Evidence before confidence.