Skip to content

Repository files navigation

Unsupervised Discovery and Predictive Modeling of Therapeutic Failure Phenotypes in the FAERS Database

This project performs end-to-end analysis of therapeutic failures reported in the FDA Adverse Event Reporting System (FAERS). It combines NLP-based clustering and machine learning classification to discover hidden phenotypes, label failure cases, and build a predictive model for real-world deployment.

Group Name: DataCraft

Group Members:

  • Aarsh Adhvaryu : 202518031
  • Harsh Jethwani : 202518055
  • Ojas Gupta : 202518057
  • Ram Sharma : 202518009

DATASET URL: https://fis.fda.gov/extensions/FPD-QDE-FAERS/FPD-QDE-FAERS.html

Project Overview

The end-to-end pipeline consists of:

  • Data Integration & Preprocessing
  • Unsupervised Failure Phenotype Discovery (Clustering)
  • Feature Engineering for Classification
  • Imbalanced Learning (SMOTE + Undersampling)
  • Supervised Model Training
  • Grid Search Optimization
  • Final Real-World Evaluation
  • Exported Artifacts for Deployment

1. Data Preprocessing

FAERS quarterly datasets (Drug, Reaction, Outcome tables) were merged and cleaned. Key preprocessing steps:

  • Merge all FAERS tables by primaryid
  • Collapse multiple reactions per case
  • Handle missing values
  • Extract severity and hospitalization signals
  • Create a binary is_failure target (based on FDA seriousness indicators)
  • Extract clean text from all_reaction_pts

2. Unsupervised Discovery of Failure Phenotypes (Clustering)

We cluster failure cases using NLP-based representation:

  • Text Embedding Pipeline
  • TF-IDF Vectorizer
  • max_features = 2000
  • ngram_range = (1, 2)
  • custom FAERS stopwords
  • Dimensionality Reduction
  • Truncated SVD → 50 components
  • KMeans Clustering
  • Searched K = 2 to 6 using silhouette score
  • Optimal or forced K = 3
Discovered Phenotypes (Clusters)
Cluster Meaning
Cluster 0 Critical Failure — life-threatening terms (death, shock, organ failure)
Cluster 1 Hospitalization Failure — emergency, pneumonia, infection
Cluster 2 Side-Effect Failure — headaches, nausea, mild symptoms

3. Classification Feature Engineering

Dropped non-predictive identifiers: primaryid, caseid, caseversion, fda_dt_parsed, all_reaction_pts, severity fields

Features were scaled using StandardScaler.

  • The phenotype label was encoded using LabelEncoder.
Artifacts saved:
  • scaler.joblib
  • label_encoder.joblib
  • training_feature_names.joblib

4. Imbalanced Learning Pipeline

To avoid data leakage, only the training split was balanced.

We used:

  • SMOTE (oversampling minority classes) — 60%
  • RandomUnderSampler — to keep 0/1/2 ratios reasonable

Saved artifacts:

  • X_train_bal.joblib
  • y_train_bal.joblib
  • X_test.joblib
  • y_test.joblib

5. Baseline Model Evaluation (Untuned)

Trained on balanced data, tested on untouched test split. Models evaluated:

  • Logistic Regression
  • Decision Tree
  • Random Forest
  • XGBoost
  • LightGBM

Scores recorded: untuned_model_scores.csv

The top 5 models were automatically selected for tuning.

6. Grid Search Optimization

A optimized grid search was run on: -Logistic Regression

  • Decision Tree
  • Random Forest
  • XGBoost
  • LightGBM

Evaluation metrics:

  • F1-Macro
  • Accuracy
  • PR-AUC

Each model produced a confusion matrix and classification report.

Results saved:

  • tuned_model_scores.csv
  • best_tuned_.joblib

The best model (by F1-Macro) was selected as:

LightGBM (Tuned) Saved as:

  • best_classifier_final.joblib

7. Final Real-World Evaluation

Evaluated tuned model on unseen, real-world FAERS data: Final Score:

  • Accuracy: 0.9601
  • F1-Macro: 0.8277 Performance Highlights
  • Critical_Failure: F1 ≈ 0.98
  • Hospitalization_Failure: F1 ≈ 0.95
  • SideEffect_Failure: F1 ≈ 0.56

This confirms:

  • Strong generalization
  • No data leakage
  • High real-world reliability

The full predictions were exported: final_evaluation_predictions.csv

About

Unsupervised Discovery and Predictive Modeling of Therapeutic Failure Phenotypes in the FAERS Database

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages