Final Capstone Project Baseball(A)IQ APP
Designed and built a regression pipeline predicting MLB players' next-season batting average from prior performance, engineering a player-season panel dataset from the Lahman Baseball Database (1962–2023) with a time-aware train/test split to prevent data leakage Engineered a sabermetrically-grounded feature set (BABIP, walk/strikeout rate, age curve, positional context) to capture predictive signal beyond raw batting average alone Trained and compared 5 model configurations (Linear Regression, Random Forest, Gradient Boosting) across algorithm and feature-set variations, tracking experiments with MLflow and selecting the best model by MAE/R² [Tuned Linear Regression- MAE: 0.0234, RMSE: 0.0298, R2: 0.2247] Built a natural-language interface (Streamlit + LLM API) supporting both real-player lookups and hypothetical "what-if" scenarios, with structured input parsing, clarifying questions for missing data, and graceful handling of out-of-scope queries Applied production MLOps practices: config-driven training (YAML), DVC-tracked datasets, a pytest suite covering preprocessing, model validation, and interface logic

