Aspiring Data & AI Professional | Turning Data into Actionable Insights | AI & Analytics
Building skills in data analytics, AI-powered data extraction, and transforming complex information into actionable insights through my experience at Pfizer.
Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.
AI & MLPythonDocument IntelligencePresentation Skills
About me
Building skills in data analytics, AI-powered data extraction, and transforming complex information into actionable insights through my experience at Pfizer.
I’m an aspiring Data & AI professional passionate about using technology and data to solve problems and uncover meaningful insights. Through my experiences at Pfizer and UMass Lowell, I’ve been developing hands-on skills in AI-powered data extraction, data analytics, and problem-solving. I enjoy learning new technologies, working with data, and exploring how AI can transform complex information into actionable insights. I’m excited to continue growing and connect with professionals who share an interest in data, technology, and innovation.
Externships
Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
Pfizer
Experience
Information Analyst Intern
Mosaic Lowell · May 2024 - Aug 2024
Education
University of Massachusetts at Lowell
Bachelors of Science in Business Administration · Class of 2027
Skills
Data AnalysisAI TechnologiesDocument InsightsData ExtractionProject Management
Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.
AI & MLPythonDocument IntelligencePresentation Skills
Overview
The work prototyped an AI-driven pipeline that converted scanned pharmaceutical PDFs into machine-readable outputs. Deliverables included Colab notebooks for data cleaning and image preprocessing, PyMuPDF word-level extracts with bounding boxes and accuracy notes, LLM tokenization documentation, a comparative OCR analysis that favored PaddleOCR for layout preservation, and a working RAG pipeline
I produced runnable Colab notebooks and analyses: LLM tokenization documentation, cleaned structured data and preprocessed images, PyMuPDF word-level extracts with bounding boxes and validation recommendations, a comparative OCR analysis identifying PaddleOCR as best at preserving layout, and a working RAG pipeline with hybrid retrieval.