Pfizer logoPfizer Externship via Extern
Shrihan Reddy Nomula portrait

Data science student building practical AI tools

Sophomore at UC Berkeley creating prototypes that turn PDFs into structured data using OCR, LLMs, and RAG.

Work samples

Portfolio items include an externship prototype that automated enterprise PDF workflows and other data science projects showcasing model pipelines, OCR integration, and extraction logic.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · ✅ Verified by Extern · ⏱️ In progress

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills

About me

Sophomore at UC Berkeley creating prototypes that turn PDFs into structured data using OCR, LLMs, and RAG.

I am a sophomore at UC Berkeley studying Data Science. I have some hands-on experience building AI prototypes, including an externship project that used OCR, large language models, and retrieval-augmented generation to extract data and surface insights from enterprise PDFs. I am exploring data science and applied machine learning as I shape my early career.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

In progress

Experience

Machine Learning Development Engineer

Presonance · Jul 2026 – Present

AI Engineer Extern

Externship — Pfizer · Jul 2026 – Present

Exoplanet Data Research Assistant

NASA Jet Propulsion Laboratory · Apr 2024 – May 2026

ML/AI Researcher

Independent Research — Federated Learning | Advisor: Prof. Ron Mahabir · Apr 2024 – May 2025

Data Science & Analytics Intern

California EPA — Department of Toxic Substances Control · Jun 2024 – Aug 2024

Education

University of California, Berkeley

B.A. Data Science (Emphasis in Robotics) & Astrophysics; GPA: 3.9 · Class of 2029

Skills

OCR integrationLarge language models (LLMs)Retrieval-augmented generation (RAG)Data extraction from PDFsData Science fundamentalsPython

✅ Verified by Extern · ⏱️ In progress

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The work prototyped AI methods for document intelligence on a supplier documentation package. The project cataloged document types and their dates, flagged inconsistent date formats and split pages, documented hedged language and repeated boilerplate, and recorded parsing and validation challenges for automated extraction.

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I produced a document catalogue with identified types and dates, cleaned and binarized an OCR-ready image with measured noise and contrast improvements, extracted Manufacture and Expiration dates with bounding boxes from a 3-page SDF, and produced an OCR comparison notebook with field-level scoring and diagnosed layout issues.

Project breakdown

I examined a supplier documentation package, identified each document type and its dates, flagged inconsistent date formats and split documents across pages, noted hedged language and repeated boilerplate, and documented parsing and validation challenges for AI extraction.

The module tasked cleaning a noisy scanned photo for OCR. I applied grayscale, median blur, Non-Local Means denoising, CLAHE contrast enhancement, and Otsu thresholding, measured noise (28.7→0.3) and contrast (31.7→50.6), and produced a binarized cleaned image.

I processed a 3-page SDF with PyMuPDF, saved each word with page and bbox, grouped words into lines, matched labels with regex, and extracted Manufacture and Expiration dates with their bounding boxes; I reported parsing issues and proposed normalization and validation steps.

Google Docs
Access

The task evaluated OCR on a Certificate of Quality. I ran Tesseract on a 300 DPI page, converted output to text, confidences and boxes, and scored 16 key fields. I diagnosed rotated-stamp and column-merge issues, and produced a notebook that standardised outputs and a field-by-field score…

Google Docs
Access
View all work