Pfizer logoCurrently an Extern @Pfizer·🟢 Open to work
Melissa G portrait

Senior studying, building applied AI projects for document workflows

I prototype AI document tools that combine OCR, LLMs, and retrieval to automate PDF workflows and turn messy files into structured data.

Work samples

Portfolio includes a Pfizer externship prototype using OCR, LLMs, and RAG to extract insights from enterprise PDFs and automate document workflows.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · ✅ Verified by Extern · ⏱️ In progress

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills
  • Dropbox Clone
    Dropbox Clone

    Built a full-stack cloud storage application using React, Node.js, Express.js, and SQL. Implemented secure file upload, storage, and user management features. Used Git for version control and collaborative development.

    ReactNode.jsExpress.jsSQLGit

About me

I prototype AI document tools that combine OCR, LLMs, and retrieval to automate PDF workflows and turn messy files into structured data.

I am a senior at Indiana Wesleyan University working toward a bachelor’s degree. I have some practical project experience and I'm building a portfolio focused on applied AI tools for real-world documents, including an externship project with Pfizer.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Pfizer

Experience

Receptionist

CDC Tax · 2024–Present

Server

Yolk · 2021–2024

Bartender

Sugar Factory · 2018–2020

Education

Indiana Wesleyan University

Bachelor of Science in Healthcare Administration

Kenzie Academy

Diploma in Full Stack Software Development

Skills

OCRLarge language models (LLMs)Retrieval-augmented generation (RAG)Prototype developmentDocument data extractionPDF workflow automationPythonReact

✅ Verified by Extern · ⏱️ In progress

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The project prototyped AI-powered document intelligence for enterprise PDF workflows, combining OCR, large language models, and retrieval-augmented generation. The work produced a unified technology reference, OCR design notes, and prototype artifacts for automated document extraction and question-answering. Deliverables included a written summary of how large language models process text and how

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I summarized how large language models predict tokens based on large text corpora, described their token-based prediction process, and documented the role of human feedback in refining model suggestions.

Project breakdown

I summarized how large language models trained on large text corpora learned patterns to predict next words, described their token-based prediction process, and noted that human feedback helped refine future suggestions.

The module required cleaning structured and unstructured document data and improving scanned images for OCR. I processed text (whitespace, encoding, punctuation, case, contractions, HTML) and reordered steps for correct results, and applied denoising, grayscale, contrast, and thresholding to a…

The module required extracting text and bounding boxes from multi-page SDFs. I used PyMuPDF to iterate pages and get_text("words"), handled table and column layouts, and produced a document of extraction results noting challenges and suggestions for locating manufacturing and expiration dates.

Google Docs
Access

I ran Tesseract, PaddleOCR, and EasyOCR on a supplier PDF, debugged installs and extraction scripts, and compared outputs. EasyOCR extracted lot numbers, product codes, manufacture and expiration dates most reliably; PaddleOCR required the most setup.

Google Docs
Access

The project required ingesting a PDF, applying chunking and sentence-transformer embeddings, and combining vector and BM25 retrieval. I debugged async issues and produced a working pipeline with hybrid retrieval and documented design choices.

Google Docs
Access
View all works

Dropbox Clone

Built a full-stack cloud storage application using React, Node.js, Express.js, and SQL. Implemented secure file upload, storage, and user management features. Used Git for version control and collaborative development.

ReactNode.jsExpress.jsSQLGit
View all works