Pfizer logoCurrently an Extern @Pfizer
Tanya Kari portrait

Aspiring Biomedical Engineer

I'm a Texas A&M sophomore focused on Biomedical Engineering, excited about merging technology with healthcare innovation.

Work samples

Explore my portfolio featuring innovative projects, including AI-driven document intelligence solutions developed during my externship at Pfizer.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · ✅ Verified by Extern · ⏱️ In progress

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills

About me

I'm a Texas A&M sophomore focused on Biomedical Engineering, excited about merging technology with healthcare innovation.

I am a sophomore at Texas A&M University, pursuing a Bachelor's degree in Biomedical Engineering. With a passion for innovative technology, I am eager to explore opportunities that blend engineering and healthcare.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Pfizer

Experience

Medical Scribe

ScribeAmerica · May 2026

Research Assistant

Lele Lab - Texas A&M Biomedical Engineering · Aug 2025 - May 2026

Research Assistant

University of Hawaii - Cancer Center · May 2025 - Aug 2025

Education

Texas A&M University College Station

B.S. Biomedical Engineering · Class of 2028

Skills

Biomedical EngineeringAI-Powered SolutionsDocument IntelligenceOCRLLMsRAGBiotech

✅ Verified by Extern · ⏱️ In progress

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The work prototyped AI document-intelligence techniques for pharmaceutical PDFs, combining explanations of LLM mechanics with hands-on extraction and OCR comparisons. Deliverables included a written summary of how LLMs predict tokens and train, Colab notebooks for data and image preprocessing, JSON outputs of extracted fields from SDF PDFs, and a comparative OCR evaluation that identified

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I produced runnable Colab notebooks, JSON extraction outputs from SDF PDFs, a technical summary of LLM internals, and a quantitative OCR comparison that identified the most accurate engine for scanned supplier documents.

Project breakdown

The submission summarized how LLMs predicted tokens using numerical representations, described Transformer self-attention, and outlined pretraining, supervised fine-tuning, and reinforcement learning from human feedback to improve responses.

The work showed Python notebooks in Google Colab that cleaned tabular data with Pandas, normalized nested JSON into flat tables, performed text cleaning and normalization, and applied image preprocessing (denoising, thresholding, contrast) to improve OCR input.

The project processed SDF PDFs with PyMuPDF, extracted every word with bounding boxes, and then used keyword anchors and proximity logic to locate fields. The submission produced JSON word output and listed extracted fields including manufacture and expiration dates.

Google Docs
Access

I ran Tesseract, PaddleOCR, and EasyOCR on a scanned supplier PDF, recorded text-region counts, per-engine confidence scores, and detected key pharmaceutical terms. I reported that PaddleOCR had the highest accuracy and confidence, with metrics and a comparison table.

Google Docs
Access
View all works