May Manna Wanya portrait
Currently an Extern @Pfizer

Aspiring Professional in AI & Data Extraction

I’m on a journey to officially start my career while honing my technical skills in AI and data extraction, currently externing at Pfizer.

Work samples

Explore my projects focused on AI-powered solutions and data extraction, reflecting my commitment to skill development and innovation.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · ✅ Verified by Extern · ⏱️ In progress

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills

About me

I’m on a journey to officially start my career while honing my technical skills in AI and data extraction, currently externing at Pfizer.

I am an aspiring professional with a Master's degree, eager to officially launch my career. Currently, I am gaining hands-on experience through an externship at Pfizer, where I'm focused on enhancing my technical skills in AI and data extraction.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Pfizer

Experience

MS Research Assistant

GUO LAB, DEPARTMENT OF BIOINFORMATICS AND GENOMICS- UNC, CHARLOTTE · JANUARY- DECEMBER 2025

Pharmacist/Manager

CITY PHARMACY LIMITED · 7TH NOVEMEBER 2022 – APRIL 2024

Pharmacist

PARADISE PRIVATE HOSPITAL · 7TH JUNE 2018 - 30TH APRIL 2021

TB Community Treatment Support Supervisor

WORLD VISION PNG · 7TH OCTOBER 2021- NOVEMBER 4TH 2022

Education

University of Papua New Guinea, School of Medicine & Health Sciences

Bachelor of Pharmacy · Class of 2016

University of North Carolina-Charlotte

Master of Science in Bioinformatics · Class of 2026

Skills

AI-powered solutionsData extractionTechnical skills developmentDocument insights

✅ Verified by Extern · ⏱️ In progress

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The work produced technical writing and runnable prototypes that demonstrated AI techniques for processing pharmaceutical documents. It documented LLM internals, including tokenization, embeddings, multi-head transformer attention, repeated layers, and gradient-descent training. The deliverables included a Google Colab notebook with data-cleaning and image-preprocessing examples and code for

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I delivered a concise technical summary of LLM internals plus a Google Colab notebook containing runnable data-cleaning and image-preprocessing examples, and code that extracted words and bounding boxes from SDF PDFs with notes on extraction quality.

Project breakdown

I summarized LLM internals: tokenization to integer IDs, embeddings as high-dimensional vectors, transformer attention with multiple heads, repeated layers, and gradient-descent training. The submission produced a concise, technical explanation of how LLMs learn and predict text.

The module provided a Google Colab notebook that showed Python setup, basic data structures, control flow, functions, and an input-to-output example. I executed code cells and included runnable examples for data cleaning and image preprocessing.

The module involved extracting words and bounding boxes from SDF PDFs with PyMuPDF, reporting which pages and layouts extracted cleanly, noting dropped elements (logos, addresses), coding a separate bounding-box extractor, and recommending regex, label proximity, and bbox rules to locate date…

Google Docs
Access

The assignment compared three OCR engines on a scanned pharmaceutical supplier PDF. I ran Tesseract, PaddleOCR and EasyOCR, inspected outputs and found PaddleOCR preserved layout and bounding boxes best, reporting its higher confidence and setup trade-offs.

Google Docs
Access
View all works

Investigation of protein secondary structure prediction

Investigation of protein secondary structure prediction. Guo Lab, Department of Bioinformatics and Genomics, UNC Charlotte, 2025.

View all works