Pfizer logoPfizer Externship via Extern
Samuel Bissou portrait

Aspiring Computer Scientist

I am a junior at Huston-Tillotson University, diving into the world of AI and data extraction at Pfizer.

Work samples

Discover my work in AI and data extraction, showcasing the skills I am developing during my studies and externship.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · ✅ Verified by Extern · ⏱️ In progress

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills

About me

I am a junior at Huston-Tillotson University, diving into the world of AI and data extraction at Pfizer.

I am a junior at Huston-Tillotson University pursuing a Bachelor's Degree in Computer Science. Currently, I am gaining valuable experience through my externship at Pfizer, focusing on AI-powered document insights and data extraction.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

In progress

Skills

AIData ExtractionDocument InsightsComputer Science

✅ Verified by Extern · ⏱️ In progress

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The externship prototyped an AI pipeline that extracted and structured data from enterprise PDF documents. Work combined optical character recognition, language models, and retrieval-augmented generation to process document text and produce working prototypes and analyses. Deliverables included prototype outputs and a documented review of model behavior on real pharmaceutical documents.

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I reviewed a vendor document packet and produced a detailed audit identifying six-plus inconsistent date formats, repeated boilerplate text that could cause de-duplication issues, a redacted address field, and fields likely to produce automated extraction errors.

Project breakdown

I reviewed a sample vendor document packet, noted six-plus inconsistent date formats, repeated boilerplate text that could cause de-duplication issues, a redacted address field, and highlighted which fields risked automated extraction errors.

The module provided a noisy low-resolution photograph. I applied upscaling, three-stage denoising (non-local means, median, non-local means), CLAHE contrast, and adaptive Gaussian thresholding. Laplacian variance metrics tracked noise removal and edge recovery.

I processed a 3-page SDF with PyMuPDF get_text("words"), grouped 617 word tokens by page, sorted by (y0,x0), rejoined fragmented multi-word values, normalized inconsistent spellings and symbols, and produced JSON records with word bboxes to enable label-to-value matching.

Google Docs
Access

A scanned supplier page showed page-segmentation faults, so I ran Tesseract with two --psm modes, merged their outputs by position, and produced a combined extract that recovered 60 words missing from one pass.

View all work