HJ@hyunwoo-jee
hyunwoo Jee portrait

Aspiring Computer Scientist

I'm a Senior at UIUC specializing in Computer Science, ready to tackle the tech industry's challenges.

Work samples

Discover my projects and experiences in Computer Science, showcasing my skills and growth.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · Sep 2026 · ✅ Verified by Extern

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills

About me

I'm a Senior at UIUC specializing in Computer Science, ready to tackle the tech industry's challenges.

I am a Senior at the University of Illinois Urbana-Champaign pursuing a Bachelor's Degree in Computer Science. I am eager to strengthen my resume and enhance my competitiveness for future roles in tech.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Completed · Sep 2026

Skills

Computer ScienceAI-Powered Document InsightsData ExtractionProblem SolvingTeam Collaboration

✅ Verified by Extern

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The externship prototyped AI document-intelligence workflows for pharmaceutical PDFs, combining OCR, image pre-processing, embedding-based retrieval, and RAG experiments. Deliverables included OCR engine comparisons, PyMuPDF extractions with bounding boxes, RAG tests using all-MiniLM-L6-v2 plus reranking, and a blob-segmentation pipeline that labeled pages with doc_type and doc_id.

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I extracted 617 words with bounding boxes from multi-page SDFs, compared Tesseract, PaddleOCR, and EasyOCR with recorded confidences and setup notes, implemented hybrid keyword-plus-vector RAG with reranking using all-MiniLM-L6-v2, and built a pipeline that segmented a 10-page blob PDF and assigned doc_type and doc_id per page.

Project breakdown

View all work