Changqi Sun portrait
Pfizer Externship via Extern

AI for document insights and data extraction

I build models and pipelines that extract structured data and surface insights from documents, developed during a Pfizer externship.

Work samples

Portfolio of externship work showing models, extraction pipelines, and summaries built to convert document content into structured outputs and insights.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    Pfizer logo

    Pfizer · ✅ Verified by Extern · ⏱️ In progress

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills

About me

I build models and pipelines that extract structured data and surface insights from documents, developed during a Pfizer externship.

I am Changqi Sun. I completed the Pfizer Advanced: AI-Powered Document Insights & Data Extraction externship, working on AI methods for extracting and summarizing information from documents and turning document content into structured data.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

In progress

Education

western university

mechanical and materials engineering · Class of 2028

Skills

Document summarizationData extractionAI model developmentPipeline implementationText preprocessingStructured data mappingAndroid DevelopmentAutomation

✅ Verified by Extern · ⏱️ In progress

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

Overview

The externship prototyped an AI document-insights pipeline combining OCR, large language models, and retrieval-augmented generation to process enterprise PDFs. The work reviewed foundational LLM concepts and implemented components of a document-intelligence workflow. Deliverables included research notes and technical explanations of how LLMs operate.

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

What I've accomplished

I produced a technical summary that explained LLM tokenization, Transformer attention mechanics, sequential token prediction, plus descriptions of pre-training, fine-tuning, and human-feedback training steps.

Project breakdown

The submission summarized how LLMs tokenized text, used Transformer attention to weigh context, and predicted tokens sequentially. It described pre-training on large datasets and subsequent fine-tuning and human feedback to improve instruction following.

The module required cleaning structured and unstructured data and improving OCR inputs. I cleaned tabular and JSON data with pandas, implemented text-standardization steps, and applied grayscale, denoising, CLAHE, and adaptive thresholding to produce a binarized image for OCR.

View all work