<Julia_Tan_-_Aspiring_Data_Scientist>

/* I blend technical expertise with a focus on impactful data solutions, driven to enhance my skills and career in AI and data science. */

🟢 Open to work
username.png
Currently an Extern @Pfizer

<Work_samples>

Explore my portfolio showcasing diverse projects in machine learning, data visualization, and predictive analytics that drive real-world impact.

  • Long-Term Financial Stability Prediction
    [01]Freelance ProjectMar 2026
    Long-Term Financial Stability Prediction

    Created a hybrid stacked model with Scikit-learn and XGBoost to predict 3 factors of long-term financial stability and presented actionable strategies to improve financial stability, specifically lowering debt.

    Scikit-learnXGBoostMatplotlib
  • Mind the Gap: Predicting TTC Subway Delays with a Machine Learning Approach
    [02]Freelance ProjectFeb 2025
    Mind the Gap: Predicting TTC Subway Delays with a Machine Learning Approach

    Developed a predictive model with Python to forecast TTC subway delays to present 3 actionable recommendations to several stakeholders on strategies to improve handling TTC delays.

    TensorFlowPythonMatplotlib

<About_me>

I blend technical expertise with a focus on impactful data solutions, driven to enhance my skills and career in AI and data science.

I'm Julia Tan, a junior pursuing a Bachelor's in Actuarial Science at the University of Toronto. I served as a Technical/Operations Lead at Riipen Labs, with experience as an Artificial Intelligence Researcher at Algoverse. I'm passionate about harnessing data to drive decisions and improve systems.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Pfizer

Experience

Technical/Operations Lead

Riipen Labs · Mar 2026 - Apr 2026

Artificial Intelligence Researcher

Algoverse · Jun 2025 - Jan 2026

Education

University of Toronto

B.Sc. Actuarial Science and Economics · Class of 2028

Skills

PythonRSQLScikit-learnTensorFlowPyTorchMachine LearningDeep LearningLarge Language ModelsGenerative AIStatistical AnalysisData ModelingData Visualization
Back_to_works

<Long-Term_Financial_Stability_Prediction>

Created a hybrid stacked model with Scikit-learn and XGBoost to predict 3 factors of long-term financial stability and presented actionable strategies to improve financial stability, specifically lowering debt.

Scikit-learnXGBoostMatplotlib

/Overview

Created a hybrid stacked model with Scikit-learn and XGBoost to predict 3 factors of long-term financial stability and presented actionable strategies to improve financial stability, specifically lowering debt.

View_all_works
Back_to_works

<Mind_the_Gap:_Predicting_TTC_Subway_Delays_with_a_Machine_Learning_Approach>

Developed a predictive model with Python to forecast TTC subway delays to present 3 actionable recommendations to several stakeholders on strategies to improve handling TTC delays.

TensorFlowPythonMatplotlib

/Overview

Developed a predictive model with Python to forecast TTC subway delays to present 3 actionable recommendations to several stakeholders on strategies to improve handling TTC delays.

View_all_works
Back_to_works

<Financial_Modelling_Case_Competition>

Developed a stock price prediction linear regression model to derive 6 trading and hedging strategies. Ranked top 5 out of several teams for effective hedging strategies and predictions.

Scikit-learnMatplotlibLSTMARIMA

/Overview

Developed a stock price prediction linear regression model to derive 6 trading and hedging strategies. Ranked top 5 out of several teams for effective hedging strategies and predictions.

View_all_works

✅ Verified by Extern · ⏱️ In progress

<Pfizer_Advanced:_AI-Powered_Document_Insights_&_Data_Extraction_Externship>

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

/Overview

The externship prototyped AI-powered document intelligence workflows for enterprise PDFs. The work combined OCR, large language models, and retrieval-augmented generation to process and extract information from documents. Deliverables included research, architectural notes, and proof-of-concept extraction pipelines.

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

//What I've accomplished

I produced a technical summary that explained tokenization, embedding formation, next-token prediction mechanics, weight updates during training, and how models select the highest-probability output token.

///Project breakdown

The submission summarized how LLMs tokenized input, converted tokens to numeric embeddings, learned word relationships by next-token prediction during training, adjusted weights from errors, and produced the highest-probability token as output.

During the externship I loaded raw document files in Colab, cleaned structured and unstructured text with Pandas and regex, processed JSON into normalized tables, and applied OpenCV denoising and histogram equalization to enhance a scanned document image.

The project processed SDF PDFs to extract text and bounding boxes, identified where tables and special symbols were missed, and noted that table structure and date formatting were not preserved. The deliverable summarized these extraction results.

Google Docs
Access
View_all_works