Outamation logoCurrently an Extern @Outamation
Harsh Dayanand Koli portrait

Aspiring Data Scientist

I’m a graduate student specializing in Data Analytics, ready to leverage my skills in the real world.

Works samples

Explore my journey through data analytics projects and my externship, showcasing my skills and potential in the field.

  • Outamation Advanced AI-Powered Document Insights and Data Extraction Externship
    Outamation logo

    Outamation · Jul 2026 · ✅ Verified by Extern

    Outamation Advanced AI-Powered Document Insights and Data Extraction Externship

    Think AI is all about chatbots and image generators? Think again. In this externship, you’ll explore how AI is transforming massive, messy document workflows into smart, searchable data. From building Python scripts to testing open-source models, you’ll turn unstructured files into usable insights—and walk away with projects, skills, and code that employers love. Top-performing externs may get the opportunity to interview with the host company for an internship!

    AI & MLData AnalysisPythonGoogle ColabLlama IndexPresentation Skills
  • Credit Risk & Fraud Detection: Predictive Modelling

    ·

    Credit Risk & Fraud Detection: Predictive Modelling

    Collected and cleaned transactional financial datasets; engineered domain-relevant features and applied resampling strategies (SMOTE) to handle class imbalance. Performed statistical analysis across risk segments and threshold settings, monitoring model performance and identifying patterns and correlations in complex datasets. Communicated model outputs and risk findings in a clear, business-readable format for non-technical stakeholders.

    Pythonscikit-learnSQLXGBoostpandas

About me

I’m a graduate student specializing in Data Analytics, ready to leverage my skills in the real world.

I am a graduate student pursuing a Master's Degree in Data Science, with a focus on Data Analytics. Although I have little work experience, I am eager to kickstart my career in this dynamic field.

Externships

Outamation Advanced AI-Powered Document Insights and Data Extraction Externship

Outamation

Experience

AI & Data Engineering Extern

Extern, Inc. (Outamation Project) · December 2025 – Present

Data Analyst Intern

Trainity · September 2023 – December 2023

Education

University of Essex

Postgraduate Diploma in Data Science

University of Mumbai

BE in Data Engineering, 2:1 · Class of 2024

Skills

Data AnalyticsData ScienceAI-Powered Document Insights
Back to works

Outamation Advanced AI-Powered Document Insights and Data Extraction Externship

Think AI is all about chatbots and image generators? Think again. In this externship, you’ll explore how AI is transforming massive, messy document workflows into smart, searchable data. From building Python scripts to testing open-source models, you’ll turn unstructured files into usable insights—and walk away with projects, skills, and code that employers love. Top-performing externs may get the opportunity to interview with the host company for an internship!

AI & MLData AnalysisPythonGoogle ColabLlama IndexPresentation Skills
Outamation logo

Overview

During my externship with Outamation, I developed advanced skills in AI-powered document insights and data extraction. I worked on projects that transformed unstructured document workflows into structured, searchable data, employing machine learning, Python, and various OCR technologies. My hands-on experience with real-world mortgage documents equipped me with the technical foundation and

What I did

Throughout the externship, I engaged in multiple projects that honed my abilities in document processing and AI integration. Below are the key accomplishments from each module I completed:

1

Project 1: How AI Reads and Understands Mortgage Documents

I explored the principles of Machine Learning, Deep Learning, and NLP to understand how AI reads mortgage documents. Additionally, I applied Computer Vision and OCR techniques, analyzing a real mortgage 'blob' file to evaluate AI model performance on structured and unstructured data.

3

Project 3: Data Extraction from Documents Using Python

I converted digital and scanned PDFs into machine-readable formats using Python. By employing tools like PyMuPDF and pdfplumber, I extracted key information from multi-page mortgage files, applying heuristics such as regex patterns to identify essential fields.

5

Project 5: Introduction to Retrieval-Augmented Generation (RAG)

I designed a Retrieval-Augmented Generation (RAG) pipeline using LlamaIndex to enhance information retrieval from large document sets. This project improved my understanding of how AI can efficiently access relevant data in extensive text collections.

7

Project 7: Blob Processing, Classification, and Routing

I developed a system to process and classify unstructured mortgage file blobs. By using layout-based clues and rule-based methods, I separated documents and built a routing function to apply appropriate extraction logic for each identified type.

9

Project 9: Final Integration, Testing, and Evaluation

It’s time to bring everything together. In this final project, you’ll integrate all components into a complete, end-to-end document intelligence system. You’ll run your pipeline on a full mortgage blob, from preprocessing and OCR to extraction, RAG retrieval, and user interaction. Then, you’ll rigorously test the system’s performance, evaluating accuracy, speed, and reliability across different document types. Finally, you’ll package your work into a clear, professional demo—including a Colab repo, documentation, and walkthrough video—that showcases your system’s real-world potential. By the end, you’ll have a fully functional AI-powered document automation platform ready to present to peers, mentors, or even hiring managers.

2

Project 2: Learn to Work with Python for AI-Powered Document Processing

I utilized Python to process and clean mortgage documents, focusing on data preparation techniques. I worked in Google Colab, gaining proficiency in Python basics and enhancing OCR accuracy through image processing methods to ensure data readiness for AI automation.

4

Project 4: Advanced OCR Comparison and Layout-Aware Extraction

I compared three OCR engines—Tesseract, PaddleOCR, and EasyOCR—by extracting text from complex mortgage documents. I learned to clean text and utilized layout-aware tools to maintain formatting, ultimately recommending the best OCR tool for automation based on performance.

6

Project 6: Advanced RAG and Open-Source Experiments

I advanced my RAG pipeline by implementing improved chunking techniques and metadata-based filtering. I evaluated the performance of open-source models, comparing them with proprietary options to optimize the pipeline for real-world mortgage document retrieval.

8

Project 8: Gradio Chatbot with RAG Integration

I created an interactive chatbot using Gradio that integrated my RAG pipeline. This project allowed users to query documents via a web interface, enhancing the usability of my retrieval engine and making it easier to switch between different AI models.

View all works
Back to works

Credit Risk & Fraud Detection: Predictive Modelling

Collected and cleaned transactional financial datasets; engineered domain-relevant features and applied resampling strategies (SMOTE) to handle class imbalance. Performed statistical analysis across risk segments and threshold settings, monitoring model performance and identifying patterns and correlations in complex datasets. Communicated model outputs and risk findings in a clear, business-readable format for non-technical stakeholders.

Pythonscikit-learnSQLXGBoostpandas

Overview

Collected and cleaned transactional financial datasets; engineered domain-relevant features and applied resampling strategies (SMOTE) to handle class imbalance. Performed statistical analysis across risk segments and threshold settings, monitoring model performance and identifying patterns and correlations in complex datasets. Communicated model outputs and risk findings in a clear, business-readable format for non-technical stakeholders.

View all works
Back to works

Customer Purchase Prediction & Recommendation System

Analysed large transactional customer datasets to identify purchasing behaviour patterns and segment customers using statistical and classification techniques. Cleaned and validated behavioural data, applying feature transformation to prepare datasets for modelling and reporting. Translated model outputs into clear pricing and targeting recommendations for non-technical business stakeholders, supported by visual summaries.

PythonSupervised LearningpandasExcel

Overview

Analysed large transactional customer datasets to identify purchasing behaviour patterns and segment customers using statistical and classification techniques. Cleaned and validated behavioural data, applying feature transformation to prepare datasets for modelling and reporting. Translated model outputs into clear pricing and targeting recommendations for non-technical business stakeholders, supported by visual summaries.

View all works