Hi! My Name Is

Aayush Sheth

Software engineer moving into data science and applied machine learning

About Me

I'm a software engineer with five years of experience building payments and data infrastructure at Interac, IBM and Scotiabank. I'm now finishing a Master of Science in Computer Science at Georgia Tech, specializing in machine learning.

I'm most interested in data science and applied machine learning, particularly fraud and risk problems in fintech, where my payments background applies directly. The projects below are where I've taken models from raw data through to evaluation.

Education

Expected 2027

Master of Science in Computer Science (Machine Learning Specialization)

Georgia Institute of Technology

2016-2020

Honors Bachelor of Computer Science

Wilfrid Laurier University

Skills

Machine Learning

Building and evaluating predictive models (scikit-learn, XGBoost, PyTorch) for classification, feature engineering, and reinforcement learning (Q-learning); comfortable in pandas/NumPy for data prep and AUC-ROC/Sharpe ratio for evaluation

Software & Data Engineering

  • Python
  • Java
  • SQL
  • Bash
  • FastAPI
  • Flask
  • Spring Boot
  • Apache Spark
  • Kafka
  • PostgreSQL
  • Oracle DB
  • Redis

Infrastructure

  • AWS (Glue, EMR, Lambda, S3, RDS)
  • Docker
  • Kubernetes (OpenShift)
  • Terraform
  • Jenkins

Experience

August 2025 - Present

Software Engineer, E-Transfers (Core Infrastructure)

Interac

  • Leading migration of core fraud detection capabilities from legacy platform to a new microservices-based fraud engine
  • Built an asynchronous Kafka retry pipeline to offload heavy risk calculation compute, reducing synchronous transaction path execution to sub-50ms and strictly maintaining enterprise SLAs
  • Designed and implemented AWS Glue and Lambda data ingestion pipelines to migrate over 3.5M daily settlement streams to AWS S3, guaranteeing 100% state consistency during cross-region syncs
March 2022 - April 2025

Data Engineer, Data Streaming Platform

IBM

  • Engineered a dynamic, load-aware routing layer across 30+ core APIs, seamlessly shifting read/write traffic between MongoDB and Oracle to slash mean database response times by 25%
  • Wrote custom Kafka consumer-lag tracking and partition tuning algorithms for multi-terabyte ingestion pipelines, reducing message processing latency by 30% during batch workloads
  • Architected CI/CD pipelines in Jenkins and OpenShift to automate the deployment of 50+ high-scale data connectors, establishing real-time data ingestion pipelines for downstream PyTorch training runs
January - August 2021

Software Engineer, High Value Payments

Scotiabank

  • Wrote deterministic Java rules to augment automated risk classification engines, reducing false-positive transaction blocks and raising payment success rates from 86% to 90%
  • Built a parallelized Python ETL pipeline to clean and structure training datasets for machine learning models, dropping preparation runtimes from 60 minutes to 5 minutes
  • Created real-time performance dashboards and automated alerting using Python and SQL, cutting operational incident response times for high-value transactional failures by 40%

Projects

Optimizing Steal Decisions in Baseball Using Machine Learning

React, D3.js, FastAPI, PostgreSQL, scikit-learn, XGBoost

A live Go/No-Go dashboard for MLB stolen-base decisions. Two models per situation, one for attempt likelihood and one for success likelihood, feed an expected-value engine alongside a 24-state Markov run-expectancy matrix.

Steal decision dashboard showing a Go recommendation, a risk meter, an expected-value chart and a comparison of seven models
The decision dashboard: enter the game situation and it returns a Go/No-Go call with the expected runs for holding versus attempting.

Built on about 2M play-by-play events from Retrosheet joined with Statcast player metrics (2015 to 2025), covering about 387K steal opportunities. Trained on 2015 to 2020 and tested on 2021 to 2025.

Compared seven classifiers by AUC-ROC, since accuracy misleads when only 7% of opportunities end in an attempt. XGBoost scored best on attempt prediction at 0.804, and a Random Forest (0.789) powers the live dashboard. Success prediction sits around 0.58 across models, in line with prior research.

Built with a team of four at Georgia Tech. I worked across the React/D3 frontend, the data pipeline, and model development and evaluation.

Algorithmic Trading Strategy Evaluation (Q-Learning)

Python, NumPy, pandas

Framed single-stock trading as a reinforcement learning problem: a Q-Learning agent with Dyna-Q planning over discretized MACD, Momentum and Stochastic Oscillator states (1000 states, 3 actions).

On JPM in-sample (2008 to 2009), the learner returned about 55% cumulatively, against 28.8% for a hand-built rules strategy and 1.2% for buy-and-hold. Its Sharpe ratio fell from 3.45 to 0.73 as simulated market impact rose from 0% to 5%.

In-sample chart for JPM, 2008 to 2009: the Q-Learning strategy ends near 1.55 times starting value, the manual strategy near 1.29 and buy-and-hold near 1.01
In sample, 2008 to 2009: the learner (blue) pulls well ahead of the manual strategy (red) and buy-and-hold (purple).

Out of sample (2010 to 2011) that edge mostly disappeared. All three strategies lost money, and the learner finished only slightly ahead of buy-and-hold. I attribute this to overfitting: the training window was the post-2008 recovery, a strongly trending market, and the test window was choppier with shorter-lived trends.

Out-of-sample chart for JPM, 2010 to 2011: all three strategies end below starting value, with the Q-Learning strategy slightly above the manual strategy and buy-and-hold
Out of sample, 2010 to 2011: all three end below where they started.

Contact

The best way to reach me is on LinkedIn: linkedin.com/in/aayusheth

My code is on GitHub: github.com/aayush4249