Data Scientist II · Applied AI Systems thinking · Open to FT roles & consulting

Souvik
Samanta

Data Scientist II → Applied AI Systems Engineer

I build production AI systems at scale — combining deep data science expertise with backend engineering, distributed systems thinking, and AI infrastructure. My view: ML that can't survive production systems constraints isn't finished work. That shows up in churn prediction across millions of learners at Coursera, an agentic AI workflow connecting LLMs to real business data, and 6+ years building end-to-end — including founding a full-stack fintech trading platform.

30%
Marketing Cost Saved
20%
Lead Quality Uplift
15%
Revenue Growth
6+
Years in Data Science
Souvik Samanta — Senior Data Scientist

Data science + systems engineering, applied to production AI

I'm a Senior Data Scientist with 6+ years across EdTech, travel, real estate, and fintech. Currently at Coursera, I own lifecycle analytics and CRM measurement across millions of learners — including an agentic AI workflow that connects an LLM directly to production data so stakeholders can self-serve root-cause analysis. I think of myself as an Applied AI Systems Engineer: someone who brings both ML judgment and systems rigor to AI that has to actually survive production.

Why this framing? Most specialists go deep in one dimension. I bring together backend engineering (FastAPI, PostgreSQL, Kubernetes), distributed systems thinking (CAP theorem, event streaming, consistency), AI infrastructure (RAG, multi-agent orchestration, evaluation pipelines), product analytics (A/B testing, experimentation, business impact), and production discipline (CI/CD, observability, incident response).

Proof point: Co-founded and built Quantbot Securities from scratch during COVID — full-stack: ML signal models, async FastAPI backend, Kubernetes orchestration, real-time WebSocket trading infrastructure. 5+ years operation, real customers, real trading profits.

Before data science, structural engineer at M. N. Dastur — learned systems thinking, rigorous validation, and the cost of errors. MBA in Finance & Analytics, IMI New Delhi.

Machine Learning
XGBoost LightGBM Scikit-learn Random Forest K-Means ARIMA Prophet
Languages & Data
Python SQL pandas NumPy Bayesian Stats
Data Engineering
Databricks BigQuery Airflow Docker Kubernetes Google Cloud
BI & Backend
Power BI Looker Braze FastAPI Django REST

GitHub Contributions

Production ML, lifecycle analytics, and experimentation at Coursera alongside personal projects — all under one account, including private repo activity.

souviksamanta95 Includes private contributions — Coursera, side projects, and this portfolio
Loading contributions…

The journey so far

Six years across EdTech, travel, real estate, and fintech — building ML systems and data strategies that moved real business needles.

Sept 2024 — Present
Data Scientist II — Lifecycle Analytics & ML
Coursera  ·  EdTech · Remote
  • Built and deployed churn prediction models (XGBoost/LightGBM) on Databricks, identifying at-risk learners across millions of users
  • Developed customer segmentation using K-Means and RFM analysis to personalise CRM campaigns across email, push, and in-app channels
  • Designed and executed A/B tests, holdout experiments, and lift studies using Bayesian and frequentist frameworks
  • Automated ML feature pipelines on Databricks for model scoring and campaign audience generation at scale
  • Thought partner to CRM and marketing leadership — translating model outputs into targeting and channel strategy
Millions of users Bayesian A/B framework
Jan 2021 — Jan 2026 Parallel · Personal Project
Co-Founder & ML Engineer
Quantbot Securities Pvt. Ltd.  ·  Fintech · Remote

Run independently, evenings/weekends, alongside full-time roles below (not a sequential job) — pivoted into mySerenity.in after this platform was shut down for regulatory reasons.

  • Co-founded and built an automated copy-trading platform from scratch during COVID — real customers, real revenue, 5 years of operation
  • Built ML models for trade signal generation, price movement prediction, and portfolio risk analysis
  • Designed a high-availability distributed backend using FastAPI, Django, Docker, and Kubernetes with async broker API integrations
  • Led full backend architecture, cloud infrastructure, and all engineering decisions end-to-end as a solo technical founder
  • Managed end-to-end: client relationships, team of 4, compliance, invoicing, and product roadmap
Auto-scaling cloud infra Real trading profits
Sept 2022 — Aug 2024
Business Analyst — Customer Insights & ML Analytics
Travelopia  ·  Travel · Remote
  • Built propensity-to-purchase and lead scoring models (Logistic Regression, Random Forest, XGBoost) improving targeting precision significantly
  • Developed LTV forecasting and time-series demand models (ARIMA, Prophet) for revenue planning and inventory decisions
  • Built automated data pipelines using Docker and Google Cloud with near-real-time refreshes for ML feature stores and reporting
  • Designed Power BI dashboards integrating multi-source data for cross-brand visibility used by senior leadership weekly
30% marketing cost reduction 20% lead quality uplift
Nov 2021 — Aug 2022
Business Analyst
NoBroker Technologies  ·  PropTech · Bangalore
  • Identified high-converting customer cohorts using segmentation and funnel analytics, directly reshaping growth team targeting strategy
  • Delivered analytical insights on growth and sales performance driving 15% revenue improvement
  • Automated reporting pipelines using Python (pandas, NumPy), cutting report turnaround by 30%
  • Designed interactive dashboards for sales and marketing visibility in Google Data Studio
15% revenue improvement 30% faster reporting
Apr 2021 — Oct 2021
Senior Business Associate — Analytics & Strategy
Tech Mahindra  ·  IT Services · Kolkata
  • Supported client-facing data analytics projects and campaign optimisation initiatives in the Big Data and Analytics domain
  • Built tracking systems for process monitoring and performance reporting
  • First enterprise analytics role — foundation for understanding how data decisions are made at scale
Aug 2017 — Jul 2018
Structural Engineer
M. N. Dastur & Co.  ·  Engineering · Kolkata
  • Structural design and load analysis for large-scale steel plant projects using STAAD Pro and CAD
  • Built precision in calculations and rigorous validation habits — the same discipline that defines good data science
  • The foundation for thinking in systems, validating assumptions, and communicating technical findings to non-technical stakeholders
⚡ Core ML
XGBoost LightGBM Random Forest Logistic Reg. K-Means Scikit-learn
📈 Forecasting & Stats
A/B Testing Bayesian Stats ARIMA Prophet Hypothesis Testing LTV Modelling
🛢 Data Engineering
Databricks BigQuery Airflow Docker Kubernetes Google Cloud
📊 BI & Languages
Python SQL Power BI Looker Braze
🔧 Backend & APIs
FastAPI Django REST Async Processing CI/CD

Impact that speaks in numbers

6+ years of shipped ML systems: from churn prediction at Coursera (millions of users) to founding mySerenity.in (guardrailed mental health AI, pivoted from Quantbot's trading platform) as a personal project alongside full-time work. Every project combines data science depth with systems engineering lessons learned.

Coursera ML · Lifecycle Analytics

Churn Prediction at Millions Scale

Impact: Built and deployed churn prediction models (XGBoost/LightGBM) on Databricks identifying at-risk learners across millions of users, powering personalized retention interventions (email, push, in-app). Developed K-Means + RFM segmentation for multi-channel CRM personalization. Designed Bayesian + frequentist A/B testing framework for campaign optimization. Automated ML feature pipelines for real-time campaign audience scoring.

Millions of users 3 targeting channels Bayesian A/B framework
XGBoost LightGBM Databricks K-Means Feature Stores Braze Looker
Travelopia Predictive Modelling · Data Eng

Lead Scoring & LTV Forecasting

Impact: Built propensity-to-purchase models (Logistic Regression → XGBoost) reducing low-propensity marketing spend by 30% while improving lead quality by 20%. Developed LTV forecasting (ARIMA, Prophet) for revenue planning & inventory allocation. Built automated data pipelines on Google Cloud (BigQuery, Cloud Run, Docker) with near-real-time feature store refreshes. Designed Power BI dashboards for senior leadership visibility.

30% cost reduction 20% quality uplift <15% MAPE on LTV
Logistic Regression Random Forest XGBoost ARIMA Prophet BigQuery Google Cloud Docker Power BI
mySerenity.in Healthtech · Personal Project

Mental Health Platform with ARIA (Founder)

Impact: An independent personal project, built and run in parallel with full-time employment. Same company/entity as Quantbot below — pivoted to a new identity after Quantbot's copy-trading platform was shut down due to regulatory constraints. Development started January 2026; live at myserenity.in since June 2026 — a platform connecting individuals with mental health professionals (psychiatrists and psychologists), with a community feature, secure consent-based video consultations, billing, and prescription management. At its core is ARIA, a guardrailed, specialized mental health chatbot capable of long conversations with patients that summarizes each conversation to help clinicians analyze patient state efficiently between sessions. Built the full stack solo: Django backend, Next.js frontend, PostgreSQL + Redis, real-time video via Janus (WebRTC), and private S3 buckets storing consent-based call recordings.

Live since June 2026 Guardrailed clinical AI Solo full-stack build
Django Next.js PostgreSQL Redis Janus (WebRTC) Private S3 Guardrailed LLM Chatbot Consent-Based Recording
Quantbot Securities Fintech · Personal Project

Automated Copy-Trading Platform (Founder, shut down)

Impact: An independent personal project, built and run in parallel with full-time employment. Co-founded & built end-to-end cloud platform from scratch during COVID, serving real customers with real trading profits. Led all backend architecture + ML engineering — trade signal generation (LSTM, Transformer concepts), price movement prediction, portfolio risk models. Designed distributed system: FastAPI async request handling, Django ORM, PostgreSQL transactions, Redis caching, WebSocket real-time streaming. Containerized with Docker, orchestrated via Kubernetes with auto-scaling. Real-time broker API integrations. Shut down in early 2026 due to regulatory constraints on retail algorithmic copy-trading — the same company/entity then pivoted into mySerenity.in above.

Real trading profits 5 years operation Founder's full ownership
FastAPI Django PostgreSQL Redis Kubernetes Docker WebSocket LSTM/Transformers Async APIs ML Risk Models
NoBroker Technologies Growth Analytics · Funnel

Funnel Analytics & Cohort Segmentation

Impact: Identified high-converting cohorts invisible to volume-only metrics. Built segmentation layer surfacing true conversion paths, directly reshaping growth team targeting and driving 15% revenue improvement. Automated Python pipelines (Pandas, NumPy) reducing reporting turnaround from weeks to days (30% faster). Built interactive Google Data Studio dashboards for daily channel performance tracking.

15% revenue growth 30% faster reporting Cohort-driven strategy
Python Pandas NumPy SQL Cohort Analysis Segmentation Google Data Studio

Now & Next

Now (2024–2025): Full-time at Coursera + selective consulting on customer analytics, lifecycle ML, and experimentation. Next (2025–2027): Deepening systems engineering + AI infrastructure while maintaining high-impact data science work. Looking for roles that reward both technical depth and business impact.


Availability
Open to full-time roles
15–20 hrs/week consulting
Remote · IST (UTC+5:30)
Async-friendly for US/EU teams
Retainers & project-based
Book a Free 30-min Call
🎯
Customer Analytics
Churn prediction, lead scoring, propensity models, LTV forecasting, and segmentation — built for targeting precision and retention, not just model accuracy.
📊
CRM & Lifecycle Analytics
End-to-end measurement of campaigns across email, push, and in-app. Currently doing this at scale for a global EdTech platform with millions of users.
🧪
Experimentation
A/B tests, holdout experiments, and lift studies designed to be statistically rigorous and business-relevant. Teams can actually act on the results.
⚙️
Data Engineering
Automated pipelines on Databricks, BigQuery, Google Cloud, and Docker. ML feature stores, ETL workflows, and real-time data infrastructure built to last.
📈
Dashboards & BI
Power BI, Looker, and Google Data Studio — built for executive decision-making, not operational visibility. Senior leaders should understand it in under 30 seconds.
🤝
Data Strategy
Identify what to measure, build easy-to-read frameworks, and coach teams to make confident data-driven decisions. For non-technical leaders who need clarity fast.
🧠
AI Systems & LLM Infrastructure
RAG systems, multi-agent orchestration, evaluation frameworks, and observability. Building the next generation of AI-powered products that work reliably in production.
Backend & Systems Engineering
FastAPI, PostgreSQL, Redis, Kubernetes, async architectures. Full-stack capability: design scalable APIs, deploy containerized services, monitor reliability at scale.

What colleagues say

Verified LinkedIn recommendations from people who've worked with me directly.

"

I had the pleasure of working with Souvik at Travelopia for over two years. He consistently showcased exceptional skills in ML modeling using Python, delivering impressive results on complex projects. Beyond his technical expertise, Souvik is a remarkable team player, fostering collaboration and encouraging open dialogue among team members. His positive attitude and strong work ethic greatly enhanced our team dynamic. I am confident that Souvik will be a valuable asset to any organization he joins, bringing both expertise and a collaborative spirit to the table.

TP
Tejash Popate
Business Analyst · Travelopia
Worked on the same team
View on LinkedIn
"

I had the pleasure of working with Souvik for more than 1.5 years at Travelopia, and I can confidently attest to his exceptional skills and contributions. Souvik was an invaluable asset to our team, delivering high-quality work on various analytics projects, including building dashboards in Power BI, analysis using Python, and developing machine learning and factor-based models. His dedication, strong work ethic, and technical expertise enabled him to tackle complex projects with ease. What impressed me most about Souvik was his willingness to support and collaborate with other team members, fostering a spirit of teamwork and knowledge sharing. I highly recommend Souvik for any future roles.

JK
Jyoti Kumar
Analytics Leader · Data Science & Predictive Modelling
Managed Souvik directly
View on LinkedIn

Deep dives & interview notes

Interactive study modules I'm building for myself while preparing for Applied AI Systems Engineer roles — staff/senior-level depth on ML fundamentals and agentic AI, published here as I go.

Module 01 Live

Agentic AI: architecture & failure modes

What an agent actually is, the think→act→observe loop, memory vs. context window vs. vector DB, and the failure modes that separate a systems engineer's intuition from a framework demo.

Agent architecture MCP RAG Evaluation
Open module
Module 02 Live

Shallow ML depth: XGBoost, LightGBM & logistic regression

How gradient boosting actually learns, level-wise vs. leaf-wise tree growth, marketing vs. fraud tuning side by side, and the caveats (leakage, calibration, monotonicity) that separate senior from junior.

XGBoost LightGBM Calibration Fraud vs. marketing
Open module
Module 03 In progress

Time series: ARIMA, SARIMA & forecasting gotchas

Stationarity, autocorrelation, when ARIMA/SARIMA/Prophet actually apply, and where forecasting quietly breaks in production.

ARIMA SARIMA Prophet
Module 04 In progress

Deep learning: when, why, and the gotchas

When DL actually beats shallow ML, core architecture vocabulary, and the assumptions that trip people up in interviews.

Architectures Attention Embeddings
Module 05 In progress

ML pipelines: deployment, versioning & drift

Deployment patterns, model versioning, drift detection, and the monitoring stack senior/staff interviewers probe for.

MLflow Drift detection Monitoring

Writing

Career lessons, personal stories, and the occasional data take — things I've lived, not just read about.

Loading posts…

Have a problem worth solving?

Whether you're looking for a senior data scientist to join your team, or a consultant to help you make sense of your data — I'm happy to talk.

✓ Message sent! I usually reply within 24 hours.