About

Hey there, I'm Ankur Patel

I am a diligent researcher with a passion in data science, motivation in healthcare, and ambition in business. A B.S. in Biomedical Engineering from UAB and industrial research experience has geared me towards a data science career. On an average day, you will probably find me working on client projects and consulting tasks, while weekends are reserved for attending meetups or participating in hackathons. Other than that, I enjoy exploring new cities, tasting new flavors and dishes from different cuisines, exercising at the gym, and watching an assortment of movies. I believe the best way to grow is to open the mind to learn and understand every part of the seen and unseen world. The key is to stay curious.

Checkout my resume

Activities besides tech

Contact me here!

Email: p.ankur.715@gmail.com

Phone: 917-721-1355

Experience

Data Engineer - Sky I.T. Group

  • Designed and maintained automated data pipelines with Python and SQL on GCP BigQuery, integrating multi-source retail POS, sales, inventory, and product data, which enabled timely analytics and reporting for business stakeholders.
  • Developed Python parsing and transformation workflows that standardized heterogeneous client data into consistent schemas and business-defined grains, improving data consistency for downstream analytics.
  • Implemented automated data validation, reconciliation, and monitoring that improved production data quality.
  • Investigated production pipeline failures using logs and root-cause analysis, then added validation checks that reduced pipeline downtime and prevented recurring errors.
  • Developed backend Python data-processing services supporting an LLM-powered analytics chatbot, providing trusted sales and inventory data for natural-language queries.
  • Partnered with analysts and business stakeholders to translate reporting and analytical requirements into data models and pipelines, delivering accessible data products that reduced report generation time and supported faster decision-making.

Jan 2024 - Jul 2026

Data Analytics Engineer - Greater New York Insurance Companies

  • Designed and optimized end-to-end data pipelines for insurance claims, policy, underwriting, violations, customer, and geographic data, delivering trusted datasets for BI reporting, predictive analytics, and executive decision-making using Python, R, SQL, Snowflake, Tableau, and SQL Server.
  • Improved data quality, consistency, and performance through validation, documentation, monitoring, and optimization of production data workflows.
  • Partnered with business stakeholders to build Tableau dashboards and Snowflake data models that enabled faster underwriting decisions and more accurate claims and pricing analysis.
  • Collaborated with data scientists to prepare feature-engineered datasets to support predictive models using DataRobot, SparkBeyond, and internal machine learning that improved risk-scoring accuracy for underwriting.

Aug 2020 - Oct 2023

Data Engineer - Executive Business Solutions Corp, NYC

  • Developed Python ETL pipelines that integrated external real-estate, U.S. Census, and public-health data via APIs into AWS-hosted SQL databases, enabling automated analytical reporting and reducing manual data-gathering effort.
  • Designed a SQL Server data warehouse for retail clients and built customer-segmentation, purchasing-analysis, and supply-chain KPI datasets, which provided business leaders with actionable insights for strategic decision-making.
  • Built automated data-processing workflows with Python and AWS EC2 that ingested, transformed, and prepared external datasets, streamlining the data preparation stage and cutting processing time.
  • Developed time-series forecasting and predictive-analytics models in Python (using pandas and scikit-learn) for financial and business datasets, delivering forecasts that supported budgeting and planning activities.
  • Delivered hands-on training in Python, SQL, Tableau, R, Excel, data analytics, and machine learning to business users and staff, increasing their self-service reporting capabilities and adoption of analytical tools.

Oct 2018 - Mar 2020

Independent Study – Business Assessment Scores Using Yelp Reviews - Verisk Analytics, NJ

  • Developed an assessment score using business and reviews json files from Yelp challenge dataset.
  • Created LDA model for topic modeling and stars after inspecting topics using pyLDAvis, and a domain specific lexicon for scores that were normalized across NAICS Codes and social networking sites.
  • Deployed model for user interaction using Flask API.

Sept 2019 - Feb 2020

Forecasting & Business Analysis Intern - Novo Nordisk, NJ

  • Processed pharmaceutical sales, demand, and product-volume data using SQL and dbt on AWS, supporting FP&A forecasting and business planning and enabling more accurate forecast inputs for the team.
  • Developed Python ETL workflows to transform and map enterprise data from multiple databases into standardized inputs for forecasting models, reducing manual data preparation.
  • Analyzed forecast data and built a Tableau Forecast Tracker to visualize Anchor Budget and Rolling Estimate forecasts, helping stakeholders monitor forecast performance and identify variances.

Jun 2019 - Aug 2019

Data Analyst - InSpirAVE, NYC

  • Analyzed e-commerce data of products and prices of selected merchant partners, compared with competitors, and collaborated with other e-commerce businesses.

Jan 2019 - Feb 2019

Researcher II - UAB Medicine, AL

  • Assisted in scientific research in diabetes to identify roles of bodyweight, genotypes, glucose/insulin/pyruvate tolerances, and proteins in insulin secretion.
  • Analyzed the data using Excel, VBA, and Python, and analyzed big genomic data using R.

May 2017 - Aug 2018

Research Assistant - Dynamic Biosciences LLC, AL

  • Performed analytical testing procedures for breweries and reported results to customers.
  • Managed the company’s sales of laboratory equipment.
  • Assisted in column chromatography and performed a mass spectrometry assay to measure purity.
  • Created graphs of the data in Graphical Analysis software.

Aug 2016 - Dec 2016, Jan 2015 - May 2015

Training

ELCHackathon 2019 - ELC Recycle Rewards Kiosk

Group won first place for developing an Amazon Echo app to interact with customers to offer rewards and an option to donate for recycling ELC products, by using an infrared sensor and LED display to identify the item. The software used were Vision API, Cloud SQL, and Python.

Sep 2019

Redline Hackathon 2019 - Seeker

Group won 2nd place for developing a SaaS app called Seeker that uses Google Cloud Video Intelligence API for object detection and time intervals, and Yuuvis for storage and Search Service API to jump to time of searched object.

Apr 2019

HackFest 2019 - SaferWay

The first routing app that provides the safest routes by using ArcGIS API and city wide traffic statistics. The app works by predicting the probability of vehicle crashes at a certain location given several inputs like precipitation chance and amount, temperature, time of day, past crashes, and sinuosity of the road.

Apr 2019

Global Legal Hackathon 2019 - Violence Reporter

Competing against 6000 participants from 24 countries, team created an audio and image classification app to notify users of real-time and location-based crimes for evacuation and for lawyers’ use during claim cases.

Feb 2019

NYC Grand Hack 2018 - MobileAI

MobileAI phone app of image or video classification to assist visually impaired users in independent mobility through real-time navigation by capturing location from the pre-trained model.

Nov 2018

Education

Bachelor of Science in Biomedical Engineering - University of Alabama at Birmingham

Birmingham, Alabama

Aug 2012 - May 2015

Master of Science in Data Science - Saint Peter's University

Jersey City, New Jersey

Sep 2018 - Feb 2020

Portfolio

Checkout a few of my projects

Mar 2020

Finance Analytics

Analyzed time-series data of stocks for stock returns, volatility calculations, OLS, back testing, predictive, along with quantopian research.

Python

View Project

Apr 2019

AWS Redshift Cluster Analysis

Extracted data from ZAGI database in PostgreSQL server, load into AWS S3, then Redshift using Python

Program: PostgreSQL, Python, S3, Redshift

View Project

Oct 2019

Heart Disease Risk Factors

Employed machine learning, ensemble, and statistical analysis for feature selection and tuned models for predicting heart disease with 93% accuracy

Program: R

View Project

May 2019

Twitter Sentiment Analysis

Analyzed tweets of 2020 Democratic Presidential Candidates from Apr 4-11 uniformly by state, and tweets from the candidates' personal Twitter accounts from Nov 1 - Mar 31

Program: R

View Project

Feb 2019

Neural Networks

Multilayer perceptron (feed forward neural network) with hidden layers on iris dataset; CNN for image classification in MNIST, SVHN, and pet dataset

Program: Python

View Project

The Truth

A day in the life of a data scientist

Hours awake 0
Hours asleep 0
Cups of coffee 0