Key accomplishments

Interactive Web Application Development for Visualizing Epigenomic Functional Maps    New York City, New York  

Collaboration with NYU Langone Health, Capstone Project                                     10/2023-12/2023

  • Implemented Python script to batch convert 30GBdata into CSV format, developing with a MySQL database, and reduced the data retrieval time with indexing method, which enhanced query efficiency by 100x, from 14s to 0.14s
  • Migrated MySQL database from local to AmazonEC2, within S3 buckets, which improved user accessibility, and deployed website online, increasing user engagement by 25x, from 20 views to 500 in 2 months
  • Visualized results with diverse Python packages and charts like Matplotliband Seaborn, which enhanced UX scores
  • Leveraged the lightweight Flaskframework to process asynchronous requests via JSON format, returning with JPEG files to the frontend, and enabled seamless and effective rendering procedure, which around 7s

 

New York City Department of Environmental Protection                                  Kingston, New York

Bureau of Water Supply, Summer Graduate Intern (Data Scientist)                                06/2023-08/2023

  • Extracted350k+ rows of water quality data via API and internal database using SQL, preprocessing the data through linear interpolation and step functions; Loaded the data into Aquarius for future analytics and visualization
  • Fitted with Extra-Treesmodel in Python for feature selection across 28 features, and implemented XGBoost regressors through Grid Search tuning method to predict Total Phosphorus levels, which resulted in 83 adj-R2 score
  • Provided data-driven insights based on machine learning models, which influenced 10M+ NYC residents’ water quality; Improved water quality by guiding NYC water policy based on the identification of key spatio-temporal factors among 28+predictors; Visualized modeling prediction results via 3+ Power BI dashboards with 15+ pages

 

Henan Junyou Digital Technology Co., Ltd.                                               Zhengzhou, China

Data Scientist Intern                                                                  02/2022-04/2022

  • Conducted E-commerce platform data integration by matching 2k+ category names based onCloud word vector API and cosine similarity to facilitate the development of future RDBMS
  • Optimized internal databases by renaming 7M+commodity names through keyword extraction, which reduced data redundancy, achieving 50% storage space savings
  • ImplementedSMOTE algorithm in Python to process 4M+ imbalanced data, which helped for future data modeling
  • Developed a logistic regression model to predict price reductions for 5M+ products with 0.7 AUC and 92% accuracy, and conducted significant tests to identify factors impacting price reductions

 

Music Recommender System based on Implicit Feedback                                     03/2023-05/2023

  • Transformed 200M+music listening records with 8k+ users from ListenBrainz into Parquet files, and employed Spark SQL window functions and bucketing for user-based data splitting
  • Developeda parallel ALS recommender system using Spark on NYU High Performance Computing clusters and tuned hyperparameters, achieving Precision of 0.19 and NDCG of 0.21 on the top 100 item predictions for each user
  • AppliedPython’s Annoy library for fast search, achieving an 85% runtime improvement (from 2245s to 337s)

 

Bank Credit Card Customer Churn Warning Based on Multi-class Logistic Regression Model     02/2022-05/2022                      

  • Conducted optimal k-value selection (k=3) using the elbow method and implemented a robust K-prototype clustering algorithm in Python to efficiently cluster 10k+credit card customer data with mixed attributes
  • Processedthe categorical data via One-Hot encoding, and performed feature selection based on the Extra-Trees model to identify the top 10 significant features for insightful analysis
  • Developed a highly accurate 3-classification logistic regression model based on the One-vs-Rest method to predict credit card customer churn, achieving an outstanding accuracy rate of 99.67%

Role 1
Data scientist
Role 2
Data mining and analysis
Test Score
AI, Machine learning, Data Science
75%

Unspecified
Any
Graduate
United States
Growth
NA
Hybrid

Education


Experience


Languages

English,
Chinese