Key accomplishments
Interactive Web Application Development for Visualizing Epigenomic Functional Maps    New York City, New York Â
Collaboration with NYU Langone Health, Capstone Project                                     10/2023-12/2023
- Implemented Python script to batch convert 30GBdata into CSV format, developing with a MySQL database, and reduced the data retrieval time with indexing method, which enhanced query efficiency by 100x, from 14s to 0.14s
- Migrated MySQL database from local to AmazonEC2, within S3Â buckets, which improved user accessibility, and deployed website online, increasing user engagement by 25x, from 20Â views to 500Â in 2 months
- Visualized results with diverse Python packages and charts like Matplotliband Seaborn, which enhanced UX scores
- Leveraged the lightweight Flaskframework to process asynchronous requests via JSON format, returning with JPEG files to the frontend, and enabled seamless and effective rendering procedure, which around 7s
Â
New York City Department of Environmental Protection                                  Kingston, New York
Bureau of Water Supply, Summer Graduate Intern (Data Scientist)Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â 06/2023-08/2023
- Extracted350k+ rows of water quality data via API and internal database using SQL, preprocessing the data through linear interpolation and step functions; Loaded the data into Aquarius for future analytics and visualization
- Fitted with Extra-Treesmodel in Python for feature selection across 28 features, and implemented XGBoost regressors through Grid Search tuning method to predict Total Phosphorus levels, which resulted in 83 adj-R2 score
- Provided data-driven insights based on machine learning models, which influenced 10M+ NYC residents’ water quality; Improved water quality by guiding NYC water policy based on the identification of key spatio-temporal factors among 28+predictors; Visualized modeling prediction results via 3+ Power BI dashboards with 15+ pages
Â
Henan Junyou Digital Technology Co., Ltd. Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Â Zhengzhou, China
Data Scientist Intern                                                                  02/2022-04/2022
- Conducted E-commerce platform data integration by matching 2k+ category names based onCloud word vector APIÂ and cosine similarity to facilitate the development of future RDBMS
- Optimized internal databases by renaming 7M+commodity names through keyword extraction, which reduced data redundancy, achieving 50% storage space savings
- ImplementedSMOTE algorithm in Python to process 4M+ imbalanced data, which helped for future data modeling
- Developed a logistic regression model to predict price reductions for 5M+ products with 0.7 AUC and 92% accuracy, and conducted significant tests to identify factors impacting price reductions
Music Recommender System based on Implicit Feedback                                     03/2023-05/2023
- Transformed 200M+music listening records with 8k+ users from ListenBrainz into Parquet files, and employed Spark SQL window functions and bucketing for user-based data splitting
- Developeda parallel ALS recommender system using Spark on NYU High Performance Computing clusters and tuned hyperparameters, achieving Precision of 0.19 and NDCG of 0.21 on the top 100 item predictions for each user
- AppliedPython’s Annoy library for fast search, achieving an 85% runtime improvement (from 2245s to 337s)
Bank Credit Card Customer Churn Warning Based on Multi-class Logistic Regression Model     02/2022-05/2022                     Â
- Conducted optimal k-value selection (k=3) using the elbow method and implemented a robust K-prototype clustering algorithm in Python to efficiently cluster 10k+credit card customer data with mixed attributes
- Processedthe categorical data via One-Hot encoding, and performed feature selection based on the Extra-Trees model to identify the top 10 significant features for insightful analysis
- Developed a highly accurate 3-classification logistic regression model based on the One-vs-Rest method to predict credit card customer churn, achieving an outstanding accuracy rate of 99.67%
Education
-
September 1 2022 - Present
New York University
Master of Science in Data Science
-
September 1 2018 - June 30 2022
Yunnan University, China
Bachelor of Engineering in Data Science
Experience
-
June 5 2023 - August 11 2023
New York City Department of Environmental Protection
Summer Graduate Intern (Data Scientist)
-
February 14 2022 - April 18 2022
Henan Junyou Digital Technology Co., Ltd.
Data Scientist Intern
