App LogoApp name
DS

Deep Sharma

Open to work5 years experience

Data Scientist

RemoteTarget Roles: Backend Engineer • Full-Stack Developer • Frontend Specialist

Data Scientist with 5 years delivering production ML systems for Fortune 500 retail and pharmaceutical clients.

GitHub

Standing Rank

Rank Not Available

Developer Badges

Skills & Technologies

36 skills
pythonsqlpysparkjavascriptpytorchscikit-learntableaurecommendaion systemspricing & promotion analyticsdemand forecastingfeature engineeringa/b testingdrift detectionmodel monitoringcausal analysisrag pipelineshybrid dense+sparse retrievalrerankingchromadbpineconemilvuslangchaindocument processingawssagemakers3ec2azure databricksmlflowdockerci/cdjenkinskafkasparksnowflakerest apis

Work Experience

Data Scientist

Thoughtworks

Jan 2021 - Dec 2023
  • Owned the data pipeline and an XGBoost-based scoring engine within a McKinsey-led Periscope Promotion Advisor engagement for Ulta Beauty and CVS Pharmacy, scoring each SKU's price sensitivity and promotion affinity across 6 major product categories and 30,000+ products; adoption reduced uncoordinated promotions by 2.5% within 3 months.
  • Built a product recommendation and promotion-targeting system using ALS and embedding-based models, serving weekly loyalty-member sessions at scale for both clients; A/B testing showed a 7.5% lift in cross-category purchase rate over the prior rules-based system.
  • Analyzed transaction and cohort data to surface churn signals and lifecycle-value trends, informing targeted campaigns that shifted spend toward higher-margin SKUs.
  • Built distributed Spark/PySpark feature pipelines supporting daily model retraining on retail interaction data, cutting recommendation staleness from 24 hours to 4 hours.
  • Productionized ML services with MLflow, CI/CD, and Docker, including model versioning and rollback, and added drift and data-quality monitoring

Data Analyst

Planview

Jul 2020 - Dec 2020
  • Validated and quality-assured data pipelines and business-rule logic for Planview's portfolio and resource-planning platform, identifying and resolving data-flow inconsistencies that were causing a 30% error rate in quarterly leadership reports.
  • Tested and certified the data layer supporting what-if scenario planning, ensuring resource-allocation models returned accurate, consistent outputs across edge cases and stakeholder workflows.
  • Served as the data integrity point-of-contact between product, engineering, and analytics teams, documenting validation findings and aligning teams on correct business rules for capacity planning reducing back-and-forth cycles on data discrepancies.

Senior Consultant Data

Deloitte

Dec 2018 - Jun 2020
  • Managed and analyzed large-scale financial and pricing datasets (contracting, chargebacks, revenue recognition) for pharmaceutical clients, supporting data-driven decision-making.
  • Designed Tableau dashboards translating pricing and financial performance into executive-ready views for senior stakeholders, cutting manual slide prep per reporting cycle from 6 hours to 2 hours.
  • Built and automated data ingestion and validation pipelines (REST APIs, SQL) to ensure data accuracy, consistency, and reliability across reporting systems.
  • Validated and monitored large-scale ETL pipelines (Kafka, Talend, HDFS, S3) across millions of records, ensuring business rule integrity and data completeness, improved downstream model reliability by 15%.
  • Processed and transformed large datasets using Spark/PySpark for scalable analytics and model development workflows.
  • Collaborated with data engineering teams to optimize batch and streaming pipelines, improving performance and resolving data quality gaps.
  • Mentored 3–4 junior consultants, reviewing analytical work, defining tasks, and ensuring adherence to data and modeling best practices.

Software Developer

LTI, Mindtree

Jul 2013 - Nov 2018
  • Validated OTP generation for a banking client and automated tests for encryption and decryption.
  • Built backend components supporting analytics workflows for media consumption and engagement metrics.
  • Developed APIs and data-driven features (C#, SQL) contributing to reliable reporting and monitoring.
  • Improved data access performance and quality through optimized database and service-layer logic.
  • Designed APIs and database logic ensuring high data integrity and system performance.

Projects

Pricing & Promotion Decision-Support Agent (LangGraph)

LangGraphFastAPILLM
  • Designed and built a LangGraph-based reasoning agent that answers retail pricing questions by autonomously calling tools (price elasticity lookup, sales trend analysis, promotion simulation) and reasoning over multi-step results before recommending a discount depth.
  • Built a custom LangGraph StateGraph (agent → tool-call routing → tool execution → loop) rather than a prebuilt agent framework, to demonstrate full control over agent reasoning flow and tool orchestration.
  • Implemented a promotion simulation tool projecting unit lift, revenue, and margin dollar impact across discount depths, enabling the agent to explicitly weigh volume gain against margin erosion mirroring real category-management tradeoffs.
  • Deployed the agent via FastAPI with a live reasoning-trace UI, making each tool call and intermediate result visible rather than surfacing only a final answer.
  • Wrote a stubbed-LLM test harness to validate agent routing and tool-execution logic independent of live model calls, supporting fast iteration and CI-style regression checks.

Hybrid RAG Search Platform

FastAPIStreamlitMilvusFlashrankTinyBERT
  • Built a full-stack hybrid-retrieval RAG chatbot (FastAPI + Streamlit) answering questions over a ~2,000-page document set, combining dense + sparse retrieval (Milvus, BM42), multi-query expansion, and a reranking layer (Flashrank + TinyBERT).
  • Evaluated the pipeline on faithfulness, context utilization, and correctness metrics, using results to iterate on chunking strategy and prompt design.

Recommendation System on Public Retail Data (Home Depot dataset)

ALS
  • Built a collaborative-filtering + embedding-based (ALS) recommendation system on a public retail interaction dataset, evaluated with Precision / Recall against a popularity baseline.
  • Designed an offline A/B testing simulation framework to estimate CTR and revenue-lift impact of the model versus baseline.

Leaderboard Standings

Leaderboard Position Pending

Global test scores, peer standing percentiles, and algorithm leaderboard ranks are updated dynamically.

Assessment Highlights

Assessments Not Completed

Coding evaluations, system assessment results, and conceptual score badges will appear here after taking a test.

AI Collaboration Score

AI Collaboration Score Pending

Developer coding behavior, assistant cooperation, and AI pair-programming indicators are evaluated during live coding sessions.

Role Compatibility Profile

Role Compatibility Analysis Pending

Custom matching reports, candidate role compatibility percentiles, and core engineer strength profiles are processed once conceptual code screenings are complete.

Achievements

McCoy Fellowship of Distinction

McCoy Fellowship of Distinction

Texas State Merit based Scholarship recipient

Texas State Merit based Scholarship recipient

About Details

Professional Bio

Data Scientist with 5 years delivering production ML systems for Fortune 500 retail and pharmaceutical clients. Experienced in building recommendation and pricing-analytics systems, and shipping LLM-based retrieval and agentic reasoning systems from prototype to production.

Master of Science in Data Analytics & Information Systems (AI/ML)

Texas State University, San Marcos, Tx (2025 - 2026)

Bachelor of Engineering in Computer Science & Engineering

RGPV University ( - 2013)

Languages: English, Hindi