App LogoApp name
JL

Josh Lloyd

Open to work10 years experience

Data Engineer

Mandeville, LATarget Roles: Backend Engineer • Full-Stack Developer • Frontend Specialist

Data Engineer | AI/ML, Cloud Infrastructure, Data Platforms

GitHub

Standing Rank

Rank Not Available

Developer Badges

Skills & Technologies

70 skills
pythonsqljavascriptnode.jsbashrjavadagstermeltanosingerdbtsparkpysparketleltreverse-etldata governancesnowplowmatillionawslambdaecssagemakerkinesiss3gcpbigquerygcsdockerkubernetesterraformterraspacepulumisnowflakepostgresqlmysqldynamodbredshiftduckdbneo4jpandasnumpyscikit-learnclearmlmatplotlibseabornplotlyjupyterdomosisensecube.jsamazon quicksighthubspotzoho crmhighlevelmarketomicrosoft teamszoomjiraconfluencemicrosoft officefhirhl7langchainragmcphipaagdprci/cdtdd

Work Experience

Founder

The Data Dude - Data Dude Intelligence

Sep 2025 - Present Remote
  • Integrated HubSpot, Zoho CRM, HighLevel, and CSV data into one Amazon QuickSight master sales dashboard.
  • Signed 10 full-time consulting clients in 3 months across multiple industries with weekly strategy sessions.
  • Delivered BI and strategy consulting to small business owners across retail, services, and professional sectors.
  • Delivered a Meltano, GCP BigQuery, DuckDB, and Preset data platform for a marketing agency client.

Principal Data Engineer

Widen & Acquia

Apr 2021 - Aug 2025 Remote
  • Cut Matillion ETL/ELT costs 90% by replacing the managed instance with Dagster and Meltano/Singer.
  • Architected an internal AI gateway with RAG and a unified data access layer for all product LLM calls.
  • Reduced Snowflake extraction and load time 95% by building a novel open-source data platform.
  • Built 10+ Singer taps and contributed 10+ times to Dagster, Meltano, dbt, and Cube.js.
  • Led implementation of an MCP server enabling LLM self-service querying across enterprise data.
  • Rebuilt Cube.js on FIPS-compliant, zero-vulnerability Kubernetes infrastructure meeting FedRAMP requirements.
  • Designed and built 2 data platforms from scratch managing 55+ TB with 99% uptime.
  • Ingested data from 40+ sources including APIs, Postgres/MySQL/DynamoDB, AWS S3, and GCP GCS.
  • Produced 100+ dashboards and charts across the company using Domo, Sisense, Matplotlib, Seaborn, and Plotly.
  • Conducted 500+ code reviews for cross-team collaborators enforcing quality and engineering standards.
  • Led Acquia data team technology strategy and standards that improved productivity, efficiency, and reliability.
  • Implemented enterprise data governance program that increased data reliability and trust company-wide.
  • Guided Snowplow real-time behavioral tracking evaluation and engineering team rollout.
  • Maintained a 200+ page data team handbook in Confluence on team dependencies, tech debt, and best practices.
  • Facilitated Agile delivery as Scrum Master in JIRA; prioritized backlog, tech debt, and cross-project dependencies.
  • Authored cross-functional documentation and meeting notes in Google Docs for 5+ cross-departmental stakeholder groups.
  • Ran cluster and statistical analyses in Python and Pandas that informed UX improvements and user satisfaction gains.
  • Analyzed data with SageMaker, Python, scikit-learn, and ClearML to deliver insights on customer churn.
  • Aligned data strategy and infrastructure to business objectives via scope, milestone, and success-metric communication.
  • Established frameworks for governance, code standards, testing/alerting, and security/privacy.
  • Mentored junior engineers and established data engineering best practices across the team.
  • Optimized pipelines, data models, queries, and databases for performance across large datasets.
  • Curated, labeled, and automated pipelines for recommendation machine learning model training.
  • Built daily pipelines in Python, SQL, Node.js, Docker, Kubernetes, AWS, Snowflake, Domo, and Sisense.
  • Built warehouse, compute, and network infrastructure from scratch with Terraform/Terraspace and Pulumi.
  • Owned enterprise data platform product roadmap and coordinated across all internal departments.
  • Identified data anomalies through EDA with Pandas and Plotly that informed UX improvement projects and customer churn.

Co-Founder / Treasurer / Data Architect

Yuniku & Doona Dev

Jan 2017 - Present Provo, UT / Córdoba, Argentina
  • Built fault-tolerant, scalable AWS data integrations for startup clients across healthcare and SaaS.
  • Saved $10K+ annually by optimizing cloud infrastructure while preserving reliability and performance.
  • Managed $100K and $180K annual revenue, payroll, and accounting in years 1 and 2 respectively.
  • Acquired a partner software development company to expand delivery capacity and year-two revenue.
  • Architected virtual pipelines for real-time FHIR and HL7 API access from EHRs, payers, and providers.
  • Maintained MySQL, PostgreSQL, Neo4j, and DynamoDB databases across full DevOps pipelines.
  • Containerized workloads with custom Docker images across dev, stage, and production CI/CD environments.
  • Built a LangChain RAG proof of concept using Selenium, Mistral/Ollama, and Hugging Face models.
  • Delivered a Meltano, GCP BigQuery, DuckDB, and Preset data platform for a marketing agency client.

Data Engineer (Contract)

The Church of Jesus Christ of Latter-day Saints (via Oasis)

Jan 2020 - Sep 2021 Riverton, UT
  • Doubled online-referral baptisms by maintaining APIs supporting production machine learning models.
  • Ran an A/B test with 100,000+ participants to focus marketing on high-value audiences.
  • Built Docker, dbt, Spark/Pyspark, and AWS data apps supporting marketing to millions of online visitors.
  • Led 3–6 person teams adopting and troubleshooting AWS SageMaker for ML engineering workloads.
  • Curated a multi-TB Redshift warehouse supporting 30+ applied, predictive, and prescriptive analytic use cases.
  • Automated hundreds of deployments with Terraform and Azure DevOps continuous delivery pipelines.
  • Ingested and reverse-ETL'd data from Facebook, Google, and Marketo APIs for marketing targeting.
  • Coordinated daily with team and manager via Microsoft Teams and Zoom across remote delivery workstreams.
  • Designed ETL/ELT pipelines in Python, SQL, JavaScript, Terraform, Bash, and AWS Lambda/ECS.
  • Sourced and quality-scored data for 10+ projects across disparate internal customer groups.
  • Documented 1,000+ data fields and processes to accelerate engineer and analyst onboarding.
  • Processed 10,000+ daily Kinesis records via AWS Lambda for real-time ML predictions.
  • Implemented TDD and CI/CD unit testing with dbt and Python across pipeline development.
  • Resolved batch and real-time production incidents as on-call data engineer.
  • Conducted EDA with Pandas, NumPy, Matplotlib, Seaborn, Plotly, and Jupyter Notebooks to develop algorithm.
  • Automated Google Analytics feeds and charts to improve paid-media performance by 5%+.

Director of Operations

OptoQuest (Cleveland Clinic company)

Sep 2014 - Dec 2019 Cleveland, OH
  • Raised $1M in Cleveland Clinic seed funding as a primary member of the OptoQuest executive team.
  • Reduced surgery prediction error ~10% through EDA, Python scikit-learn, and SQL-driven research.
  • Supported 500+ surgeries via AWS cloud application built and operated by a managed Agile team.
  • Deployed Python ML algorithms into C# production environment for clinical simulation workflows.
  • Recruited an 8+ physician customer advisory panel that expanded clinical awareness and adoption.
  • Published 10+ professional and academic reports through partner and internal collaborations.
  • Delivered quarterly board reports with Excel financial forecasts on product, clinical, and ops metrics.
  • Served as Scrum Master in JIRA for an Agile team, tracking tasks, tech debt, and project dependencies.
  • Built and maintained multivariate linear regression models in Excel for the surgery simulation product.
  • Generated professional reports and contract reviews in MS Office for clinical stakeholders and regulatory documentation.
  • Rebuilt AWS infrastructure and data pipelines to AWS Well-Architected Framework standards.
  • Deployed Docker-based Linux Python service running the core simulation engine.
  • Authored HIPAA, GDPR, and ISO 13485 compliant quality and data management documentation.
  • Directed product strategy, marketing, and board operations as effective CEO of the startup.

Projects

Projects Not Populated

Personal applications, open-source work, and code repos will show here.

Leaderboard Standings

Leaderboard Position Pending

Global test scores, peer standing percentiles, and algorithm leaderboard ranks are updated dynamically.

Assessment Highlights

Assessments Not Completed

Coding evaluations, system assessment results, and conceptual score badges will appear here after taking a test.

AI Collaboration Score

AI Collaboration Score Pending

Developer coding behavior, assistant cooperation, and AI pair-programming indicators are evaluated during live coding sessions.

Role Compatibility Profile

Role Compatibility Analysis Pending

Custom matching reports, candidate role compatibility percentiles, and core engineer strength profiles are processed once conceptual code screenings are complete.

Achievements

Crocker Innovation Fellow

Crocker Innovation Fellow (BYU, 2012) - 1 of 20 selected from 150+ applicants.

Data Scientist in Python Certification

Data Scientist in Python, DataQuest.io (170+ assignments, 200+ hours).

AWS Solutions Architect Associate

AWS Solutions Architect Associate.

About Details

Professional Bio

Data Engineer and Founder with extensive experience in building scalable data platforms, AI/ML pipelines, and cloud infrastructure. Proven track record in leading data strategy, optimizing ETL/ELT processes, and delivering actionable insights for diverse industries.

Master of Science in Entrepreneurial Biotechnology

Case Western Reserve University ( - 2015)

Bachelor of Science in Molecular Biology (Minor in Business)

Brigham Young University ( - 2014)

Languages: English, Hindi