App LogoApp name
YV

Yuvraj Verma

Open to work4 years experience

AI/ML Engineer

Noida, IndiaTarget Roles: Backend Engineer • Full-Stack Developer • Frontend Specialist

AI/ML Engineer | Production GenAI, Agentic Systems & Industrial Computer Vision

GitHub

Standing Rank

Rank Not Available

Developer Badges

No badges earned yet

Skills & Technologies

39 skills
pythonraglangchainlanggraphllamaindexmcp serversmulti-agent orchestrationprompt engineeringlangsmithyolosam segmentationvlmsdeepsorttensorrtvllmtritonlivekitvapitwiliowhisperelevenlabsstable diffusionawsazuredockerjenkinsgithub actionsfastapiflasksocket.iocelerypostgresqlpgvectorbigquerysnowflakeelasticsearchmongodbredissupabase

Work Experience

Software Engineer – AI/ML Python

HestaBit Technologies Private Limited

Nov 2022 - Present Noida, India
  • Involved in end-to-end delivery of 10+ production-grade AI projects across enterprises and startups, spanning Multimodal AI, Computer Vision, NLP, and Machine Learning, from requirement gathering through development, deployment, and post-launch support.
  • Partnered with clients across 6+ countries during presales and discovery, translating business requirements into technical architecture before driving development and delivery.
  • Specialized in multi-agent orchestration systems, LLM quantization, inference optimization, and deep learning architectures, consistently achieving 90%+ accuracy and performance benchmarks in production.

Associate Data Scientist

Genterpretr

Jul 2022 - Nov 2022 Remote, India
  • Led data collection, cleaning, and model development for OncopretR, a genetic-profile-driven therapy recommendation engine delivering clinical, molecular, and drug-response insights to oncologists and researchers.
  • Streamlined the OncopretR/OncoBase data-processing API, improving system efficiency by 10% and enabling faster insight delivery across the platform.
  • Trained and deployed ML and deep learning models across recommendation systems, predictive analytics, and classification use cases, improving model accuracy by 15–20% through iterative feature engineering and hyperparameter tuning.
  • Built a recommendation engine achieving measurable lift in prediction relevance, and delivered classification models for real-world clinical use cases under production-style evaluation.

Projects

QBit – AI-Powered Enterprise Federated Query & Intelligent Data Access Platform

PythonAgentic AIAI AgentsMulti-Agent SystemsMCP ServersMicrosoft Foundry Agent ServiceAzure OpenAIMicrosoft Entra IDSpice RuntimeDataHubFederated QueryingSnowflakeDatabricksBigQueryRAGSaaS ArchitectureCloud SecurityREST APIs
  • Led Agentic AI workflows using MCP Servers and Microsoft Foundry Agent Service with Azure OpenAI for intelligent query planning across structured data and chat sources like Slack and Teams.
  • Built federated query architecture using Spice Runtime and DataHub cataloging across Snowflake, Databricks, and BigQuery, secured with Microsoft Entra ID authentication.
  • Implemented secure AI-powered data discovery, governance, and role-based analytics across domains, achieving ~80% query relevance.

HestaVoice – AI-Powered Voice Agent Platform for Enterprise Conversational Automation

PythonFastAPILiveKitTwilioVoice AIConversational AILLMsSTTTTSAI AgentsCRM IntegrationsWebSocketsPostgreSQLCloud Infrastructure
  • Architected real-time voice infrastructure on LiveKit, integrating multiple STT, TTS, and LLM providers plus multiple telephony providers, enabling flexible, provider-agnostic deployment for enterprise clients.
  • Built an intuitive workflow and campaign builder with native CRM integrations, empowering non-technical users to launch production-ready voice agents in minutes.
  • Delivered low-latency, real-time conversational experiences at scale, powering enterprise deployments across multiple concurrent tenants.

JAI OD – AI-Driven Video Intelligence Platform

PythonYOLO (v8–v11)DeepSORTGPT-4o VisionGemini 2.5 ProVLMsOpenCVDocker ComposeRTSPFastAPISocket.IOTesla T4MultiprocessingGPU Inference
  • Owned backend architecture and YOLO–VLM pipeline integration across GPU clusters, delivering self-healing, production-grade surveillance infrastructure for real-time monitoring.
  • Designed end-to-end pipeline: overhead fisheye capture → dual-YOLO detection → DeepSORT tracking → VLM verification (GPT-4o Vision / Gemini 2.5 Pro) → ROI alerting, deployed across multiple Jindal Steel plants.
  • Delivered automated plate counting and crane lift-state detection, replacing hours/day of manual tallying with real-time AI monitoring across concurrent camera streams.
  • Owned self-healing service architecture (Docker Compose, RTSP recovery, multiprocessing workers) and server procurement spec for dual-GPU, 64-core deployment nodes.

Synthexa – AI-Powered Document Query System

PythonFastAPIAdvanced RAGMultimodal RAGQwen2InternVLElasticsearchHybrid SearchSemantic SearchRe-rankingVector SearchSentence TransformersCeleryvLLMNVIDIA A100DockerLinux
  • Developed Advanced and Multimodal RAG pipelines using hybrid search, semantic retrieval, reranking, and SQL agents for structured Excel retrieval.
  • Optimized indexing and deployment, cutting indexing time ~34% and infrastructure cost ~25% through self- hosted, cloud-free architecture.
  • Integrated Qwen2, InternVL, and embedding models across text, tables, charts, and visuals, achieving ~92% retrieval accuracy and ~87% answer relevance.

OtherHalf – AI Anime Companion Bot

PythonGenerative AILLMsMultimodal AIYi-34BWhisperTTSAWQGPTQvLLMTensorRT-LLMNVIDIA TritonFAISSFastAPIWebSocketsPostgreSQLAWS EC2SageMakerCognito
  • Fine-tuned and deployed Yi-34B models with AWQ/GPTQ quantization, optimizing inference via vLLM, TensorRT- LLM, and NVIDIA Triton to achieve time-to-first-token under 0.2 seconds.
  • Optimized Whisper STT to under 0.5 seconds response time for 20-second audio clips, achieving 95% transcription accuracy, with full STT-to-LLM-to-TTS pipeline latency under 2 seconds.
  • Built multimodal reasoning, long-term memory, and vector retrieval systems, driving real-time conversational services with strong personalization through semantic memory.
  • Scaled to 200,000+ downloads with 95%+ user satisfaction rating.

EY Genie – AI Project Management Assistant

PythonLangChainMulti-Agent SystemsLLMsPrompt EngineeringEvaluation HarnessFastAPIReal-time Dashboards
  • Engineered multi-agent orchestration with a prompt evaluation harness, measurably reducing the hallucination/error rate of generated responses.
  • Cut report-generation time from minutes to seconds, lifting team productivity by 30%.

Leaderboard Standings

Leaderboard Position Pending

Global test scores, peer standing percentiles, and algorithm leaderboard ranks are updated dynamically.

Assessment Highlights

Assessments Not Completed

Coding evaluations, system assessment results, and conceptual score badges will appear here after taking a test.

AI Collaboration Score

AI Collaboration Score Pending

Developer coding behavior, assistant cooperation, and AI pair-programming indicators are evaluated during live coding sessions.

Role Compatibility Profile

Role Compatibility Analysis Pending

Custom matching reports, candidate role compatibility percentiles, and core engineer strength profiles are processed once conceptual code screenings are complete.

Achievements

User Scale and Satisfaction

Scaled OtherHalf AI companion to 200,000+ downloads with 95%+ user satisfaction rating.

Production Performance Benchmarks

Achieved 90%+ accuracy and performance benchmarks in production for multi-agent orchestration and LLM systems.

About Details

Professional Bio

Certified AI/ML Engineer with 4+ years architecting and deploying intelligent systems across Generative AI, Agentic AI, and Computer Vision. Proven expertise in building voice agents, multimodal RAG pipelines, and multi-agent orchestration systems for enterprise-scale production solutions.

B.Tech in Computer Science and Engineering

Sir Chhotu Ram Institute of Engineering and Technology, Meerut (2018 - 2022)

Languages: English, Hindi