App LogoApp name
DD

darshan d

Open to work4 years experience

Site Reliability Engineer

BangaloreTarget Roles: Backend Engineer • Full-Stack Developer • Frontend Specialist

Site Reliability Engineer with nearly 4 years of experience in cloud infrastructure and CI/CD.

GitHub

Standing Rank

Rank Not Available

Developer Badges

No badges earned yet

Skills & Technologies

19 skills
awspythonlinuxdockerkuberneteshelmterraformgithub actionsjenkinsargo cdprometheusgrafanaopentelemetryelk stacksplunkdynatracegitjiraservicenow

Work Experience

SRE-2

Bread Financial

Apr 2025 - Jul 2025 Bangalore
  • Handled traffic spike by configuring Auto Scaling (target tracking on CPU/ALB requests) , ensured system stability under load.
  • Performed Linux-level troubleshooting during production incidents by analyzing CPU , memory, disk I/O, process states, and network issues using tools like Top, Htop, Vmstat, Iostat, Netstat, and Journalctl.
  • Fixed latency spike (200ms -> 5s) using OpenTelemetry tracing , found slow DB query, optimized query, reduced latency to <1s.
  • Diagnosed and resolved pod failures, container crashes, and deployment issues in production Kubernetes clusters, ensuring high availability and zero downtime.

SRE

Genpact

Mar 2022 - Mar 2025 Bangalore
  • Improved service uptime to 99.9% by fixing recurring failures and adding proactive monitoring.
  • Fixed DB “too many connections” issue by implementing connection pooling, timeouts, and autoscaling DB instances
  • Implemented RBAC and channel-based access control for Slack-triggered deployments, restricting production releases to approved teams.
  • Managed AWS services including VPC, EC2, IAM, Route 53, Load Balancers, and ECS to build scalable and secure cloud solutions.
  • Built end-to-end monitoring using Prometheus, Grafana, OpenTelemetry.
  • Troubleshoot CrashLoopBackOff by analyzing pod logs and config, identified incorrect env variable/DB config, fixed deployment, restored service.
  • Automated on-call tasks using Python (API integrations, Alertmanager -> Jira, log parsing, token rotation).
  • Identified the root cause of continuously restarting pods due to OOMKilled; fixed memory leak, updated resource requests/limits -> eliminated restarts.
  • Managed remote state handling for multiple development teams using Terraform workspaces.
  • Resolved AWS access denied (ec2:RunInstances) by analyzing IAM policies, identified explicit deny, corrected permissions, restored infra provisioning.

Projects

Projects Not Populated

Personal applications, open-source work, and code repos will show here.

Leaderboard Standings

Leaderboard Position Pending

Global test scores, peer standing percentiles, and algorithm leaderboard ranks are updated dynamically.

Assessment Highlights

Assessments Not Completed

Coding evaluations, system assessment results, and conceptual score badges will appear here after taking a test.

AI Collaboration Score

AI Collaboration Score Pending

Developer coding behavior, assistant cooperation, and AI pair-programming indicators are evaluated during live coding sessions.

Role Compatibility Profile

Role Compatibility Analysis Pending

Custom matching reports, candidate role compatibility percentiles, and core engineer strength profiles are processed once conceptual code screenings are complete.

Achievements

Service Uptime Improvement

Improved service uptime to 99.9% by fixing recurring failures and adding proactive monitoring.

Container Optimization

Reduced container image size by 40% using multi-stage Docker builds, improving deployment speed.

About Details

Professional Bio

Dynamic and results-driven Site Reliability Engineer with nearly 4 years of experience in automating infrastructure, enhancing system reliability, and streamlining CI/CD pipelines. Seeking to contribute to scalable, secure, and high-performing environments.

Master of Computer Applications (MCA)

RV College of Engineering, Bengaluru (2019 - 2022)

Languages: English, Hindi
darshan d - Profile | Swiftcruit