App LogoApp name
NK

Naveen Kumar M

Open to work0 years experience

Cloud & DevOps Engineer

Bangalore, IndiaTarget Roles: Backend Engineer • Full-Stack Developer • Frontend Specialist

Cloud & DevOps Engineer | AWS | Azure | GCP | Kubernetes | Terraform | CI/CD | Observability

Standing Rank

Rank Not Available

Developer Badges

Skills & Technologies

24 skills
awsazuregcpterraformansiblekubernetesdockerhelmopenshiftjenkinsgithub actionsgitlab ciazure devopsprometheusgrafanaaws cloudwatchazure monitorelk stackpythonbashgitgithubgitlabhashicorp vault

GitHub Activity Evidence

GitHub Footprint Not Connected

Recruiters value verified code contributions. Linking a GitHub account displays live heatmaps, active contribution metrics, repository languages, and push activity tags.

Work Experience

Work Experience Not Populated

Professional job history, career timeline, and previous industry roles will show here once linked or added.

Projects

Cloud Observability & Alerting Platform

PrometheusGrafanaAWS EC2KubernetesPythonAWS CloudWatchAzure Monitor
  • Configured Prometheus to scrape and store real-time metrics from AWS EC2 instances, Kubernetes pods, and application endpoints, enabling continuous health monitoring across multi-cloud infrastructure.
  • Built Grafana dashboards with custom panels tracking CPU utilization, memory consumption, request latency, and error rates — enabling proactive incident detection before user impact occurred.
  • Implemented alert threshold tuning in Prometheus Alertmanager for CPU (>85%), memory (>80%), and pod crash-loop events — reducing false positive alerts by ~35% and improving on-call signal quality.
  • Integrated AWS CloudWatch and Azure Monitor alongside Prometheus to provide unified observability across multi- cloud environments, ensuring no blind spots in production systems.
  • Automated incident response runbooks using Python scripts to trigger remediation workflows on alert — reducing mean time to recovery (MTTR) by ~40% for common operational issues.
  • Documented monitoring procedures, alerting standards, SLO/SLA definitions, and dashboard best practices to enable consistent observability across the platform engineering team.

Scalable DevOps SaaS Deployment Platform

AWSDockerKubernetesEKSJenkinsGitHub ActionsTerraformAnsible
  • Deployed a full-stack application on AWS using Docker and Kubernetes (EKS) with zero-downtime rolling deployments, achieving high availability and fault tolerance across distributed nodes.
  • Built and maintained CI/CD pipelines using Jenkins and GitHub Actions, automating build, test, and deployment workflows — reducing release cycle time by ~40% and eliminating manual deployment errors.
  • Provisioned all cloud infrastructure using Terraform IaC, enabling repeatable and version-controlled deployments across EC2, S3, Lambda, VPC, and CloudWatch with least-privilege IAM policies enforced by default.
  • Configured Ansible playbooks and roles for post-provision OS configuration management, automating application setup and patching across all managed cloud hosts.
  • Responded to and remediated operational incidents including pod failures, resource exhaustion, and deployment rollbacks — applying SRE incident response practices and producing RCA documentation for each event.
  • Collaborated across engineering and operations workflows to define key reliability indicators (SLOs/KPIs) and align the monitoring strategy with business availability requirements.

Resilience Mesh — Multi-Cloud Hybrid Architecture

AWSAzureGCPTerraformPrometheusGrafanaKubernetesGKEAKSPythonBash
  • Engineered a highly available hybrid infrastructure spanning AWS, Azure, and GCP using Terraform workspaces for consistent multi-cloud provisioning, eliminating single points of failure and achieving 99.9% platform uptime.
  • Implemented cross-cloud monitoring using Prometheus federation and Grafana multi-source dashboards, providing a single-pane-of-glass view of infrastructure health across all cloud providers.
  • Configured Kubernetes auto-scaling policies based on custom Prometheus metrics on GKE and AKS, enabling the platform to dynamically scale during traffic spikes and maintain reliability under load.
  • Automated routine operational tasks using Python and Bash scripts — environment health checks, log rotation, stale resource cleanup — saving ~5 hours/week of manual SRE toil.
  • Implemented RBAC and IAM least-privilege access controls across all cloud environments via Terraform-managed policy resources, reducing the blast radius of any security or operational incident.
  • Configured cross-cloud VPC networking, firewall rules, and private connectivity to ensure secure and segmented communication between workloads across AWS, Azure, and GCP.

Leaderboard Standings

Leaderboard Position Pending

Global test scores, peer standing percentiles, and algorithm leaderboard ranks are updated dynamically.

Assessment Highlights

Assessments Not Completed

Coding evaluations, system assessment results, and conceptual score badges will appear here after taking a test.

AI Collaboration Score

AI Collaboration Score Pending

Developer coding behavior, assistant cooperation, and AI pair-programming indicators are evaluated during live coding sessions.

Role Compatibility Profile

Role Compatibility Analysis Pending

Custom matching reports, candidate role compatibility percentiles, and core engineer strength profiles are processed once conceptual code screenings are complete.

Achievements

Achievements Not Earned

Special rewards, developer badges, system recognition certificates, and conceptual milestones will display here.

About Details

Professional Bio

2025 Computer Science graduate with strong hands-on experience in cloud infrastructure engineering, DevOps automation, and site reliability practices across AWS, Azure, and GCP. Committed to automation-first, reliability-focused engineering aligned with modern Cloud and DevOps best practices.

Bachelor of Technology in Computer Science

Presidency University, Bangalore (2021 - 2025)

Languages: English, Hindi