Shubham Kumar's Resume

Shubham Kumar

Cloud Platform Engineer / Site Reliability Engineer with a DevOps foundation, now increasingly focused on AI Platform Operations (AWS Bedrock, RAG, FinOps, and cost optimization).

Gurugram, India (IST)

Shubham Kumar's profile picture

About

Cloud Platform Engineer / Site Reliability Engineer with 7+ years of experience across DevOps, Kubernetes, and multi-cloud infrastructure (AWS, Azure, GCP) — building observable, automated, production-grade platforms. That foundation is still how I approach systems design today.

Over the past year I've been increasingly focused on AI Platform Operations: building and operating production AI tooling on AWS Bedrock with Claude models, including a human-gated RAG auto-triage agent and an AI-assisted incident RCA tool, alongside FinOps cost governance and Zero Trust identity (Keycloak, OIDC/OAuth2).

Work Experience

SingleStore

Feb 2026 - Present

Cloud Platform Engineer

  • Designed and built ATLAS, an internal operational-intelligence platform (Airflow, SingleStore, Next.js) used daily by Support, Engineering, and Leadership to track SLA risk and recurring issues.
  • Built production AI tooling on AWS Bedrock (Claude 3/3.5) — an AI-assisted incident RCA tool on Grafana MCP and a guardrailed RAG auto-triage agent, both human-gated by design.
  • Implemented GPU-backed autoscaling on EKS using Karpenter, cutting ML infrastructure cost by 35%+ (utilization ~25% to 65%), measured via Kubecost against real AWS billing.
  • Own the production Keycloak identity platform (Zero Trust, OIDC/OAuth2) serving ~150 daily internal users.
  • Contribute Go backend code to an internal multi-cloud cost-governance (FinOps) platform.
  • AWS Bedrock
  • FinOps
  • Kubernetes
  • Terraform
  • Zero Trust

AirFi Aviation Solutions

Aug 2025 - Feb 2026

Senior DevOps Engineer

  • Led an uptime initiative that raised platform availability from 97.8% to 99.95%.
  • Cut incident resolution time by 45%+ through centralized observability and standardized runbooks.
  • Mentored junior engineers and set incident-response and IaC standards adopted across every team on the shared platform.
  • DevOps
  • Observability
  • Incident Response
  • Mentorship

AirFi Aviation Solutions

Oct 2023 - Jul 2025

DevOps Engineer

  • Developed and deployed automation across a fleet of 8,000+ embedded IFE (in-flight entertainment) devices — including firmware rollout pipelines that cut release time by 40%, and telemetry-based PMIC monitoring with secure LTE-based diagnostics that reduced MTTR by 35%.
  • Built DISCO, an internal Python-based tool for processing onboard infotainment box log data at scale — pulling and parsing logs from AWS S3 for fleet-wide diagnostics.
  • Operated and optimized AWS and Azure Kubernetes environments for production workloads; automated infrastructure changes with Terraform and CI-driven workflows.
  • Implemented monitoring and alerting improvements that reduced production outages by 40%.
  • CI/CD
  • Terraform
  • Ansible
  • Automation
  • Python
  • Embedded Systems

Innoitus

Jun 2023 - Sep 2023

Site Reliability Engineer

  • Improved observability and alert quality through custom tooling and hands-on monitoring improvements.
  • Reduced critical incident frequency by 35% through proactive monitoring and reliability practices.
  • Improved incident response times by 30% with better alerting and on-call workflows.
  • SRE
  • Monitoring

Amazon

Oct 2021 - Jun 2023

Quality Analyst

  • Built Jenkins pipelines integrating Prometheus and Grafana dashboards for better pipeline and environment visibility.
  • Managed AWS-based environments with a focus on scalability, uptime, and dependable delivery workflows.
  • Administered Kubernetes workloads with resource optimization across QA and production-adjacent systems.
  • AWS
  • Jenkins
  • Prometheus
  • Grafana
  • Kubernetes

Extreme Soft Management

Apr 2019 - Aug 2021

Site Reliability Engineer

  • Operated and maintained production infrastructure on Google Cloud Platform (GCP), introducing automation for repetitive operational tasks.
  • Led a year-long GCP-to-AWS migration, modernizing the deployment stack end-to-end.
  • Automated workflows that saved 80+ engineering hours per month across recurring processes.
  • SRE
  • GCP
  • AWS Migration
  • Automation

Education

RNS Institute of Technology

2014 - 2018
B.E in Mechanical Engineering

Skills

Tools

Kubernetes
Terraform
AWS
Azure
Google Cloud Platform
Jenkins
Ansible
Docker
GitHub
Argo CD
Apache Kafka
Linux
AWS Bedrock / LLM Ops
FinOps / Cost Optimization
Keycloak / Zero Trust IAM

Framework And Runtime

Node.js
Next.js
AWS Lambda

Programming Language

Python
Go
TypeScript
Bash

Databases

MongoDB
PostgreSQL
Elasticsearch

Side projects

ATLAS — Operational Intelligence Platform

Internal operational-intelligence platform used daily by Support, Engineering, and Leadership at SingleStore to track SLA risk and recurring issues. Built on Airflow for orchestration, SingleStore for storage, and a Next.js frontend.

  • Airflow
  • SingleStore
  • Next.js
  • TypeScript

AI-Assisted Incident RCA & RAG Auto-Triage

Production AI tooling on AWS Bedrock (Claude 3/3.5): an incident RCA tool built on Grafana MCP, and a guardrailed RAG auto-triage agent — both deliberately human-gated rather than fully autonomous.

  • AWS Bedrock
  • Claude
  • Grafana MCP
  • RAG
  • Python

GPU Autoscaling & ML Infra Cost Optimization

Implemented GPU-backed autoscaling on EKS using Karpenter, cutting ML infrastructure cost by 35%+ and raising utilization from ~25% to ~65%, measured via Kubecost against real AWS billing.

  • Karpenter
  • EKS
  • Kubernetes
  • Kubecost
  • AWS

k8s-gitops-platform

Production Kubernetes platform on AWS with automated cluster provisioning (Terraform) and GitOps deployments (Argo CD). Full observability stack — Prometheus, Grafana, Loki, and Tempo — for metrics, logs, and traces.

  • Kubernetes
  • Terraform
  • AWS
  • Argo CD
  • Prometheus
  • Grafana
  • Loki
  • Tempo

Europe Job Hunter

A job search tool combining a Go REST API backend, a React frontend, and Python services, with Gemini AI integrated to assist with job matching.

  • Go
  • React
  • Python
  • Gemini AI