Kaushik Kumaran

Site Reliability Engineer · IBM

SRE by title, systems archaeologist by nature, with 5+ years across cloud infrastructure and site reliability. I dig through config drift and dependency chains to find the toil worth killing. Currently building agentic infrastructure to make the next page unnecessary.

Education

Boston University

MS, Computer Information Systems

SRM Institute of Science and Technology

B.Tech, Computer Science Engineering

Professional Experience

IBM

Site Reliability Engineer · TLS Top Performer ’25

Austin, Texas
  • Eliminated 1,500+ hours of annual toil across 7,000+ VMs. Built a multi-API Go CLI that accelerated retrieval by 93% and cut maintenance preparation from one hour to under five minutes.
  • Caught 30+ production defects before customer impact. Co-built an Instana-driven synthetic monitoring, chaos, and regression framework across 30 datacenters.
  • Made production AAP reconstructible from Git across 9,372 hosts, 159 job templates, and 181 schedules. Generated 16,728 lines of dependency-ordered Config-as-Code from live AWX state.
  • Removed failed-primary and operator-workstation dependencies from Phase 1 DR recovery. Built a self-bootstrapping AAP job that restores configuration directly from Git.
  • Enabled one Config-as-Code baseline to support warm standby across us-south and eu-de. Added approval-gated site overlays for schedules, credentials, and regional settings.
  • Turned AAP configuration drift into an auditable defect. Scheduled live-to-Git checks and scrubbed Secrets Manager identifiers across 23 custom credential types.
  • Architected AI data infrastructure around a Python ingestion engine, MCP server, Qdrant, watsonx embeddings, and a dynamic schema for context-aware SRE retrieval.
  • Expanded fleet observability across 30 datacenters and 12,000+ VMs. Built Prometheus and Grafana dashboards backed by a custom Instana API scraping pipeline.
  • Standardized DevSecOps delivery across 20+ repositories. Generated Tekton pipeline configuration, centralized secret scanning, and automated post-merge Box uploads across nine repositories.
  • Reduced OpenShift deployment time per datacenter by 66%. Replaced 23 manual Helm configuration files with one dynamic automation source.
  • Eliminated maintenance-alert false positives across the global estate. Connected ServiceNow change logic to Instana maintenance windows through an API-driven suppression workflow.
PythonGoOpenShiftAnsible/AAPTektonIBM PowerVS

Searce Inc

Cloud Engineer

Bangalore, India
  • Improved monitoring reliability by 40% and reduced server crashes by 50% by migrating self-managed Prometheus to Google Managed Prometheus and resolving 10+ critical configuration issues.
  • Executed end-to-end cloud migrations for enterprise clients using Velostrata and v5 Migrate, modernizing legacy environments into cloud-native architectures.
  • Improved operational efficiency by 40% with Python automation and infrastructure optimizations that reduced recurring overhead and cloud spend.
GCPDockerPythonPrometheusGrafanaTerraformGKE

Software Developer · Norinco Pvt Ltd

Developed a scalable NestJS and TypeScript backend with PostgreSQL, improving query time by 40%. Supported client projects across the development lifecycle and built an authenticated React Native logistics application.

Software Development Engineer · Magnox Technologies

Designed a scalable React Native EdTech platform that reduced the projected implementation workforce cost by 90%.

Software Development Engineer · My Equation (formerly Tech Analogy)

Built React and Next.js products in TypeScript, improved SEO by 50%, and deployed services on AWS EC2. Supervised a team of 10 developers while tracking backlogs and delivery through an agile workflow.

Research

Research Intern

Geriatric health systems

Developed an all-in-one geriatric health solution with a centralized, authorization-controlled reporting system. Reduced the required assessment parameters by 55%.

Research Developer

Smart-campus waste management

Designed an IoT and mobile system for live waste-bin status across a 250-acre campus, combining machine-learning optimization, visualization, and real-time authority notifications.

Research Intern

Cardiac-risk prediction

Built a medical-data model for coronary-artery-disease prediction with 95% accuracy while reducing required parameters by 55%. The resulting paper was accepted to IEEE ICESIC 2022.

Publications

Coronary Artery Disease Prediction using Machine Learning Algorithms

Accepted to IEEE ICESIC 2022 Conference · 2022

Machine-learning research on coronary-artery-disease prediction with a smaller clinical feature set.

Read ↗

Efficient Waste Management System using IoT and Smart Credit System for College Campuses

Accepted to Springer CCIS · 2023

An IoT and smart-credit architecture for live waste monitoring and optimized collection across university campuses.

Read ↗

Engineering Projects

Kubernetes security · AI infrastructure

Argus

View code ↗
  • Architected a multi-layer Kubernetes observability system capturing network and kernel signals through eBPF for real-time threat analysis.
  • Built a threat-detection and AI-evaluation pipeline that correlates low-level telemetry with the MITRE ATT&CK framework.
  • Implemented human-governed remediation with Cilium and Falco for pod isolation and process termination.
  • Designed chaos threat-injection workflows to evaluate detection accuracy and reasoning robustness under failure conditions.
  • Delivered an end-to-end platform with 71 passing tests and a React control interface.

Sentinel

2026

A LangGraph multi-agent orchestrator coordinating Argus for security and Phoenix for chaos and regression testing.

LangGraphPythonKubernetesFastAPI

Phoenix

2026

A Kubernetes chaos and self-healing agent that injects failures, diagnoses incidents, and gates remediation behind human approval.

PythonChaos MeshKubernetesClaude API

Open Source

dcgm-exporter · NVIDIA

NVIDIA's official Prometheus exporter for GPU telemetry, built on DCGM — the metrics backbone behind most GPU-fleet observability stacks.

Contribution ↗

LangGraph · LangChain

LangChain's low-level orchestration runtime for building stateful, resilient agents as graphs — 39k+ stars and the backbone under most production LangChain agent stacks.

Contribution ↗

Volcano · CNCF

A CNCF batch-scheduling system for Kubernetes purpose-built for AI/ML training and HPC workloads — gang scheduling, queues, and GPU-aware fair-share at cluster scale.

Contribution ↗

Tornjak · CNCF · SPIFFE

A CNCF SPIFFE/SPIRE management plane — the UI and API layer operators use to broker human access and administer one or more SPIRE workload-identity deployments.

Contribution ↗

Omnigent · Databricks

Databricks' open-source AI agent meta-harness — 9k+ stars, orchestrating Claude Code, Codex, Cursor, and custom agents behind one policy and sandboxing layer.

Contribution ↗

Temporal · Temporal

The open-source durable execution engine for workflow orchestration — 22k+ stars, running production reliability infrastructure at OpenAI, NVIDIA, Snap, and DoorDash.

Contribution ↗

Technical Skills

Languages

Go · Python · TypeScript · Shell · SQL · GraphQL

Cloud & Infrastructure

Kubernetes · OpenShift · IBM PowerVS · GCP · AWS · Terraform · Ansible/AAP · Helm · Tekton · ServiceNow

Observability

Prometheus · Grafana · Instana · ELK Stack

AI/Agentic Systems

LangGraph · FastAPI · Qdrant · watsonx embeddings · Claude API

Security

eBPF · Cilium · Falco · Kyverno

Product Engineering

React · Node.js · NestJS · PostgreSQL · MongoDB · Neo4j · React Native

Leadership

Head of Web/App Development

IEEE SRM Student Branch

Led 100+ members and conducted 10+ workshops for 300+ people.

Technical Executive

Alexa Developers SRMIST

Ran AWS fundamentals workshops and built the team site backend.

Web Developer

GenY SRMIST

Built a file-uploader application and a MERN CRM.

Certifications