Boston University
MS, Computer Information Systems
Site Reliability Engineer · IBM
SRE by title, systems archaeologist by nature, with 5+ years across cloud infrastructure and site reliability. I dig through config drift and dependency chains to find the toil worth killing. Currently building agentic infrastructure to make the next page unnecessary.
MS, Computer Information Systems
B.Tech, Computer Science Engineering
Site Reliability Engineer · TLS Top Performer ’25
Cloud Engineer
Developed a scalable NestJS and TypeScript backend with PostgreSQL, improving query time by 40%. Supported client projects across the development lifecycle and built an authenticated React Native logistics application.
Designed a scalable React Native EdTech platform that reduced the projected implementation workforce cost by 90%.
Built React and Next.js products in TypeScript, improved SEO by 50%, and deployed services on AWS EC2. Supervised a team of 10 developers while tracking backlogs and delivery through an agile workflow.
Developed an all-in-one geriatric health solution with a centralized, authorization-controlled reporting system. Reduced the required assessment parameters by 55%.
Designed an IoT and mobile system for live waste-bin status across a 250-acre campus, combining machine-learning optimization, visualization, and real-time authority notifications.
Built a medical-data model for coronary-artery-disease prediction with 95% accuracy while reducing required parameters by 55%. The resulting paper was accepted to IEEE ICESIC 2022.
Machine-learning research on coronary-artery-disease prediction with a smaller clinical feature set.
An IoT and smart-credit architecture for live waste monitoring and optimized collection across university campuses.
A LangGraph multi-agent orchestrator coordinating Argus for security and Phoenix for chaos and regression testing.
A Kubernetes chaos and self-healing agent that injects failures, diagnoses incidents, and gates remediation behind human approval.
NVIDIA's official Prometheus exporter for GPU telemetry, built on DCGM — the metrics backbone behind most GPU-fleet observability stacks.
LangChain's low-level orchestration runtime for building stateful, resilient agents as graphs — 39k+ stars and the backbone under most production LangChain agent stacks.
A CNCF batch-scheduling system for Kubernetes purpose-built for AI/ML training and HPC workloads — gang scheduling, queues, and GPU-aware fair-share at cluster scale.
A CNCF SPIFFE/SPIRE management plane — the UI and API layer operators use to broker human access and administer one or more SPIRE workload-identity deployments.
Databricks' open-source AI agent meta-harness — 9k+ stars, orchestrating Claude Code, Codex, Cursor, and custom agents behind one policy and sandboxing layer.
The open-source durable execution engine for workflow orchestration — 22k+ stars, running production reliability infrastructure at OpenAI, NVIDIA, Snap, and DoorDash.
Go · Python · TypeScript · Shell · SQL · GraphQL
Kubernetes · OpenShift · IBM PowerVS · GCP · AWS · Terraform · Ansible/AAP · Helm · Tekton · ServiceNow
Prometheus · Grafana · Instana · ELK Stack
LangGraph · FastAPI · Qdrant · watsonx embeddings · Claude API
eBPF · Cilium · Falco · Kyverno
React · Node.js · NestJS · PostgreSQL · MongoDB · Neo4j · React Native
Led 100+ members and conducted 10+ workshops for 300+ people.
Ran AWS fundamentals workshops and built the team site backend.
Built a file-uploader application and a MERN CRM.