Own and evolve our cloud infrastructure across multi-region production environments, end-to-end.
Lead our GitOps deployment model - designing and maintaining declarative, automated deployment workflows with zero manual gates.
Build, maintain, and optimize CI/CD pipelines with a strong focus on developer experience, reliability, and speed.
Initiate, implement, and champion an AI-first DevOps & SRE ecosystem - identifying opportunities, building AI agents and intelligent automation, and driving their adoption across engineering operations.
Develop automation frameworks for provisioning, scaling, observability, and incident response, leveraging AI-powered tooling and agentic workflows to reduce toil.
Operate and improve our observability platform: metrics, logs, alerting, dashboards, SLOs/SLIs, and on-call tooling.
Champion zero-trust secrets management and credential-less authentication patterns across the stack.
Partner with architects and engineering leadership on cloud cost optimization, availability, and performance.
Build internal tooling and automation that multiplies engineering velocity across the organization.
Qualifications
5+ years of hands-on DevOps experience in a SaaS product environment - Must.
Demonstrated initiative in applying AI to engineering operations - designing and building AI agents, agentic workflows, LLM-powered automation, or Model Context Protocol (MCP) integrations that reduced operational toil and improved production reliability, quality, or velocity - Must.
Strong scripting and programming skills - Python and Bash for automation, tooling, and AI agent development; Go is a plus.
Strong motivation to continuously learn and adopt emerging technologies, and to share that knowledge across the team.
Deep, hands-on AWS expertise; multi-cloud (AWS, GCP, Azure) experience is a strong plus - Must.
Strong understanding of containers and orchestration - Docker, Kubernetes, including workloads, networking, service mesh (Istio), Helm/Kustomize, and autoscaling (KEDA, HPA, VPA).