Skip to content
← Back to job listings

Senior Staff DevOps Engineer

jobgether · India

Software DevelopmentRemoteExternal listingfull-time3 days ago

About The Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Staff DevOps Engineer based in India.

This is a high-impact technical leadership role within a global Infrastructure Platform team responsible for operating large-scale, cloud-native SaaS <infrastructure.You> will shape the architecture, reliability, security, and scalability of platforms running across AWS, Azure, and GCP.A major focus will be Kubernetes leadership, including EKS, AKS, service mesh, GitOps, and cloud-native deployment <patterns.You> will also help advance AI-assisted and agentic automation across DevOps and SRE workflows.The role offers significant organizational influence through architecture decisions, technical standards, mentoring, and cross-functional <leadership.You>’ll work with globally distributed engineering teams across India, the US, EMEA, and APAC in a fast-paced 24×7 environment.This is an individual-contributor leadership position with substantial scope to improve reliability, automation, security, and engineering productivity.

Accountabilities

Lead cloud infrastructure strategy: Design, build, and evolve resilient, secure, scalable, and cost-efficient multi-account cloud infrastructure, primarily across AWS while supporting Azure and GCP environments.

Drive Kubernetes excellence: Serve as a technical authority for production Kubernetes, including EKS and AKS cluster architecture, upgrades, networking, storage, capacity planning, security, multi-tenancy, and troubleshooting.

Own service mesh capabilities: Design and operate production service mesh solutions, implementing secure service-to-service communication, mTLS, traffic management, observability, resilience, and progressive delivery.

Establish platform standards: Define and promote best practices for Kubernetes, cloud networking, IAM, secrets management, infrastructure governance, workload design, Helm, deployment patterns, and security controls.

Advance Infrastructure as Code and GitOps: Automate infrastructure provisioning, deployment, monitoring, incident response, and capacity management using Terraform, CI/CD, GitOps, and related platform engineering practices.

Strengthen reliability and operations: Improve SLIs, SLOs, error budgets, observability, runbooks, incident response, on-call practices, post-incident reviews, and systemic remediation across a 24×7 SaaS environment.

Support security and compliance: Help implement secure platform controls and maintain infrastructure aligned with requirements such as PCI DSS, including audit readiness, governance, and secure-by-design practices.

Champion AI-enabled operations: Introduce LLM-based tooling, AI coding assistants, and agentic workflows to improve infrastructure development, incident triage, root-cause analysis, deployment validation, compliance checks, and operational efficiency.

Build safe AI automation: Establish appropriate guardrails, observability, cost controls, and human-in-the-loop practices for production AI and agentic workflows, including solutions built with services such as Amazon Bedrock.

Provide technical leadership: Influence engineers across teams and geographies without direct management responsibility, drive consensus on architecture, contribute to critical escalations, and raise engineering standards.

Mentor engineering talent: Coach Staff, Senior, and mid-level engineers while documenting reusable patterns, sharing technical knowledge, and encouraging stronger engineering practices.

Lead strategic initiatives: Drive cross-functional projects that improve uptime, deployment velocity, cloud consistency, cost efficiency, operational toil, and engineering productivity.

Requirements

Experience: 12+ years working in 24×7 production operations and highly available SaaS or cloud environments, with prior experience as a technical lead or Staff+ individual contributor in a global engineering organization.

Cloud expertise: 5+ years of hands-on experience with multi-account AWS infrastructure, including AWS Organizations, Account Factory, guardrails, SCPs, landing zones, networking, IAM, and cross-account connectivity.

Kubernetes: 5+ years of production Kubernetes experience at scale, with deep expertise in EKS and/or AKS, cluster operations, networking, storage, security, workload management, and performance optimization.

Infrastructure as Code: 5+ years of Terraform experience managing infrastructure across multiple AWS accounts and regions.

CI/CD & GitOps: 5+ years designing and implementing CI/CD pipelines for Terraform, Kubernetes, and microservices, plus practical experience with GitOps platforms such as ArgoCD, Kargo, or Flux.

Programming & systems: Strong Python, Go, or similar programming skills combined with advanced shell scripting and solid knowledge of Linux, networking, distributed systems, and production troubleshooting.

Service mesh: Hands-on experience implementing and operating a production service mesh such as Istio, Linkerd, or AWS App Mesh.

Observability: Experience with monitoring and logging technologies such as Prometheus, Grafana, OpenSearch, or equivalent platforms.

SRE practices: Strong understanding of SLIs, SLOs, error budgets, incident management, observability, reliability engineering, and operational excellence.

AI & automation: Experience applying AI tools on AWS or equivalent platforms to improve engineering productivity, automation, or operational efficiency; experience with Amazon Bedrock, LLM agents, or agentic workflows is highly valuable.

Security & compliance: Experience designing secure cloud platforms and familiarity with regulated enterprise environments; knowledge of PCI DSS and related compliance practices is preferred.

Regional infrastructure: Experience supporting data sovereignty or regional cloud deployments is a plus, particularly across markets with specific residency requirements.

Leadership: Strong interpersonal and communication skills, with the ability to influence engineers and stakeholders across teams, time zones, and organizational boundaries.

Education: Bachelor’s or Master’s degree in Computer Science or a related technical discipline, or equivalent practical experience.

Working style: Self-directed, collaborative, adaptable, quality-focused, and comfortable operating in a complex, fast-moving environment where technical decisions have broad organizational impact.

Benefits

  • Fully remote position for candidates based in India.
  • Opportunity to work on large-scale, cloud-native infrastructure supporting a global enterprise SaaS platform.
  • Significant technical influence and ownership without requiring people management.
  • Exposure to AWS, Azure, GCP, Kubernetes, service mesh, GitOps, and Infrastructure as Code at scale.
  • Opportunity to lead the adoption of AI-assisted and agentic DevOps workflows, including LLM-powered automation.
  • Collaboration with globally distributed engineering teams across India, US, EMEA, and APAC.
  • Scope to shape platform architecture, engineering standards, reliability practices, and long-term infrastructure strategy.
  • Opportunities to mentor experienced engineers and establish yourself as a subject-matter expert across the broader engineering organization.
  • Participation in strategic initiatives focused on automation, reliability, security, cost efficiency, and engineering productivity.
  • Inclusive workplace committed to equal opportunity and reasonable accommodations throughout the hiring process.
  • Opportunity to contribute to a modern engineering environment focused on innovation, operational excellence, and continuous improvement.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing