Skip to content
← Back to job listings

Platform Engineer & Cloud Ops Engineer

Sutherland · Remote, TS, India

Software DevelopmentRemoteImported listingfull-time1 day ago

About The Role

KEY RESPONSIBILITIES

Platform Architecture & Strategy (Senior-focused)

  • Lead the architecture and deployment of complex multi-cloud solutions (GCP/AWS), including networking, compute, storage, and multi-environment design.
  • Define platform standards and reusable patterns (e.g., Terraform modules, cluster blueprints).
  • Evaluate emerging cloud technologies to enhance cloud strategy and roadmap.
  • Drive cost optimization and FinOps practices, including right-sizing and governance.
  • Own platform reliability, scalability, capacity planning, and disaster recovery design.

Infrastructure Provisioning & Automation

  • Deploy, configure, and manage cloud infrastructure across GCP and AWS.
  • Write and maintain infrastructure as code (Terraform primary; CloudFormation/ARM where applicable).
  • Maintain multi-environment infrastructure consistency (dev, staging, prod).
  • Automate provisioning, configuration, and operational tasks to reduce manual toil.

Kubernetes & Container Platform

  • Design, operate, and support production-grade Kubernetes clusters (GKE preferred).
  • Manage upgrades, autoscaling, node pools, namespaces, and RBAC policies.
  • Own Helm/Kustomize standards for application deployment.
  • Support or manage service mesh (Istio) for traffic management, mTLS, observability, and security.
  • Define and promote golden paths for safe and consistent app deployment.

CI/CD, Monitoring & Operational Support

  • Build and maintain GitLab CI/CD pipelines for infrastructure and application delivery.
  • Monitor platform health using Datadog — metrics, logs, traces, SLOs, and alerts.
  • Troubleshoot issues, support incident response, and lead post-incident reviews (senior).
  • Maintain runbooks and dashboards; lead or support on-call rotations.

Leadership & Collaboration (Senior-focused)

  • Mentor junior and mid-level engineers, set technical direction, and review designs.
  • Collaborate effectively with global, cross-functional, and on/offshore teams.
  • Communicate architecture decisions clearly to technical and business stakeholders.

TECH STACK
Required:

  • Cloud Platforms: GCP (Compute Engine, GKE, VPC, Storage, IAM, Load Balancing), AWS (EC2, EKS, VPC, S3, IAM)
  • Kubernetes: GKE architecture, upgrades, Helm/Kustomize, autoscaling, RBAC
  • Infrastructure as Code: Terraform (multi-environment, reusable modules, remote state)
  • CI/CD: GitLab pipeline development and support
  • Monitoring: Datadog (metrics, logs, traces, alerting, SLOs)
  • Service Mesh: Istio (traffic management, mTLS, observability)

Good to have:

  • GitOps tools (ArgoCD, Flux)
  • Cloud FinOps tooling and cost optimization experience
  • Scripting languages (Python, Go, Bash)
  • Vault, Packer, service catalogs, and self-service platforms
  • Experience in regulated environments (HIPAA, SOC 2, ISO 27001)

REQUIREMENTS

Must have:

  • Platform Engineer: 3+ years in cloud infrastructure/platform or DevOps engineering.
  • Senior Platform Engineer: 8+ years, with leadership or architect-level responsibilities.
  • Deep, hands-on expertise with GCP and/or AWS cloud platforms.
  • Strong Kubernetes production experience, preferably GKE.
  • Expert-level Terraform skills with reusable modules and multi-env IaC.
  • Proven CI/CD automation experience at scale.
  • Ability to troubleshoot complex cloud environment issues.
  • Collaborative teamwork mindset with clear communication skills.

Nice to have:

  • Certifications: Google Professional Cloud Architect, AWS Solutions Architect Professional, Certified Kubernetes Administrator (CKA).
  • Experience with GitOps and progressive delivery techniques.
  • Prior FinOps or cloud cost optimization role experience.
  • Scripting for automation and Linux system administration background.

HOW SUCCESS IS MEASURED

  • Platform reliability: uptime and SLO achievement for shared services.
  • Automation coverage and manual toil reduction.
  • Adoption and enforcement of platform standards and self-service deployment paths.
  • Cloud cost efficiency achieved through optimization.
  • Timely and successful delivery of platform initiatives and projects.
  • Mean time to recovery (MTTR) for platform-impacting incidents.

All your information will be kept confidential according to EEO guidelines.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing