Skip to content
← Back to job listings

Senior Platform Architect

wizdaa · Remote

Software DevelopmentRemoteImported listingcontractabout 14 hours ago

About The Role

We're looking for a Platform Architect who can set the standard for how we build, ship, and operate reliable cloud platforms at scale. You sit at the intersection of platform engineering and SRE. You'll own the path from infrastructure design to reliable production services, bringing DevOps rigor to complex systems.

This is not a ticket-processing role, and it's not a research role. You'll tackle hard problems: platform reliability, scalability, cost efficiency, deployment automation, and workload operations. You'll have the scope to solve them properly. Senior professionals here identify problems before they're asked and raise the ceiling on what the platform can do.

What you will work on

  • Build and operate scalable backend and AI infrastructure, supporting real-time and batch workloads with a focus on performance, reliability, and multi-tenant architecture.
  • Design and maintain deployment workflows across services and environments, including versioning, staged rollouts, automated releases, monitoring, and safe rollback strategies.
  • Build and operate LLM and agentic systems in production, integrating model providers, APIs, gateways, tools, and external services while managing rate limits, reliability, guardrails, and graceful degradation.
  • Develop reusable services, APIs, automation, and data pipelines that support AI-powered products and internal platform capabilities.
  • Extend infrastructure-as-code across the platform using Terraform and reusable patterns to provision and manage cloud services consistently across projects and environments.
  • Maintain GitOps-based deployment workflows using tools such as ArgoCD, improving automation and consistency across environments and tenants.
  • Run distributed workloads on Kubernetes (GKE), managing scaling, workload placement, tenant isolation, service reliability, and infrastructure capacity.
  • Improve platform observability and reliability through metrics, logging, tracing, SLOs, alerting, incident response practices, and operational tooling.
  • Identify performance and infrastructure cost improvements across cloud services, compute resources, APIs, and AI workloads.
  • Use agentic coding and AI development tools to accelerate engineering work, including scaffolding services, generating and reviewing infrastructure and application code, debugging, and automating repetitive workflows.

What you won’t find here

A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing