Platform Engineer
sunset · New York
About The Role
About SunsetAt its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of <businesses.In> 2025, we had a unique insight: the data every company generates each day through collaboration, communication, and building is some of the most valuable training data in the world. Public and synthetic data can only get frontier models so far, so the next generation of model progress depends on real, proprietary data grounded in how actual businesses operate. We are a primary source of it, partnering directly with the frontier AI labs building what comes next.Why Join Sunset NowWe have scaled from $0 to a multi-eight-figure run rate in a matter of monthsWe have raised from top-tier investors, including Floodgate, Afore, Ludlow, and Hustle FundWe are small enough that you will carry outsized responsibility and grow as quickly as the company doesYou will partner with and build for some of the fastest and most important companies in the worldYou will help build a massive, category-defining business from the ground floorThe RoleSunset operates customer-facing SaaS products, connector and ingestion services, asynchronous workers, high-volume data pipelines, model-backed systems, review tools, and customer-delivery paths. These workloads have different shapes, but they need a coherent foundation for infrastructure, delivery, observability, recovery, access, and <cost.You> will build and operate the shared platform that lets our product, data, and AI teams ship reliable, secure, observable, and cost-aware systems without manual infrastructure work or operational risk growing linearly. You will write software and infrastructure, improve real engineering workflows, lead through incidents, and create paved roads teams can use without waiting on you.This is not a deployment-operator or internal-IT role. Product, data, and ML teams remain responsible for the systems they build. You will give them the runtime, delivery, visibility, recovery, and operating patterns to own those systems well. You will partner closely with our Security Lead, but you will not be expected to run the entire security or compliance program.Problems You Might OwnMake several workload shapes feel like one coherent platformCreate a small set of supported patterns for customer-facing services, connectors, scheduled jobs, data-processing pipelines, model-backed workloads, and evaluation runs. Define the contracts for environments, compute, state, networking, delivery, secrets, telemetry, failure handling, and recovery without forcing every workload into an inappropriate stack.Turn delivery and operations into product-quality experiencesMake it straightforward for an engineer to create an environment, ship a safe change, understand a failed deploy or job, get the right access, recover a system, and know who owns the result. Build useful self-service and escape hatches while making unsupported paths and exceptions explicit.Make reliability visible from customer request to completed workloadConnect service, queue, job, pipeline, and model telemetry to the outcome that matters. Establish practical objectives, alerts, incident mechanics, replay and recovery paths, and reviews that remove recurring failure classes instead of only documenting them.Make infrastructure cost and control evidence part of normal operationExpose cost and capacity in workload-relevant units, then improve them without hiding reliability, security, quality, or developer time. Work with Security to implement least privilege, secrets, logging, backup, deployment, and audit controls whose evidence comes from the systems that actually enforce them.What You'll DoEstablish Sunset's current platform, workload, reliability, ownership, toil, recovery, cost, and technical-control baselineBuild reusable infrastructure-as-code modules, runtime templates, deployment workflows, environment contracts, and operational toolingCreate supported paths for customer-facing services, asynchronous and batch jobs, data pipelines, and model-backed workloadsImprove deploy safety, workload visibility, backup and recovery, incident response, replay, rollback, and durable remediationWork with engineering teams to define useful service and pipeline objectives, ownership, escalation, and recovery pathsBuild self-service for common infrastructure, environment, access, deploy, debugging, and recovery work without becoming a central approval queueMake cloud and vendor cost understandable by service and workload and improve efficiency within explicit reliability and security boundsPartner with Security on cloud identity, secrets, isolation, audit logging, vulnerability response, incident readiness, and automated control evidenceSupport employees and contractors through bounded access, safe environments, release controls, documentation, and timely removal of authorityUse AI tools deeply for platform engineering and operations while verifying generated code, plans, queries, state changes, and incident conclusionsWhat Success Looks LikeSunset's environments, runtimes, deploy paths, service and pipeline owners, reliability risks, recovery gaps, manual work, and infrastructure costs are visible and prioritizedOne consequential failure or toil class is materially reduced in your first 90 days, and another team can use the resulting paved road without case-by-case helpProduct, data, and AI teams can ship and understand their systems faster while retaining clear operating ownershipPriority services and pipelines have useful objectives, actionable telemetry, tested recovery paths, and incident learning that removes recurring failuresCommon platform work becomes self-service while exceptions remain explicit, owned, monitored, and time-boundedCloud cost and capacity are understandable in workload-relevant units and improve without hidden reliability, security, or developer-time regressionsSecurity and customer-trust evidence becomes easier to produce because it reflects current technical controlsYou Might Thrive Here IfYou have personally owned production cloud infrastructure and delivery or reliability systems across multiple services, including an asynchronous, batch-data, or model-backed workloadYou are a strong software engineer who is comfortable changing application, platform, and infrastructure code and operating the result in productionYou can reason from user impact through dependencies, state, telemetry, incident response, recovery, and durable remediationYou have built paved roads other engineers adopted because they made real work easier, not because a platform team required themYou understand both long-running services and high-volume or scheduled workloads and know where their reliability models should differYou can make pragmatic tradeoffs among delivery speed, least privilege, isolation, recovery, developer experience, and unit costYou are effective in an early-stage environment where the first step is often to establish ownership and a trustworthy baselineYou can lead calmly through ambiguous incidents, communicate clearly, and leave the system and operating model stronger afterwardYou use modern AI engineering tools fluently and verify generated infrastructure, queries, code, and operational conclusions before they affect productionThis Role May Not Be for You IfYou want a deployment or cloud-administration role where product teams hand systems to you to operate permanentlyYou prefer designing a platform in isolation to learning how engineers, services, pipelines, and customer deliveries actually workYou measure platform success by migration, ticket, dashboard, or uptime counts without connecting them to adoption, reliability, recovery, and user impactYou want to standardize every workload on one stack regardless of its state, scale, failure, or recovery requirementsYou do not want AI tools to be part of your daily engineering and operational workflowBonusExperience as an early platform or SRE hire at a fast-growing companyExperience with AWS, Terraform, container runtimes, workflow orchestration, and observability systemsExperience with high-volume data processing, model serving, evaluation jobs, GPU workloads, or machine-learning platformsExperience improving developer environments, preview systems, CI/CD, progressive delivery, or internal developer platformsExperience with replayable pipelines, backup and restore, disaster recovery, capacity planning, or cloud-cost allocationExperience implementing technical controls and automated evidence for SOC 2 or enterprise customer requirements
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
