Senior DevOps Engineer
solvd · Remote, Argentina
About The Role
Solvd Inc. is a rapidly growing AI-native consulting and technology services firm delivering enterprise transformation across cloud, data, software engineering, and artificial intelligence. We work with industry-leading organizations to design, build, and operationalize technology solutions that drive measurable business outcomes.
Following the acquisition of Tooploox, a premier AI and product development company, Solvd now offers true end-to-end delivery—from strategic advisory and solution design to custom AI development and enterprise-scale implementation. Our capability centers combine deep technical expertise, proven delivery methodologies, and sector-specific knowledge to address complex business challenges quickly and effectively.
WHAT YOU'LL WORK ON
- Design, implement, and maintain secure, reliable, and scalable cloud infrastructure.
- Build and improve CI/CD pipelines that support frequent, repeatable, and low-risk deployments.
- Automate infrastructure provisioning, configuration, deployment, and environment management.
- Maintain infrastructure as code using tools such as Terraform, CloudFormation, Pulumi, or comparable technologies.
- Improve consistency across development, testing, staging, and production environments.
- Build and maintain containerized application environments and orchestration capabilities.
- Implement and improve monitoring, logging, tracing, alerting, dashboards, and production-health reporting.
- Partner with engineers to establish SLIs, SLOs, and actionable operational metrics.
- Improve application and infrastructure resiliency, scalability, availability, and disaster recovery readiness.
- Manage secrets, certificates, identity, permissions, network controls, and cloud security configurations.
- Support vulnerability management, dependency security, patching, and infrastructure hardening.
- Improve deployment strategies including automated rollback, blue-green deployment, canary releases.
- Diagnose and resolve production incidents involving infrastructure, networking, application deployment, performance, capacity, or cloud services.
- Participate in incident response, post-incident reviews, root-cause analysis, and corrective-action planning.
- Reduce cloud waste and improve infrastructure cost visibility.
- Build self-service tools and reusable deployment patterns for application engineering teams.
- Document infrastructure architecture, operational procedures, recovery processes, and troubleshooting guidance.
- Participate in production support and an appropriate on-call rotation.
- Use AI-assisted engineering tools to accelerate scripting, troubleshooting, documentation, and infrastructure analysis.
WHAT YOU BRING
- Approximately 5+ years of experience in DevOps, SRE, Cloud Engineering, Platform Engineering, or closely related role.
- Strong experience operating production workloads in AWS.
- Strong hands-on experience with infrastructure as code (Terraform, CloudFormation, Pulumi).
- Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CircleCI, AWS CodePipeline).
- Strong knowledge of Docker and containerized application delivery.
- Experience with Kubernetes, Amazon ECS, or another container orchestration platform.
- Strong understanding of cloud networking (VPCs, subnets, routing, load balancers, DNS, firewalls, gateways, private connectivity).
- Experience with cloud IAM, role-based access, least-privilege design, and secrets management.
- Experience implementing logging, metrics, dashboards, alerts, and distributed tracing.
- Familiarity with observability platforms (Datadog, New Relic, Grafana, Prometheus, CloudWatch, OpenTelemetry).
- Strong Linux and command-line skills.
- Scripting experience with Python, Bash, JavaScript, or comparable language.
- Understanding of application deployment patterns for modern web applications, APIs, background workers, and event-driven systems.
- Experience supporting relational databases, backups, restore processes, high availability.
- Knowledge of incident-management practices, root-cause analysis, capacity planning, and production-readiness reviews.
- Working knowledge of cloud security, vulnerability remediation, encryption, certificates, key management, and compliance-oriented controls.
NICE TO HAVE
- Experience supporting a multi-tenant SaaS platform.
- Experience with serverless infrastructure and event-driven cloud services.
- Experience with PostgreSQL and managed database services.
- Experience with CDNs, media delivery, large-file processing, video workloads, or storage-intensive applications.
- Experience with security and compliance frameworks relevant to enterprise SaaS.
- Experience building internal developer platforms or self-service engineering capabilities.
- Experience with automated performance, resilience, chaos, or disaster-recovery testing.
- Experience improving cloud cost allocation, forecasting, and optimization.
- Experience supporting globally distributed engineering teams.
When you join Solvd, you'll…
- Shape real-world AI-driven projects across key industries, working with clients from startup innovation to enterprise transformation.
- Be part of a global team with equal opportunities for collaboration across continents and cultures.
- Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.
- Ready to make an impact?
- If you're excited to build things that matter, champion responsible AI, and grow with some of the industry’s sharpest minds. Apply today and let’s innovate together.
- Solvd is an equal opportunity employer.
This listing was posted by a verified recruiter at solvd. Report this listing
JobSpring