Founding AI Platform Engineer (MLOps / Backend)
jobgether · Brazil
About The Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Founding AI Platform Engineer (MLOps / Backend) based in Brazil.
This is a foundational engineering opportunity for someone who wants to shape the systems that turn ML and GenAI capabilities into reliable, production-ready <products.You> will operate at the intersection of backend engineering, cloud infrastructure, MLOps, platform reliability, and product delivery.The role offers broad ownership across model lifecycle management, deployment, serving, experimentation, observability, and CI/CD.You will build the production infrastructure that enables AI systems to scale safely while remaining performant, maintainable, secure, and cost-efficient.Working closely with ML and product teams, you will turn ambiguous technical challenges into practical, durable <solutions.As> an early platform engineer, your decisions will directly influence engineering standards, tooling, reliability practices, and the future scalability of the platform.
Accountabilities
- Build and maintain infrastructure and tooling for training, evaluating, deploying, serving, and monitoring ML models and GenAI services.
- Develop and operate production backend services, APIs, and pipelines supporting recommendations, agent workflows, and customer-facing integrations.
- Improve CI/CD pipelines, automated testing, release processes, rollback strategies, and environment management.
- Establish comprehensive observability across application health, model behavior, agent quality, latency, costs, and operational failure modes.
- Build reproducibility and lifecycle management practices for models, prompts, datasets, configurations, and software releases.
- Support experimentation and measurement infrastructure that enables ML and product teams to evaluate changes reliably.
- Strengthen platform reliability, scalability, security, performance, and cost efficiency across the technology stack.
- Troubleshoot production issues end-to-end and convert recurring operational problems into long-term engineering improvements.
- Collaborate closely with ML, product, and engineering teams to move ambiguous initiatives from concept to completion.
- Establish engineering standards and platform practices that can support future growth and increasing system complexity.
- Identify platform, reliability, and scaling risks early and proactively address them before they affect customers or delivery.
Requirements
- Strong software engineering background with experience building, deploying, and operating production systems.
- Proven experience with backend services, cloud infrastructure, CI/CD, automated testing, observability, and engineering automation.
- Strong proficiency in Python and the ability to work effectively across backend services, infrastructure, tooling, and operational workflows.
- Good understanding of reliability, performance, maintainability, scalability, and infrastructure cost tradeoffs.
- Ability to collaborate effectively with ML and product teams and independently drive ambiguous technical work to completion.
- Strong ownership mentality, attention to detail, and a practical approach focused on simplifying and strengthening systems.
- Experience with MLOps workflows covering model training, evaluation, deployment, and monitoring is highly advantageous.
- Experience serving machine-learning models or LLM-powered applications in production is a strong plus.
- Familiarity with experimentation platforms, event pipelines, analytics instrumentation, or feature delivery platforms is beneficial.
- Experience with agent evaluation, prompt versioning, retrieval and search infrastructure, or vector-backed systems is an advantage.
- Experience supporting customer-facing APIs or SaaS platform infrastructure is preferred.
- Strong troubleshooting, communication, and cross-functional collaboration skills.
Benefits
- Opportunity to take foundational ownership of an AI platform and its engineering standards.
- Broad technical scope spanning backend engineering, cloud infrastructure, MLOps, GenAI, observability, and reliability.
- Direct opportunity to influence how ML and GenAI capabilities are brought into production.
- High level of autonomy and ownership in a growing technology environment.
- Opportunity to work closely with ML, product, and engineering teams on high-impact systems.
- Ability to shape scalable infrastructure, deployment practices, and platform architecture from an early stage.
- Fully remote work environment.
- Full-time position within the IT function.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring