Skip to content
← Back to job listings

Associate Director- AI Engineering

Trinity Partners India LLP · Bangalore, India

Software DevelopmentExternal listingfull-timeabout 1 hour ago

About The Role

About Trinity & This Role

Trinity Life Sciences is a leading analytics and technology partner to the biotech and pharma industry, helping clients solve their hardest commercialisation challenges. AI Foundations is the internal team that makes AI adoption fast, safe, and repeatable across Trinity — owning the platforms, standards, and reusable infrastructure that every product team builds on.

Trinity is seeking a Associate Director to own the platform choices, cloud infrastructure setup, and tooling that underpin Trinity's AI capabilities. In practice, cloud infrastructure at Trinity may be provisioned and governed in collaboration with a central IT or cloud operations team. The Platform Engineer must be equally effective whether building and owning infrastructure directly or operating expertly within a shared, IT-governed environment — in either case, the responsibility for making the AI platform work well for product teams sits squarely with this role.

The right person cares deeply about building platforms that engineers want to use: reliable, well-documented, and designed to accelerate product delivery. They are as comfortable navigating an enterprise IT process to get the right infrastructure approved as they are writing Terraform to provision it themselves.

Key Responsibilities

Cloud Infrastructure: Design, Setup & Governance

The approach to infrastructure provisioning will depend on Trinity's operating model — this role spans the full range.

  • Where infrastructure ownership sits with AI Foundations : design, build, and maintain cloud infrastructure for AI and agentic workloads across AWS, Azure , and/or Snowflake — making deliberate, well-justified choices about services, topology, networking, and cost architecture using Terraform, Pulumi , or AWS CDK/CloudFormation .
  • Where infrastructure is IT-governed : work effectively within centrally provisioned environments — understand what has been set up, identify gaps or misconfigurations that affect AI workloads, and work with IT to get the right resources allocated, correctly sized, and properly secured. Be the bridge between AI Foundations' technical requirements and the IT team's provisioning processes.
  • Regardless of who provisions: own the configuration, connectivity, and correct operation of the infrastructure AI Foundations uses — network topology, IAM roles and permissions, service endpoints, inter-service connectivity, and environment isolation.
  • Design and document infrastructure for AI-specific requirements: GPU/accelerated compute provisioning, model serving endpoints, high-throughput inference scaling, vector store hosting, and low-latency retrieval pipelines.
  • Own cloud cost management — rightsizing, tagging strategies, reserved/spot instance recommendations, and per-workload cost visibility — and translate cost implications into clear decisions for engineering and leadership.
  • Infrastructure as Code & Automation
  • Where the team has the latitude to build infrastructure directly, IaC is the standard. Where infrastructure is IT-managed, automation still applies to the layers above it.
  • Manage infrastructure through code — Terraform, Pulumi , or AWS CDK — wherever the team has provisioning ownership; ensure all environments are reproducible, version-controlled, and auditable.
  • Build and maintain CI/CD pipelines for infrastructure changes with automated validation, drift detection, and safe rollout processes.
  • Automate environment configuration and application-layer setup above the infrastructure layer — even in IT-governed setups, there is significant automation value in Kubernetes namespace setup, service account configuration, secrets injection, and deployment scaffolding.
  • Implement infrastructure monitoring and alerting: health checks, cost anomaly detection, capacity alerts, and runbooks for common failure scenarios.

Navigating Enterprise IT in a Governed Environment

Many enterprise AI teams operate in environments where cloud accounts, networking, and core services are set up and controlled by a central IT or cloud platform team. This is a common and legitimate model — and navigating it effectively is a real skill.

  • Understand Trinity's IT governance model for cloud provisioning — how requests are raised, approved, and fulfilled — and become the person who knows how to move quickly within it without cutting corners.
  • Translate AI Foundations' infrastructure requirements into well-specified requests that IT teams can act on: clear resource specifications, security requirements, tagging conventions, and acceptance criteria.
  • Proactively identify provisioning gaps or mismatches between what IT has set up and what agentic workloads actually need — and drive resolution rather than working around them.
  • Build relationships with IT and cloud ops counterparts so that AI Foundations is not waiting in a queue but is a known, trusted team whose requirements are understood and prioritised appropriately.
  • Maintain a clear picture of what Trinity's AI platform looks like end-to-end — even when different parts are owned by different teams — so there are no surprises at deployment time.

Container Orchestration & Deployment

  • Own Trinity's container platform for AI workloads — Docker image standards, base image management, and Kubernetes cluster configuration ( EKS, AKS , or equivalent) including node pools, autoscaling, resource quotas, and namespace governance.
  • Design deployment patterns for agentic AI workloads: stateless agent workers, async job queues, model serving deployments with rolling updates, and sidecar patterns for observability and policy enforcement.
  • Build and maintain Helm charts or equivalent deployment manifests; manage cluster upgrades and platform stability across AI Foundations' deployments.
  • Support engineers in containerising their workloads correctly — the subject matter expert on Docker and Kubernetes best practices within the team.

Platform Adoption in Product Development

Building or configuring the platform is necessary but not sufficient. The real measure of success is whether Trinity's product teams adopt it well and build on it confidently.

  • Work closely with AI Engineers and Data Scientists to ensure platform choices are reflected in how products are built — not just in documentation that nobody reads.
  • Define and promote platform-level patterns for how agentic workloads should be deployed, scaled, monitored, and updated — making the right way the easy way.
  • Identify friction where engineers are working around the platform rather than with it; fix the root cause rather than asking engineers to adapt.
  • Contribute to Trinity's internal developer platform: self-service tooling, environment templates, and shared infrastructure modules that accelerate onboarding and reduce duplicated setup work across teams.
  • Participate in architecture reviews to ensure new product workloads are platform-compatible from the design stage — not retrofitted at deployment time.

Security, Compliance & Reliability

  • Implement cloud security controls aligned with Trinity's enterprise security standards: network segmentation, IAM policies with least-privilege principles, RBAC at cluster and service levels, secrets management ( AWS Secrets Manager, Azure Key Vault, HashiCorp Vault ), and encryption at rest and in transit.
  • Ensure infrastructure and platform configurations meet compliance requirements relevant to Trinity's life sciences client base — data residency, access audit trails, and environment isolation between workloads where required.
  • Design for reliability: multi-AZ deployments, disaster recovery runbooks, backup strategies, and SLA-aware infrastructure choices that match the criticality of the workloads they support.
  • Build and maintain observability infrastructure — centralised logging ( ELK, CloudWatch, Azure Monitor ), metrics ( Prometheus/Grafana ), distributed tracing, and alerting — so the team has full visibility into platform and application health.

Platform Strategy & Vendor Evaluation

  • Evaluate cloud services, managed AI platforms, and infrastructure tooling with well-evidenced views on build vs buy vs managed service for each layer of the stack.
  • Track the evolution of the cloud AI platform landscape ( AWS Bedrock, Azure AI Foundry, Snowflake Cortex, Databricks ) and advise the team on when new managed capabilities should replace custom infrastructure.
  • Contribute to Trinity's infrastructure roadmap — translating near-term product needs and longer-term platform ambitions into a coherent, prioritised plan.

What We Are Looking For

  • 12+ years of cloud infrastructure or platform engineering in production environments, with hands-on ownership of complex, multi-service cloud architectures.
  • Deep expertise in at least one major cloud provider ( AWS preferred; Azure or GCP considered) — able to design architectures, debug failures, and optimise costs at scale.
  • Experience operating in both self-managed and IT-governed cloud environments — comfortable provisioning infrastructure directly and equally comfortable navigating enterprise governance processes to get the right resources allocated.
  • Infrastructure as Code skills — Terraform, Pulumi, or AWS CDK — and experience managing infrastructure through CI/CD pipelines; working knowledge of IaC even in primarily IT-managed setups.
  • Proficient with Docker and Kubernetes in production — cluster management, workload deployment, autoscaling, and day-2 operations on EKS, AKS, or equivalent.
  • Solid understanding of cloud security fundamentals: IAM, RBAC, network security groups, secrets management, and compliance controls in enterprise environments.
  • Experience with AI or ML workloads on cloud infrastructure — GPU provisioning, model serving, vector databases, or managed AI platforms — is highly valued.
  • Strong SQL and scripting skills ( Python, Bash ); experience with observability tooling (Prometheus, Grafana, ELK, CloudWatch).
  • Degree in Computer Science, Engineering, or related field from IITs, NITs, BITS, or comparable. Strong applied track record considered equally.

The Kind of Achievements We Respect

  • "Designed and built a multi-tenant EKS platform for ML inference workloads. Reduced model deployment time from 2 days to under 30 minutes. Adopted by 5 engineering teams."
  • "In an IT-governed environment, identified and resolved 12 infrastructure misconfigurations blocking ML workload deployment. Established a requirements specification process with the IT team that cut provisioning turnaround from 3 weeks to 4 days."
  • "Migrated a manually provisioned AWS environment to fully Terraform-managed infrastructure. Eliminated configuration drift and reduced environment setup time by 80%."
  • "Implemented GPU autoscaling for LLM inference on EKS with spot instance fallback. Reduced inference infrastructure cost by 45% at equivalent throughput."

Why This Role Matters

Every agentic AI system Trinity builds runs on the infrastructure and platform this role owns — or navigates. Getting the platform right, whether built from scratch or configured within an enterprise IT model, determines whether AI Foundations can deliver at the speed and quality Trinity's growth requires. If you find platform engineering most satisfying when it enables others to build great things — and you are equally at home with a Terraform plan or a ticket to IT — this is the right role for you.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing