Skip to content
← Back to job listings

Senior Technical Anchor

Ford Global Career Site · Chennai, Tamil Nadu, India

External listingfull-timeabout 3 hours ago

About The Role

This is a hybrid role with a mixed Anchoring of Data Engineering and Machine Learning leadership. You will be the technical authority responsible for scaling ingestion pipelines across diverse formats (PDFs, images, Exl, Doc, Video and audio, raw text, logs) and integrating state-of-the-art ML/NLP/Vision extraction models into robust, enterprise-grade production pipelines.

As a Technical Anchor, you will balance hands-on architecture and coding with technical mentorship, driving engineering standards, MLOps best practices, Platform reliability, and cross-team alignment.

Key Responsibilities

  1. Technical Anchor & Architecture Leadership
  • Direct the end-to-end technical strategy, design, and architecture for the unstructured data ingestion and ML extraction platform.
  • Serve as the primary technical contact and subject matter expert across engineering teams, product managers, and enterprise stakeholders.
  • Lead architectural reviews, define design patterns, establish coding standards, and enforce data security, privacy, and governance protocols.
  • Mentor and coach senior and mid-level software engineers, data engineers, and ML engineers to foster technical excellence.
  1. Unstructured Data Engineering & Pipeline Infrastructure
  • Architect scalable, fault-tolerant batch and real-time streaming ingestion pipelines for processing complex unstructured data sources at scale. (PDFs, PPT, Excel, Word, Video and Audio Formats)
  • Maintain robust data orchestration, transformation, parsing, chunking, deduplication, and metadata enrichment strategies.
  • Build automated data quality gates, validation frameworks, and observability pipelines to ensure high data integrity.
  1. ML Extraction & AI Integration
  • Integrate advanced Machine Learning, Computer Vision, OCR, NLP, and LLM-based models into production extraction pipelines (e.g., document intelligence, entity extraction, semantic parsing).
  • Establish scalable model serving and inference patterns (low-latency streaming inference and high-throughput batch extraction).
  • Partner with Data Science and ML research teams to optimize model deployment, quantization, prompt engineering, fine-tuning workflows, and output validation.
  1. Platform Reliability
  • Drive platform scalability, performance tuning, infrastructure cost optimization, and high availability across cloud and hybrid environments.

Required Qualifications & Technical Experience

Experience: 8+ years of progressive experience in Software Engineering, Data Engineering, or ML Engineering, with at least 2+ years as a Technical Anchor, Lead Engineer, or Principal Architect.

  • Cloud Platform Expertise: Extensive hands-on architectural experience with Google Cloud Platform (GCP), specifically:
  • Serverless Compute: Cloud Run, Cloud Functions, Cloud Scheduler.
  • Messaging & Orchestration: GCP Pub/Sub, Cloud Tasks, Workflow Orchestration.
  • Storage & Databases: Cloud Storage (GCS), Cloud SQL (PostgreSQL/MySQL), BigQuery.
  • Data Engineering & Pipeline Skills:
  • Advanced proficiency in Python (primary), PySpark/Spark, SQL, and REST/gRPC microservice design.
  • Proven track record building high-throughput streaming and batch pipelines.
  • Deep experience designing distributed system fault-tolerance (Multi-Tier DLQ, exponential backoff, state machines).
  • Machine Learning & Extraction Stack:
  • Hands-on experience with NLP, Layout Parsing, Computer Vision, and OCR (e.g., LayoutLM, Tesseract, OpenCV, PyTorch/TensorFlow, Hugging Face).
  • Direct experience deploying and managing LLM inference integrations, Multi-Model Architecture, prompt engineering pipelines, and RAG architectures.
  • Experience with Vector Databases (e.g., Vertex AI Vector Search, Pgvector) and unstructured text chunking/embedding strategies.
  • DevOps & Infrastructure-as-Code:
  • Strong expertise in Docker containerization, Kubernetes, CI/CD pipelines (GitHub Actions, Tekton, Cloud Build), and Infrastructure-as-Code (Terraform).

Preferred Qualifications

  • Experience interfacing data extraction workloads with High-Performance Computing (HPC) clusters or GPU accelerated environments.
  • Knowledge of Human-in-the-Loop (HITL) annotation, active learning workflows, and data validation frameworks.
  • Automotive, enterprise manufacturing, or large-scale document automation domain knowledge.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing