← Back to job listings
FG
Data Engineer
Ford Global Career Site · India
About The Role
We are looking for someone with advanced Python programming skills who applies robust software engineering principles to data problems. You will collaborate closely with Data Scientists, ML Engineers, and Product Managers to build the scalable, automated pipelines required to train, deploy, and monitor machine learning models in production.
Key Responsibilities
- GCP Pipeline Development: Design, build, and maintain highly scalable ETL/ELT data pipelines using Python and GCP-native data processing tools (e.g., Cloud Run, Cloud Functions).
- AI/ML Infrastructure Support: Engineer feature stores, robust data feeds specifically optimized for machine learning training and inference. Work closely with ML Engineers to operationalize models using Vertex AI.
- Data Integration & Ingestion: Write clean, modular Python code to ingest data from diverse sources (APIs, streaming platforms, on-prem databases) into BigQuery and Google Cloud Storage (GCS).
- System Optimization: Optimize BigQuery architecture, partition/cluster tables, and tune complex SQL queries to ensure performance and cost-efficiency at a massive scale.
- Software Engineering Best Practices: Champion best practices in Python development, including version control (Git), CI/CD pipelines (Cloud Build / GitHub Actions), code reviews, and comprehensive unit/integration testing.
- Data Quality & Governance: Implement robust data quality checks, alerting, and monitoring to ensure the data feeding our AI models is accurate and reliable.
Required Qualifications
- Degree: Bachelor’s or Master’s degree in Computer Science, Engineering, Mathematics, or a related technical field (or equivalent practical experience).
- Experience: 4 to 6 years of professional experience in Data Engineering, Software Engineering, or a closely related field.
- Advanced Python: Deep expertise in Python programming. You should be highly comfortable with:
- Data processing and ML-adjacent libraries (e.g., PySpark, Pandas, NumPy).
- API development
- Writing efficient and production-grade code.
- GCP Mastery: Proven, hands-on experience designing and operating data architectures on Google Cloud Platform. Must have strong experience with:
- BigQuery (advanced SQL, architecture, and optimization).
- Google Cloud Storage (GCS) .
- Compute/Serverless (Cloud Functions, Cloud Run).
- AI/ML Acumen: Experience working alongside Data Science teams. A strong understanding of the ML lifecycle, feature engineering, and the data requirements for model training and deployment.
- MLOps: Understanding of MLOps principles, model registry, and continuous training pipelines.
Preferred Qualifications
- Vertex AI: Direct experience interacting with or deploying pipelines using Google Cloud's Vertex AI platform.
- Streaming Technologies: Familiarity with real-time data processing using Google Cloud Pub/Sub and streaming Dataflow jobs.
- Infrastructure as Code: Experience managing GCP resources using Terraform.
- Containerization: Proficiency with Docker.
Similar roles you might like
See all →0A
Lead AI Engineer
0063 Applied Materials India Private Limited
Salary not disclosedPosted today
IA
Lead Consultant - Snowflake and Cortex AI Engineer
IN12 AstraZeneca India Pvt Ltd Company
Salary not disclosedPosted today
IA
IT Senior Business Analyst
IN12 AstraZeneca India Pvt Ltd Company
Salary not disclosedPosted today
6S
UI LMTS - AI Engineering
611 salesforce.com India Private Limited
Salary not disclosedPosted today
G
Consultant - Data and Analytics Advisory 4A
Genpact
Salary not disclosedPosted today
1S
LLM Data Scientist, Senior Associate
1685 SS CORP SVCS MUMBAI PVT LTD
Salary not disclosedPosted today
II
Senior ML Engineer
IN01 IND Fresenius Medical Care India Private Limited
Salary not disclosedPosted today
AP
Staff data analyst
A3R0-GTCI Private Limited
Salary not disclosedPosted today
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
