Skip to content
← Back to job listings

Senior Data Engineer - Real World Data & Healthcare Analytics

Zifo · Chennai, Tamil Nadu, India

Data Science / AI / Machine LearningExternal listingfull-time43 minutes ago

About The Role

We are looking for a highly skilled Data Engineer with Real World Data (RWD) experience to build and manage end-to-end data pipelines for multimodal healthcare datasets. The ideal candidate will work at the intersection of data engineering, analytics, and healthcare research, transforming complex healthcare data into analysis-ready assets that support advanced analytics, AI/ML initiatives, and evidence-generation studies. This role requires expertise in large-scale healthcare data processing, data harmonization, cloud platforms, and modern data engineering practices.

Responsibilities

Data Engineering & Pipeline Development

  • Design, develop, and maintain scalable ETL/ELT pipelines for large healthcare and real world datasets.
  • Build and manage data ingestion, transformation, harmonization, and analytics layers.
  • Implement data quality frameworks, governance controls, lineage tracking, and monitoring.
  • Manage data lifecycle processes across raw, curated, and analytics-ready environments.
  • Work with structured and unstructured healthcare datasets from multiple sources.

Healthcare Data Harmonization

  • Harmonize heterogeneous healthcare data sources and coding systems into standardized formats.
  • Map and transform clinical terminologies including
  • o SNOMED CT
  • o ICD-10
  • o LOINC
  • o RxNorm
  • o CPT/HCPCS
  • Support implementation of common data models such as OMOP and FHIR.

Analytics & Study Support

  • Support data feasibility assessments and data quality evaluations.
  • Collaborate with epidemiologists, biostatisticians, data scientists, and business stakeholders.
  • Develop reusable data assets, cohorts, and model-ready datasets.
  • Enable advanced analytics and AI/ML use cases through reliable data engineering practices.

Application & Platform Development

  • Contribute to analyst-facing applications, dashboards, and self-service data products.
  • Support development of data products using modern workflow automation and AI assisted engineering approaches.
  • Provide guidance on efficient querying and optimization of large longitudinal datasets.

Requirements

Data Engineering

  • Strong experience with Python, SQL, Spark / PySpark
  • Experience building production-grade ETL/ELT pipelines.
  • Strong understanding of data modelling concepts - Star schema, Snowflake schema, Normalization and denormalization
  • Experience with metadata management, lineage, monitoring, and data governance.

Platforms & Technologies

Experience in one or more of the following - Palantir Foundry, Databricks, Snowflake, AWS or equivalent cloud platforms, HPC environments

  • Containerized workloads Software Engineering Practices
  • Git
  • CI/CD pipelines
  • Unit testing and automation
  • Performance monitoring and optimization

Domain Expertise

Candidates should have working knowledge of Healthcare Real World Data (RWD), Claims data, Electronic Health Records (EHR), Registries, Patient-reported outcomes, Wearables and digital health datasets

Understanding of study feasibility, observational research, and healthcare analytics workflows is highly desirable.

AI & Automation Experience

Preferred experience with Large Language Models (LLMs), AI-assisted data engineering, Agentic workflows, Data profiling and automated data quality assessments, integration of ML outputs into production data pipelines

Qualification / Requirement

  • 4-8 years of experience in data engineering, healthcare analytics, or real-world data platforms.
  • Experience working with large-scale healthcare datasets in regulated environments.
  • Strong problem-solving and analytical skills.
  • Excellent stakeholder communication capabilities.
  • Formal educational qualifications are flexible; relevant experience and expertise are valued.

Nice to Have

  • Experience with multimodal healthcare datasets (clinical, omics, imaging, genomics, proteomics, microbiome, etc.).
  • Hands-on experience implementing OMOP/FHIR at scale.
  • Experience building self-service applications and data products for business users.
  • Familiarity with federated data networks and data quality frameworks.

What We're Looking For

  • Systems thinker who can work with complex and evolving datasets.
  • Strong collaboration skills across technical and business teams.
  • Agile mindset with a focus on delivery.
  • Commitment to data privacy and ethics.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing