Skip to content
← Back to job listings

DE&A - Sr. Data Engineer - Databricks

Fa Etvl Saasfaprod1 · Pune, Maharashtra, India

Data Science / AI / Machine LearningSenior LevelExternal listingfull-time6 days ago

About The Role

We are looking for a hands-on Senior Data Engineer with 7 to 9 years of experience to design, build, and optimize large-scale data pipelines and lakehouse solutions on Databricks. The ideal candidate is a strong individual contributor who stays current with the latest Databricks platform capabilities — including Unity Catalog, Lakeflow (Delta Live Tables), Lakebase, and Databricks' expanding AI/agent tooling — and can apply them to solve real-world data engineering problems at scale.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines on Databricks using PySpark, Spark SQL, and Delta Lake for batch and streaming workloads.
  • Build and manage declarative pipelines using Lakeflow / Delta Live Tables (DLT), including expectations, data quality checks, and change data capture (CDC).
  • Implement and manage data governance, access control, lineage, and data sharing using Unity Catalog across multiple workspaces and clouds.
  • Optimize Spark jobs and Databricks clusters for performance and cost, leveraging Photon, serverless compute, auto-scaling, and job clustering best practices.
  • Design and implement medallion architecture (bronze/silver/gold) data models and lakehouse patterns for analytics and ML consumption.
  • Work with Databricks Workflows (Jobs) to orchestrate multi-task pipelines, including dependency management, retries, and monitoring/alerting.
  • Integrate Databricks with cloud-native services (AWS/Azure/GCP) such as S3/ADLS/GCS, Kafka/Event Hubs/Kinesis, Glue/ADF, and IAM/Entra ID for secure, automated data flows.
  • Apply CI/CD practices for Databricks using Databricks Asset Bundles (DABs), Repos, and Git integration; automate deployments across dev/test/prod.
  • Evaluate and adopt newer Databricks capabilities — Lakebase (serverless Postgres on the lakehouse), Unity Catalog Metrics, Genie/Agent Bricks, Mosaic AI, and real-time/streaming enhancements — and recommend where they add value to existing pipelines.
  • Implement data quality, testing, and observability frameworks (e.g., Great Expectations, DLT expectations, Lakehouse Monitoring) to ensure trustworthy, production-grade data.
  • Collaborate with data scientists, analysts, and business stakeholders to understand requirements and translate them into robust, reusable data engineering solutions.
  • Mentor junior engineers, participate in code reviews, and contribute to engineering best practices, coding standards, and documentation.
  • Troubleshoot production data pipeline issues, perform root-cause analysis, and drive continuous improvement in reliability and performance.

Required Skills & Experience

  • 7–9 years of overall experience in Data Engineering, with at least 3–4 years of hands-on, production experience on the Databricks platform.
  • Strong programming skills in Python and/or Scala, with deep hands-on expertise in PySpark and Spark SQL.
  • Solid experience with Delta Lake (ACID transactions, time travel, schema evolution, optimize/vacuum/Z-ordering, liquid clustering).
  • Hands-on experience with Lakeflow / Delta Live Tables (DLT) for building declarative, quality-controlled pipelines.
  • Working knowledge of Unity Catalog for centralized governance, fine-grained access control, data lineage, and cross-workspace data sharing.
  • Experience with Databricks Workflows/Jobs for pipeline orchestration, scheduling, and monitoring.
  • Proficiency with at least one major cloud platform (AWS, Azure, or GCP) and its native storage/compute/security services.
  • Experience with streaming technologies such as Structured Streaming, Kafka, Event Hubs, or Kinesis.
  • Strong SQL skills, including performance tuning, partitioning strategies, and query optimization on large datasets.
  • Familiarity with CI/CD for data platforms — Databricks Asset Bundles, Git-based version control, Terraform, and automated testing/deployment pipelines.
  • Understanding of data modeling concepts (dimensional modeling, medallion/lakehouse architecture) and data warehousing fundamentals.
  • Demonstrated ability to stay current with the Databricks product roadmap (e.g., Unity Catalog enhancements, Lakebase, Genie/Agent Bricks, Mosaic AI, Lakehouse Monitoring, serverless compute) and apply relevant updates to existing systems.
  • Strong analytical, debugging, and performance-tuning skills across the Databricks/Spark stack.
  • Excellent communication skills with the ability to work directly with cross-functional stakeholders and, where applicable, mentor junior team members.

Good to Have

  • Databricks Certified Data Engineer Associate/Professional or Databricks Certified Associate/Professional Developer for Apache Spark certification.
  • Exposure to MLOps/MLflow, Mosaic AI, or Databricks' agent/GenAI tooling (Agent Bricks, Genie, AI/BI dashboards).
  • Experience with Lakebase or other Postgres/OLTP-on-lakehouse patterns for operational analytics use cases.
  • Experience with dbt, Airflow, or similar orchestration/transformation tools alongside Databricks.
  • Prior experience in a regulated or high-governance data environment (finance, healthcare, or similar).
  • Contributions to internal frameworks, reusable pipeline templates, or engineering best-practice documentation.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing