Skip to content
← Back to job listings

Data Engineer

hudsonmanpower (recruitee) · New Jersey, United States

Data Science / AI / Machine LearningEntry LevelExternal listingcontract2 days ago

About The Role

Senior Data Engineer – 8+ Years Experience

Job Summary

We are looking for an experienced Senior Data Engineer with 8+ years of hands-on experience in designing, developing, and maintaining scalable data platforms, data pipelines, and analytics solutions. The ideal candidate will have strong expertise in Python/SQL, ETL/ELT, cloud data platforms, data warehousing, distributed data processing, orchestration, and data architecture .

The candidate will work closely with Data Scientists, BI Developers, Software Engineers, Product Managers, and business stakeholders to build reliable, secure, high-performance data solutions that support business-critical analytics and AI/ML initiatives.

Key Responsibilities

  • Design, develop, and maintain scalable and reliable batch and real-time data pipelines .
  • Build robust ETL/ELT workflows to ingest, transform, validate, and distribute data from multiple sources.
  • Develop highly optimized and complex SQL queries, stored procedures, and data transformations .
  • Design and implement data warehouses, data lakes, lakehouses, and dimensional data models .
  • Work with large datasets using distributed processing technologies such as Apache Spark/PySpark .
  • Develop data pipelines using orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms .
  • Implement data solutions on major cloud platforms such as AWS, Azure, or GCP .
  • Design and optimize cloud data platforms and services such as Amazon Redshift, Snowflake, Databricks, Azure Synapse, BigQuery, or equivalent technologies .
  • Implement data quality, data validation, reconciliation, monitoring, and observability frameworks.
  • Develop solutions for incremental data processing, CDC, slowly changing dimensions, partitioning, and performance optimization .
  • Build and maintain real-time/streaming data pipelines using technologies such as Kafka, Kinesis, or equivalent tools.
  • Implement appropriate data security, governance, access control, encryption, and compliance practices.
  • Collaborate with data architects to translate business requirements into scalable technical solutions.
  • Perform performance tuning of data pipelines, databases, Spark jobs, and cloud data workloads.
  • Establish and maintain CI/CD practices for data engineering workflows.
  • Write unit, integration, and data-quality tests to ensure reliability of production pipelines.
  • Troubleshoot production data issues and participate in incident resolution and root-cause analysis.
  • Conduct code reviews and promote engineering best practices across the data engineering team.
  • Mentor junior and mid-level data engineers and provide technical leadership.
  • Document data architecture, pipeline designs, data models, operational procedures, and technical decisions.
  • Stay current with emerging technologies in cloud, big data, data engineering, data platforms, and AI/ML .

Required Technical Skills

Programming & Database

  • Strong proficiency in Python .
  • Advanced SQL skills.
  • Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle .
  • Experience with NoSQL databases such as MongoDB, DynamoDB, Cassandra, or similar is advantageous.
  • Strong understanding of database design, indexing, query optimization, and transaction management.

Big Data & Distributed Processing

  • Strong experience with Apache Spark / PySpark .
  • Experience with Hadoop ecosystem technologies is desirable.
  • Understanding of distributed computing, partitioning, parallel processing, and performance optimization.

Data Engineering & ETL

  • Extensive experience building ETL/ELT pipelines .
  • Experience with tools such as:
  • Apache Airflow
  • Azure Data Factory
  • AWS Glue
  • dbt
  • Informatica
  • Talend
  • SSIS
  • Experience handling structured, semi-structured, and unstructured data.

Cloud Technologies

Strong experience with at least one major cloud platform

AWS

  • S3
  • Glue
  • EMR
  • Redshift
  • Lambda
  • Kinesis
  • Athena
  • IAM

Azure

  • Azure Data Factory
  • Azure Data Lake Storage
  • Azure Databricks
  • Azure Synapse Analytics
  • Azure Functions
  • Event Hubs
  • Key Vault

GCP

  • BigQuery
  • Cloud Storage
  • Dataflow
  • Dataproc
  • Pub/Sub
  • Cloud Composer

Data Warehousing & Lakehouse

  • Strong understanding of data warehouse architecture .
  • Experience with Snowflake, Databricks, Redshift, Synapse, BigQuery , or equivalent.
  • Expertise in:
  • Star and Snowflake schemas
  • Fact and dimension tables
  • Slowly Changing Dimensions (SCD)
  • Data marts
  • Data lakes
  • Lakehouse architecture
  • Partitioning and clustering
  • Data modeling

Streaming & Real-Time Data

  • Experience with Apache Kafka or equivalent streaming platforms.
  • Understanding of producers, consumers, topics, partitions, offsets, consumer groups, and schema management.
  • Experience developing real-time or near-real-time data processing pipelines.

DevOps & Engineering Practices

  • Experience with Git/GitHub/GitLab/Bitbucket .
  • Experience with CI/CD pipelines .
  • Knowledge of Docker and Kubernetes is desirable.
  • Experience with Infrastructure as Code tools such as Terraform is advantageous.
  • Familiarity with automated testing, deployment, monitoring, and observability.

Data Governance & Security

  • Understanding of data governance, metadata management, lineage, data cataloging, and data quality .
  • Experience implementing role-based access control and secure data access.
  • Knowledge of privacy and compliance requirements such as GDPR, CCPA, HIPAA, or equivalent regulations , depending on business requirements.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing