
Data Engineer
hudsonmanpower (recruitee) · New Jersey, United States
About The Role
Senior Data Engineer – 8+ Years Experience
Job Summary
We are looking for an experienced Senior Data Engineer with 8+ years of hands-on experience in designing, developing, and maintaining scalable data platforms, data pipelines, and analytics solutions. The ideal candidate will have strong expertise in Python/SQL, ETL/ELT, cloud data platforms, data warehousing, distributed data processing, orchestration, and data architecture .
The candidate will work closely with Data Scientists, BI Developers, Software Engineers, Product Managers, and business stakeholders to build reliable, secure, high-performance data solutions that support business-critical analytics and AI/ML initiatives.
Key Responsibilities
- Design, develop, and maintain scalable and reliable batch and real-time data pipelines .
- Build robust ETL/ELT workflows to ingest, transform, validate, and distribute data from multiple sources.
- Develop highly optimized and complex SQL queries, stored procedures, and data transformations .
- Design and implement data warehouses, data lakes, lakehouses, and dimensional data models .
- Work with large datasets using distributed processing technologies such as Apache Spark/PySpark .
- Develop data pipelines using orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms .
- Implement data solutions on major cloud platforms such as AWS, Azure, or GCP .
- Design and optimize cloud data platforms and services such as Amazon Redshift, Snowflake, Databricks, Azure Synapse, BigQuery, or equivalent technologies .
- Implement data quality, data validation, reconciliation, monitoring, and observability frameworks.
- Develop solutions for incremental data processing, CDC, slowly changing dimensions, partitioning, and performance optimization .
- Build and maintain real-time/streaming data pipelines using technologies such as Kafka, Kinesis, or equivalent tools.
- Implement appropriate data security, governance, access control, encryption, and compliance practices.
- Collaborate with data architects to translate business requirements into scalable technical solutions.
- Perform performance tuning of data pipelines, databases, Spark jobs, and cloud data workloads.
- Establish and maintain CI/CD practices for data engineering workflows.
- Write unit, integration, and data-quality tests to ensure reliability of production pipelines.
- Troubleshoot production data issues and participate in incident resolution and root-cause analysis.
- Conduct code reviews and promote engineering best practices across the data engineering team.
- Mentor junior and mid-level data engineers and provide technical leadership.
- Document data architecture, pipeline designs, data models, operational procedures, and technical decisions.
- Stay current with emerging technologies in cloud, big data, data engineering, data platforms, and AI/ML .
Required Technical Skills
Programming & Database
- Strong proficiency in Python .
- Advanced SQL skills.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle .
- Experience with NoSQL databases such as MongoDB, DynamoDB, Cassandra, or similar is advantageous.
- Strong understanding of database design, indexing, query optimization, and transaction management.
Big Data & Distributed Processing
- Strong experience with Apache Spark / PySpark .
- Experience with Hadoop ecosystem technologies is desirable.
- Understanding of distributed computing, partitioning, parallel processing, and performance optimization.
Data Engineering & ETL
- Extensive experience building ETL/ELT pipelines .
- Experience with tools such as:
- Apache Airflow
- Azure Data Factory
- AWS Glue
- dbt
- Informatica
- Talend
- SSIS
- Experience handling structured, semi-structured, and unstructured data.
Cloud Technologies
Strong experience with at least one major cloud platform
AWS
- S3
- Glue
- EMR
- Redshift
- Lambda
- Kinesis
- Athena
- IAM
Azure
- Azure Data Factory
- Azure Data Lake Storage
- Azure Databricks
- Azure Synapse Analytics
- Azure Functions
- Event Hubs
- Key Vault
GCP
- BigQuery
- Cloud Storage
- Dataflow
- Dataproc
- Pub/Sub
- Cloud Composer
Data Warehousing & Lakehouse
- Strong understanding of data warehouse architecture .
- Experience with Snowflake, Databricks, Redshift, Synapse, BigQuery , or equivalent.
- Expertise in:
- Star and Snowflake schemas
- Fact and dimension tables
- Slowly Changing Dimensions (SCD)
- Data marts
- Data lakes
- Lakehouse architecture
- Partitioning and clustering
- Data modeling
Streaming & Real-Time Data
- Experience with Apache Kafka or equivalent streaming platforms.
- Understanding of producers, consumers, topics, partitions, offsets, consumer groups, and schema management.
- Experience developing real-time or near-real-time data processing pipelines.
DevOps & Engineering Practices
- Experience with Git/GitHub/GitLab/Bitbucket .
- Experience with CI/CD pipelines .
- Knowledge of Docker and Kubernetes is desirable.
- Experience with Infrastructure as Code tools such as Terraform is advantageous.
- Familiarity with automated testing, deployment, monitoring, and observability.
Data Governance & Security
- Understanding of data governance, metadata management, lineage, data cataloging, and data quality .
- Experience implementing role-based access control and secure data access.
- Knowledge of privacy and compliance requirements such as GDPR, CCPA, HIPAA, or equivalent regulations , depending on business requirements.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring