Lead Data Engineer
MobiKwik · Gurgaon, Haryana, India
About The Role
About the Role As a Lead Data Engineer at MobiKwik, you will be responsible for designing, building, and optimizing scalable data pipelines and infrastructure that fuel analytics, product innovation, and business insights. You will work closely with data scientists, analysts, and product managers to ensure reliable, high-quality, and accessible data across the organization. This is a hands-on technical leadership role where you’ll both code and mentor, ensuring best practices in data engineering are followed. Key Responsibilities 1. Data Infrastructure: Build and maintain scalable, secure, and high-performance data platforms on AWS (Glue, Redshift, S3, EMR, Athena). Ensure systems are reliable, cost-efficient, and meet business SLAs. 2. Data Pipelines: Design and implement ETL/ELT workflows to ingest, process, and transform large structured and unstructured datasets from multiple sources (transactional DBs, APIs, logs, Kafka). Enable both batch and real-time data streaming solutions. 3. Data Modeling & Warehousing: Develop efficient data models and schemas to support analytics, BI, and reporting. Optimize storage and queries in Redshift, Hive, or MySQL. 4. Performance & Optimization: Continuously tune data processing frameworks (PySpark, Glue, Kafka, Hive) to ensure performance, scalability, and cost-effectiveness. 5. Collaboration: Partner with data scientists, product managers, and analysts to translate requirements into scalable data solutions. Ensure data availability and accessibility for experimentation, insights, and decision-making. 6. Quality & Monitoring: Implement robust monitoring, alerting, and data validation frameworks to ensure data integrity, reliability, and availability. 7. Team Mentorship: Guide junior engineers with technical reviews, coding best practices, and architecture discussions. Foster a culture of innovation, ownership, and continuous improvement. Requirements 7–10 years of experience in Data Engineering, with hands-on expertise in: ● Python, PySpark, SQL, Shell scripting ● AWS data stack (Glue, Redshift, Athena, EMR, S3) ● Streaming technologies (Kafka, Kinesis) ● ETL/ELT frameworks and data orchestration tools (Airflow, Step Functions). ● Strong understanding of databases (MySQL, MongoDB, Hive, NoSQL) and query optimization. ● Proven experience designing and scaling data pipelines for large volumes of data. ● Excellent problem-solving, debugging, and performance-tuning skills. ● Ability to mentor, influence, and lead while remaining hands-on with code.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
