Skip to content
← Back to job listings

Senior Data Engineer

Sigma Software · Remote, Masovian Voivodeship, Poland

Data Science / AI / Machine LearningRemoteImported listingfull-time1 day ago

About The Role

  • Write and defend diagnostic SQL queries against large-scale production datasets
  • Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
  • Harmonize fields across independently designed datasets and maintain versioned field mappings
  • Develop point-in-time-correct feature tables and aggregation pipelines
  • Design and maintain conversion and labeling pipelines with delayed label handling
  • Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
  • Build experimentation infrastructure including traffic splitting and reporting pipelines
  • Perform large-scale historical backfills and safe reprocessing after mapping changes
  • Implement data isolation and safe-aggregation controls for advertiser data protection
  • Develop automated data quality validation frameworks
  • Collaborate closely with Customer engineers and prepare operational documentation
  • Contribute to architecture discussions and platform scalability improvements
  • 5+ years of experience in Data Engineering
  • At least 2 years of experience working with production ML or large-scale analytics pipelines
  • Expert-level SQL skills including window functions and incremental processing patterns
  • Strong Python skills for production-grade pipeline development
  • Hands-on experience with Spark or PySpark
  • Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
  • Experience working with cloud data warehouses at scale, preferably BigQuery
  • Strong understanding of data modeling and point-in-time correctness
  • Experience working with event-driven or clickstream datasets at very large scale
  • Experience supporting business-critical production pipelines
  • Upper-Intermediate English level or higher

WILL BE A PLUS

  • Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
  • Experience building streaming or near-real-time ingestion systems
  • Understanding of feature stores, train/serve skew, and label leakage prevention
  • Experience in AdTech or auction-based environments
  • Experience handling delayed or incomplete labels in ML systems
  • Experience with dbt or similar transformation frameworks
  • Experience delivering solutions into Customer-owned infrastructure
  • Knowledge of GDPR/CCPA-related privacy engineering practices
  • Experience with experimentation infrastructure and statistical validation pipelines
  • Experience working in hybrid cloud/on-prem Linux environments
  • Terraform and Kubernetes experience
  • Experience optimizing warehouse cost and performance

PERSONAL PROFILE

  • Strong analytical and problem-solving skills
  • Ownership-oriented mindset
  • Ability to work independently in a client-facing environment
  • Strong communication and documentation skills
  • Comfortable working in a fast-paced engineering environment
  • Collaborative and proactive attitud

This is an external listing. JobSpring does not represent or verify the employer. Report this listing