Data Engineer
bureau · Bangalore
About The Role
About BureauBureau is a unified risk decisioning platform for Compliance, Fraud, and Transaction risks. Our platform is a single decision-making engine, powered by a 1 billion+ identity knowledge graph. Over 150 Banks, fintechs, retailers, and digital platforms use Bureau to verify identities faster and stop fraud earlier globally.Bureau has raised $50M+ from renowned Silicon Valley and global investors including Sorenson Capital and PayPal Ventures and is expanding rapidly from APAC to Americas, Europe, and beyond.Why Bureau?Bureau is building the infrastructure that makes digital identities and transactions safe and trustworthy for billions of people. The mission is big, the problems are complex, and the impact is real.We hire people who want that level of responsibility. People who move fast, build systems from scratch, and care deeply about turning strategy into execution. If you want predictability or narrow scope, this won't be your place. If you want to shape how a scaling global company operates—keep reading.What You'll DoBuild and maintain batch and streaming pipelines that ingest, clean, and transform data from device SDKs, internal services, partner APIs, and third-party data providersWrite and own backend services and RESTful APIs that serve data to internal teams and to real-time decisioning pathsDevelop Airflow DAGs powering reporting, compliance analytics, model training, and feature pipelines, and keep them healthy day to dayWrite Spark jobs and SQL transformations against our data lake, and tune them when they get slow or expensiveAdd monitoring, alerting, and data quality checks to the pipelines you own so problems surface before a stakeholder noticesDebug production issues across the stack: a Kafka consumer lagging, a schema change breaking downstream, a query that got 10x slower after a data volume jumpWork with the infrastructure your pipelines and services run on — containers, deployments, cluster configs — and help keep it stable and cost-saneContribute to our identity graph work, helping model relationships between entities to surface fraud rings and hidden linkagesWrite documentation and tests, participate in design reviews, and help keep our schemas and contracts sane as the system growsWhat You'll BringMust have1–3 years of professional software engineering experience, with meaningful exposure to data-intensive systemsStrong programming skills in Python, Java, or Scala, and the ability to write production-quality, tested codeStrong SQL: joins, window functions, aggregations, and enough of a mental model of query execution to know why something is slowWorking understanding of databases, including the difference between OLTP and OLAP systems and when each is appropriateHands-on experience with at least one distributed data processing framework (Spark preferred) or a genuine willingness to ramp up quicklyExperience building or maintaining backend services and REST APIsFamiliarity with a major cloud platform (AWS preferred) and core services like S3, EC2, and managed databasesSolid computer science fundamentals: data structures, concurrency, and the basics of distributed systemsComfort with Git, code review, and CI/CDNice to haveInfrastructure and systems knowledge — Docker, Kubernetes, Terraform or similar IaC, and a working sense of how services get deployed, scaled, and monitored in productionExperience running or tuning distributed workloads: cluster sizing, resource tuning, or tracking down a bottleneck between compute and storageExposure to Kafka or MSK, or any streaming/event-driven systemExperience with Airflow or a similar orchestration toolExposure to EMR, Athena, ClickHouse, Databricks, or SnowflakeAwareness of lakehouse table formats such as Iceberg, Delta Lake, or HudiFamiliarity with observability tooling — Prometheus, Grafana, Datadog, or equivalentAny experience with graph databases (Neo4j, TigerGraph, Amazon Neptune)Interest in or exposure to fraud, risk, identity, fintech, or other high-stakes real-time domainsExposure to ML workflows: feature pipelines, training data preparation, or model servingWhat We Look ForYou debug rather than guess, and you can explain what actually went wrongYou ask what the data is for before deciding how to model itYou're curious about the layer below the one you work in — how your code actually runs, where it fails, what it costsYou're comfortable being new to a tool and getting productive in it quicklyYou care that numbers are right, because at Bureau a wrong number is a wrong risk decisionOur CultureWe hire self-motivated people and get out of their wayWe value performance, not hours workedSpeed, ownership, and impact matter mostCompensationCompetitive salary + potential equityHealth benefits, flexible PTO, learning budget
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing