
IN_Senior Associate_Data Engineer + AI_One Consulting_Advisory_Kolkata
PwC Asia · Kolkata Y-14, Kolkata, West Bengal, India
About The Role
Job Description & Summary: · We are looking for a Senior Data Engineer with deep, hands-on expertise in Microsoft Fabric and the broader Azure data ecosystem to design, build, and operate enterprise-scale analytics platforms. You will own end-to-end data solutions — from real-time ingestion and medallion-architecture lakehouses to governed semantic models and business-facing dashboards — for large enterprise clients, including regulated industries such as energy and utilities. · This is a builder role. You will write production code (Python/PySpark, T-SQL, KQL), automate deployments through CI/CD, and work directly with architects and business stakeholders to turn requirements into reliable, observable data products. Job Position Title: Senior Associate Responsibilities: · Architecture & Platform Design · Define end-to-end solution architecture for enterprise analytics platforms on Microsoft Fabric and Azure - ingestion, storage, compute, serving, and consumption layers — with clear rationale and documented trade-offs. · Design workspace topology, capacity strategy, and environment separation (dev/test/prod), including naming standards, item organization, and promotion paths. · Establish platform-level design patterns: medallion layering standards, hot-path vs. cold-path (Lambda-style) architectures, metadata/control-table-driven orchestration frameworks, and reusable pipeline templates. · Make and defend technology selection decisions (e.g., Warehouse vs. Lakehouse, Eventhouse vs. Warehouse for a workload, ADF vs. Fabric pipelines, native vs. custom alerting) based on cost, latency, governance, and skills fit. · Design for security and compliance by default: Entra ID RBAC models, workspace/item-level permissions, service-principal boundaries, network and data-residency considerations, and Purview-integrated governance. · Define non-functional requirements and design against them — scalability, capacity/CU consumption, refresh SLAs, disaster recovery, and observability/monitoring architecture. · Produce and maintain architecture artifacts: high-level and detailed design documents, data-flow diagrams, decision records (ADRs), and review-board presentations. · Conduct design and code reviews, set engineering standards for the team, and guide build teams through implementation of the approved architecture. · Data Platform Engineering (Microsoft Fabric) · Design and implement medallion architectures (Bronze/Silver/Gold) using Fabric Lakehouses, Data Warehouses, and OneLake with Delta/Parquet as the storage foundation. · Build and orchestrate Fabric Data Pipelines and Notebooks (PySpark/Python) for batch ELT, including metadata-driven/control-table-based orchestration frameworks. · Develop Real-Time Intelligence (RTI) solutions: Eventstreams for streaming ingestion, Eventhouse/KQL databases for hot-path analytics, materialized views, and update policies. · Implement Data Activator (Activator) rules and alerting — threshold, anomaly, and spike detection with actions into Microsoft Teams, email, and downstream systems. · Write performant T-SQL against Fabric Warehouses and SQL analytics endpoints, and KQL against Eventhouses. · Semantic Modeling & BI · Design and maintain enterprise semantic models (star schemas, DAX measures, incremental refresh, RLS) consumed by Power BI reports and dashboards. · Build business-facing reports (consumption analytics, operational monitoring, KPI dashboards) and manage their lifecycle across dev/test/prod workspaces. · Prevent and remediate model/report drift through source-controlled definitions (TMDL/PBIP). · Azure Integration & Migration · Integrate Fabric with the wider Azure estate: Azure Data Factory (including ADF-to-Fabric pipeline migration), Azure Data Lake Storage, Azure Key Vault, Azure Functions, and Azure OpenAI for AI-enriched data workloads. · Migrate legacy ETL frameworks (ADF, SSIS, Synapse) into Fabric-native equivalents with parity validation. · AI & Microsoft Foundry · Build AI-enabled data solutions with Microsoft Foundry (Azure AI Foundry): model catalog and deployments (GPT-4o and other frontier models), prompt flow, and evaluation of model outputs for quality and safety. · Develop agentic and RAG (retrieval-augmented generation) workloads using Foundry Agent Service / Azure OpenAI with enterprise data sources — Fabric lakehouses, Azure AI Search indexes, and vector stores (pgvector, Eventhouse). · Ground AI applications in governed data: connect Foundry projects to OneLake/Fabric data, apply responsible AI practices (content filters, evaluations, red-team findings), and monitor deployed models for cost, latency, and drift. · Operationalize AI workloads alongside data pipelines — embedding generation, enrichment/classification steps in ELT, and LLM-as-judge evaluation harnesses — with the same CI/CD and observability rigor as the rest of the platform. · DevOps, CI/CD & Automation · Implement Git integration for Fabric workspaces and manage promotion via deployment pipelines and the Fabric REST APIs (item CRUD, service-principal automation). · Build headless deployment tooling in Python (e.g., workspace provisioning, Variable Library configuration, notebook/pipeline deployment) and CI/CD in Azure Pipelines or GitHub Actions. · Authenticate and automate using Entra ID service principals, managed identities, and least-privilege RBAC. · Governance, Quality & Operations · Implement data governance with Microsoft Purview: cataloging, lineage, classification, sensitivity labels, and glossary seeding via API/automation. · Establish data-quality checks, reconciliation (hot vs. cold path), monitoring, and runbooks; verify end-to-end flows with automated smoke/E2E tests. · Document architecture, control tables, and operational procedures to an audit-ready standard. Mandatory skill sets: o 5+ years of professional data engineering experience on the Microsoft data stack (Azure and/or Fabric), with at least 1–2 years hands-on with Microsoft Fabric in real project delivery. o Demonstrated architecture and platform design experience: designing end-to-end data platform solutions (workspace/environment topology, medallion layering, hot/cold path patterns, security and RBAC models) and producing design documents and decision records that guide build teams. o Strong programming skills in Python (including PySpark) for data processing and platform automation. o Advanced T-SQL: dimensional modeling, stored procedures, performance tuning, warehouse workload patterns. o Working proficiency in KQL (Kusto) for streaming/telemetry analytics. o Proven experience with lakehouse/medallion architecture, Delta Lake, and Parquet. o Hands-on Power BI / semantic model experience: star schemas, DAX, row-level security, deployment across environments. o Experience with Azure Data Factory or Fabric Data Pipelines for orchestration, including parameterized, metadata-driven designs. o CI/CD experience with Azure DevOps Pipelines or GitHub Actions, plus Git-based source control for data artifacts. o Experience with Entra ID (Azure AD) authentication patterns — service principals, OAuth2/MSAL, managed identities — and secure secret management (Key Vault). o Strong debugging and verification discipline: ability to prove a pipeline works end-to-end, reconcile data across layers, and document evidence. Preferred skill sets: - Azure Databricks and its associated tech stack: Spark on Databricks (batch and Structured Streaming), Delta Lake / Delta Live Tables, Unity Catalog for governance, Databricks Workflows/Jobs, notebooks and Repos, cluster/compute management, MLflow for experiment tracking, and Databricks SQL warehouses. Experience integrating or migrating between Databricks and Fabric (e.g., shared Delta tables via OneLake shortcuts, Unity Catalog mirroring) is especially valuable. - Real-Time Intelligence depth: Eventstream, Eventhouse, update policies, materialized views, Activator alert rules. - Microsoft Purview implementation experience (catalog, lineage, APIs). - Experience calling Fabric REST APIs for headless/programmatic deployment at scale. - Hands-on Microsoft Foundry (Azure AI Foundry) experience: model deployments, prompt flow, Foundry Agent Service, AI evaluations, and integrating Foundry projects with enterprise data platforms. - Exposure to Azure OpenAI / AI-augmented data workloads (RAG, embeddings, pgvector, or GraphRAG-style knowledge graphs) and Azure AI Search. - Full-stack awareness: Next.js/React/TypeScript apps consuming data APIs, MSAL browser auth. - PostgreSQL (pgvector, graph extensions) or other polyglot persistence experience. - Domain experience in energy & utilities (AMI/smart-meter data, grid events, consumption billing) or another regulated industry. - Testing tooling: pytest, Playwright, or similar automated verification frameworks. Certifications (Good to Have) · Microsoft Fabric & Data - DP-600 — Fabric Analytics Engineer Associate - DP-700 — Fabric Data Engineer Associate - DP-203 — Azure Data Engineer Associate - PL-300 — Power BI Data Analyst Associate · Azure Architecture & AI - AZ-305 — Azure Solutions Architect Expert - AI-102 — Azure AI Engineer Associate (Azure OpenAI / Foundry workloads) - AZ-204 — Azure Developer Associate - AZ-400 — DevOps Engineer Expert · Databricks - Databricks Certified Data Engineer Associate / Professional - Databricks Certified Associate Developer for Apache Spark - Databricks Certified Machine Learning Associate · Fabric certifications (DP-600/DP-700) carry the most weight given the platform focus; the rest strengthen the profile but are not screening criteria. Years of experience required: 5 to 7 Years Education qualification: Bachelor's degree in Computer Science, IT, or a related field. Location Kolkata
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
