Sr. Azure Data Engineer | Accenture | Serving Notice | Immediate Joiner · Bengaluru, Karnataka, India
Senior Azure Data Engineer passionate about designing and modernizing enterprise-scale Azure data platforms that transform raw data into trusted, business-ready insights. My expertise includes Azure Synapse Analytics, Azure Data Factory, Azure Databricks, Spark, Delta Lake, and modern ELT architectures, with experience building scalable, high-performance solutions across complex enterprise ecosystems.
My work spans cloud migration, data warehousing, metadata-driven ELT pipelines, data modeling, performance optimization, data governance, CI/CD collaboration, and production support, with a strong focus on delivering reliable, secure, and scalable analytics platforms. More recently, I've been exploring the intersection of Data Engineering and AI, building solutions using RAG, Databricks Vector Search, and Agentic AI to automate troubleshooting, enhance knowledge discovery, and improve engineering productivity.
I enjoy solving complex engineering problems, learning emerging technologies, and taking ownership from design through production. I'm always looking for opportunities to build scalable solutions, optimize performance, and create meaningful business impact through data.
Outside of work, you'll usually find me watching movies (especially horror), reading a good book or spending time with dogs.
Contact details — on requestProof of Work — on request
Experience
Data Engineering Senior Analyst
Accenture · Nov 2023 – Present
Not yet confirmed
Engineered Azure-based enterprise data platforms using Azure Data Factory, Azure Databricks, Azure Synapse Analytics, ADLS Gen2, PySpark, Spark SQL and Delta Lake, supporting large-scale retail domains including Customer, Stores, Supply Chain, Pricing, Inventory and Master Data.
Designed and implemented end-to-end ELT pipelines by integrating 12+ heterogeneous data sources (databases, APIs, flat files and enterprise applications) into a centralized Unity Catalog-governed Delta Lake platform, utilizing Time Travel to support data versioning and ensure secure, high-quality datasets for analytics and reporting.
Optimized a 200+ TB enterprise data platform containing structured and semi-structured data (Delta, Parquet, CSV, JSON) using Spark and Delta Lake techniques including partitioning, caching, broadcast joins, Deletion Vectors, OPTIMIZE, Z-Ordering, Auto Compaction, Schema Evolution, indexing and query tuning, improving pipeline performance by 50% and query performance by 30%.
Designed scalable enterprise data warehouse solutions using Azure Synapse Analytics, implementing dimensional modeling, incremental loading, Serverless SQL Pools and optimized storage strategies to improve scalability while reducing cloud storage and compute costs.
Owned and supported 60+ production data pipelines, implementing monitoring, SLA tracking and alerting using Azure Monitor and Log Analytics, ensuring high platform availability, proactive incident detection and faster root cause resolution.
Led cloud migration initiatives from legacy platforms to Azure, including schema design, historical backfilling, partitioning, data validation and system decommissioning, modernizing enterprise reporting and analytics platforms.
Delivered optimized PySpark/T-SQL transformations, stored procedures, views, functions, indexing strategies and performance tuning to improve data processing efficiency and support high-volume analytical workloads.
Managed Git-based source control, branching, pull requests and Azure Data Factory CI/CD processes, while collaborating with the Azure DevOps team to enable automated deployments.
Collaborated with BI, Data Science, Strategy, Product and Business teams in Agile/Scrum (JIRA/Confluence) to deliver trusted datasets, improve dashboard performance and enable self-service analytics.
Designed and developed a Databricks-based Agentic AI application using RAG to help developers troubleshoot pipeline failures, perform RCA and lineage tracing across pipelines/code/logs - reducing debugging time by 70%.
Architected a hybrid AI solution integrating Databricks Vector Search with live SQL execution, automating enterprise knowledge ingestion, indexing and semantic retrieval while enabling non-technical users to query enterprise data and engineering knowledge through natural language.
Led a team of 5 Data Engineers, providing technical guidance, code reviews, production support and delivery planning across multiple data platform initiatives.
TechStar
Accenture · Dec 2022 – Present
Not yet confirmed
Recognized in the Accenture NA Market with this prestigious award given to top 0.01% of the high performing technology consultants and additional rewards worth 450k INR.
Data Engineering Analyst
Accenture · Oct 2021 – Oct 2023
Not yet confirmed
Engineered Azure-based enterprise data platforms using Azure Data Factory, Azure Databricks, Azure Synapse Analytics, ADLS Gen2, PySpark, Spark SQL and Delta Lake, supporting large-scale retail domains including Customer, Stores, Supply Chain, Pricing, Inventory and Master Data.
Designed and implemented end-to-end ELT pipelines by integrating 12+ heterogeneous data sources (databases, APIs, flat files and enterprise applications) into a centralized Unity Catalog-governed Delta Lake platform, utilizing Time Travel to support data versioning and ensure secure, high-quality datasets for analytics and reporting.
Optimized a 200+ TB enterprise data platform containing structured and semi-structured data (Delta, Parquet, CSV, JSON) using Spark and Delta Lake techniques including partitioning, caching, broadcast joins, Deletion Vectors, OPTIMIZE, Z-Ordering, Auto Compaction, Schema Evolution, indexing and query tuning, improving pipeline performance by 50% and query performance by 30%.
Designed scalable enterprise data warehouse solutions using Azure Synapse Analytics, implementing dimensional modeling, incremental loading, Serverless SQL Pools and optimized storage strategies to improve scalability while reducing cloud storage and compute costs.
Owned and supported 60+ production data pipelines, implementing monitoring, SLA tracking and alerting using Azure Monitor and Log Analytics, ensuring high platform availability, proactive incident detection and faster root cause resolution.
Led cloud migration initiatives from legacy platforms to Azure, including schema design, historical backfilling, partitioning, data validation and system decommissioning, modernizing enterprise reporting and analytics platforms.
Delivered optimized PySpark/T-SQL transformations, stored procedures, views, functions, indexing strategies and performance tuning to improve data processing efficiency and support high-volume analytical workloads.
Managed Git-based source control, branching, pull requests and Azure Data Factory CI/CD processes, while collaborating with the Azure DevOps team to enable automated deployments.
Collaborated with BI, Data Science, Strategy, Product and Business teams in Agile/Scrum (JIRA/Confluence) to deliver trusted datasets, improve dashboard performance and enable self-service analytics.
Designed and developed a Databricks-based Agentic AI application using RAG to help developers troubleshoot pipeline failures, perform RCA and lineage tracing across pipelines/code/logs - reducing debugging time by 70%.
Architected a hybrid AI solution integrating Databricks Vector Search with live SQL execution, automating enterprise knowledge ingestion, indexing and semantic retrieval while enabling non-technical users to query enterprise data and engineering knowledge through natural language.
Led a team of 5 Data Engineers, providing technical guidance, code reviews, production support and delivery planning across multiple data platform initiatives.
Data Engineering Associate
Accenture · Oct 2019 – Oct 2021
Not yet confirmed
Engineered Azure-based enterprise data platforms using Azure Data Factory, Azure Databricks, Azure Synapse Analytics, ADLS Gen2, PySpark, Spark SQL and Delta Lake, supporting large-scale retail domains including Customer, Stores, Supply Chain, Pricing, Inventory and Master Data.
Designed and implemented end-to-end ELT pipelines by integrating 12+ heterogeneous data sources (databases, APIs, flat files and enterprise applications) into a centralized Unity Catalog-governed Delta Lake platform, utilizing Time Travel to support data versioning and ensure secure, high-quality datasets for analytics and reporting.
Optimized a 200+ TB enterprise data platform containing structured and semi-structured data (Delta, Parquet, CSV, JSON) using Spark and Delta Lake techniques including partitioning, caching, broadcast joins, Deletion Vectors, OPTIMIZE, Z-Ordering, Auto Compaction, Schema Evolution, indexing and query tuning, improving pipeline performance by 50% and query performance by 30%.
Designed scalable enterprise data warehouse solutions using Azure Synapse Analytics, implementing dimensional modeling, incremental loading, Serverless SQL Pools and optimized storage strategies to improve scalability while reducing cloud storage and compute costs.
Owned and supported 60+ production data pipelines, implementing monitoring, SLA tracking and alerting using Azure Monitor and Log Analytics, ensuring high platform availability, proactive incident detection and faster root cause resolution.
Led cloud migration initiatives from legacy platforms to Azure, including schema design, historical backfilling, partitioning, data validation and system decommissioning, modernizing enterprise reporting and analytics platforms.
Delivered optimized PySpark/T-SQL transformations, stored procedures, views, functions, indexing strategies and performance tuning to improve data processing efficiency and support high-volume analytical workloads.
Managed Git-based source control, branching, pull requests and Azure Data Factory CI/CD processes, while collaborating with the Azure DevOps team to enable automated deployments.
Collaborated with BI, Data Science, Strategy, Product and Business teams in Agile/Scrum (JIRA/Confluence) to deliver trusted datasets, improve dashboard performance and enable self-service analytics.
Designed and developed a Databricks-based Agentic AI application using RAG to help developers troubleshoot pipeline failures, perform RCA and lineage tracing across pipelines/code/logs - reducing debugging time by 70%.
Architected a hybrid AI solution integrating Databricks Vector Search with live SQL execution, automating enterprise knowledge ingestion, indexing and semantic retrieval while enabling non-technical users to query enterprise data and engineering knowledge through natural language.
Led a team of 5 Data Engineers, providing technical guidance, code reviews, production support and delivery planning across multiple data platform initiatives.
U
Business Analyst (Marketing & Strategy)
UCIM · Apr 2019 – Sep 2019
Not yet confirmed
Created Excel pivot charts and Tableau dashboards to present project data and KPI metrics to VCs and investors, enhancing data visibility.
Collaborated with startups at UCIM to support fundraising and marketing initiatives, strengthening their market presence.
Business Analyst
Exeliq Consulting Inc. · Apr 2019 – Sep 2019
Not yet confirmed
Intern
Almora · Jan 2019 – Mar 2019
Not yet confirmed
Research Intern
Airport Authority of India, Udaipur · Oct 2018 – Dec 2018
Not yet confirmed
Collaborated with a team of highly skilled Ph.
D.
Researchers.
Thereafter research was conducted and published the first research paper on Passenger Screening Algorithm, as part of my major project using Kaggle dataset.
SDE Intern
DigiCommSols Private Limited · Apr 2018 – Aug 2018
Not yet confirmed
Operations Intern
ConneXTech · Oct 2017 – Mar 2018
Not yet confirmed
Research Intern
Delhi Metro Rail Corporation Ltd · Jun 2017 – Jul 2017
Not yet confirmed
Did a multivariate analysis of 10000+ values.
Spatial Data Mining for Noida City Metro Line Expansion.
Teaching Assistant
Coding Blocks · Jan 2017 – Apr 2017
Not yet confirmed
Skills 0 proven through work
Also works with
Data StagingData OrganizationData VisualizationTableauMicrosoft ExcelData ValidationTaking OwnershipReal-Time Dashboard DevelopmentMedallion Architecture DesignTechnical TroubleshootingWork Planning and PrioritizationTechnical Input ProvisionTeam LeadershipData ShufflingDatabase Query OptimizationAzure Data Lake UsageObject Storage ManagementStar Schema DesignData TransformationTabular Data Cleaning & ValidationProcess ImprovementData MonitoringAzure Synapse Analytics UsageAzure Data Factory Pipeline DevelopmentDatabase IntegrationETL Pipeline Development
Proof of Work
Proof of Work
Sarthak shares this with people who ask. You'll hear back either way.
Education
Executive Programme in Strategy and Management, Sponsored by Accenture
Indian Institute of Management, Calcutta
BTech Information Technology
Maharaja Agrasen Institute Of Technology, Delhi · 2015 — 2019