Rajat Mehta
Data Engineer
Data Engineer
I'm a Data Engineer at Accenture, currently working as an Advanced App Engineering Analyst with 1+ years of experience building production-grade data pipelines. My work spans the modern data stack — SQL, PySpark, and Databricks for large-scale data processing, AWS for cloud infrastructure, and Airflow and DBT for orchestration and transformation — handling datasets with millions of rows. I also bring experience with Snowflake for cloud data warehousing and Kafka for real-time data streaming. I enjoy solving problems at the intersection of scale, reliability, and performance in data engineering.
Contributed to a production data pipeline for a healthcare client, working across ingestion, transformation, and orchestration to help deliver the client's first consolidated Customer Data Platform (CDP) as part of a cross-functional data engineering team
Partnered on ingesting and mapping 10 plus fragmented source systems into a Unified Customer Profile using Databricks and PySpark, consolidating 900,000 plus rows of siloed healthcare data into a single queryable data model
Built PySpark transformation logic for data cleansing, deduplication, and restructuring within a multi-layer pipeline, improving consistency and usability of unified customer data for downstream analytics
Applied advanced SQL (CTEs, window functions) to develop Calculated Insights used for business reporting and customer segmentation
Collaborated on pipeline orchestration and scheduling in Databricks to support reliable, repeatable data refreshes across the team-owned pipeline
Validated pipeline accuracy and data quality with QA stakeholders using Agile and Jira sprint methodology; participated in SDLC processes including software validation and test execution
Technologies / Skills Used : Databricks, PySpark, SQL, ETL, Data Unification, Customer Data Platform, Data Quality, Agile, Jira, SDLC, QA Testing
Databases : MySQL, PostgreSQL, AWS Redshift, BigQuery, MongoDB, SQL Server
Frameworks & Libraries : Apache Airflow, Docker, Node.js, React, Express.js, dbt, Power BI
Programming Languages : Python, SQL (Advanced), C++, JavaScript, Shell Scripting
Soft Skills : Agile / Scrum, Jira, Stakeholder Management, Technical Documentation, Risk & Controls
Tools & Platforms : AWS (S3, EMR, Redshift, Glue, Lambda, Athena, EC2, Step Functions), GCP (BigQuery, GCS), Databricks, Snowflake, Azure Data Factory