Technology & Tools: Python, AWS Glue, Amazon S3, Athena, CloudFormation, Apache Spark,
Spark SQL, Azure, Azure Data Factory, ADLS Gen2, Azure Databricks, Azure Synapse Analytics,
Airflow, Amazon Redshift, PySpark, MySQL
Role: Data Engineer
• Developed ETL processes using AWS Glue, PySpark, Apache Spark, and Spark SQL for
large-scale publishing and product data processing.
• Managed Data Lake storage in S3 and analytical workloads in Athena and Redshift,
ensuring data quality and consistency across pipeline stages.
• Implemented orchestration using Apache Airflow and cloud-native services and
optimized transformations for distributed data processing.
• Worked with MySQL and performed statistical analysis on inventory and product data.
• Used Azure Data Factory, ADLS Gen2, Azure Databricks, and Azure Synapse Analytics
for ingestion, storage, processing, transformation, and analytics.