Job Description:
Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark.
Develop PySpark applications for processing and transforming large volumes of structured and semi-structured data.
Write efficient Python scripts and reusable functions for data processing, validation, and automation.
Develop complex SQL queries for data extraction, transformation, validation, and analysis.
Work with Apache Spark, Spark SQL, DataFrames, Hive, and Hadoop for distributed data processing.
Perform data cleansing, transformation, validation, reconciliation, and quality checks.
Optimize Spark/PySpark jobs using techniques such as partitioning, caching, repartitioning, and broadcast joins.
Develop and maintain data pipelines for ingestion from multiple source systems.
Implement error handling, retry mechanisms, and monitoring to ensure reliable data pipelines.
By continuing you agree to our Terms & Privacy Policy.