Job Title:
PySpark Developer Location:
Chennai / Bangalore / Hyderabad / Pune Notice Period:
Immediate to 30 Days
Job Description : We are seeking a skilled
PySpark Developer
with strong experience in Python, PySpark, SQL, and Data Warehousing concepts. The ideal candidate will be responsible for designing, developing, and optimizing large-scale data processing pipelines and ETL solutions using Spark-based technologies.
Key Responsibilities : Design, develop, and maintain scalable ETL/ELT pipelines using PySpark. Build and optimize Spark jobs for performance, reliability, and scalability. Process and transform large datasets using Spark SQL, DataFrames, and RDDs. Develop batch and real-time data processing solutions. Integrate data pipelines with Hive, HDFS, Snowflake, Redshift, and other data platforms. Collaborate with data engineers, analysts, and business stakeholders. Monitor data pipelines, troubleshoot issues, and ensure SLA compliance. Follow coding best practices, version control, and CI/CD processes. Work with Hadoop ecosystem tools and cloud platforms such as AWS, Azure, or GCP.
Required Skills: Strong hands-on experience in
Python and PySpark Expertise in
Spark SQL, DataFrames, and RDDs Good knowledge of
Hadoop (Hive, HDFS, YARN) Strong
SQL
and query optimization skills Experience with
Data Warehousing concepts Knowledge of
Parquet, Avro, JSON
data formats Experience with
Git
version control Familiarity with
Airflow, Oozie, or similar scheduling tools Exposure to
AWS, Azure, or GCP
is an added advantage