Primary Skill: PySpark, Hive, Python, SQL
Secondary: Unix, Kafka, Java
Roles & Responsibilities:
We are seeking a highly experienced Hadoop Spark Developer with 10+ years of expertise in Big Data technologies, including PySpark, Hadoop Ecosystem, Hive, and Python.
The ideal candidate will be responsible for designing, developing, optimizing, and maintaining large-scale data processing solutions.
Experience with Microsoft Copilot for AI-assisted development and productivity enhancement is highly desirable. The developer should hold a bachelor's or master's degree.
The candidate should possess strong analytical skills, hands-on experience in distributed data processing, and the ability to work closely with business stakeholders, architects, and data engineering teams.
  • Design, develop, and maintain scalable Big Data solutions using Hadoop and Spark.
  • Build and optimize ETL/ELT pipelines using PySpark, Hive, and Python.
  • Process and analyze large datasets in distributed environments.
  • Develop high-performance Spark jobs and optimize existing workloads.
  • Create and manage Hive tables, partitions, views, and complex queries.
  • Implement data quality, data validation, and reconciliation frameworks.
  • Perform code reviews and ensure adherence to coding standards and best practices.
  • Utilize Microsoft Copilot to accelerate development, automate code generation, troubleshooting, documentation, and testing activities.
  • Strong experience building both batch and real-time streaming applications with Kafka.
  • Collaborate with Data Architects, Data Scientists, Business Analysts, and DevOps teams.
  • Troubleshoot production issues and perform root cause analysis.
  • Design data ingestion frameworks for structured, semi-structured, and unstructured data.
  • Participate in Agile ceremonies including sprint planning, estimation, and retrospectives.
  • Mentor junior developers and provide technical leadership.
Similar jobs

Hadoop Developer with Java

Apply Now
Back to search page