Strong hands-on experience in Python for data engineering and application development. Extensive experience with AWS cloud services, including S3, EMR, Glue, Lambda, IAM, EC2, ECS/EKS, CloudWatch, and Redshift. Strong expertise in Apache Spark for large-scale batch data processing. Hands-on experience with Apache Flink for real-time stream processing and event-driven data pipelines. Experience designing and implementing batch and streaming data architectures. Strong knowledge of data modeling, data warehousing, and data lake/lakehouse concepts. Experience with ETL/ELT frameworks and data integration. Strong SQL skills and experience with relational and NoSQL databases. Experience with Apache Kafka or similar messaging/event streaming platformsStrong understanding of distributed computing and big data technologies. Experience with Docker, Kubernetes, and CI/CD pipelinesHands-on experience with Git and Agile development methodologies. Design and implement scalable, secure, and high-performance data architecture solutions on AWS. Build and optimize batch processing pipelines using Apache Spark. Develop real-time streaming data solutions using Apache Flink. Design end-to-end data ingestion, transformation, and processing pipelines. Define data models, governance standards, and architectural best practices. Python, Apache Spark, AWS, Apache Flink, batch and streaming data architectures, ETL/ELT, Apache Kafka, Docker, Kubernetes, CI/CD pipelines.
By continuing you agree to our Terms & Privacy Policy.