Position : Data Engineer with AWS & Python
Location : Mclean, VA (Onsite)
Term : W2 Role
Job Summary:
We are looking for an experienced Data Engineer with strong AWS and Python expertise to design, develop, and maintain scalable data pipelines and cloud-based data solutions. The ideal candidate will have hands-on experience with AWS data services, Python programming, ETL/ELT development, data integration, and modern data engineering practices.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT workflows using Python and AWS services.
- Build and optimize data ingestion and transformation pipelines for structured and unstructured data.
- Develop reusable Python scripts and applications for data processing, automation, and integration.
- Work with AWS services such as S3, Glue, Lambda, Redshift, EMR, Athena, Kinesis, and CloudWatch.
- Implement data processing solutions using PySpark and distributed computing frameworks.
- Develop data models and optimize data storage and retrieval processes.
- Perform data quality checks, validation, reconciliation, and error handling.
- Optimize data pipelines for performance, scalability, reliability, and cost efficiency.
- Integrate data from APIs, databases, files, and other enterprise data sources.
- Implement monitoring, logging, and alerting for data pipelines.
- Collaborate with Data Scientists, Data Analysts, Architects, and business stakeholders.
- Follow best practices for CI/CD, version control, testing, security, and documentation.
- Troubleshoot production data issues and provide timely resolution.
Required Skills / Must Have
- 5+ years of experience in Data Engineering or a related field.
- Strong hands-on experience with Python for data engineering and automation.
- Strong experience with AWS cloud services, particularly:
- Amazon S3
- AWS Glue
- AWS Lambda
- Amazon Redshift
- Amazon Athena
- Strong knowledge of SQL and relational databases.
- Experience developing ETL/ELT data pipelines.
- Experience with PySpark/Spark and distributed data processing.
- Strong understanding of data warehousing and data lake concepts.
- Experience with Git and CI/CD pipelines.
- Strong troubleshooting and analytical skills.
Preferred / Nice to Have
- Experience with AWS EMR, Kinesis, Step Functions, or MWAA/Airflow.
- Experience with Terraform or CloudFormation.
- Knowledge of Databricks, Snowflake, or Delta Lake.
- Experience working with streaming data pipelines.
- Knowledge of Docker/Kubernetes.
- Experience with REST APIs and third-party data integrations.
- Knowledge of AWS IAM, encryption, and cloud security best practices.
- Experience with Agile/Scrum methodology.
Education
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Equivalent professional experience may be considered.
Cloud BC Labs Inc is a digital transformation organization aimed at creating seamless solutions for clients to effectively manage their business operations. The company specializes in Business and Management Consulting, AI/ML, Data Analytics & Visualization, Cloud Data Warehouse Migration, Snowflake Implementation, Informatica Implementation & Upgrade, Staffing Services and Data Management Solutions