StreamSets Developer
Location: Remote
Job Type: Long Term Contract
Position Overview
We are looking for an experienced StreamSets Developer to design, develop, implement, and support enterprise-scale data integration and data ingestion pipelines. The ideal candidate will have strong hands-on experience with StreamSets Data Collector, ETL/ELT, real-time data integration, APIs, databases, cloud platforms, and data engineering technologies.
This is a Remote opportunity.
Key Responsibilities
Design, develop, configure, and maintain data ingestion pipelines using StreamSets.
Build scalable ETL/ELT pipelines for batch and real-time data integration.
Develop and maintain StreamSets Data Collector (SDC) pipelines and data flows.
Integrate data from multiple sources including databases, APIs, files, cloud platforms, and enterprise applications.
Configure StreamSets origins, processors, destinations, and pipeline parameters.
Develop data transformation, cleansing, enrichment, validation, and routing logic.
Monitor pipeline performance and troubleshoot data processing and integration issues.
Implement error handling, data quality checks, logging, and recovery mechanisms.
Work with structured and semi-structured data including JSON, XML, CSV, Avro, and Parquet.
Integrate StreamSets with relational and NoSQL databases.
Develop integrations with cloud-based data platforms and data warehouses.
Optimize pipelines for performance, scalability, reliability, and throughput.
Required Skills
Strong hands-on experience with StreamSets Data Collector (SDC).
Experience designing and developing StreamSets pipelines.
Strong ETL/ELT and data integration experience.
Strong SQL skills.
Experience with relational databases such as Oracle, SQL Server, PostgreSQL, or MySQL.
Experience working with APIs and web services.
Strong understanding of REST APIs.
Experience with JSON, XML, CSV, Avro, and Parquet.
Preferred Skills
Experience with StreamSets Transformer and StreamSets Control Hub.
Experience with cloud platforms such as AWS, Azure, or GCP.
Experience with cloud data platforms such as Snowflake, Databricks, or BigQuery.
Experience with Kafka and event-driven data integration.
Experience with Hadoop/HDFS or other distributed data platforms.
Experience with Python or Java.
Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab, or Azure DevOps.
Experience with Docker/Kubernetes is a plus.
Experience working with data warehouses and data lakes.
Knowledge of data governance, metadata management, and data quality.
Qualifications
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred.
5+ years of experience in data engineering, ETL, data integration, or related technologies.
Strong StreamSets development experience is required.
Excellent analytical, troubleshooting, communication, and problem-solving skills.
By continuing you agree to our Terms & Privacy Policy.