Data Engineer
Location: Charlotte, NC - Hybrid (2 days/week, Tues/Wed)
Duration: 2+ Years
Experience: 7 11 years
Role Overview
Build and optimize scalable cloud-based geospatial data solutions using AWS, PySpark, and Python. The role focuses on large-scale data pipelines, spatial analytics, and map-centric applications supporting transportation, urban planning, and environmental initiatives.
Key Responsibilities
-
Design and maintain scalable geospatial data pipelines on AWS.
-
Develop high-performance PySpark workflows for large-scale spatial datasets.
-
Build reusable Python modules for spatial computations, transformations, and business rules.
-
Configure and optimize AWS storage, compute, and orchestration services.
-
Perform spatial analytics, queries, indexing, and performance tuning for maps, dashboards, and APIs.
-
Implement data quality, validation, and monitoring processes for geospatial data.
-
Integrate external sources such as satellite imagery, open data, and partner feeds.
-
Troubleshoot production pipelines and drive continuous improvements in reliability and performance.
-
Collaborate with data scientists, domain experts, and business stakeholders.
Top 3 Required Skills
-
Python Strong programming and geospatial development
-
PySpark Large-scale distributed data processing & optimization
-
AWS Production-grade cloud data engineering
Preferred Experience
-
Strong understanding of geospatial concepts including projections, vector/raster data, spatial joins, topology, and spatial indexing.
-
Experience building production-grade data pipelines and cloud solutions.
-
Strong troubleshooting, communication, and technical documentation skills.
Ideal Profile: 7 11 year hands-on Data Engineer with strong Python, PySpark, and AWS expertise, ideally with experience in
geospatial/location intelligence data and spatial analytics.