We are searching for an experienced Data Ops/Data Engineer at our Grayston Drive Sandton facility.

Primary Duties And Responsibilities Job Purpose

We are looking for an autonomous, logic-driven Data Engineer who wants to step up, take ownership of a massive dataset, and bridge the gap between data pipelines and server environments. Your core focus will be engineering robust ETL pipelines using Apache Airflow and Python. Simultaneously, you will take ownership of our local and lab LAMP/CPanel server environments. You do not need to be a veteran Sys Admin on day one, you will have the backing of our corporate Blue Label Telecoms Group IT support team for deep infrastructure emergencies. However, you must be a fearless, self-driven problem solver who is excited to dive into server logs, optimize large databases, and build the future data foundations for our Machine Learning platform.

Key Responsibilities Data Engineering & Pipeline Orchestration (Core Focus) Design, build, and maintain automated ETL pipelines using Python and Apache Airflow. Ingest, aggregate, and reconcile massive datasets (100 M+ rows monthly) from telecom operator logs, matching delivery rates against financial campaign reporting. Architect and automate the monthly update of our "Golden Customer Record," integrating opt-in and multi-channel customer engagement data. Build and optimize data pipelines to feed clean, structured data into Open Data Accelerators for our machine Learning platform. Database Administration & Infrastructure Ownership Manage, monitor, and optimize our LAMP stack (Linux, Apache, My SQL/Maria DB, PHP) and CPanel environments across lab/cloud and local setups. Implement aggressive database partitioning, indexing, and tuning to ensure high-volume imports do not impact active production systems. Design and execute automated data archiving strategies to keep operational databases fast while maintaining strict historical records for compliance. Work alongside Group IT support to monitor server uptime, resource allocation, and ensure hardware/ infrastructure stability. Compliance & Governance Ensure all data pipelines and storage architecture strictly adhere to POPIA compliance and data. Governance standards, safely handling and isolating Personal Identifiable Information (PII). Technical Skills Familiarity with Linux command-line operations. Basic server administration skills. Experience with CPanel management. Exposure to PHP development. Exposure to shell scripting (Bash). Experience working with telecommunications data. Experience with high-volume marketing platforms. Experience with financial reconciliation systems. Ability to manage and process high-volume datasets. Computer Science degree, diploma, or equivalent practical experience in data-related environments. Behavioral Skills Demonstrated ability to handle large and complex datasets effectively. Strong analytical and problem-solving capability. Attention to detail when working with data and reconciliation processes. Ability to learn and adapt to new technologies and systems. Self-driven approach supported by a proven track record of delivering results in data-intensive environments Education & Experience Education: 3 to 5 years of experience in Data Engineering, Database Administration, or Backend Software Engineering. Strong proficiency in Python and advanced, highly optimized SQL (writing raw, efficient queries for large datasets are non-negotiable). Practical experience building and maintaining Directed Acyclic Graphs (DAGs) in Apache Airflow. Deep understanding of relational database design, indexing strategies, and table partitioning.

Our company provides equal employment opportunities (EE) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability or genetics. #J-18808-Ljbffr

Similar jobs

Data ops/ data engineer

Apply Now
Back to search page