Architect for RAG: Design and scale pipelines for Retrieval-Augmented Generation (RAG), transforming large volumes of unstructured IT logs and documentation into optimized vector embeddings.
Scale vector infrastructure: Oversee the health and performance of vector databases (Pinecone, Milvus, Weaviate), ensuring sub-second retrieval speeds for agentic reasoning loops.
Engineer semantic layers: Build knowledge graphs and semantic layers beyond simple ETL to provide agents with the necessary context for navigating complex infrastructure puzzles.
Automate data excellence: Build automated guardrails to detect noise, bias, or PII before it reaches the model.
Bridge raw, messy data sources and deep technical AI work, identifying and resolving quality issues at the source.
Progress to production: Build, deploy, and maintain CI/CD pipelines for data infrastructure, ensuring that the context window remains fresh and reliable.
Requirements
Expertise in data mining, data storage, and ETL processes.
Experience in data pipelines development and tooling (Glue, Databricks, Synapse, Dataproc).
Experience with relational and NoSQL databases (PostgreSQL, DB2, MongoDB).
Excellent problem‑solving, analytical, and critical thinking skills.
Ability to manage multiple projects simultaneously while maintaining a high level of attention to detail.
Ability to communicate with both technical and non‑technical colleagues, translating technical requirements from business needs.
Experience as a Data Engineer and/or in cloud modernization (preferred).
Experience in data modelling to create conceptual models of how data connects and is used in business processes (preferred).
Professional certification (e.g., Open Certified Technical Specialist with Data Engineering Specialization) (preferred).
Cloud platform certification (e.g., AWS Certified Data Analytics – Specialty, Elastic Certified Engineer, Google Cloud Professional Data Engineer, Microsoft Certified: Azure Data Engineer Associate) (preferred).
Understanding of social coding and integrated development environments (GitHub, Visual Studio) (preferred).
Degree in a scientific discipline (Computer Science, Software Engineering, Information Technology) (preferred).
Hard Skills
Data Mining
ETL Processes
Data Pipeline Development
Data Modeling
Relational Databases
NoSQL Databases
Vector Embeddings
Automated Data Guardrails
CI/CD Pipelines
Knowledge Graphs
Soft Skills
Problem‑Solving
Analytical Thinking
Attention to Detail
Communication Skills
Project Management
Certifications & Qualifications
Open Certified Technical Specialist with Data Engineering Specialization
AWS Certified Data Analytics – Specialty
Elastic Certified Engineer
Google Cloud Professional Data Engineer
Microsoft Certified: Azure Data Engineer Associate