Responsibilities

  • Architect for RAG: Design and scale pipelines for Retrieval-Augmented Generation (RAG), transforming large volumes of unstructured IT logs and documentation into optimized vector embeddings.
  • Scale vector infrastructure: Oversee the health and performance of vector databases (Pinecone, Milvus, Weaviate), ensuring sub-second retrieval speeds for agentic reasoning loops.
  • Engineer semantic layers: Build knowledge graphs and semantic layers beyond simple ETL to provide agents with the necessary context for navigating complex infrastructure puzzles.
  • Automate data excellence: Build automated guardrails to detect noise, bias, or PII before it reaches the model.
  • Bridge raw, messy data sources and deep technical AI work, identifying and resolving quality issues at the source.
  • Progress to production: Build, deploy, and maintain CI/CD pipelines for data infrastructure, ensuring that the context window remains fresh and reliable.

Requirements

  • Expertise in data mining, data storage, and ETL processes.
  • Experience in data pipelines development and tooling (Glue, Databricks, Synapse, Dataproc).
  • Experience with relational and NoSQL databases (PostgreSQL, DB2, MongoDB).
  • Excellent problem‑solving, analytical, and critical thinking skills.
  • Ability to manage multiple projects simultaneously while maintaining a high level of attention to detail.
  • Ability to communicate with both technical and non‑technical colleagues, translating technical requirements from business needs.
  • Experience as a Data Engineer and/or in cloud modernization (preferred).
  • Experience in data modelling to create conceptual models of how data connects and is used in business processes (preferred).
  • Professional certification (e.g., Open Certified Technical Specialist with Data Engineering Specialization) (preferred).
  • Cloud platform certification (e.g., AWS Certified Data Analytics – Specialty, Elastic Certified Engineer, Google Cloud Professional Data Engineer, Microsoft Certified: Azure Data Engineer Associate) (preferred).
  • Understanding of social coding and integrated development environments (GitHub, Visual Studio) (preferred).
  • Degree in a scientific discipline (Computer Science, Software Engineering, Information Technology) (preferred).

Hard Skills

  • Data Mining
  • ETL Processes
  • Data Pipeline Development
  • Data Modeling
  • Relational Databases
  • NoSQL Databases
  • Vector Embeddings
  • Automated Data Guardrails
  • CI/CD Pipelines
  • Knowledge Graphs

Soft Skills

  • Problem‑Solving
  • Analytical Thinking
  • Attention to Detail
  • Communication Skills
  • Project Management

Certifications & Qualifications

  • Open Certified Technical Specialist with Data Engineering Specialization
  • AWS Certified Data Analytics – Specialty
  • Elastic Certified Engineer
  • Google Cloud Professional Data Engineer
  • Microsoft Certified: Azure Data Engineer Associate
#J-18808-Ljbffr
Similar jobs

Data Engineer

Apply Now
Back to search page