Job Title: Lead Databricks Data Engineer
Location: Remote
Duration: 6 months+

Purpose:

Build and operate the pipelines that land, conform and curate sources into the lakehouse. Each integration follows the same pattern - ingest to bronze, design, build silver, build gold - so this role adds throughput against a queue of integrations that is currently the binding constraint on coverage.

Key responsibilities:

  • Build ingestion into the bronze layer for assigned sources: gateway and observability logs, productivity tool admin APIs, AI-enabled SaaS usage, hyperscaler billing exports and reference data. Land raw and untransformed, on a scheduled refresh, replayable if the downstream design changes.
  • Work to the shared bronze landing contract so each tool is ingested once and serves both this program and the parallel productivity initiative, rather than being integrated twice.
  • Build the silver layer: typed, deduplicated and conformed to the canonical dimensions, refreshed independently of any downstream publication schedule.
  • Build gold marts carrying attribution method, attribution level, cost basis and provisional status alongside cost and usage.
  • Implement the attribution and allocation logic designed by the analysts, including precedence resolution and ratio-based splitting of shared endpoint cost.
  • Work within Unity Catalog governance - shared bronze and silver, separate gold marts with a recorded owner per dataset - including permissions, lineage and cataloging.
  • Implement data quality rules and monitoring: completeness, freshness and tag-coverage checks with alerting, so pipeline problems surface before they reach a divisional invoice.
  • Manage the volume impact of enabling caller-identity data in the cost and usage report, which multiplies row counts by the number of calling identities per model.
  • Work to the per-source cadence - daily where controls and anomaly detection depend on it, monthly where they do not - within the team's existing CI/CD and promotion practices.

Essential skills and experience:

  • Advanced Databricks engineering: Delta Lake, medallion architecture, Databricks Workflows, Auto Loader and incremental ingestion patterns.
  • Unity Catalog to a governance standard - catalogs, schemas, permissions, lineage - not merely as a place tables happen to live.
  • Strong Python and PySpark, and strong SQL. Notebook-based development.
  • Ingestion from REST APIs including pagination, throttling, incremental watermarks and credential handling, plus cloud object storage across AWS, Azure and GCP.
  • Performance and cost optimization of Spark workloads: partitioning, clustering, file sizing and cluster configuration.
  • CI/CD for Databricks - asset bundles or equivalent - and Git-based development workflow.
  • Able to work to an existing catalog structure and coding standard rather than introducing a parallel approach.

Desirable:

  • Databricks Genie familiarity, including preparing semantic context so natural-language querying returns trustworthy answers.
  • Experience with cloud billing data at volume.
  • Prior work on a shared platform where another team owned adjacent datasets in the same catalog.

About TekNinjas

TekNinjas is a global IT staffing partner placing skilled professionals with leading enterprises and system integrators across the US, Canada, UK, India, and the EU. We focus on the right fit - not just the fast fill - and stay engaged with our consultants well beyond the start date.

We're the right partners in your success.

TekNinjas is an Equal Opportunity Employer.



Lead Databricks Data Engineer

Apply Now
Back to search page