Job Title: Lead Databricks Data Engineer
Location: Remote
Duration: 6 months+
Purpose:
Build and operate the pipelines that land, conform and curate sources into the lakehouse. Each integration follows the same pattern - ingest to bronze, design, build silver, build gold - so this role adds throughput against a queue of integrations that is currently the binding constraint on coverage.
Key responsibilities:
- Build ingestion into the bronze layer for assigned sources: gateway and observability logs, productivity tool admin APIs, AI-enabled SaaS usage, hyperscaler billing exports and reference data. Land raw and untransformed, on a scheduled refresh, replayable if the downstream design changes.
- Work to the shared bronze landing contract so each tool is ingested once and serves both this program and the parallel productivity initiative, rather than being integrated twice.
- Build the silver layer: typed, deduplicated and conformed to the canonical dimensions, refreshed independently of any downstream publication schedule.
- Build gold marts carrying attribution method, attribution level, cost basis and provisional status alongside cost and usage.
- Implement the attribution and allocation logic designed by the analysts, including precedence resolution and ratio-based splitting of shared endpoint cost.
- Work within Unity Catalog governance - shared bronze and silver, separate gold marts with a recorded owner per dataset - including permissions, lineage and cataloging.
- Implement data quality rules and monitoring: completeness, freshness and tag-coverage checks with alerting, so pipeline problems surface before they reach a divisional invoice.
- Manage the volume impact of enabling caller-identity data in the cost and usage report, which multiplies row counts by the number of calling identities per model.
- Work to the per-source cadence - daily where controls and anomaly detection depend on it, monthly where they do not - within the team's existing CI/CD and promotion practices.
Essential skills and experience:
- Advanced Databricks engineering: Delta Lake, medallion architecture, Databricks Workflows, Auto Loader and incremental ingestion patterns.
- Unity Catalog to a governance standard - catalogs, schemas, permissions, lineage - not merely as a place tables happen to live.
- Strong Python and PySpark, and strong SQL. Notebook-based development.
- Ingestion from REST APIs including pagination, throttling, incremental watermarks and credential handling, plus cloud object storage across AWS, Azure and GCP.
- Performance and cost optimization of Spark workloads: partitioning, clustering, file sizing and cluster configuration.
- CI/CD for Databricks - asset bundles or equivalent - and Git-based development workflow.
- Able to work to an existing catalog structure and coding standard rather than introducing a parallel approach.
Desirable:
- Databricks Genie familiarity, including preparing semantic context so natural-language querying returns trustworthy answers.
- Experience with cloud billing data at volume.
- Prior work on a shared platform where another team owned adjacent datasets in the same catalog.
About TekNinjas
TekNinjas is a global IT staffing partner placing skilled professionals with leading enterprises and system integrators across the US, Canada, UK, India, and the EU. We focus on the right fit - not just the fast fill - and stay engaged with our consultants well beyond the start date.
We're the right partners in your success.
TekNinjas is an Equal Opportunity Employer.