Job Description The Senior Data Engineer will lead the end-to-end migration of complex structured finance datasets from a legacy proprietary platform to Databricks . This is a highly independent, architecture-focused role. The engineer will own data modeling, pipeline architecture, migration strategy, data validation, performance optimization, and production cutover. The source platform contains significant business logic within a large, sparsely documented legacy codebase. The ideal candidate must be able to reverse-engineer legacy code, recover undocumented business rules, establish financial data parity, and build scalable production data solutions in Databricks .
Responsibilities - Lead migration of structured finance datasets from legacy platforms to Databricks.
- Reverse-engineer undocumented business logic from legacy SQL, stored procedures, C++, Java, or custom code.
- Design scalable data models for complex financial instruments and time-series data.
- Implement point-in-time processing, as-of joins, bitemporal history, late-arriving data, and corporate action adjustments.
- Build scalable ingestion, transformation, and publishing pipelines using Databricks, Spark, PySpark, SQL, and Delta Lake .
- Establish data quality, reconciliation, parity validation, lineage, and audit frameworks.
- Optimize Databricks/Spark workloads for performance and cloud cost.
- Apply partitioning, clustering, Z-Ordering, file compaction, and query optimization techniques.
- Work with quantitative, risk, product, and business teams to translate financial requirements into technical solutions.
- Document architectural decisions, trade-offs, data models, and migration strategies.
- Lead or support controlled production cutover from the legacy platform.
Required Skills & Qualifications - 8+ years of Data Engineering experience with demonstrated architectural ownership.
- Strong production experience with Databricks .
- Expert-level experience with: Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, Databricks Workflows / Job Orchestration, Spark performance tuning
- Strong Financial Services / Structured Finance experience.
- Experience with ABS, MBS, mortgage/loan-level data, cash flow analytics, or similar financial datasets .
- Strong understanding of financial time-series processing, including: Point-in-time accuracy, As-of joins, Bitemporal data, Late-arriving data, Historical data processing, Corporate actions
- Proven ability to understand and reverse-engineer complex legacy production codebases with limited documentation.
- Experience building data reconciliation and parity validation frameworks .
- Advanced Python and SQL .
- Strong shell scripting skills.
- Experience with Git, CI/CD, Terraform , or similar DevOps/IaC tooling.
- Ability to quantify business impact, such as performance improvements, cost savings, migration scale, processing volumes, or data parity results.
Preferred Qualifications - Experience with Delta Live Tables / Lakeflow .
- Strong Unity Catalog and data governance experience.
- Experience with FinOps / Databricks cost optimization .
- Experience taking a legacy data platform completely offline following successful migration and parity validation.
- Experience presenting at Databricks Data + AI Summit, AWS re, Snowflake Summit, FINOS , or other major data engineering conferences.
- Published technical blogs, whitepapers, or research related to data engineering, financial data, or financial modeling.
For applications and inquiries, contact:[email protected]