Job Summary
The Lead Data Engineer will focus approximately 70% on data engineering and 30% on analytical responsibilities, combining deep Databricks and PySpark expertise with strong analytical capabilities and finance or payroll domain experience. The role will design and maintain scalable cloud-native data platforms, production ETL pipelines, and large-scale data processing solutions while supporting exploratory analysis, data science workflows, and financial research. The engineer will work closely with data scientists, economists, and business stakeholders to translate complex analytical requirements into reliable, scalable, and maintainable solutions.
Key Responsibilities
• Design, develop, and maintain scalable data pipelines for ingestion, transformation, and distribution of payroll, macroeconomic, and financial datasets.
• Build and support Databricks-based platforms for financial research and analytical workloads.
• Implement ETL/ELT frameworks using PySpark and Delta Lake for structured and unstructured data.
• Develop data models and data marts optimized for analytics, reporting, and machine learning use cases.
• Ensure data quality, consistency, lineage, governance, security, and observability across data assets.
• Optimize processing for large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
• Translate business and analytical requirements into scalable and maintainable data solutions.
• Perform exploratory data analysis, including profiling datasets and identifying distributions, outliers, missing patterns, and data drift.
• Translate data scientist logic into efficient PySpark implementations, including cross-sectional metrics and time-windowed aggregations.
• Build validation dashboards and exploratory notebooks to verify pipeline outputs and data quality.
• Support feature engineering through complex aggregation and transformation logic at scale.
• Independently validate analytical outputs, investigate anomalies, and identify results requiring further analysis.
• Perform ad hoc analytical work using pandas and NumPy alongside PySpark.
• Contribute to lakehouse architecture, medallion architecture, data mesh concepts, catalog design, and technical tradeoff discussions.
• Implement CI/CD pipelines using Databricks Asset Bundles, Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
• Manage Unity Catalog governance, access patterns, and schema design.
• Leverage AI-assisted development tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent platforms.
• Review AI-generated code for correctness, performance, scalability, security, and maintainability.
Required Qualifications
• Bachelor’s or Master’s degree in Computer Science, Data Engineering, Information Systems, Statistics, Economics, Finance, or a related field.
• 5+ years of experience in Data Engineering or Data Platform development.
• Finance or payroll domain experience, including familiarity with payroll data structures, pay-period logic, compensation and deduction relationships, or financial-services data environments.
• Experience working with billions of records, multi-terabyte datasets, and large-scale time-series data.
• Strong Databricks expertise, including Unity Catalog, Delta Lake, Databricks Workflows, and Databricks Asset Bundles or equivalent deployment frameworks.
• Strong understanding of Delta Lake optimization, clustering, change data feed, and versioning.
• Strong Python, SQL, PySpark, data modeling, and ETL/ELT development skills.
• Experience with data quality and validation frameworks.
• Strong analytical fluency, including exploratory data analysis, basic statistical concepts, distributions, correlations, time-series patterns, and feature engineering.
• Proficiency with pandas and NumPy for analytical work.
• Experience participating in architecture-level design decisions and evaluating technical tradeoffs.
• Experience implementing CI/CD using tools such as Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
• Experience with AI-assisted development tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent.
• Strong understanding of data governance, data quality, and metadata management.
Preferred Qualifications
• Experience with macroeconomic, capital markets, or financial-services data.
• Exposure to lakehouse patterns, data mesh concepts, and medallion architecture.
• Experience with Kafka or Structured Streaming for streaming and event-driven pipelines.
• Experience processing Census or geographic data, including TIGER and FIPS codes.
• Infrastructure-as-code experience with Terraform, CDK, or similar technologies.
• Databricks certifications at the Associate or Professional level.
• Experience migrating legacy data platforms such as Glue, EMR, or HDInsight to Databricks.
• Experience supporting machine learning and AI-driven analytics solutions.
• Data visualization experience using Power BI, Tableau, Databricks Dashboards, or similar platforms.
• Scala experience.
• Experience with SQL Server, PostgreSQL, Delta Tables, and NoSQL databases.