About the Role
We're looking for a Senior Data Engineer who treats data integrity like a security issue. You'll own resilient ingestion pipelines from hostile third-party APIs and streaming sources, build a bulletproof medallion architecture in Databricks, and develop the kind of entity resolution logic that turns messy free-text into trustworthy analytics. This role reports to the Data Engineering lead and directly impacts the reliability of our observability and business intelligence platform.
Key Responsibilities
- Build and maintain REST API ingestion pipelines from third-party vendors (voice-AI, telephony platforms) that handle divergent pagination schemes, row caps, and schema drift — implementing contract checks and field-rename detection that alert rather than silently propagate nulls
- Own end-to-end streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and Auto Loader, including checkpoint management, deduplication logic, and source-parity reconciliation to detect late arrivals, duplicates, and schema mismatches
- Design and maintain Delta Lake medallion pipelines (bronze → silver → gold) using Spark SQL and PySpark with idempotent MERGE upserts, Delta Live Tables, and materialized-view constraints that enforce data quality at every layer
- Develop entity resolution logic that normalizes high-cardinality free-text fields (practice names, provider identifiers, payer names) using regex, fuzzy matching, canonical dictionaries, and confidence-scored crosswalks
- Establish and maintain reconciliation workflows that trace dashboard discrepancies back to raw event streams, documenting where data quality issues enter the pipeline and building alerts for drift detection
- Collaborate with analytics and product teams to define precise metric definitions and maintain semantic layer rigor, ensuring metrics are reproducible and auditable
Requirements
- 5+ years as a data engineer with hands-on production experience building streaming and batch pipelines at scale
- Deep expertise with Databricks (Delta Lake, Spark SQL, PySpark, Delta Live Tables) and experience designing medallion/multi-layer architectures
- Proven experience with streaming ingestion frameworks (Structured Streaming, Kafka, Event Hubs) including checkpoint management, watermarking, and deduplication
- Strong SQL and Python skills; comfort writing complex transformations and debugging data lineage issues across multiple pipeline stages
- Track record of building resilient data pipelines that handle real-world messiness: schema drift, late-arriving data, duplicates, and API inconsistencies
- Nice to have: Experience with entity resolution, fuzzy matching, or data quality frameworks (Great Expectations, dbt tests, or custom reconciliation logic)
- Nice to have: Familiarity with Azure ecosystem (Event Hubs, Data Factory, Synapse) or REST API pagination patterns and rate-limiting strategies
Required Skills & Tech Stack
Azure (Required)Apache Kafka (Required)SQL (Required)Etl Pipeline Design (Required)Apache Spark (Required)Delta Lake (Required)Python (Required)Pyspark (Required)Data Warehousing (Required)
Job Overview
Salary
Competitive Salary
Work Mode
onsite
Experience
senior
