Hybrid
São Paulo, SP, Brazil
Salary Range
Contractor
Experience Level
Senior
Requirements
Desired Skills
Tasks and Responsibilities
Show originalAbout the Role
We are seeking a Senior Data Engineer to technically lead data initiatives at a national scale. This position focuses primarily on entity resolution, master data governance, and reconciliation of third-party sources, ensuring data quality and observability across the entire production chain.
You will serve as the technical reference for the team, conducting code reviews focused on correctness, cost, and maintainability, while applying AI-assisted engineering with guardrails and traceability.
Key Responsibilities
- Define single source of truth rules, survival logic, stable identifiers, and manual overrides for license data, deeds, tax records, and commercial data.
- Reconcile conflicting third-party sources at a national scale and implement observability to detect deviations in logic.
- Act as a data model steward, onboarding new sources and evolving models when they no longer meet business needs.
- Ensure end-to-end data quality: testing, monitoring, alerting, lineage, and root cause analysis of incidents.
- Conduct analyses to guide the roadmap, prioritizing source acquisitions and datasets based on business impact.
- Deploy data science models to production, defining data contracts and monitoring drift.
- Serve as the team's technical reference for code reviews and architectural trade-offs.
- Apply AI-assisted engineering using agents, maintaining guardrails and traceability.
What We're Looking For
- Solid SQL and Python skills applied throughout the entire data lifecycle, from ingestion to serving.
- Hands-on experience with Databricks in production: notebooks, jobs, workflows, Delta Lake, and PySpark.
- Deep expertise in data modeling: dimensional modeling, slowly changing dimensions, grain definition, and layered architectures like Medallion.
- Experience with entity resolution and data reconciliation.
- Knowledge of data quality and observability: testing, monitoring, alerting, and lineage.
- Familiarity with transformation and orchestration frameworks such as dbt, Airflow, and Databricks Workflows, along with version control and CI/CD.
- Experience delivering data products that feed production applications, plus the ability to translate ambiguous business questions into actionable metrics.
Differentiators
- Domain-oriented data product architecture in a lakehouse environment.
- AI architectures applied to data: pipelines for RAG, MCP for standardized agent access, and quality assurance in non-deterministic agentic systems.
- Advanced Databricks ecosystem: Unity Catalog and Delta Live Tables.
- Data governance: access control, cataloging, and lineage.
- Advanced MLOps: model serving, drift monitoring, and retraining workflows.
Tech Stack
Python, SQL, PySpark, Databricks, Delta Lake, dbt, AWS, Terraform, Postgres, GitHub Actions, and AI agents in daily development workflows.
Share job:
Share job: