GeekHunter Logo

Plans

Login

EN

EN

Databuild

Mid-Level Data Engineer (Remote, Contract) - Airflow | Spark | Python | SQL

Show original

Remote

(Anywhere)

Salary Range

Not informed

Experience Level

Mid level

Requirements

3+ years of experience in the career
SQL
Python
Analise de dados com uso de Python
Apache Airflow
Desenvolvimento de sistema com Python e MySQL

Tasks and Responsibilities

Show original

Tasks and Responsibilities

We are a fast-growing Data Science and Data Engineering consultancy, and we are expanding our team to deliver large, complex, high-impact projects.

Here, AI is not just a “concept”: it is part of the daily work in automations, pipelines, observability, quality, and productivity, always with a focus on robustness and scale for enterprise clients.

If you enjoy building things that actually work — with data arriving on time, reliable orchestration, governance, performance, and traceability — come with us. 👇

✅ What you will do

  • Build and evolve data pipelines (batch and, when necessary, near real-time)
  • Orchestrate workflows with Apache Airflow: resilient, idempotent, and easy-to-maintain DAGs, with well-handled dependencies, retries, and backfills
  • Develop and optimize distributed processing with Spark (PySpark) in environments such as Databricks
  • Work with Python and SQL for ingestion, transformation, and validation
  • Implement data quality, monitoring/alerts, logs, and observability
  • Collaborate with Data Science, Analytics, and Product teams
  • Document decisions and support engineering standards (Git, CI/CD, code review)

🎯 Requirements

  • Hands-on experience with Apache Airflow in production: creating and maintaining DAGs, operators, sensors, and hooks, scheduling, retries, backfill, and troubleshooting executions
  • Solid Spark / PySpark (transformations, optimization, efficient read/write)
  • Python applied to data engineering (ingestion, API integration, transformation, and testing)
  • Strong SQL (joins, CTEs, window functions, modeling, and performance)
  • Clear understanding of data modeling and governance
  • Plus: Databricks (Jobs, notebooks, clusters, Delta Lake/medallion architecture, Unity Catalog), managed Airflow or Airflow on Kubernetes (MWAA, Cloud Composer, Astronomer), cloud (Azure, AWS, or GCP), dbt, and streaming (Kafka, Spark Structured Streaming)

💡 What you will find here

  • Challenging projects with large companies and real data
  • A fast-growing environment with room for ownership and growth
  • Team culture: collaboration, autonomy, delivery, and quality
  • Modern stack, with real encouragement for automation and best practices

Share job:

Share job: