Hybrid
São Paulo, SP, Brazil
Salary Range
Not informed
Experience Level
Mid level
Requirements
Tasks and Responsibilities
Show originalAct from raw data through to delivery for analytics, building and maintaining ETL/ELT pipelines with Python and SQL, as well as collecting external data via web scraping when necessary. The focus is on quality, traceability, automation, and compliance, avoiding fragile pipelines.
Responsibilities
- Develop and maintain ETL/ELT pipelines with Python and SQL.
- Process and standardize data, including cleaning, deduplication, validations, business rules, and enrichments.
- Implement web scraping from public sources, with rate limiting, reprocessing, versioning, logging, and handling changes on the website.
- Structure collection routines with fault tolerance, for example retries, backoff, and error capture.
- Create data layers for analytical consumption, with documentation and a data dictionary.
- Monitor pipelines, create alerts, quality metrics, audit, and reconciliation.
- Work with stakeholders to translate business needs into data and technical specifications.
- Ensure good governance practices, privacy, and proper use of data, especially for external sources.
Requirements
- Solid experience with Python for data manipulation, automation, and integration.
- Experience with SQL and transformation routines.
- Experience with data pipelines, ETL/ELT, and publishing to a data warehouse or data lake.
- Practical experience with web scraping, including HTML parsing and pagination mechanisms.
- Knowledge of Git and good development practices, testing, and documentation.
Share job:
Share job: