GeekHunter Logo

Solutions

Recruitment Software

Post jobs, import profiles, add screening questions, and manage hiring all in one place with personalized questions and control the entire recruitment process.

Recruitment Service

Let our expert team handle the key stages of tech recruitment for you with our specialized team acting in the main stages of recruitment.

Talent Pool

Top-tier Brazilian tech professionals, pre-vetted and ready for action, all pre-selected and ready for new opportunities.

Our Plans

Discover the perfect plan for your needs.

Login

English

EN

Nava Technology for Business

Senior Site Reliability Engineer - SRE (Affirmative Action for Women)

Show original

Hybrid

São Paulo, SP, Brazil

Salary Range

Not informed

Experience Level

Senior

Requirements

5+ years of experience in the career

Tasks and Responsibilities

Show original

We are looking for a Site Reliability Engineer (SRE) to act as the guardian of reliability, stability, and performance of our products and services. If you enjoy working with critical environments, data-driven decisions, and a blameless culture, this role may be for you.



🎯 Role Mission

Ensure that our systems operate with high reliability, efficiency, and predictability, balancing delivery speed and operational robustness. The SRE will be a key piece in the technical maturity evolution of the squad and in sustaining critical services.


The professional will work on a rotating on-call scale, responding to incidents within defined SLAs, conducting rapid stabilizations, participating in blameless postmortems, and proposing continuous improvements to reduce recurrence. On-call follows internal compensation policies.


Main Responsibilities


Reliability and Governance

  • Define, maintain, and evolve SLIs and SLOs for critical APIs
  • Manage error budgets and support release decisions
  • Act as a reference in balancing agility and stability

Observability and Operations

  • Implement and evolve monitoring, metrics, logs, and tracing
  • Ensure actionable alerts and efficient dashboards
  • Lead or support incident response and war rooms

Incident Management

  • Structure and execute blameless incident response processes
  • Conduct postmortems and ensure corrective actions
  • Act in reducing MTTA, MTTR, and recurrence

Automation and Toil Reduction

  • Automate repetitive tasks and operational flows
  • Create runbooks, automations, and CI/CD improvements
  • Standardize rollout, rollback, and resilience testing processes
  • Infrastructure and Performance
  • Work with Kubernetes/EKS, Azure DevOps, Kafka, and databases

Required Requirements

  • Experience in Engineering, Infra, Platform, or SRE/DevOps
  • Experience with SLO, SLI, error budget, and incident management
  • Strong troubleshooting skills and RCA (Root Cause Analysis)
  • Technologies
  • Kubernetes/EKS, Azure DevOps
  • Observability: Prometheus, Grafana, ELK, CloudWatch, X-Ray
  • Kafka, Oracle, MySQL
  • Operational security and IAM
  • Languages and Automation
  • Bash, PowerShell, Python
  • Ansible, Terraform, Helm
  • Differentiator: .NET Framework and .NET Core

Availability to work in the hybrid model in the Vila Olímpia region of São Paulo, 1 to 2 times a week, is required.



In addition to being a Great Place to Work certified company, you will find at NAVA:

✅ Career opportunities 🚀

✅ Freedom to write your own code 🏆

✅ Diversity and different ways of seeing the world 🌈

✅ Communities that encourage everyone's growth 📚

✅ In-Company Training 💻

✅ An incredible team 😎 ✅ Company engaged in the UN Global Compact 💪🏼

✅ Innovative projects 💡

✅ High rating on Glassdoor 📣


See all jobs at Nava Technology for Business

Share job:

Share job: