GeekHunter Logo

Solutions

Recruitment Software

Post jobs, import profiles, add screening questions, and manage hiring all in one place with personalized questions and control the entire recruitment process.

Recruitment Service

Let our expert team handle the key stages of tech recruitment for you with our specialized team acting in the main stages of recruitment.

Talent Pool

Top-tier Brazilian tech professionals, pre-vetted and ready for action, all pre-selected and ready for new opportunities.

Our Plans

Discover the perfect plan for your needs.

Login

English

EN

ZG Soluções

Senior SRE Analyst

Show original

Remote

(Anywhere)

Salary Range

BRL

Full Time Employee

Experience Level

Senior

Requirements

5+ years of experience in the career
Terraform
Ansible
Virtualization
Kubernetes
Docker
Linux

Desired Skills

Automação & CI/CD
PostgreSQL
Prometheus, Grafana, ELK, Datadog

Tasks and Responsibilities

Show original

We are seeking a Senior Infrastructure and Reliability professional with proven experience in critical production environments.

Our environments are hybrid: workloads on public cloud, workloads in data centers with virtualization, and workloads within our clients' private infrastructure. This requires someone who masters the foundation (networking, virtualization, and Linux) and automates everything that supports this foundation.

By joining us, you will be responsible for provisioning infrastructure in an automated manner, operating our production Kubernetes clusters, and ensuring the resilience and observability of applications across all these environments.

We evaluate technical depth and diagnostic capability. We do not evaluate visibility, public speaking, or social media presence.


RESPONSIBILITIES AND DUTIES

• Automate infrastructure provisioning and configuration using Terraform and Ansible;

• Operate and evolve production Kubernetes clusters, including networking, storage, RBAC, and update strategies;

• Build and maintain lean, reproducible Docker images and containers;

• Maintain connectivity between our environments and those of our clients: routing, VPN, DNS, TLS, firewall, and reverse proxy;

• Administer and scale the virtualization layer and the Linux systems supporting the platform;

• Maintain and evolve the architecture with a focus on resilience and high availability;

• Evolve the CI/CD pipeline and delivery via GitOps (ArgoCD);

• Promote observability and monitoring with Prometheus, Grafana, and Zabbix, reducing alert noise;

• Lead incident analysis to root cause and transform findings into documented structural fixes;

• Support the operation of Java applications and PostgreSQL databases in production;

• Promote security management across servers and services in the ecosystem.


REQUIREMENTS AND QUALIFICATIONS

The requirements below are non-negotiable for this position. We expect you to be able to explain each in depth, with examples of your work:

• Solid experience as an SRE, DevOps, or Infrastructure engineer in critical production environments, with direct involvement in high-impact incidents;

• Mastery of networking: TCP/IP, routing, DNS, NAT, TLS/SSL, VPN, firewall, and reverse proxy with NGINX;

• Experience managing virtualization in production, preferably Proxmox (VMware, KVM, or equivalents are also considered);

• Advanced Linux operations and automation with Shell/Bash;

• Provisioning automation with Terraform and Ansible in real-world, long-term use;

• Management of production Kubernetes clusters, with the ability to investigate and resolve workload, traffic, and node failures;

• Building, updating, and managing Docker images and containers;

• Experience with CI/CD processes and delivery via GitOps (ArgoCD or equivalent);

• Experience with monitoring and observability tools (Prometheus, Grafana, Zabbix);

• Version control with Git as a daily practice;

• Autonomy to take on ill-defined problems, investigate them thoroughly, and deliver solutions;

• Clear communication and high-quality technical documentation, especially during incidents.


YOU WILL STAND OUT IF YOU HAVE

• Experience in regulated environments (healthcare, finance, or public sector) and familiarity with LGPD (Brazilian General Data Protection Law);

• Hands-on experience troubleshooting distributed systems and intermittent failures;

• Experience with middleware, Java application servers, message queues, and API Gateways;

• Hands-on experience with hardening, secrets management, and vulnerability management;

• PostgreSQL expertise at the level of tuning, replication, and recovery;

• Certifications that validate your practice (CKA/CKS, RHCE, LPIC, networking, or cloud);

• Ability to write your own tools when none exist (Python, Go, or equivalent);

• Use of AI in daily technical work to build tools, automate tasks, and accelerate problem resolution.


THIS ROLE IS NOT FOR YOU IF

• Your Kubernetes experience is limited to applying manifests written by others;

• You treat networking as a black box and escalate anything outside the application to third parties;

• You provision via graphical consoles and do not version control infrastructure.


See all jobs at ZG Soluções

Share job:

Share job: