Geekhunter Logo

Solutions

Recruitment Software

Post jobs, import profiles, add screening questions, and manage hiring all in one place with personalized questions and control the entire recruitment process.

Recruitment Service

Let our expert team handle the key stages of tech recruitment for you with our specialized team acting in the main stages of recruitment.

Talent Pool

Top-tier Brazilian tech professionals, pre-vetted and ready for action, all pre-selected and ready for new opportunities.

Our Plans

Discover the perfect plan for your needs.

Login

English

EN

Dell Technologies | Radancy

Senior Site Reliability Engineer (SRE) – Server

Show original

On-site

Eldorado do Sul, RS, Brazil

Salary Range

Not informed

Full Time Employee

Experience Level

Senior

Requirements

5+ years of experience in the career
Inglês advanced
Ansible
Oracle Linux
Git
Windows Server
Observabilidade
VMware vSphere
GitOps
Monitoramento
SRE
RHEL
CI/CD
Linux
Python
vCenter

Desired Skills

Zabbix
Kickstart
Packer
vROps
LLM
GitLab CI
Dynatrace
PCI-DSS
SOX
NetBox
ISO 27001
Preseed
Agentic AI
ServiceNow CMDB

Tasks and Responsibilities

Show original

Join us as a Senior Site Reliability Engineer (SRE) on our Infrastructure Engineering team in Eldorado do Sul/RS to do the best work of your career and generate a profound social impact.


As a Senior Site Reliability Engineer (SRE), you will be responsible for designing, developing, and managing automations for the entire compute infrastructure lifecycle, ensuring scalability, availability, and reliability of services. You will work with Compute Engineering, Platform Automation, and Observability teams on a corporate platform based on automation, supporting critical workloads and driving operational excellence through modern SRE practices and large-scale automation.



What you will do


  • Design, develop, and maintain automations for the entire compute infrastructure lifecycle, including server discovery, bare-metal provisioning, operating system deployment, firmware and patch management, hypervisor configuration, decommissioning, and automated diagnostic and remediation processes using Ansible, Python, and CI/CD pipelines.
  • Define and operate reliability practices for compute services by creating and maintaining SLIs, SLOs, and error budgets, in addition to enhancing observability, reducing monitoring noise, conducting post-incident analyses, and transforming recurring failures into automated solutions.
  • Collaborate with Compute Engineering, Platform Automation, and Observability teams to define operational requirements, integrate automations into the corporate AI and automation platform, ensure operational readiness criteria, and contribute to policies-as-code aimed at safe and scalable operation of compute platforms.
  • Lead continuous improvement initiatives by measuring automation coverage, increasing autonomous resolution rates, maintaining structured knowledge bases, identifying opportunities to eliminate manual activities, and evaluating new tools for integration into the corporate ecosystem.
  • Participate in the Compute tower on-call rotation, acting as an escalation point when automated mechanisms fail to resolve critical incidents within defined service levels, and leveraging previously consolidated diagnostic information to accelerate resolution.


Essential Requirements


  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related fields, with experience in SRE, Infrastructure Engineering, Server Engineering, or Large-Scale Platform Operations.
  • Advanced or fluent English, with the ability to collaborate effectively with global teams.
  • Advanced hands-on experience in at least two of the following areas: VMware vSphere/vCenter, bare-metal server platforms (Dell PowerEdge, HPE, Cisco UCS, or equivalents), Linux operating systems (Oracle Linux, RHEL, or equivalents), or Windows Server environments.
  • Proven experience in developing infrastructure automations using Ansible, Python, Git, CI/CD pipelines, and GitOps practices, including automated management of the compute platform lifecycle.
  • Solid knowledge in observability, infrastructure monitoring, system reliability, troubleshooting, incident management, root cause analysis, and resolution of complex failures involving hardware, hypervisors, operating systems, and distributed environments.


Desirable Requirements


  • Experience with vROps, Zabbix, Dynatrace, ServiceNow CMDB, NetBox, GitLab CI, image management tools (Kickstart, Preseed, Packer), firmware automation, formal Site Reliability Engineering (SRE) practices, and environments regulated by compliance requirements such as SOX, PCI-DSS, or ISO 27001.
  • Familiarity with corporate AI-assisted automation platforms, agentic AI solutions, integration of LLMs into operational processes, and advanced ecosystems of corporate observability and automation.
See all jobs at Dell Technologies | Radancy

Share job:

Share job: