On-site
Eldorado do Sul, RS, Brazil
Salary Range
Full Time Employee
Experience Level
Senior
Requirements
Desired Skills
Tasks and Responsibilities
Show originalJoin us as a Senior Site Reliability Engineer (SRE) – Storage on our Infrastructure Engineering team in Eldorado do Sul/RS to do the best work of your career and drive profound social impact.
You will be responsible for designing, developing, and managing automations for the entire lifecycle of storage services, ensuring scalability, availability, performance, and reliability of corporate environments. You will work with Storage Engineering, Platform Automation, Observability, and Infrastructure teams to deliver a highly automated operation, supporting critical workloads through modern SRE practices, automation, and reliability engineering.
What You Will Do
- Design, develop, and maintain automations for the entire lifecycle of storage services, including volume and LUN provisioning, snapshot management and replication, capacity monitoring, firmware management, decommissioning, and automated remediation using Ansible, REST APIs, and CI/CD pipelines.
- Define and operate reliability practices for storage services by creating and maintaining SLIs, SLOs, and error budgets, enhancing observability, reducing operational noise, conducting post-incident analyses, and transforming recurring issues into automated solutions.
- Collaborate with Storage Engineering, Platform Automation, and Observability teams to establish operational requirements, integrate automations into the corporate AI platform, define operational readiness criteria, and contribute to policies-as-code aimed at protecting the availability, durability, and capacity of environments.
- Lead continuous improvement initiatives by measuring automation coverage, increasing autonomous resolution rates, maintaining a structured knowledge base, identifying automation opportunities, and evaluating new tools for integration into the corporate ecosystem.
- Participate in the Storage tower on-call rotation, acting as an escalation point when automated mechanisms fail to resolve critical incidents within established service levels, and leveraging pre-consolidated diagnostics to accelerate decision-making and resolution.
Essential Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related fields, with experience in Storage Engineering, SRE, Storage Operations, or large-scale corporate Infrastructure.
- Advanced or fluent English, with the ability to collaborate effectively with global teams.
- Advanced hands-on experience in at least two of the following areas: SAN/Block storage, NAS/File Storage, Object Storage, or equivalent corporate platforms, including solutions from Dell, Pure Storage, NetApp, HPE, or similar.
- Proven experience developing automations for storage platforms using Ansible, Python, REST APIs, Git, CI/CD pipelines, and GitOps practices, including provisioning, replication, capacity management, and environment lifecycle.
- Solid knowledge of observability, monitoring, service reliability, incident analysis, troubleshooting, and resolving complex failures involving performance, replication, capacity, controllers, disks, and distributed storage architectures.
Desired Requirements
- Experience with multi-vendor storage environments, monitoring platforms such as Zabbix or vROps, ServiceNow CMDB, GitLab CI, cloud storage services, hybrid architectures, and formal Site Reliability Engineering (SRE) practices.
- Familiarity with corporate AI-assisted automation platforms, agentic AI solutions, integration of LLMs into operational processes, environments regulated by compliance requirements (SOX, PCI-DSS, ISO 27001), and generation of evidence for audits.
Share job:
Share job: