D
Remote
(Anywhere)
Salary Range
BRL
Full Time Employee
BRL
Contractor
Experience Level
Senior
Requirements
5+ years of experience in the career
Kubernetes
Python
Streaming
Ambiente cloud (GCP, AWS, Azure).
Tasks and Responsibilities
Show originalRESPONSIBILITIES
- Lead the construction and evolution of AI systems in production, combining batch and streaming data pipelines, backend services, and deployment of machine learning and LLM models.
- Design and implement robust pipelines for ingestion, transformation, and serving of structured, semi-structured, and unstructured data, with a focus on generative AI applications, RAG, and advanced analytics.
- Develop scalable backend services for real-time and batch inference, using REST/gRPC/GraphQL APIs, ensuring versioning, observability, and SLAs.
- Design event-driven distributed architectures for intelligent applications, using tools like Kafka, Pub/Sub, Kinesis, or SQS/SNS.
- Define standards and best practices for relational, non-relational, and vector databases, including embedding modeling, feature stores, and retrieval strategies for RAG.
- Operate AI workloads in the cloud and Kubernetes, including GPU optimization, autoscaling, scheduling, and cost control.
- Implement MLOps and LLOps pipelines, including CI/CD for models, drift monitoring, dataset versioning, and experiment tracking.
- Ensure governance, quality, traceability, and security of data and models throughout their lifecycle.
- Act in the design and optimization of data and AI architecture, prioritizing performance, reliability, cost-effectiveness, and security.
- Anticipate technical risks in pipelines, inference, and infrastructure, proposing scalable and resilient solutions.
- Support the construction of internal data and AI platforms (data platform / AI platform), promoting standardization and reuse of components.
- Collaborate with multifunctional product, engineering, and data science squads, leading strategic technical decisions.
- Mentor mid-level and junior engineers, disseminating best practices in software, data, and AI engineering.
- Actively contribute to the corporate AI strategy, aligning technology with business objectives and product evolution.
MANDATORY REQUIREMENTS
- Minimum of 5 years of experience in Data Engineering, Backend, ML Engineering, or AI Engineering.
- Solid experience in developing batch and streaming pipelines in production.
- Proven experience in deploying and operating ML or LLM models in production.
- Experience in distributed and event-driven architecture.
- Mastery of Python, SQL, and API development.
- Experience with containers and Kubernetes.
- Experience in cloud (AWS, GCP, or Azure) and data governance best practices.
- Experience in system observability (structured logs, metrics, tracing).
- Experience in optimizing pipeline and service performance.
- Strong technical communication skills with technical and non-technical stakeholders.
DESIRED REQUIREMENTS
- Experience with LLM and RAG frameworks (LangChain, LlamaIndex, etc.).
- Experience with Ray, Spark, Flint, Beam, or distributed processing systems.
- Experience with experiment tracking and MLOps platforms (ClearML, MLflow, Vertex AI, SageMaker).
- Advanced knowledge in infrastructure as code (Terraform, Helm, CDK).
- Experience deploying scalable applications on Kubernetes.
- Experience with vector databases and semantic search systems.
- Experience with GPU workloads and cost optimization in the cloud.
- Track record of technical leadership in complex projects or internal platforms.
ESSENTIAL SOFT SKILLS
- Clear and effective communication at all levels of the organization.
- Ability to lead complex technical discussions and architectural decisions.
- Mentorship and technical development skills for the team.
- Systemic vision to assess risks, trade-offs, and opportunities.
- Autonomy and initiative in solving complex problems.
- Platform ownership mindset, focused on reliability, cost, and developer experience.
VALUED PROJECTS / ACHIEVEMENTS
- Implementation of AI systems in production with high business impact.
- Construction of streaming pipelines or internal data/AI platforms.
- Deployment of RAG systems with corporate data.
- Proven optimization of infrastructure costs or GPU usage.
- Creation of scalable and resilient distributed architectures.
- Definition of organizational standards for data or AI engineering.
Share job:
Share job: