Company Description
Ubisoft is a world leader in video games, with teams spread across the globe creating original and memorable gaming experiences, from Assassin's Creed to Rainbow Six to Just Dance and many more. We believe that diversity of perspectives drives progress for both players and teams. If you are passionate about innovation and want to push the boundaries of entertainment, join our adventure and help us create the unknown!
Job Description
As a Site Reliability Specialist at Ubisoft Montreal, you will join the IT Games and Studios team and contribute to improving the availability, reliability, and performance of platforms and services essential to game development at Ubisoft.
You will collaborate with development teams, cloud specialists, and infrastructure teams to design resilient solutions, improve observability practices, and support operational excellence in production environments. Through automation, continuous improvement, and reliability-focused initiatives, you will help ensure the stability and efficiency of services used across the organization.
What you will do
- Collaborate with service teams to define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Design and implement automation solutions to improve operational efficiency and service reliability
- Document technical solutions and support their integration into internal platforms and services
- Ensure alignment of technical solutions with quality standards, engineering practices, and IT guidelines
- Work closely with development, infrastructure, and platform teams to improve operational consistency and system reliability
- Support observability practices, including monitoring, logging, alerting, and incident management
- Participate in root cause analyses and continuous improvement initiatives following service incidents
- Optimize deployment processes and operations through automation and infrastructure improvements
- Support the evolution and maintenance of cloud environments and services
- Share your knowledge and contribute to service reliability initiatives
Qualifications
What you bring to the team
- Experience in infrastructure engineering, automation, and DevOps practices
- Knowledge of GitLab and GitLab CI/CD for deployments and automation
- Proficiency in programming and scripting languages such as Python, Bash, and Go
- Experience with Terraform, Infrastructure as Code (IaC) practices, and Kubernetes (K8s) in public cloud environments such as Amazon Web Services (AWS) or Microsoft Azure
- Familiarity with configuration management tools such as Ansible or Chef
- Experience with observability and monitoring platforms such as Prometheus and Grafana
- Interest in AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar solutions that promote development, automation, and operational efficiency
- Ability to collaborate effectively with technical and non-technical partners while contributing to problem-solving and continuous improvement
