Skip to main content

Specialist – Tooling and Infrastructure Reliability

Auto-translated from French · original: Spécialiste – Fiabilité des outils et de l’infrastructure

4 × 8hr days80% payOnsite · Montreal, Canada

Company Description

Ubisoft is a world leader in video games, with teams across the globe creating original and memorable gaming experiences, from Assassin’s Creed to Rainbow Six, Just Dance, and many more. We believe that diverse perspectives help both players and teams grow. If you are passionate about innovation and want to push the boundaries of entertainment, join our adventure and help us create the unknown!

Job Description

The incumbent ensures the ongoing viability, stability, and performance of the operational tools and infrastructure that support the development of GaaS games. They design, develop, and operate tools and pipelines (build, configuration, versioning, deployment, publishing) to simplify, optimize, and automate development processes. They educate and support teams on testing, quality, security, and automation before go-live, and promote best practices aimed at delivering a reliable and high-performance gaming experience.

Responsibilities

  • Support development teams in technological and tooling choices to improve the visibility, control, and robustness of internal and external services.
  • Educate, support, and guide development teams in improving continuous integration and deployment systems.
  • Research, integrate, and develop technologies that improve reliability, performance, and productivity.
  • Design, operate, and take ownership of build, configuration, versioning, and publishing pipelines (including packaging, signatures, SBOM, artifacts).
  • Implement and support CI/CD tooling (automated testing, quality, security), IaC, and secure, reproducible, and controlled deployments.
  • Maintain tooling products to provide exemplary service quality to the project (internal SLOs).
  • Implement and maintain game deployment guides and document the implementation as well as the technical specifications of network and server infrastructures.
  • Collaborate with development teams to diagnose and fix anomalies and outages related to online services.
  • Implement and maintain incident management processes.
  • Manage the Cloud using appropriate tools.
  • Develop tools and processes that facilitate the deployment of services by developers in a secure and controlled manner.
  • Define and track SLAs/SLOs/SLIs, deploy observability (logs, metrics, traces), manage capacity, and contribute to FinOps initiatives.

Qualifications

Education

  • University degree in Computer Science, Computer Engineering, or any other relevant field.

Experience

  • 5 to 8 years of experience in software development and systems administration.
  • Experience in infrastructure automation (Cloud).
  • Experience in managing high-throughput systems.
  • Experience in designing resilient, scalable, and redundant architectures.
  • Experience in software development and optimization.

Skills and Knowledge

  • Excellent analytical and synthesis skills.
  • Ability to solve complex problems.
  • Ability to adapt quickly to change.
  • Ability to work under pressure.
  • Very good knowledge of distributed systems.
  • Very good knowledge of Linux and Windows systems administration.
  • Languages: Python, Go, C#, or C++.
  • CI/CD (GitLab, GitHub, Azure DevOps), IaC (Terraform, CloudFormation), containers and orchestration (Docker, Kubernetes).
  • Observability: Prometheus/Grafana, ELK/EFK, OpenTelemetry (or equivalents).
  • Cloud: AWS, Azure, GCP; databases; networks (DNS, CDN, load balancing, TLS).
  • Assets: Unreal Engine 5 (or similar engine), DevOps methodology, experience in infrastructure automation.