DevOps/MLOps Engineer - INTL India

Post Date

Jul 18, 2026

Location

Vancouver,
British Columbia

ZIP/Postal Code

V5T 1
Canada
Sep 19, 2026 Insight Global

Job Type

Contract-to-perm

Category

Software Engineering

Req #

VAN-79890fe7-abbb-47e2-ba78-0f5a1777a7c3

Pay Rate

$18 - $22 (hourly estimate)

Who Can Apply

  • Candidates must be legally authorized to work in Canada

Job Description

Insight Global is seeking an seeking an AI DevOps / MLOps Engineer to join a platform engineering team for a AAA gaming company focused on maintaining and supporting production AI and cloud-native applications. This is a hands-on operational role responsible for keeping Kubernetes-based services stable, secure, and reliable across global engineering teams.

Key Responsibilities:
- Maintain and support Kubernetes-based application deployments across development and production environments.
- Manage and troubleshoot GitFlow-based CI/CD pipelines used by multiple engineering teams.
- Support containerized workloads using Docker, Helm Charts, ArgoCD, and Kubernetes.
- Perform platform maintenance activities including security upgrades, patching, dependency updates, and infrastructure improvements.
- Partner with engineering teams across North America and Eastern Europe to resolve production issues and operational challenges.
- Troubleshoot application, infrastructure, networking, and deployment issues across Kubernetes clusters.
- Support AI services, model evaluation workflows, RAG pipelines, automation services, and related cloud-native applications.
- Improve reliability through monitoring, alerting, health checks, readiness probes, retries, scaling strategies, and rollback procedures.
- Build and maintain operational dashboards, logging, metrics, alerts, and runbooks.
- Investigate incidents, identify recurring failure patterns, and implement operational improvements to increase platform resiliency.
- Provide day-to-day support for engineering teams during active troubleshooting scenarios, serving as a technical resource when production issues arise.
- Collaborate with platform, AI, and operations teams to establish operational best practices for running production services.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Required Skills & Experience

- 5+ years of experience in DevOps, MLOps, Site Reliability Engineering (SRE), Platform Engineering, or Cloud Engineering, with strong hands-on production Kubernetes experience.
- 2-4 years of experience building, deploying, and supporting containerized applications using Docker, Helm, ArgoCD, CI/CD pipelines, and GitFlow methodologies.
- Experience with Kubeflow or similar machine learning workflow orchestration platforms.
- Experience supporting applications and services within Azure environments, including familiarity with Azure OpenAI, Azure AI Search, RAG architectures, and LLM-powered services.
- Experience operating production systems with monitoring, alerting, incident response, operational support, observability tooling (logs, metrics, dashboards, alerts), and runbook creation.
- Experience with OpenCity or similar internal developer platforms and cloud enablement tools.
- Strong troubleshooting skills across Kubernetes, networking, infrastructure, secrets management, cloud dependencies, and application services.
- Understanding of reliability engineering principles including scaling, health probes, retries, timeouts, failover, recovery, rollback strategies, and graceful degradation.

Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.