Job Description
We are seeking a Senior Site Reliability Engineer (SRE) to join a high-impact Platform Engineering team focused on building scalable cloud infrastructure, reusable platform capabilities, and automation frameworks that enable software engineers to develop and deploy applications at scale. This is not a traditional operations role—the team focuses on platform engineering, Infrastructure as Code, developer enablement, and reliability through software and automation.
You will design and build reusable cloud infrastructure across AWS and Azure, develop Infrastructure as Code from the ground up, create automation and self-service platform capabilities, enhance observability using Datadog, and contribute to modern CI/CD practices. The ideal candidate has a strong software engineering mindset and enjoys building platforms that improve engineering productivity while supporting highly available production environments.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
• 6+ years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, or Cloud Engineering.
• Hands-on experience designing, building, and supporting production cloud environments across both AWS and Azure.
• Strong experience building reusable Infrastructure as Code using Terraform, including custom modules, shared frameworks, and cloud platform components (AWS CDK and/or CloudFormation preferred).
• Experience designing and supporting Kubernetes-based platforms in production environments.
• Strong software engineering mindset with experience building automation using languages such as Go, Python, PowerShell, TypeScript, or similar.
• Strong experience with Datadog (or similar observability platforms), including monitoring, troubleshooting production issues, and improving platform reliability.
• Experience building modern CI/CD pipelines and developer enablement capabilities through platform engineering.
Nice to Have Skills & Experience
• Experience leveraging AI tools, AI agents, or LLM-powered workflows to improve engineering productivity or platform automation.
• Experience working in regulated industries such as healthcare or medical devices.
• Experience with serverless technologies (AWS Lambda, Azure Functions).
• AWS, Azure, Terraform, or Kubernetes certifications.
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.