Job Description
Insight Global is seeking an AWS Site Reliability Engineer (SRE) / Observability Engineer for a leading technology-focused client. This individual will play a critical role in maintaining the reliability, scalability, and operational intelligence of cloud-based applications and infrastructure while supporting emerging AI-enabled operational workflows.
This position sits at the intersection of AWS cloud operations, observability engineering, incident management, and AI Operations. The ideal candidate is a highly analytical problem solver who can quickly dive into monitoring platforms, investigate logs, analyze metrics and alerts, identify root causes, and drive rapid issue resolution. They will leverage tools such as AWS, Terraform, Datadog, Splunk, Dynatrace, GitLab, and modern observability platforms to improve system health, operational visibility, and service reliability.
In addition to traditional SRE responsibilities, this individual will help support a growing AI Operations ecosystem by assisting with agentic AI workflows, participating in human-in-the-loop validation processes, evaluating AI-generated outputs, and refining AI prompts to improve operational effectiveness. The team is looking for someone who can thrive in a fast-paced environment, rapidly understand complex systems, determine operational impact, and continuously improve observability, monitoring, and AI-assisted operational processes.
This is an exciting opportunity for an engineer who enjoys solving production challenges, strengthening observability practices, driving incident response initiatives, and helping organizations scale next-generation AI-powered operational capabilities.
-Monitor and support AWS application and infrastructure environments
-Plan and execute application and infrastructure configuration changes
-Respond to production incidents, critical outages, and operational emergencies
-Triage application and infrastructure issues using monitoring and observability platforms
-Analyze logs, metrics, traces, and alerts to identify root cause and operational impact
-Determine whether issues are system-related, application-related, infrastructure-related, or AI workflow-related
-Lead incident response efforts and coordinate cross-functional resolution activities
-Develop and maintain postmortems, operational documentation, and technical runbooks
-Partner with software engineering teams to improve reliability, resiliency, and performance
-Enhance monitoring, logging, alerting, and observability capabilities across the environment
-Drive faster issue detection and resolution through monitoring best practices
-Support emerging agentic AI solutions being deployed across operational workflows
-Perform human-in-the-loop validation of AI-generated recommendations and outputs
-Review, refine, and optimize AI prompts to improve accuracy and operational effectiveness
-Collaborate with AI, development, and operations teams to improve operational intelligence
-Develop automation solutions that reduce manual effort and improve operational efficiency
-Participate in system design reviews, capacity planning initiatives, and architectural discussions
-Drive continuous improvement efforts focused on service reliability, observability maturity, and operational excellence
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Required Skills & Experience
Terraform expertise for Infrastructure as Code (IaC)
Strong AWS experience, particularly with EC2 and cloud-native services
Experience supporting AWS environments, including API Gateway technologies
Strong experience with Datadog monitoring and observability
Experience with enterprise monitoring and observability platforms such as Datadog, Dynatrace, and Splunk
GitLab and source control experience using Git
Strong troubleshooting, incident response, and production support experience
Strong log analysis experience within enterprise monitoring environments
Ability to quickly triage production issues and determine root cause using observability tools
Observability engineering experience, including identifying performance bottlenecks and system issues
Experience leading incident response efforts and supporting application teams
Experience creating and facilitating blameless postmortems
Ability to improve logging strategies, operational runbooks, and reduce technical debt
Experience supporting high-availability production applications and infrastructure
Experience supporting AI-powered operational workflows with human oversight and validation
Strong analytical mindset with the ability to rapidly understand system behavior and operational impact
Strong collaboration and communication skills when working across development, operations, and platform teams
Nice to Have Skills & Experience
Experience with Agentic AI systems or AI-enabled operational platforms
Prompt engineering experience, including creating and optimizing AI prompts
Experience working with Anthropic Claude or comparable enterprise AI models
Experience evaluating and validating AI-generated outputs
Experience supporting AI observability and AI Operations initiatives
API development experience
Experience building and supporting microservices architectures
Front-end development experience
Automation and platform engineering experience
Capacity planning and system design consulting experience
Experience promoting observability best practices across engineering organizations
Previous experience in large-scale cloud environments
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.