Job Description

Job Description
A large healthcare client of ours is seeking a Site Reliability Engineer to join a fast-paced operations team responsible for maintaining platform stability, managing critical incidents, and supporting enterprise applications. This individual will play a key role in incident response, service reliability, cross-functional coordination, and operational excellence.
The ideal candidate brings a strong operations mindset, excels under pressure, and can effectively lead technical incident management efforts while communicating with both technical and business stakeholders.

Responsibilities
Lead and coordinate response efforts for P1/P2 production incidents.
Serve as Incident Commander during major outages and service disruptions.
Drive incident triage and engage appropriate engineering, infrastructure, and business teams.
Monitor application and platform health to ensure system availability and reliability
Facilitate root cause analysis activiti...

Ready to Apply?

Take the next step in your AI career. Submit your application to Insight Global today.

Submit Application