plaintext JOB SUMMARY Senior Azure Site Reliability Engineer
Key Responsibilities Act as the primary support lead for high-severity production outages and critical business escalations. Coordinate incident response activities and bridge calls to restore service for mission-critical healthcare applications. Perform troubleshooting and root cause analysis across Azure infrastructure, databases, and application services. Execute recovery and remediation activities to minimize disruption to patient care and business operations. Monitor and maintain the health, availability, and performance of Azure App Services, Virtual Machines, and Azure Functions. Analyze telemetry and system performance data using Azure Monitor, Log Analytics, and Application Insights. Support and optimize Azure SQL databases, enterprise data pipelines, and cloud storage solutions. Deploy system patches, hotfixes, and updates in accordance with security and operational standards. Audit system logs, user access, and security controls to ensure ongoing compliance and data protection. Develop and maintain automation scripts to reduce manual support activities and improve operational efficiency. Create and maintain technical documentation, support runbooks, disaster recovery procedures, and compliance artifacts. Collaborate with engineering, infrastructure, security, and application teams to drive service reliability and continuous improvement.
Required Qualifications 5+ years of experience in Production Support, Site Reliability Engineering (SRE), or a related support engineering function, including 2+ years supporting healthcare environments. Hands-on experience managing and supporting production systems within Microsoft Azure. Strong expertise in Azure SQL, T-SQL, and healthcare data exchange standards such as HL7 and FHIR. Advanced scripting and automation skills using PowerShell, Azure CLI, Bash, or Python. Ability to serve as the primary point of contact for high-severity production incidents and critical escalations. Experience leading incident response efforts, coordinating bridge calls, and driving rapid resolution within strict SLA requirements. Proven ability to perform root cause analysis (RCA) across infrastructure, application, and database layers. Commitment to participating in a 24/7 on-call rotation and supporting mission-critical healthcare applications. Working knowledge of HIPAA, HITRUST, and healthcare data privacy and security requirements.
Preferred Qualifications Experience supporting Epic EHR environments, including Interconnect, Bridges, Cogito, or related Epic modules. Familiarity with Azure OpenAI Services, AI-powered applications, and Large Language Model (LLM) operations. Knowledge of monitoring AI application performance, including prompt/response latency and operational health metrics. Experience working with vector databases or AI-driven search and retrieval platforms. Exposure to enterprise disaster recovery planning, compliance audits, and healthcare regulatory environments.
This field is requiredPlease enter valid emailId.
This field is requiredPlease enter valid cell phone.
This field is requiredPlease enter valid first name.
This field is requiredPlease enter valid last name.