Submit Resume

Site Reliability Engineer

  • ONTARIO, Toronto

  • 07/22/2026

  • Contract

  • Active

Job Description:

  • Job Summary
    We are seeking a Site Reliability Engineer to support Cyber Data Risk & Resilience by ensuring the reliability, availability, performance, and operational visibility of critical cybersecurity platforms and services. This role is responsible for maintaining production systems, building observability solutions, automating operational processes, supporting incident response, enhancing executive dashboards, and driving continuous improvements across cloud-based and distributed environments.

    Key Responsibilities
    • Maintain and improve the reliability, availability, scalability, and performance of cybersecurity platforms and supporting infrastructure.
    • Monitor system health, identify operational risks, respond to incidents, and drive timely resolution of service-impacting issues.
    • Instrument infrastructure, applications, APIs, databases, cloud components, and data pipelines for end-to-end observability.
    • Design, build, and enhance monitoring, alerting, logging, tracing, and dashboard solutions across distributed systems.
    • Develop actionable alerts that reduce noise, improve signal quality, and accelerate incident response.
    • Define and monitor SLIs, SLOs, SLAs, error budgets, latency, throughput, availability, and operational risk metrics.
    • Build, maintain, and improve operational dashboards for engineering, operations, cybersecurity, risk, and executive leadership.
    • Continuously enhance executive dashboards with service health, reliability trends, incident reporting, and operational performance metrics.
    • Partner with engineering, cloud, infrastructure, application, and cybersecurity teams to improve platform reliability.
    • Participate in incident response, root cause analysis, post-incident reviews, and problem management activities.
    • Automate operational tasks, health checks, reporting, deployment validation, and recovery procedures.
    • Support CI/CD pipelines, DevOps processes, release readiness, rollback validation, and production support activities.
    • Contribute to resiliency engineering initiatives including capacity planning, performance tuning, disaster recovery, failover testing, and resilience validation.
    • Ensure monitoring, dashboards, alerting, and operational processes comply with enterprise security, governance, and risk standards.
    • Perform other duties as assigned.

    Required Qualifications
    • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience.
    • 10+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Systems Engineering, Software Engineering, or Production Operations.
    • Experience supporting highly available, distributed, cloud-based, or mission-critical technology platforms.
    • Strong experience with monitoring, observability, logging, tracing, dashboards, and service health reporting.
    • Experience instrumenting applications, infrastructure, cloud services, APIs, databases, and distributed systems.
    • Strong understanding of SRE concepts including SLIs, SLOs, SLAs, error budgets, capacity planning, and incident management.
    • Experience designing meaningful monitoring and actionable alerting strategies.
    • Strong scripting or programming skills using Python, Java, Bash, PowerShell, or similar languages.
    • Experience with cloud platforms including AWS, Azure, or GCP.
    • Experience with Infrastructure as Code tools such as Terraform.
    • Experience supporting CI/CD pipelines, DevOps workflows, release management, and production environments.
    • Experience troubleshooting distributed systems, REST APIs, event-driven architectures, messaging platforms, and service integrations.
    • Familiarity with relational and NoSQL databases such as PostgreSQL, Microsoft SQL Server, MongoDB, or similar technologies.
    • Strong analytical, troubleshooting, problem-solving, and communication skills.

    Preferred Qualifications
    • Experience supporting cybersecurity, risk, resilience, or enterprise security platforms.
    • Experience with observability platforms such as Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch, or OpenTelemetry.
    • Experience building executive-level dashboards, operational scorecards, and reliability reporting.
    • Experience with automated health checks, synthetic monitoring, dependency mapping, and operational runbooks.
    • Experience with Kubernetes, Docker, serverless platforms, or cloud-native technologies.
    • Experience with Apache Kafka or other messaging platforms.
    • Familiarity with cloud security governance tools such as Azure Policy, AWS SCP, Wiz, Prisma, CloudGuard, or similar solutions.
    • Experience with AI cloud platforms such as Azure AI, AWS Bedrock, or Google Vertex AI.
    • Experience supporting Linux and Windows environments through scripting and automation.

    Required Skills
    1. Site Reliability Engineering (SRE)
    2. DevOps
    3. Cloud Platforms (AWS, Azure, GCP)
    4. Infrastructure as Code (Terraform)
    5. Monitoring & Observability
    6. Logging & Tracing
    7. Alerting
    8. Dashboard Development
    9. Incident Management
    10. Root Cause Analysis
    11. Service Reliability
    12. SLIs / SLOs / SLAs
    13. Performance Monitoring
    14. Capacity Planning
    15. Disaster Recovery
    16. CI/CD
    17. Python
    18. Java
    19. Bash
    20. PowerShell
    21. REST APIs
    22. Distributed Systems
    23. Microservices
    24. SQL Databases
    25. NoSQL Databases
    26. Automation
    27. Troubleshooting
    28. Cloud Infrastructure
    29. Operational Excellence
    30. Communication Skills
    31. Problem Solving

    Preferred Skills
    1. Splunk
    2. Grafana
    3. Prometheus
    4. Datadog
    5. Dynatrace
    6. New Relic
    7. Azure Monitor
    8. AWS CloudWatch
    9. OpenTelemetry
    10. Kubernetes
    11. Docker
    12. Apache Kafka
    13. Wiz
    14. Prisma Cloud
    15. CloudGuard
    16. Azure AI
    17. AWS Bedrock
    18. Google Vertex AI
    19. Linux Administration
    20. Windows Administration
    21. Executive Dashboard Reporting
    22. Synthetic Monitoring
    23. Service Dependency Mapping
    24. Cloud Security
    25. Cybersecurity Platforms

    Education
    Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience.

    Certifications
    Relevant cloud, DevOps, SRE, Kubernetes, or cybersecurity certifications are preferred.

.

.

.