All jobs

Merative

Principle SRE

United States · Remote

About this role

Merge posting Verbiage (add as first line in the Job Summary) : Merge medical imaging solutions, offered by Merative , combine intelligent, scalable imaging workflow tools with deep and broad expertise to help healthcare organizations improve their confidence in patient outcomes and optimize care delivery. With minimal supervision, leverages deep technical expertise to set the reliability, observability, and operational automation direction for business-critical, PHI-bearing cloud platforms. Establishes reliability standards, service level objectives, and automation practices across multiple teams, and leads their implementation through technical influence rather than direct reporting authority. Responsibilities: People Providetechnical guidance, mentorship, and leadership to engineers across development, QA, and operations teams. Interact regularly with lower and/or senior management on matters concerning multiple functional areas, departments, and/or customers. Act as a senior point of escalation for complex or high-severity production issues. Build reliability capability in others through design reviews, pairing, and blameless post-incident learning. Foster a culture and develop approaches that generate innovative ideas, products, and services. Reliability Engineering Define, publish, and govern service levelobjectives(SLOs)management framework,help to defineservice level indicators, and error budgets for the platform and its shared services. Set observability standards for metrics, logging, tracing, and alerting, and drive consistent adoption across teams. Own the technical side of the incident lifecycle: detection, response, escalation, and post-incident review; drive corrective actions to closure. Lead capacity planning, performance engineering, failure-domain isolation, and disaster recovery design. Drive resilience patterns into product architecture in partnership with development teams. Monitor and act on incoming issues from support, customers, or other stakeholders. Operational Automation Design and Implementation Design the automation strategy for platform operations and set the standards, patterns, and tooling other engineers build against. Design and implement automated remediation for recurring failure modes so routine faults are resolved without human intervention. Build andmaintaininfrastructure-as-code, environment provisioning, and deployment automation. Automate recurring operational work including patching, scaling, certificate rotation, backup and restore validation, and disaster recovery exercises. Build self-service tooling that lets development and support teams perform routine operational tasks safely and without escalation. Automate the collection of evidence and the verification of security and compliance controls that apply to PHI-bearing workloads. Identify, measure, and reduce operational toil; set and report against measurable toil-reduction targets. Process Provide guidance on company processes. Liaise with cross functional teams (development, product, program management, support, implementation, security, etc.) in delivering and supporting their projects. Support cross functional teams in resolving customer concerns. Participate in the creation and/or review and/or approval of architecture, design, and project documents. Plan, track, and deliver assigned reliability and automation initiatives, on-timeand on-budget. Report initiative status andimmediatelyescalate when work is varying from commitments. Effectively represent the platform’s reliability posture to the customer and to auditors if/as needed. Adhere to Mergemethodology, quality system requirements, and good engineering practices. Provide input into applicable budgets, including cloud consumption and tooling spend. Pursue self-development as an employee to be better at their current role as well as to grow into their next role if/asappropriate. Adhere to all applicable legal requirements. Core Competencies Organized : Demonstrates stro

Skills and categories

Site-Reliability-EngineeringPrincipal-Site-Reliability-EngineerCloud-Infrastructure-EngineeringPlatform-EngineeringDevOps-EngineerSenior-SREDeveloper

Listing provided by Himalayas. Verify availability on the source board before applying — Kartavyam aggregates public listings as they appear and does not manage this application.

Similar roles

Browse all jobs