All jobs

Coalfire

Senior Site Reliability Engineer (Fedramp)

United States · Remote

About this role

About Coalfire Coalfire is on a mission to make the world a safer place by solving our clients’ hardest cybersecurity challenges. We work at the cutting edge of technology to advise, assess, automate, and ultimately help companies navigate the ever-changing cybersecurity landscape. We are headquartered in Chicago, Illinois with offices across the U.S. and U.K., and we support clients around the world. But that’s not who we are – that’s just what we do. We are thought leaders, consultants, and cybersecurity experts, but above all else, we are a team of passionate problem-solvers who are hungry to learn, grow, and make a difference. Why Join Us FedRAMP now defines compliance as a set of Key Security Indicators (KSIs); measurable security outcomes that automation validates continuously. Holding an estate at that standard day after day is an operations job. Coalfire organizes its delivery engineering into capability-focused teams, and the Run teams keep authorized client environments available and provably compliant long after the build team leaves. As a Senior Site Reliability Engineer you own one operational capability, such as remidation, monitoring and alerting or backup and recovery, and you design how Coalfire observes a client estate and holds it in a compliant state. You own the automation and the procedure that each managed environment inherits, you carry escalation for the hardest operational problems, and you set and continually raise the engineering bar for the engineers around you. If you are driven by a desire to innovate, excel at operational excellence, and thrive in a collaborative environment, come be part of a team committed to making the world a more secure and complaint place. What You'll Do Own one operational capability for the managed estate, such as monitoring and alerting or backup and recovery, including its automation, its runbooks, and the service standard each managed environment inherits. Design how Coalfire observes a regulated cloud environment: telemetry and log pipelines, service-level objectives, alert quality, and the escalation paths behind them. Make alerts useful and actionable. Build and maintain the continuous-monitoring evidence pipeline so you can prove the authorized state each day, in machine-readable form where the framework allows it. Own backup and recovery engineering for your capability: recovery procedure you have tested, recovery objectives you can measure, and the automation that runs them during an outage. Serve as the senior escalation point in client environments. Diagnose and troubleshoot beyond the runbook, resolve the incident, then fix the runbook so the next engineer on call does not repeat the diagnosis. Lead incident and problem management for your domain, including incident command on major events, blameless post-incident review, and corrective action that closes the problem out. Automate operational toil out of the estate with infrastructure-as-code, pipelines, and scripting. Decide what Coalfire automates once for the whole estate and what stays specific to one client. Partner with Engagement Architects and the Build teams on transition into managed operations: operational readiness review, monitoring and runbook coverage, and the service commitments Coalfire can meet at go-live. Hold on-call for your capability and improve the rotation: coverage, alert actionability, and the load it places on the team. Represent operational posture in front of clients, and support renewals and expansions with a credible account of what Coalfire operates for them and at what cost. Mentor and lead Site Reliability Engineers and junior staff: review designs as well as changes, set operational standards, and grow named individuals in your domain. Author and peer review code, runbooks, operational design documentation, and the compliance artifacts that evidence continuous monitoring, inclusive of vendor best practices. What You'll Bring Automation-first mindset with deep Infrastructure

Skills and categories

Site-Reliability-EngineeringFedRAMP-ComplianceCloud-OperationsObservability-EngineeringInfrastructure-EngineeringSenior-Site-Reliability-EngineerFedRAMP-Security-EngineerFedRAMP-Compliance-EngineerSenior-Site-Reliability-Engineering-ArchitectSite-Reliability-EngineerDevOps-Site-Reliability-EngineerCybersecurity-Site-Reliability-EngineerSite-Reliability-Operations-EngineerSite-Reliability-Engineer-IISenior-FedRAMP-ConsultantDeveloper

Listing provided by Himalayas. Verify availability on the source board before applying — Kartavyam aggregates public listings as they appear and does not manage this application.

Similar roles

Browse all jobs