# Senior Site Reliability Engineer

[Heidihealth Com Au](https://gurify.com/jobs?q=Heidihealth%20Com%20Au) · London, United Kingdom · Added 4 months ago

Senior

[DevOps](https://gurify.com/jobs/devops)

[Apply on the original posting → (opens in a new tab)](https://jobs.ashbyhq.com/heidihealth.com.au/5a0d7a56-bfd0-443d-8123-29db65679a59)

## Job description

Senior Site Reliability Engineer
Company: Heidi
Location: London, London, United Kingdom / Melbourne, Victoria, Australia
Type: FULL_TIME

Who We Are
Healthcare needs a better rhythm: one that keeps care continuous and deeply human. Heidi is building an AI Care Partner that works alongside clinicians to make that possible.

We’re a team of doctors, engineers, designers, researchers, and creatives building tools that help clinicians stay focused on what matters most: their patients.

In just 18 months, Heidi has given back more than 18 million hours to healthcare professionals - supporting 73 million patient visits in 116 countries. Today, more than two million patient visits each week are powered by Heidi worldwide.

Backed by nearly $100 million in funding, we’re growing in the US, UK, Canada, and Europe, partnering with leading health systems including the NHS, Beth Israel Lahey Health, and Monash Health.

The Role
This role sits in the core Platform/SRE team that owns production. You’ll work directly on incident response, on-call, system reliability, and day-to-day operations for Heidi’s platform.

We’re open to candidates who are strong mid-level SREs ready to take on more ownership, as well as senior SREs who enjoy being hands-on in operations. The role is intentionally ops-heavy and focused on keeping real systems healthy in production.

What you’ll do
- Participate in on-call and incident response: Respond to production incidents, contribute to service restoration, and support clear communication during incidents. Over time, take increasing responsibility for leading incidents end-to-end.

- Improve operational reliability: Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements.

- Own parts of the production environment: Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing ownership as familiarity increases.

- Strengthen observability: Improve dashboards, alerts, logs, and traces so issues are detected earlier and diagnosed faster, with a strong focus on actionable signals.

- Reduce operational toil: Automate repetitive tasks, simplify runbooks, and improve tooling to make on-call and day-to-day operations easier and safer.

- Support safe change: Improve deployments, rollback mechanisms, and operational readiness to reduce the risk of incidents caused by change.

- Contribute to operational practices: Write and maintain runbooks, participate in blameless post-mortems, and help improve incident response processes over time.

- Collaborate closely with engineers: Work with product and feature teams to improve production readiness, service ownership, and reliability expectations.

What we’re looking for
- 3–6+ years in SRE, DevOps, Platform, or operations-heavy engineering roles.

- Experience supporting production systems and participating in on-call rotations.

- Comfortable debugging live systems under pressure.

- Experience operating cloud infrastructure (AWS preferred).

- Working knowledge of Kubernetes and containerised workloads.

- Infrastructure as Code experience (Terraform or similar).

- Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc).

- Scripting or automation experience (Python, Bash, or similar).

Nice to have:
- Experience leading incidents or mentoring others during on-call.

- Experience in regulated or security-sensitive environments.

- Familiarity with databases, queues, and caches in production.

- Interest in reliability practices such as SLOs, error budgets, and capacity planning.

How We Work
- We own production: The Platform/SRE team is responsible for reliability and incident response.

- Incidents are blameless: We focus on learning and improving systems, not assigning fault.

- Practical over perfect: We prioritise improvements that reduce real operational pain.

- Calm under pressure: Clear thinking and communication matter during incidents.

What do we believe in?
Heidi builds for the future of healthcare, not just the next quarter, and our goals are ambitious because the world’s health demands it. We believe in progress built through precision, pace, and ownership.

- Live Forever - Every release moves care forward: measured, safe, and built to last. Data guides us, but patients define the truth that matters.

- Practice Ownership - Decisions follow logic and proof, not hierarchy. Exceptional care demands exceptional standards in our work, our thinking, and our character.

- Small Cuts Heal Faster - Stability earns trust, speed delivers impact. Progress is about learning fast without breaking what people depend on.

- Make others better - Feedback is direct, kindness is constant, and excellence lifts everyone. Our success is measured by collective growth, not individual output.

Our mission is clear: expand the world’s capacity to care, and do it without losing the humanity that makes care worth delivering.

Why you should join Heidi 🚀
- Real product momentum. We’re not trying to generate interest, we’re channeling it.

- Equity from day one. When Heidi wins, you win. You’ll share directly in the success you help create.

- Unmatched impact. Play a pivotal role in defining and scaling customer success at a critical growth moment - all while working on a product that delivers tangible value to clinicians and patients every day.

- Work alongside world-class talent. Join a team of operators and builders who’ve scaled unicorns.

- Global reach. Help shape our international expansion as we bring Heidi to key international markets.

- Growth and balance. Enjoy a personal development budget, work from anywhere for a month, dedicated wellness days, and your birthday off to recharge.

- Flexibility that works. A hybrid environment, with 3 days in the office.

Heidi’s commitment to Diversity, Equity and Inclusion
Heidi is dedicated to creating an equitable, inclusive, and supportive work environment that brings people together from diverse backgrounds, experiences, and perspectives. Our strength is in our differences. We're proud to be an equal opportunity employer and are proud to welcome all applicants as we're committed to promoting a culture of opportunity for all.

**Live in Heidihealth Com Au’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- LO [Senior DevOps Engineer](https://gurify.com/job/senior-devops-engineer-at-localstack-08c24eb4c67f) Localstack · United Kingdom · 3 weeks ago
- LD [Lead Devops Engineer - Waracle](https://gurify.com/job/lead-devops-engineer-waracle-3366c71a88fb) Dundee, United Kingdom · 2 weeks ago
- PA [Forward Deployed Infrastructure Engineer, New Grad](https://gurify.com/job/forward-deployed-infrastructure-engineer-new-grad-at-palantir-5c2cdf2ccee8) Palantir · London, United Kingdom · 2 weeks ago
- EL [Engineering Manager - DevOps & AI Platform](https://gurify.com/job/engineering-manager-devops-ai-platform-at-elliptic-542e7113af20) Elliptic · London, United Kingdom · last week
- BL [DevOps Engineer (FTC)](https://gurify.com/job/devops-engineer-ftc-at-bluecrestcapitalmanagement-e0a576aea185) Bluecrestcapitalmanagement · London, United Kingdom · 3 weeks ago
- BL [AI DevOps Engineer](https://gurify.com/job/ai-devops-engineer-at-bluecrestcapitalmanagement-b926ee977b64) Bluecrestcapitalmanagement · London, United Kingdom · 3 weeks ago

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Senior Site Reliability Engineer","description":"\u003Cp\u003ESenior Site Reliability Engineer \u003Cbr /\u003ECompany: Heidi\u003Cbr /\u003ELocation: London, London, United Kingdom / Melbourne, Victoria, Australia\u003Cbr /\u003EType: FULL_TIME\u003C/p\u003E\u003Cp\u003EWho We Are\u003Cbr /\u003EHealthcare needs a better rhythm: one that keeps care continuous and deeply human. Heidi is building an AI Care Partner that works alongside clinicians to make that possible.\u003C/p\u003E\u003Cp\u003EWe\u2019re a team of doctors, engineers, designers, researchers, and creatives building tools that help clinicians stay focused on what matters most: their patients.\u003C/p\u003E\u003Cp\u003EIn just 18 months, Heidi has given back more than 18 million hours to healthcare professionals - supporting 73 million patient visits in 116 countries. Today, more than two million patient visits each week are powered by Heidi worldwide.\u003C/p\u003E\u003Cp\u003EBacked by nearly $100 million in funding, we\u2019re growing in the US, UK, Canada, and Europe, partnering with leading health systems including the NHS, Beth Israel Lahey Health, and Monash Health.\u003C/p\u003E\u003Cp\u003EThe Role\u003Cbr /\u003EThis role sits in the core Platform/SRE team that owns production. You\u2019ll work directly on incident response, on-call, system reliability, and day-to-day operations for Heidi\u2019s platform.\u003C/p\u003E\u003Cp\u003EWe\u2019re open to candidates who are strong mid-level SREs ready to take on more ownership, as well as senior SREs who enjoy being hands-on in operations. The role is intentionally ops-heavy and focused on keeping real systems healthy in production.\u003C/p\u003E\u003Cp\u003EWhat you\u2019ll do\u003Cbr /\u003E- Participate in on-call and incident response: Respond to production incidents, contribute to service restoration, and support clear communication during incidents. Over time, take increasing responsibility for leading incidents end-to-end.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003EImprove operational reliability: Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EOwn parts of the production environment: Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing ownership as familiarity increases.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrengthen observability: Improve dashboards, alerts, logs, and traces so issues are detected earlier and diagnosed faster, with a strong focus on actionable signals.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EReduce operational toil: Automate repetitive tasks, simplify runbooks, and improve tooling to make on-call and day-to-day operations easier and safer.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESupport safe change: Improve deployments, rollback mechanisms, and operational readiness to reduce the risk of incidents caused by change.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EContribute to operational practices: Write and maintain runbooks, participate in blameless post-mortems, and help improve incident response processes over time.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECollaborate closely with engineers: Work with product and feature teams to improve production readiness, service ownership, and reliability expectations.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EWhat we\u2019re looking for\u003Cbr /\u003E- 3\u20136\u002B years in SRE, DevOps, Platform, or operations-heavy engineering roles.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003EExperience supporting production systems and participating in on-call rotations.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EComfortable debugging live systems under pressure.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience operating cloud infrastructure (AWS preferred).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWorking knowledge of Kubernetes and containerised workloads.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EInfrastructure as Code experience (Terraform or similar).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with monitoring and alerting tools (Datadog, Prometheus, etc).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EScripting or automation experience (Python, Bash, or similar).\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003ENice to have:\u003Cbr /\u003E- Experience leading incidents or mentoring others during on-call.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003EExperience in regulated or security-sensitive environments.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with databases, queues, and caches in production.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EInterest in reliability practices such as SLOs, error budgets, and capacity planning.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EHow We Work\u003Cbr /\u003E- We own production: The Platform/SRE team is responsible for reliability and incident response.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003EIncidents are blameless: We focus on learning and improving systems, not assigning fault.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EPractical over perfect: We prioritise improvements that reduce real operational pain.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECalm under pressure: Clear thinking and communication matter during incidents.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EWhat do we believe in?\u003Cbr /\u003EHeidi builds for the future of healthcare, not just the next quarter, and our goals are ambitious because the world\u2019s health demands it. We believe in progress built through precision, pace, and ownership.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003ELive Forever - Every release moves care forward: measured, safe, and built to last. Data guides us, but patients define the truth that matters.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EPractice Ownership - Decisions follow logic and proof, not hierarchy. Exceptional care demands exceptional standards in our work, our thinking, and our character.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESmall Cuts Heal Faster - Stability earns trust, speed delivers impact. Progress is about learning fast without breaking what people depend on.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EMake others better - Feedback is direct, kindness is constant, and excellence lifts everyone. Our success is measured by collective growth, not individual output.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EOur mission is clear: expand the world\u2019s capacity to care, and do it without losing the humanity that makes care worth delivering.\u003C/p\u003E\u003Cp\u003EWhy you should join Heidi \u0026#128640;\u003Cbr /\u003E- Real product momentum. We\u2019re not trying to generate interest, we\u2019re channeling it.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003EEquity from day one. When Heidi wins, you win. You\u2019ll share directly in the success you help create.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EUnmatched impact. Play a pivotal role in defining and scaling customer success at a critical growth moment - all while working on a product that delivers tangible value to clinicians and patients every day.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWork alongside world-class talent. Join a team of operators and builders who\u2019ve scaled unicorns.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EGlobal reach. Help shape our international expansion as we bring Heidi to key international markets.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EGrowth and balance. Enjoy a personal development budget, work from anywhere for a month, dedicated wellness days, and your birthday off to recharge.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFlexibility that works. A hybrid environment, with 3 days in the office.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EHeidi\u2019s commitment to Diversity, Equity and Inclusion\u003Cbr /\u003EHeidi is dedicated to creating an equitable, inclusive, and supportive work environment that brings people together from diverse backgrounds, experiences, and perspectives. Our strength is in our differences. We\u0026#39;re proud to be an equal opportunity employer and are proud to welcome all applicants as we\u0026#39;re committed to promoting a culture of opportunity for all.\u003C/p\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"senior-site-reliability-engineer-at-heidihealth-com-au-653420e72639"},"url":"https://gurify.com/job/senior-site-reliability-engineer-at-heidihealth-com-au-653420e72639","datePosted":"2026-05-07","validThrough":"2026-11-04T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Heidihealth Com Au","sameAs":"https://jobs.ashbyhq.com/heidihealth.com.au"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"GB","addressLocality":"London"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"United Kingdom","item":"https://gurify.com/jobs/united-kingdom"},{"@type":"ListItem","position":3,"name":"Senior Site Reliability Engineer","item":"https://gurify.com/job/senior-site-reliability-engineer-at-heidihealth-com-au-653420e72639"}]}
```
