# Senior Product Manager, Observability

[Nscaleoperationsukltd](https://gurify.com/jobs?q=Nscaleoperationsukltd) · United Kingdom · Posted last week

Senior

[Product Manager](https://gurify.com/jobs/product-manager)

[Apply on the original posting → (opens in a new tab)](https://job-boards.eu.greenhouse.io/nscaleoperationsukltd/jobs/4957117101)

## Job description

### About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

### About the role

Technical Product Managers at Nscale own the definition, delivery, and ongoing evolution of a slice of the Nscale platform, partnering with engineering, design, and go-to-market to turn customer and operational problems into shippable outcomes. As a Senior Technical Product Manager for Observability, you own the platform that gives customers and internal operators real-time visibility into their GPU fleet: the telemetry pipeline that scrapes data from physical infrastructure, the aggregation and storage layer, and the observability surfaces (logs, metrics, and traces) that enable fleet management, incident response, and alerting at scale. You partner daily with Fleet Software, Network Engineering, Data Centre Operations, and customer teams to make fleet health visible, actionable, and reliable as Nscale scales from a handful of deployments to a globally distributed fleet.

### What you'll be doing

- Own the roadmap for Nscale's observability platform: the telemetry pipeline, log and metrics aggregation, trace collection, and customer facing APIs and dashboards that surface fleet health to customers and operators.

- Define how logs, metrics, and traces are captured from physical infrastructure, aggregated, and surfaced through the observability platform to enable customers to manage their fleet and handle incidents.

- Own alerting strategy and optimisation: define what matters, reduce noise, and ensure the right signal reaches the right person at the right time.

- Capture and prioritise new telemetry requirements as the fleet scales, working with engineering to extend coverage across new hardware, sites, and deployment types.

- Shadow incident reviews and site operations to turn recurring manual effort and visibility gaps into platform capabilities.

- Define and drive the metrics that matter: alert signal-to-noise ratio, time-to-detect, time-to-resolve, telemetry coverage, and platform reliability.

- Mentor junior PMs and raise the bar for PRDs, reviews, and product decisions across the team.

### What you need

- 5–8 years in product management, with a track record owning significant areas in observability, infrastructure, or operations-facing products.

- Demonstrated experience building observability stacks: you have owned a product that captures and surfaces logs, metrics, and traces at scale, and you understand the architectural and UX tradeoffs involved.

- Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry.

- Experience with deployment tooling in a data centre or infrastructure context, including provisioning workflows, networking automation, or zero-touch deployment pipelines.

- Experience building for operators and delivery teams (design engineers, project controllers, PMs, SREs, DC technicians) and a genuine appetite for their workflows.

- Strong technical fluency: you can lead architecture and trade-off discussions across telemetry pipelines, time-series storage, alerting systems, and observability integrations.

- A record of moving ambiguous operational problems to shipped outcomes that measurably improve visibility, incident response, or fleet reliability.

- Excellent written and verbal communication across engineers, operators, and executives.

### Nice to haves

- Broader observability problem domain experience across different toolsets beyond the above stack.

- Familiarity with bare-metal provisioning tools (OpenStack Ironic, MAAS, or similar) or network automation tooling (NetBox, Nautobot, or similar).Degree in CS or engineering, or prior experience as an engineer, SRE, or infrastructure operator.

- Familiarity with GPU or accelerated compute infrastructure, data centre operations, or hyperscaler-style deployment at scale.

- ITSM: Jira Service Management, ServiceNow, Zendesk, or Freshservice.

- Experience in high-growth environments where the product is being built alongside the fleet it monitors.

Join Nscale as we build a world-class AI cloud platform. If you're excited about owning the software that turns contracts into live GPU capacity, we'd love to hear from you!

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

**Live in Nscaleoperationsukltd’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- NS [Senior Product Manager, Platform](https://gurify.com/job/senior-product-manager-platform-at-nscaleoperationsukltd-c6b7ce3df4c4) Nscaleoperationsukltd · United Kingdom · 3 weeks ago
- MU [Senior Product Manager](https://gurify.com/job/senior-product-manager-at-mubi-d57d2d5584bd) Mubi · London · last week
- KR [Senior Product Manager - Charging & Tax](https://gurify.com/job/senior-product-manager-charging-tax-at-krakentech-ee9960aaf779) Krakentech · London, United Kingdom · 5 days ago
- SP [Product Manager - Customer Service Platform](https://gurify.com/job/product-manager-customer-service-platform-at-spotify-afeb641513bd) Spotify · London · last week
- BR [AI Product Manager](https://gurify.com/job/ai-product-manager-at-brunswickgroup-a2fbb1842397) Brunswickgroup · London, United Kingdom · 5 days ago
- 9F [Senior AI Product Manager](https://gurify.com/job/senior-ai-product-manager-at-9fin-4b575d21cb36) 9FIN · London · last week

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Senior Product Manager, Observability","description":"\u003Ch3\u003EAbout Nscale\u003C/h3\u003E\u003Cp\u003ENscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today\u0026#39;s AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you\u0026#39;ll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.\u003C/p\u003E\u003Ch3\u003EAbout the role\u003C/h3\u003E\u003Cp\u003ETechnical Product Managers at Nscale own the definition, delivery, and ongoing evolution of a slice of the Nscale platform, partnering with engineering, design, and go-to-market to turn customer and operational problems into shippable outcomes. As a Senior Technical Product Manager for Observability, you own the platform that gives customers and internal operators real-time visibility into their GPU fleet: the telemetry pipeline that scrapes data from physical infrastructure, the aggregation and storage layer, and the observability surfaces (logs, metrics, and traces) that enable fleet management, incident response, and alerting at scale. You partner daily with Fleet Software, Network Engineering, Data Centre Operations, and customer teams to make fleet health visible, actionable, and reliable as Nscale scales from a handful of deployments to a globally distributed fleet.\u003C/p\u003E\u003Ch3\u003EWhat you\u0026#39;ll be doing\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EOwn the roadmap for Nscale\u0026#39;s observability platform: the telemetry pipeline, log and metrics aggregation, trace collection, and customer facing APIs and dashboards that surface fleet health to customers and operators.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EDefine how logs, metrics, and traces are captured from physical infrastructure, aggregated, and surfaced through the observability platform to enable customers to manage their fleet and handle incidents.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EOwn alerting strategy and optimisation: define what matters, reduce noise, and ensure the right signal reaches the right person at the right time.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECapture and prioritise new telemetry requirements as the fleet scales, working with engineering to extend coverage across new hardware, sites, and deployment types.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EShadow incident reviews and site operations to turn recurring manual effort and visibility gaps into platform capabilities.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EDefine and drive the metrics that matter: alert signal-to-noise ratio, time-to-detect, time-to-resolve, telemetry coverage, and platform reliability.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EMentor junior PMs and raise the bar for PRDs, reviews, and product decisions across the team.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003EWhat you need\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003E5\u20138 years in product management, with a track record owning significant areas in observability, infrastructure, or operations-facing products.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EDemonstrated experience building observability stacks: you have owned a product that captures and surfaces logs, metrics, and traces at scale, and you understand the architectural and UX tradeoffs involved.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience with deployment tooling in a data centre or infrastructure context, including provisioning workflows, networking automation, or zero-touch deployment pipelines.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience building for operators and delivery teams (design engineers, project controllers, PMs, SREs, DC technicians) and a genuine appetite for their workflows.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong technical fluency: you can lead architecture and trade-off discussions across telemetry pipelines, time-series storage, alerting systems, and observability integrations.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EA record of moving ambiguous operational problems to shipped outcomes that measurably improve visibility, incident response, or fleet reliability.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExcellent written and verbal communication across engineers, operators, and executives.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003ENice to haves\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EBroader observability problem domain experience across different toolsets beyond the above stack.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with bare-metal provisioning tools (OpenStack Ironic, MAAS, or similar) or network automation tooling (NetBox, Nautobot, or similar).Degree in CS or engineering, or prior experience as an engineer, SRE, or infrastructure operator.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with GPU or accelerated compute infrastructure, data centre operations, or hyperscaler-style deployment at scale.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EITSM: Jira Service Management, ServiceNow, Zendesk, or Freshservice.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience in high-growth environments where the product is being built alongside the fleet it monitors.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EJoin Nscale as we build a world-class AI cloud platform. If you\u0026#39;re excited about owning the software that turns contracts into live GPU capacity, we\u0026#39;d love to hear from you!\u003C/p\u003E\u003Cp\u003EAt Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ\u002B community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.\u003C/p\u003E\u003Cp\u003EIf there\u2019s anything we can do to accommodate your specific situation, please let us know.\u003C/p\u003E\u003Cp\u003EThe responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.\u003C/p\u003E\u003Cp\u003EFor information on how Nscale handles candidate personal data, please see our Employee \u0026amp; Candidate Privacy Notice: Here.\u003C/p\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"senior-product-manager-observability-at-nscaleoperationsukltd-8185b8d7cfff"},"url":"https://gurify.com/job/senior-product-manager-observability-at-nscaleoperationsukltd-8185b8d7cfff","datePosted":"2026-09-01","validThrough":"2026-10-29T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Nscaleoperationsukltd","sameAs":"https://job-boards.eu.greenhouse.io/nscaleoperationsukltd"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"GB"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"United Kingdom","item":"https://gurify.com/jobs/united-kingdom"},{"@type":"ListItem","position":3,"name":"Senior Product Manager, Observability","item":"https://gurify.com/job/senior-product-manager-observability-at-nscaleoperationsukltd-8185b8d7cfff"}]}
```
