# Tech, Experimentation & Failure

[Flightstory](https://gurify.com/jobs?q=Flightstory) · LONDON · Posted last week

[Product Manager](https://gurify.com/jobs/product-manager)

[Apply on the original posting → (opens in a new tab)](https://flightstory.teamtailor.com/jobs/8163083-senior-product-manager-tech-experimentation-failure)

## Job description

### SENIOR PRODUCT MANAGER - TECH, EXPERIMENTATION & FAILURE

### COMPANY: STEVEN.COM

### REPORTING TO: CITO

### LOCATION: LONDON

### ABOUT STEVEN.COM

Steven.com is building the operating system for the creator economy, forecast to pass a trillion dollars by the early 2030's. This industry is currently held back by fragmented tools and a lack of professional infrastructure. Steven.com is the unlock. We are the end-to-end Operating System designed to scale what is irreplaceably human.

We have built proprietary technology and obsessed teams to identify and scale the highest-potential creators across our core pillars:

- Creator Media: Amplifying reach, influence, and trust

- Creator Community: Transforming audiences into connected tribes

- Creator Products: Providing creators with the infrastructure to build and back aligned products and ventures

- Powered by Creator Tech & Intelligence: A proprietary data and technology suite that fuels smarter decisions and drives innovation across the entire flywheel.

Our Experimentation & Failure team has one of the most strategically important and unusual mandates at Steven.com: increase the rate of failure. As Steven Bartlett puts it: "the path to the correct answer is out-failing your competition." This isn't a growth team or an optimisation function. It's the team that exists to make sure we learn faster than anyone else.

### ROLE MISSION

Lead the Experimentation & Failure team, reporting to the CITO. You'll out-experiment and out-fail the competition - running high-velocity, rigorous experiments across every show, creator, piece of content, and commercial bet at Steven.com, while building a team and a culture that treats deliberate failure as the primary learning mechanism. This is deeply hands-on: you'll be setting hypotheses, isolating variables, checking statistical power, reading results, and moving to the next test - not directing from a distance.

### KEY OUTCOMES

- Own and drive the experimentation agenda across every show, creator, and IP property in the FlightStory portfolio - from podcast topic selection and episode structure through to thumbnail design and social tile copy. No detail too small to test.

- Lead hypothesis formation for every experiment, with clear success, failure, and inconclusive criteria defined in advance - and enforce single-variable discipline across the board.

- Ensure every experiment is adequately powered before launch: sample sizes calculated, measurement windows defined, results interpretable by design.

- Systematically increase experimentation velocity and build the intake process that makes high-volume testing the default across FlightStory and Steven.com.

- Build institutional memory - a searchable, structured record of every experiment run, what was learned, and what was decided - as a compounding organisational asset.

- Partner with the VP of Engineering & Applied AI to apply AI tooling to experiment design, analysis, and reporting, and to ensure infrastructure supports testing at this velocity.

### CORE COMPETENCIES

- Deep, first-principles command of experimentation mathematics: statistical significance, power, sample size calculation, p-values, confidence intervals, Type I/II errors, and the difference between statistical and practical significance.

- Genuine mastery of the scientific method applied to product and content - hypothesis formation, single-variable isolation, measurement design, and result interpretation.

- A track record of building experimentation culture, not just running tests - creating an environment where the whole team experiments and failure is rewarded.

- Comfortable querying data and working shoulder-to-shoulder with engineers and data scientists at implementation depth.

- Strong written and verbal communication - able to write a hypothesis an engineer respects and explain a result a producer will act on.

- Experience operating at pace, in high-volume testing environments where speed of learning is the competitive edge.

### YOU'LL THRIVE HERE IF

- You think in hypotheses, not features.

- You isolate one variable at a time, and understand viscerally why changing five things at once makes a result meaningless.

- You give a null or inconclusive result the same intellectual respect as a win - you know how to extract the signal either way.

- You're obsessive about measurement: an experiment that can't be measured is just a change, not an experiment.

- You're deeply sceptical of your own results, and design experiments to prove yourself wrong rather than confirm what you already believe.

- You move fast, expect others to move fast, and don't wait for perfect conditions to run a test.

- You believe failure is feedback, feedback is knowledge, and knowledge is power - and you build systems to generate that knowledge at the highest possible rate.

### IDEAL BACKGROUND

- Demonstrable experience leading (not just participating in) product or content experimentation programs at a technology, media, or creator-economy company - owning the methodology, volume, and culture.

- Strong advantage: experience with algorithmic platforms and how to design experiments against platform-specific metrics; podcasting, video, social, or creator-economy background; familiarity with YouTube/Spotify/social analytics (CTR, retention, watch time, audience behaviour).

- Nice to have: experience building an experimentation platform from scratch; familiarity with causal inference beyond standard A/B testing (holdouts, quasi-experiments, diff-in-diff); experience experimenting on AI/ML systems or prompt variations in production; a background in statistics, maths, CS, economics, or a natural science; exposure to early-stage environments where you had to build the experimentation infrastructure yourself.

**Live in Flightstory’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- EX [Senior Product Manager](https://gurify.com/job/senior-product-manager-at-exclaimer-223073822a52) Exclaimer · London, United Kingdom · last week
- EB [API Connectivity](https://gurify.com/job/api-connectivity-at-ebury-a67717482e7c) Ebury · London · 6 days ago
- SP [Principal Product Manager – Content Platform](https://gurify.com/job/principal-product-manager-content-platform-at-spotify-5a8889b85bc1) Spotify · London · 5 days ago
- PP [Principal Product Manager - Waracle](https://gurify.com/job/principal-product-manager-waracle-6356c74b39db) London, United Kingdom · last week
- AP [Associate Product Manager | London](https://gurify.com/job/associate-product-manager-london-e0ae787d9295) London, United Kingdom · last week
- CC [Product Manager, Payouts](https://gurify.com/job/product-manager-payouts-at-checkout-com-a0464cfa4daa) Checkout Com · London · last week

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Tech, Experimentation \u0026 Failure","description":"\u003Ch3\u003ESENIOR PRODUCT MANAGER - TECH, EXPERIMENTATION \u0026amp; FAILURE\u003C/h3\u003E\u003Ch3\u003ECOMPANY: STEVEN.COM\u003C/h3\u003E\u003Ch3\u003EREPORTING TO: CITO\u003C/h3\u003E\u003Ch3\u003ELOCATION: LONDON\u003C/h3\u003E\u003Ch3\u003EABOUT STEVEN.COM\u003C/h3\u003E\u003Cp\u003ESteven.com is building the operating system for the creator economy, forecast to pass a trillion dollars by the early 2030\u0026#39;s. This industry is currently held back by fragmented tools and a lack of professional infrastructure. Steven.com is the unlock. We are the end-to-end Operating System designed to scale what is irreplaceably human.\u003C/p\u003E\u003Cp\u003EWe have built proprietary technology and obsessed teams to identify and scale the highest-potential creators across our core pillars:\u003C/p\u003E\u003Cul\u003E\u003Cli\u003ECreator Media: Amplifying reach, influence, and trust\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECreator Community: Transforming audiences into connected tribes\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECreator Products: Providing creators with the infrastructure to build and back aligned products and ventures\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EPowered by Creator Tech \u0026amp; Intelligence: A proprietary data and technology suite that fuels smarter decisions and drives innovation across the entire flywheel.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EOur Experimentation \u0026amp; Failure team has one of the most strategically important and unusual mandates at Steven.com: increase the rate of failure. As Steven Bartlett puts it: \u0026quot;the path to the correct answer is out-failing your competition.\u0026quot; This isn\u0026#39;t a growth team or an optimisation function. It\u0026#39;s the team that exists to make sure we learn faster than anyone else.\u003C/p\u003E\u003Ch3\u003EROLE MISSION\u003C/h3\u003E\u003Cp\u003ELead the Experimentation \u0026amp; Failure team, reporting to the CITO. You\u0026#39;ll out-experiment and out-fail the competition - running high-velocity, rigorous experiments across every show, creator, piece of content, and commercial bet at Steven.com, while building a team and a culture that treats deliberate failure as the primary learning mechanism. This is deeply hands-on: you\u0026#39;ll be setting hypotheses, isolating variables, checking statistical power, reading results, and moving to the next test - not directing from a distance.\u003C/p\u003E\u003Ch3\u003EKEY OUTCOMES\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EOwn and drive the experimentation agenda across every show, creator, and IP property in the FlightStory portfolio - from podcast topic selection and episode structure through to thumbnail design and social tile copy. No detail too small to test.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ELead hypothesis formation for every experiment, with clear success, failure, and inconclusive criteria defined in advance - and enforce single-variable discipline across the board.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EEnsure every experiment is adequately powered before launch: sample sizes calculated, measurement windows defined, results interpretable by design.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESystematically increase experimentation velocity and build the intake process that makes high-volume testing the default across FlightStory and Steven.com.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBuild institutional memory - a searchable, structured record of every experiment run, what was learned, and what was decided - as a compounding organisational asset.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EPartner with the VP of Engineering \u0026amp; Applied AI to apply AI tooling to experiment design, analysis, and reporting, and to ensure infrastructure supports testing at this velocity.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003ECORE COMPETENCIES\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EDeep, first-principles command of experimentation mathematics: statistical significance, power, sample size calculation, p-values, confidence intervals, Type I/II errors, and the difference between statistical and practical significance.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EGenuine mastery of the scientific method applied to product and content - hypothesis formation, single-variable isolation, measurement design, and result interpretation.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EA track record of building experimentation culture, not just running tests - creating an environment where the whole team experiments and failure is rewarded.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EComfortable querying data and working shoulder-to-shoulder with engineers and data scientists at implementation depth.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong written and verbal communication - able to write a hypothesis an engineer respects and explain a result a producer will act on.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience operating at pace, in high-volume testing environments where speed of learning is the competitive edge.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003EYOU\u0026#39;LL THRIVE HERE IF\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EYou think in hypotheses, not features.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EYou isolate one variable at a time, and understand viscerally why changing five things at once makes a result meaningless.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EYou give a null or inconclusive result the same intellectual respect as a win - you know how to extract the signal either way.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EYou\u0026#39;re obsessive about measurement: an experiment that can\u0026#39;t be measured is just a change, not an experiment.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EYou\u0026#39;re deeply sceptical of your own results, and design experiments to prove yourself wrong rather than confirm what you already believe.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EYou move fast, expect others to move fast, and don\u0026#39;t wait for perfect conditions to run a test.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EYou believe failure is feedback, feedback is knowledge, and knowledge is power - and you build systems to generate that knowledge at the highest possible rate.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003EIDEAL BACKGROUND\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EDemonstrable experience leading (not just participating in) product or content experimentation programs at a technology, media, or creator-economy company - owning the methodology, volume, and culture.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong advantage: experience with algorithmic platforms and how to design experiments against platform-specific metrics; podcasting, video, social, or creator-economy background; familiarity with YouTube/Spotify/social analytics (CTR, retention, watch time, audience behaviour).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ENice to have: experience building an experimentation platform from scratch; familiarity with causal inference beyond standard A/B testing (holdouts, quasi-experiments, diff-in-diff); experience experimenting on AI/ML systems or prompt variations in production; a background in statistics, maths, CS, economics, or a natural science; exposure to early-stage environments where you had to build the experimentation infrastructure yourself.\u003C/li\u003E\u003C/ul\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"tech-experimentation-failure-at-flightstory-3a6904188aa6"},"url":"https://gurify.com/job/tech-experimentation-failure-at-flightstory-3a6904188aa6","datePosted":"2026-08-03","validThrough":"2026-09-26T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Flightstory","sameAs":"https://flightstory.teamtailor.com"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"GB","addressLocality":"LONDON"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"United Kingdom","item":"https://gurify.com/jobs/united-kingdom"},{"@type":"ListItem","position":3,"name":"Tech, Experimentation \u0026 Failure","item":"https://gurify.com/job/tech-experimentation-failure-at-flightstory-3a6904188aa6"}]}
```
