# Data Engineer

[Solveintelligence](https://gurify.com/jobs?q=Solveintelligence) · London · Posted yesterday

[Data](https://gurify.com/jobs/data)

[Apply on the original posting → (opens in a new tab)](https://jobs.ashbyhq.com/solveintelligence/7e42954a-cb4a-4aa4-a317-fb0a23ded383)

## Job description

💡 ABOUT US

We're the fastest-growing startup transforming the IP industry.

- Traction: 20-30% MoM revenue growth; selling to 700+ global IP teams (DLA Piper, tech boutiques, and global enterprises).

- Proven Value: Users report 50-90% efficiency gains using our AI platform.

- Backing: Recently featured in Sifted following our $40M Series B announcement, bringing our total funding to $55M from elite investors including Y Combinator, 20VC, Visionaries, Microsoft and Thomson Reuters.

🏗️ ABOUT THE ROLE

We’re hiring a data engineer to build the ingestion and search systems behind Solve Intelligence’s AI products.

Our sources include global patent literature, case law, technical standards and contributions, scientific databases, academic papers, and content from across the web. The data spans structured records, documents, images, audio and video. You’ll work across bulk ingestion and on-demand retrieval, making this information searchable and useful in our products

You’ll own systems from source acquisition through to serving queries. The work includes:

- Large-scale ingestion. Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures.

- Document processing and data quality. Extract useful content from complex documents and other formats. Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata.

- Search and serving. Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records.

- Connecting information across sources. Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources.

- Performance engineering. Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale. Diagnose CPU, memory and storage I/O bottlenecks, and tune jobs and infrastructure for throughput, latency and cost.

You’ll work closely with our AI researchers and product engineers, with substantial freedom to choose the approach and build the systems yourself.

🛠️ WHAT YOU BRING

### Must haves:

- Strong Python and SQL, with experience designing and operating production databases.

- Solid experience building and operating production data pipelines over large, messy datasets.

- Expertise with running search systems over large document collections.

- End-to-end ownership from raw data to user-facing functionality.

- A good understanding of schema design, indexing and query optimisation.

- A track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement.

### Nice to Have:

- Experience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C++ is useful.

👋 THE FOUNDERS

You'll partner with a founding team of AI PhDs and elite systems engineers:

- Sanj (CRO): PhD in AI (Gatsby Unit, UCL), ex-Huawei R&D, former lead at Magic Carpet AI (acquired).

- Chris (CEO): PhD in AI (UCL), published researcher, ex-Dyson and Alan Turing Institute.

- Angus (CTO): MEng Computer Science, ex-Qualcomm and Coremont (Brevan Howard).

### WHAT WE OFFER

- Competitive Salary + Significant Equity: We want you to have true ownership in the success of the company.

- Founding Impact: You'll have a direct hand in how we build out the data infrastructure the rest of the product depends on.

- Support: Full visa sponsorship and private medical insurance.

- The Environment: Free meals and a seat at the table with an incredibly smart, ambitious team.

**Live in Solveintelligence’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- SU [Senior Data Engineer](https://gurify.com/job/senior-data-engineer-at-superpayments-c3938afbf460) Superpayments · London · last week
- PL [Senior Data Engineer (Commercial Analytics)](https://gurify.com/job/senior-data-engineer-commercial-analytics-at-pleo-20f49b17964f) Pleo · London · yesterday
- TR [Embedded Data Engineer - ML](https://gurify.com/job/embedded-data-engineer-ml-at-trainline-87a1c4df04fa) Trainline · London · last week
- LE [Junior Data Engineer](https://gurify.com/job/junior-data-engineer-at-lendable-d9d9cf24853a) Lendable · London · 2 days ago
- FO [Graduate Data Engineer](https://gurify.com/job/graduate-data-engineer-at-fosphamarketing-4b8b0ab6d1c9) Fosphamarketing · London · yesterday
- BL [Graduate Data Engineer](https://gurify.com/job/graduate-data-engineer-at-blenheimchalcot-9f20865a84ff) Blenheimchalcot · London · yesterday

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Data Engineer","description":"\u003Cp\u003E\u0026#128161; ABOUT US\u003C/p\u003E\u003Cp\u003EWe\u0026#39;re the fastest-growing startup transforming the IP industry.\u003C/p\u003E\u003Cul\u003E\u003Cli\u003ETraction: 20-30% MoM revenue growth; selling to 700\u002B global IP teams (DLA Piper, tech boutiques, and global enterprises).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EProven Value: Users report 50-90% efficiency gains using our AI platform.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBacking: Recently featured in Sifted following our $40M Series B announcement, bringing our total funding to $55M from elite investors including Y Combinator, 20VC, Visionaries, Microsoft and Thomson Reuters.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003E\u0026#127959;\uFE0F ABOUT THE ROLE\u003C/p\u003E\u003Cp\u003EWe\u2019re hiring a data engineer to build the ingestion and search systems behind Solve Intelligence\u2019s AI products.\u003C/p\u003E\u003Cp\u003EOur sources include global patent literature, case law, technical standards and contributions, scientific databases, academic papers, and content from across the web. The data spans structured records, documents, images, audio and video. You\u2019ll work across bulk ingestion and on-demand retrieval, making this information searchable and useful in our products\u003C/p\u003E\u003Cp\u003EYou\u2019ll own systems from source acquisition through to serving queries. The work includes:\u003C/p\u003E\u003Cul\u003E\u003Cli\u003ELarge-scale ingestion. Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EDocument processing and data quality. Extract useful content from complex documents and other formats. Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESearch and serving. Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EConnecting information across sources. Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EPerformance engineering. Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale. Diagnose CPU, memory and storage I/O bottlenecks, and tune jobs and infrastructure for throughput, latency and cost.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EYou\u2019ll work closely with our AI researchers and product engineers, with substantial freedom to choose the approach and build the systems yourself.\u003C/p\u003E\u003Cp\u003E\u0026#128736;\uFE0F WHAT YOU BRING\u003C/p\u003E\u003Ch3\u003EMust haves:\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EStrong Python and SQL, with experience designing and operating production databases.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESolid experience building and operating production data pipelines over large, messy datasets.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExpertise with running search systems over large document collections.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EEnd-to-end ownership from raw data to user-facing functionality.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EA good understanding of schema design, indexing and query optimisation.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EA track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003ENice to Have:\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EExperience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C\u002B\u002B is useful.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003E\u0026#128075; THE FOUNDERS\u003C/p\u003E\u003Cp\u003EYou\u0026#39;ll partner with a founding team of AI PhDs and elite systems engineers:\u003C/p\u003E\u003Cul\u003E\u003Cli\u003ESanj (CRO): PhD in AI (Gatsby Unit, UCL), ex-Huawei R\u0026amp;D, former lead at Magic Carpet AI (acquired).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EChris (CEO): PhD in AI (UCL), published researcher, ex-Dyson and Alan Turing Institute.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EAngus (CTO): MEng Computer Science, ex-Qualcomm and Coremont (Brevan Howard).\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003EWHAT WE OFFER\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003ECompetitive Salary \u002B Significant Equity: We want you to have true ownership in the success of the company.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFounding Impact: You\u0026#39;ll have a direct hand in how we build out the data infrastructure the rest of the product depends on.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESupport: Full visa sponsorship and private medical insurance.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EThe Environment: Free meals and a seat at the table with an incredibly smart, ambitious team.\u003C/li\u003E\u003C/ul\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"data-engineer-at-solveintelligence-ebf465f19936"},"url":"https://gurify.com/job/data-engineer-at-solveintelligence-ebf465f19936","datePosted":"2026-09-16","validThrough":"2026-11-01T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Solveintelligence","sameAs":"https://jobs.ashbyhq.com/solveintelligence"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"GB","addressLocality":"London"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"United Kingdom","item":"https://gurify.com/jobs/united-kingdom"},{"@type":"ListItem","position":3,"name":"Data Engineer","item":"https://gurify.com/job/data-engineer-at-solveintelligence-ebf465f19936"}]}
```
