# Data Engineer – Applied ML

[Similarweb](https://gurify.com/jobs?q=Similarweb) · Tel Aviv, Israel · Posted yesterday

[Data](https://gurify.com/jobs/data)

[AI & ML](https://gurify.com/jobs/ai-ml)

[Apply on the original posting → (opens in a new tab)](https://job-boards.greenhouse.io/similarweb/jobs/8239378)

## Job description

Similarweb is the leading digital intelligence platform used by over 3500 global customers. Our wide range of solutions power the digital strategies of companies like Google, eBay, and Adidas.

We help our customers succeed in today’s digital world by giving them access to data-driven insights, competitive benchmarks, strategic analysis, and more.

In 2021, we went public on the New York Stock Exchange, and we haven’t stopped growing since!

We’re looking for a Data Engineer with a strong applied ML focus to join our R&D department!

### Why is this role so important at Similarweb?

Our Retail Intelligence products help leading brands and retailers understand how their products, brands and categories perform online. Behind them is data collected from retailers and marketplaces around the world: product pages, brands and categories, each described differently by every site.

Your mission is to turn that data into a single, trusted view: classifying products into a unified taxonomy, normalizing brands and attributes, and matching the same entities across sources. And doing it at scale, across a catalog of more than a billion records that keeps growing and changing every day.

This is an applied ML role within data engineering. You’ll build with LLMs, agentic frameworks such as LangGraph, embeddings and classical ML, and ship them as production pipelines. It’s hands-on work, not research for its own sake, but it takes a real understanding of classification and NLP methods to choose the right tool for each problem and prove that it works.

### So, what will you be doing all day?

### Your daily responsibilities may include:

- Designing and building LLM-powered and ML-based pipelines that classify, normalize, structure and match product, brand and category data

- Building agentic workflows (LangGraph or similar) that automate complex data tasks end to end

- Choosing the right approach for each problem (LLMs, embeddings, fine-tuned models, classical classifiers or rules), balancing accuracy, cost and latency

- Scaling solutions to run efficiently over billions of records, using Spark, Databricks and our cloud infrastructure

- Building evaluation frameworks: ground-truth datasets, labeling processes, quality metrics and ongoing monitoring

- Taking solutions from POC to production, and owning them after launch

- Working closely with Product to define requirements and shape the roadmap

- Collaborating with data engineers, data scientists and other R&D teams on infrastructure and best practices

### This is the perfect job for someone who:

- Holds a B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics or another relevant field

- Has 4+ years of hands-on experience as a data engineer, ML engineer or data scientist, with solutions running in production

- Has strong Python skills and writes production-quality code

- Has hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings, evaluation)

- Has worked with the modern LLM stack: LLM provider APIs (OpenAI, Anthropic, etc.), LangGraph or LangChain, Hugging Face and vector stores

- Has a solid grasp of text classification and NLP methods, both classical and modern, and knows when to use each

- Has experience processing large-scale data with Spark/PySpark, Databricks or similar, on AWS or another cloud

- Understands evaluation and data quality well: precision/recall trade-offs, building ground truth, error analysis

- Is pragmatic and delivery-focused, comfortable with ambiguity, and communicates clearly with Product and business stakeholders

- Has experience with taxonomies, entity resolution or product/e-commerce data (advantage)

- Has experience with fine-tuning or deploying open-source models (advantage)

At Similarweb, collaborating with our colleagues in-office creates a more connected, unified culture. Our best work is a product of our face-to-face collaboration, with the ability to work partially from home.

### Why you’ll love being a Similarwebber:

You’ll actually love the product you work with: Our customers aren’t our only raving fans. When we asked our employees why they chose to come work at Similarweb, 99% of them said “the product.” Imagine how exciting your job is when you get to work with the most powerful digital intelligence platform in the world.

You’ll find a home for your big ideas: We encourage an open dialogue and empower employees to bring their ideas to the table. You’ll find the resources you need to take initiative and create meaningful change within the organization.

We offer competitive perks & benefits: We take your well-being seriously, and offer competitive compensation packages to all employees. We also put a strong emphasis on community, with regular team outings and happy hours.

You can grow your career in any direction you choose: Interested in becoming a VP or want to transition into a different department? Whether it’s Career Week, personalized coaching, or our ongoing learning solutions, you’ll find all the tools and opportunities you need to develop your career right here.

Diversity isn’t just a buzzword: People want to work in a place where they can be themselves. We strive to create a workplace that is reflective of the communities we serve, where everyone is empowered to bring their full, authentic selves to work. We are committed to inclusivity across race, gender, ethnicity, culture, sexual orientation, age, religion, spirituality, identity and experience. We believe our culture of equality and mutual respect also helps us better understand and serve our customers in a world that is becoming more global, more diverse, and more digital every day.

We will handle your application and information related to your application in accordance with the Applicant Privacy Policy available here: https://www.similarweb.com/corp/legal/applicant-privacy-policies/

We will handle your application and information related to your application in accordance with the Applicant Privacy Policy available here.

**Live in Similarweb’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- QU [Data Engineer](https://gurify.com/job/data-engineer-at-quicklizard-1abe5e1cf1a6) Quicklizard · Petah Tikva, Israel · today
- SI [Ops Data Engineer](https://gurify.com/job/ops-data-engineer-at-similarweb-3e6abc7a7b5f) Similarweb · Tel Aviv-Yafo, Israel · 6 weeks ago
- AL [Senior Data Engineer](https://gurify.com/job/senior-data-engineer-at-allcloud-e0332f70f766) Allcloud · Ra'anana, Israel · 3 weeks ago
- CI [Analytics Data Engineer](https://gurify.com/job/analytics-data-engineer-at-comm-it-9f06680c7994) Comm IT · Israel · 3 weeks ago
- TR [Data Engineer](https://gurify.com/job/data-engineer-at-trigo-ca95ed09ab0a) Trigo · Ramat Gan, Israel · 3 weeks ago
- VE [Data Lead- Full-Stack Data Engineer](https://gurify.com/job/data-lead-full-stack-data-engineer-at-venncity-02ec7b8acca3) Venncity · Tel Aviv District, Israel · 3 weeks ago

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Data Engineer \u2013 Applied ML","description":"\u003Cp\u003ESimilarweb is the leading digital intelligence platform used by over 3500 global customers. Our wide range of solutions power the digital strategies of companies like Google, eBay, and Adidas.\u003C/p\u003E\u003Cp\u003EWe help our customers succeed in today\u2019s digital world by giving them access to data-driven insights, competitive benchmarks, strategic analysis, and more.\u003C/p\u003E\u003Cp\u003EIn 2021, we went public on the New York Stock Exchange, and we haven\u2019t stopped growing since!\u003C/p\u003E\u003Cp\u003EWe\u2019re looking for a Data Engineer with a strong applied ML focus to join our R\u0026amp;D department!\u003C/p\u003E\u003Ch3\u003EWhy is this role so important at Similarweb?\u003C/h3\u003E\u003Cp\u003EOur Retail Intelligence products help leading brands and retailers understand how their products, brands and categories perform online. Behind them is data collected from retailers and marketplaces around the world: product pages, brands and categories, each described differently by every site.\u003C/p\u003E\u003Cp\u003EYour mission is to turn that data into a single, trusted view: classifying products into a unified taxonomy, normalizing brands and attributes, and matching the same entities across sources. And doing it at scale, across a catalog of more than a billion records that keeps growing and changing every day.\u003C/p\u003E\u003Cp\u003EThis is an applied ML role within data engineering. You\u2019ll build with LLMs, agentic frameworks such as LangGraph, embeddings and classical ML, and ship them as production pipelines. It\u2019s hands-on work, not research for its own sake, but it takes a real understanding of classification and NLP methods to choose the right tool for each problem and prove that it works.\u003C/p\u003E\u003Ch3\u003ESo, what will you be doing all day?\u003C/h3\u003E\u003Ch3\u003EYour daily responsibilities may include:\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EDesigning and building LLM-powered and ML-based pipelines that classify, normalize, structure and match product, brand and category data\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBuilding agentic workflows (LangGraph or similar) that automate complex data tasks end to end\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EChoosing the right approach for each problem (LLMs, embeddings, fine-tuned models, classical classifiers or rules), balancing accuracy, cost and latency\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EScaling solutions to run efficiently over billions of records, using Spark, Databricks and our cloud infrastructure\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBuilding evaluation frameworks: ground-truth datasets, labeling processes, quality metrics and ongoing monitoring\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ETaking solutions from POC to production, and owning them after launch\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWorking closely with Product to define requirements and shape the roadmap\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECollaborating with data engineers, data scientists and other R\u0026amp;D teams on infrastructure and best practices\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003EThis is the perfect job for someone who:\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EHolds a B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics or another relevant field\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas 4\u002B years of hands-on experience as a data engineer, ML engineer or data scientist, with solutions running in production\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas strong Python skills and writes production-quality code\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings, evaluation)\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas worked with the modern LLM stack: LLM provider APIs (OpenAI, Anthropic, etc.), LangGraph or LangChain, Hugging Face and vector stores\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas a solid grasp of text classification and NLP methods, both classical and modern, and knows when to use each\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas experience processing large-scale data with Spark/PySpark, Databricks or similar, on AWS or another cloud\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EUnderstands evaluation and data quality well: precision/recall trade-offs, building ground truth, error analysis\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EIs pragmatic and delivery-focused, comfortable with ambiguity, and communicates clearly with Product and business stakeholders\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas experience with taxonomies, entity resolution or product/e-commerce data (advantage)\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EHas experience with fine-tuning or deploying open-source models (advantage)\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EAt Similarweb, collaborating with our colleagues in-office creates a more connected, unified culture. Our best work is a product of our face-to-face collaboration, with the ability to work partially from home.\u003C/p\u003E\u003Ch3\u003EWhy you\u2019ll love being a Similarwebber:\u003C/h3\u003E\u003Cp\u003EYou\u2019ll actually love the product you work with: Our customers aren\u2019t our only raving fans. When we asked our employees why they chose to come work at Similarweb, 99% of them said \u201Cthe product.\u201D Imagine how exciting your job is when you get to work with the most powerful digital intelligence platform in the world.\u003C/p\u003E\u003Cp\u003EYou\u2019ll find a home for your big ideas: We encourage an open dialogue and empower employees to bring their ideas to the table. You\u2019ll find the resources you need to take initiative and create meaningful change within the organization.\u003C/p\u003E\u003Cp\u003EWe offer competitive perks \u0026amp; benefits: We take your well-being seriously, and offer competitive compensation packages to all employees. We also put a strong emphasis on community, with regular team outings and happy hours.\u003C/p\u003E\u003Cp\u003EYou can grow your career in any direction you choose: Interested in becoming a VP or want to transition into a different department? Whether it\u2019s Career Week, personalized coaching, or our ongoing learning solutions, you\u2019ll find all the tools and opportunities you need to develop your career right here.\u003C/p\u003E\u003Cp\u003EDiversity isn\u2019t just a buzzword: People want to work in a place where they can be themselves. We strive to create a workplace that is reflective of the communities we serve, where everyone is empowered to bring their full, authentic selves to work. We are committed to inclusivity across race, gender, ethnicity, culture, sexual orientation, age, religion, spirituality, identity and experience. We believe our culture of equality and mutual respect also helps us better understand and serve our customers in a world that is becoming more global, more diverse, and more digital every day.\u003C/p\u003E\u003Cp\u003EWe will handle your application and information related to your application in accordance with the Applicant Privacy Policy available here: https://www.similarweb.com/corp/legal/applicant-privacy-policies/\u003C/p\u003E\u003Cp\u003EWe will handle your application and information related to your application in accordance with the Applicant Privacy Policy available here.\u003C/p\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"data-engineer-applied-ml-at-similarweb-bc9ef18b9695"},"url":"https://gurify.com/job/data-engineer-applied-ml-at-similarweb-bc9ef18b9695","datePosted":"2026-09-30","validThrough":"2026-11-15T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Similarweb","sameAs":"https://job-boards.greenhouse.io/similarweb"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"IL","addressLocality":"Tel Aviv"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"Israel","item":"https://gurify.com/jobs/israel"},{"@type":"ListItem","position":3,"name":"Data Engineer \u2013 Applied ML","item":"https://gurify.com/job/data-engineer-applied-ml-at-similarweb-bc9ef18b9695"}]}
```
