# Warsaw - Careers - Vantage Point

[Vantagepointglobal](https://gurify.com/jobs?q=Vantagepointglobal) · Warsaw, Poland · Posted 9 months ago

[Data](https://gurify.com/jobs/data)

[Apply on the original posting → (opens in a new tab)](https://vantagepointglobal.teamtailor.com/jobs/6870031-data-engineer-warsaw)

## Job description

### About the Role

We are looking for a Data Engineer to design and maintain scalable data solutions that power advanced analytics and AI-driven insights. This role combines expertise in big data engineering and web scraping, enabling you to work on high-impact projects involving large, complex, and unstructured datasets.

You will architect data pipelines, enforce governance standards, and build tools for extracting and processing data from diverse sources, including websites and external vendors. If you thrive in solving complex data challenges and want to work with cutting-edge technologies, this is the role for you.

##

### What You’ll Do

- Architect, develop, and maintain high-throughput data pipelines in Databricks and AWS (Glue, EMR, Fargate, Step Functions).

- Ingest, normalize, and enrich large volumes of structured and unstructured data, including internal data, market data, vendor feeds, and alternative sources.

- Collaborate with AI Engineers, ML scientists, and software teams to translate requirements into scalable data architectures, schemas, and APIs.

- Optimize pipeline performance and cost using distributed processing techniques (Spark, Delta, Arrow) and AWS best practices (spot fleets, autoscaling).

- Enforce data governance, privacy, and lineage standards, cataloguing assets in Unity Catalog and managing PII/PCI classification.

- Build automated validation, testing, and monitoring frameworks to ensure data quality and freshness for both offline and online workloads.

- Support onboarding and integration of new external data vendors, ensuring compliance and rapid time-to-value.

- Continuously evaluate emerging GenAI tooling (vector stores, LLMOps platforms, synthetic-data generators) and drive proof-of-concepts.

- Web Scraping Focus:

- Own the creation of tools and workflows for web crawling and scraping using compliance-approved technologies.

- Test and validate scraped data for accuracy, quality, and compliance.

- Identify and resolve issues with scrapes and scale processes as needed.

##

### What’s Required

- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.

- 5+ years of experience in data engineering or a related role.

- Strong experience in Python and SQL.

- Experience with Spark or Scala and distributed data processing.

- Proficiency in building scalable, distributed data pipelines in a cloud environment.

- Familiarity with Linux/UNIX, HTTP, HTML, JavaScript, and networking concepts.

- Knowledge of web scraping tools and libraries (e.g., Requests, BeautifulSoup, Scrapy, Pandas, Selenium, Spark).

- Working knowledge of version control systems and open-source practices.

- Solid understanding of data architecture principles, data modeling, and data warehousing.

- Excellent analytical and problem-solving skills.

- Strong communication skills in English (written and spoken).

- Commitment to the highest ethical standards.

##

### Preferred Skills

- Experience extracting text from PDFs, images, and applications.

- Familiarity with system monitoring/administration tools.

- Knowledge of graph databases.

- Prior experience analysing big data sets.

Vantage Point Global is fully committed to being an Equal Opportunities, inclusive employer. We are passionate about attracting diverse talent, and welcome applications regardless of ethnicity, culture, age, gender, nationality, religion, disability, or sexual orientation.

### Things you need to know:

- To apply, you’ll need to provide us with a CV and answer a few initial questions.

- We’d like to make you aware that if you have not heard back from us within three weeks of the date of application that we will not be progressing your application.

**Live in Vantagepointglobal’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- AS [Data Engineer in Warsaw • Asana Jobs](https://gurify.com/job/data-engineer-in-warsaw-asana-jobs-7da3727e1c61) Asana · Warsaw · 5 weeks ago
- VI [Data Scientist (Growth)](https://gurify.com/job/data-scientist-growth-at-viktor-de04c550356a) Viktor · Warsaw · 4 days ago
- VI [Data Scientist (Growth)](https://gurify.com/job/data-scientist-growth-at-viktor-a590adf6f6f7) Viktor · Warsaw · 4 days ago
- WP [Senior Data Engineer](https://gurify.com/job/senior-data-engineer-at-wpp-4f5d1b84380a) WPP · Warsaw, Poland · 2 weeks ago
- WP [Data Engineer](https://gurify.com/job/data-engineer-at-wpp-7ad2c6680d01) WPP · Warsaw, Poland · 2 weeks ago
- TR [Data Engineer](https://gurify.com/job/data-engineer-at-tripledotstudios-e50f13ffd86b) Tripledotstudios · Warsaw · 3 weeks ago

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Warsaw - Careers - Vantage Point","description":"\u003Ch3\u003EAbout the Role\u003C/h3\u003E\u003Cp\u003EWe are looking for a Data Engineer to design and maintain scalable data solutions that power advanced analytics and AI-driven insights. This role combines expertise in big data engineering and web scraping, enabling you to work on high-impact projects involving large, complex, and unstructured datasets.\u003C/p\u003E\u003Cp\u003EYou will architect data pipelines, enforce governance standards, and build tools for extracting and processing data from diverse sources, including websites and external vendors. If you thrive in solving complex data challenges and want to work with cutting-edge technologies, this is the role for you.\u003C/p\u003E\u003Cp\u003E##\u003C/p\u003E\u003Ch3\u003EWhat You\u2019ll Do\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EArchitect, develop, and maintain high-throughput data pipelines in Databricks and AWS (Glue, EMR, Fargate, Step Functions).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EIngest, normalize, and enrich large volumes of structured and unstructured data, including internal data, market data, vendor feeds, and alternative sources.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECollaborate with AI Engineers, ML scientists, and software teams to translate requirements into scalable data architectures, schemas, and APIs.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EOptimize pipeline performance and cost using distributed processing techniques (Spark, Delta, Arrow) and AWS best practices (spot fleets, autoscaling).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EEnforce data governance, privacy, and lineage standards, cataloguing assets in Unity Catalog and managing PII/PCI classification.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBuild automated validation, testing, and monitoring frameworks to ensure data quality and freshness for both offline and online workloads.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESupport onboarding and integration of new external data vendors, ensuring compliance and rapid time-to-value.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EContinuously evaluate emerging GenAI tooling (vector stores, LLMOps platforms, synthetic-data generators) and drive proof-of-concepts.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWeb Scraping Focus:\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EOwn the creation of tools and workflows for web crawling and scraping using compliance-approved technologies.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ETest and validate scraped data for accuracy, quality, and compliance.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EIdentify and resolve issues with scrapes and scale processes as needed.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003E##\u003C/p\u003E\u003Ch3\u003EWhat\u2019s Required\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EBachelor\u2019s or Master\u2019s degree in Computer Science, Engineering, or related field.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003E5\u002B years of experience in data engineering or a related role.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong experience in Python and SQL.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience with Spark or Scala and distributed data processing.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EProficiency in building scalable, distributed data pipelines in a cloud environment.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with Linux/UNIX, HTTP, HTML, JavaScript, and networking concepts.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EKnowledge of web scraping tools and libraries (e.g., Requests, BeautifulSoup, Scrapy, Pandas, Selenium, Spark).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWorking knowledge of version control systems and open-source practices.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ESolid understanding of data architecture principles, data modeling, and data warehousing.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExcellent analytical and problem-solving skills.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong communication skills in English (written and spoken).\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECommitment to the highest ethical standards.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003E##\u003C/p\u003E\u003Ch3\u003EPreferred Skills\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EExperience extracting text from PDFs, images, and applications.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with system monitoring/administration tools.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EKnowledge of graph databases.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EPrior experience analysing big data sets.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EVantage Point Global is fully committed to being an Equal Opportunities, inclusive employer. We are passionate about attracting diverse talent, and welcome applications regardless of ethnicity, culture, age, gender, nationality, religion, disability, or sexual orientation.\u003C/p\u003E\u003Ch3\u003EThings you need to know:\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003ETo apply, you\u2019ll need to provide us with a CV and answer a few initial questions.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWe\u2019d like to make you aware that if you have not heard back from us within three weeks of the date of application that we will not be progressing your application.\u003C/li\u003E\u003C/ul\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"warsaw-careers-vantage-point-at-vantagepointglobal-4c48215de4b0"},"url":"https://gurify.com/job/warsaw-careers-vantage-point-at-vantagepointglobal-4c48215de4b0","datePosted":"2025-12-02","validThrough":"2026-11-06T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Vantagepointglobal","sameAs":"https://vantagepointglobal.teamtailor.com"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"PL","addressLocality":"Warsaw"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"Poland","item":"https://gurify.com/jobs/poland"},{"@type":"ListItem","position":3,"name":"Warsaw - Careers - Vantage Point","item":"https://gurify.com/job/warsaw-careers-vantage-point-at-vantagepointglobal-4c48215de4b0"}]}
```
