Job

Staff Software Engineer (Agentic Search, Crawler)

Актуальные вакансии Infrastructure & System Engineering

About the Company:

We’re partnering with a company building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. From large-scale GPU orchestration to inference optimization, the team owns the hard problems across compute, storage, networking and applied AI. The company has a global footprint with R&D hubs across Europe, the UK, North America and Israel. The team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

About the Role:

The team is looking for a Staff Software Engineer to work on the content acquisition and crawling infrastructure of a novel search engine tailored for agentic AI consumption. In this role, you will focus on building systems that discover, fetch, and continuously refresh content from the open web and other large-scale data sources.

Why this opportunity is different?

You will join one of the founding Search engineering teams at the company. Over the coming year, they’re building a world-class team to create a next-generation search platform designed specifically for AI systems. This is a rare opportunity to shape architecture, influence technical strategy, and help build a critical piece of the future AI stack from the ground up — while backed by the resources, stability, and ambition of a publicly traded AI infrastructure company.

Your responsibility will be to:

  • Design, implement, and operate web-scale crawling systems for acquiring content from the internet
  • Build ingestion workflows for internal and external data sources, including crawlers, structured feeds, and partner integrations
  • Develop crawl scheduling, prioritisation, recrawl policies, and freshness strategies
  • Build systems for URL discovery, deduplication, content extraction, and crawl orchestration
  • Ensure reliable operation of crawling infrastructure under high-throughput conditions
  • Define observability and quality metrics for crawl coverage, freshness, throughput, and content quality
  • Monitor resource usage, bandwidth consumption, and infrastructure cost
  • Collaborate with indexing and ML teams to ensure acquired content meets retrieval and ranking requirements
  • Enable safe experimentation with crawling strategies and content acquisition policies

You may be a good fit if you have:

  • Experience building backend or distributed systems
  • Deep Go or C++ expertise (experience with Python / Java is also valuable)
  • Experience in Web crawling
  • Experience with large-scale distributed systems (10k+ RPS, billions of URLs, high-throughput pipelines)
  • Understanding of web protocols (HTTP, DNS, TLS), crawling, scraping, and content extraction
  • Experience operating production systems and debugging failures in distributed environments
  • Solid understanding of scalability, fault tolerance, and resource management

It would be a plus if you have experience with:

  • Building streaming data pipelines and event-driven systems
  • Kafka, Pulsar, NATS, RabbitMQ, or similar messaging platforms
  • Designing distributed schedulers, queues, and asynchronous processing systems
  • Spark, Flink, Beam, or MapReduce
  • Ad tech, social networks, search engines, or other large-scale content platforms

Conditions & work format:

  • Hybrid in Amsterdam, London, Tel Aviv, or New York (relocation support is available)
  • Remote work format across Europe/US is possible
  • Competitive base salary and annual bonus
  • Healthcare and local benefits package
Send your CV on Telegram @dariiyah