Skip to content

Jobgether

Senior Data Engineer (Web Scraping)

Location
India, IN
Arrangement
Remote
Employment type
Full-time
Level
Senior
Posted
30 September 2026 (today)

Checked todayApplications go to the employer, never to RoleSprint

About this role

Accountabilities: • Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines.

• Design and establish scalable web-scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling.

• Build, maintain, and improve reliable production scrapers for both new and existing data sources.

• Investigate websites and determine the most appropriate acquisition method, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling.

• Evaluate build-versus-buy options for scraping infrastructure and external services, considering capabilities, reliability, cost, operational complexity, and risk.

• Ensure web-data acquisition activities appropriately account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations, escalating unclear situations when required.

• Diagnose and resolve scraping challenges related to website changes, dynamic content, authentication, sessions, rate limits, concurrency, and other operational constraints.

• Integrate scraping workloads into scalable data-platform and lakehouse architectures.

• Improve scheduling, monitoring, storage, validation, and operational support for scraping workloads.

• Use AI-assisted engineering tools where appropriate while maintaining a thorough understanding of, and accountability for, the code being delivered.

• Support production workloads through monitoring, debugging, maintenance, and continuous improvement.

• Contribute to a remote engineering environment through code reviews, documentation, ticket-based workflows, and knowledge sharing.

Requirements

• Demonstrated professional experience building and operating production web-scraping systems at scale .

• Proven ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and ongoing production support.

• Strong production-level Python engineering skills, with experience developing maintainable applications rather than standalone scripts.

• Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium .

• Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior.

• Experience addressing common scraping challenges including pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies.

• Strong understanding of data pipelines, data quality, and how collected data should be validated, stored, and consumed by downstream systems.

• Experience deploying, monitoring, and supporting production workloads in a cloud environment.

• Strong debugging, analytical, and problem-solving abilities, with the judgment to make effective engineering decisions independently.

• Comfortable working within a remote engineering team and participating in code reviews, documentation, and ticket-based development workflows.

• Experience with AWS is desirable.

• Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg , is a plus.

• Experience with PySpark or other distributed data-processing technologies is beneficial.

• Familiarity with Docker and containerized workloads is advantageous.

• Experience with Terraform or other infrastructure-as-code tools is a plus.

• Familiarity with Grafana or comparable observability platforms is desirable.

• Experience operating high-volume or distributed crawling systems is beneficial.

• Experience evaluating or operating commercial scraping, proxy, or browser-infrastructure services is a plus.

• Experience implementing automated scraper testing, canary runs, or source-drift detection is desirable.

• Exposure to legal, compliance, privacy, or data-governance processes related to web-data acquisition is advantageous.

• Strong ownership, autonomy, documentation, communication, and collaboration skills.

Benefits

• Fully remote position within a remote-first technology team.

• Opportunity to take ownership of a critical web-data acquisition capability and influence its architecture and operating standards.

• Senior, hands-on role with substantial autonomy across investigation, engineering, deployment, and production support.

• Work on scalable data pipelines and modern lakehouse architectures supporting research and data products.

• Exposure to cloud infrastructure, distributed processing, observability, browser automation, APIs, and production scraping technologies.

• Opportunity to establish reusable engineering patterns and improve the reliability and scalability of data ingestion.

• Collaboration with a distributed engineering team through code reviews, documentation, and structured workflows.

• Environment that supports independent problem-solving, technical ownership, and continuous improvement.

• Fully remote setup available across the relevant distributed team environment.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Work location

  • IN

Related jobs

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Published on Lever under the board identifier Jobgether, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 30 September 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

Browse all current openings

No credit card requiredStart free