HamburgerMenu
hirist

Merit Group - Technical Lead - Web Scraping & Data Engineering

Merit Data and Technology
10 - 16 Years
Anywhere in India/Multiple Locations

Posted on: 13/08/2026

Job Description

Experience : 12 - 18 Years


Location : PAN INDIA


Work Mode : Permanent Remote


Job Type : Full Time Employment


Notice : Immediate Joiners


Job Description :


Primary (70%) : Web scraping architecture anti-bot, distributed crawling, proxy/CAPTCHA strategy, compliance.


Secondary (30%) : Data engineering pipelines, lakehouse, orchestration, dbt.


Mandatory ownership of the technical solution and effort estimation for every new scraping proposal/RFP.


Senior profile : 9+ years overall, 5+ years in scraping, 4+ years in data engineering.


Technical Lead / Architect Web Scraping & Data Engineering :


Role Summary :


We are seeking a Technical lead for the design and delivery of large-scale web scraping and data extraction solutions, with strong supporting expertise in data engineering, the technical authority on all scraping initiatives from pre-sales solutioning and effort estimation through to architecture, build, and stabilization.


The primary mandate is web scraping : designing resilient crawlers, anti-bot strategies, distributed extraction systems, and compliance frameworks. The secondary mandate is data engineering : ensuring extracted data flows reliably into well-modelled, query-ready storage layers and downstream analytical or operational systems.


Mandatory Involvement Areas :


- Author or co-author the technical solution section of every new RFP / RFI / proposal document for scraping projects.


- Lead technical discovery calls with prospective clients to understand target sources, data SLAs, volume, and compliance constraints.


- Conduct target-site feasibility assessments anti-bot complexity, dynamic content, login walls, geo-restrictions, rate limits and document findings before commercials are committed.


- Own end-to-end effort estimation for scraping engagements : crawler build effort, infrastructure sizing, proxy and CAPTCHA cost projections, maintenance overhead, and contingency.


- Produce solution architecture diagrams, tech stack recommendations, and assumption logs as part of every proposal.


- Define and document SLAs, KPIs, and acceptance criteria proposed to the client.


- Participate in client orals, technical defence sessions, and commercial negotiations as the technical SPOC.


- Maintain an internal estimation knowledge base reusable estimation templates, complexity matrices, target-site classification, and historical actuals and continuously refine it after every project closure.


- Sign off on the technical feasibility and risk profile of every proposal before submission. No scraping proposal goes out without architect approval.


Core Responsibilities :


- Design end-to-end scraping solutions covering crawl orchestration, extraction, parsing, storage, and downstream data consumption.


- Define reference architectures for high-volume, high-velocity, and high-variety scraping use cases.


- Architect resilient systems handling JavaScript-heavy sites, CAPTCHAs, rate limiting, IP blocking, and frequent DOM changes.


- Evaluate and select frameworks, proxy networks, and anti-bot bypass strategies aligned with cost, performance, and compliance.


- Design data quality, deduplication, validation, and schema-evolution strategies for scraped data.


Data Engineering Architecture (Secondary) :


- Design downstream data pipelines that ingest, transform, and serve scraped data to analytics, ML, or operational consumers.


- Architect lakehouse / warehouse layers, define data modelling standards (dimensional, Data Vault, or hybrid), and govern schema evolution.


- Define ELT/ETL patterns, orchestration strategy, and SLA-backed data freshness commitments.


- Establish data quality, observability, lineage, and cataloguing practices across the platform.


- Recommend storage formats, partitioning, and indexing strategies for cost and query performance.


Delivery Leadership :


- Translate business requirements into technical specifications, sprint plans, and implementation roadmaps.


- Produce HLD, LLD, and Architecture Decision Records (ADRs) for each engagement.


- Provide hands-on guidance, perform code reviews, and mentor scraping and data engineers.


- Drive proof-of-concept (PoC) builds for complex or high-risk targets before full-scale rollout.


Compliance, Risk & Governance :


- Ensure all scraping work adheres to applicable laws and platform terms (GDPR, CCPA, DPDP, robots.txt, copyright).


- Define and enforce ethical scraping practices, request throttling, and PII handling guidelines.


- Conduct risk assessments for each target source and recommend mitigations.


Performance, Cost & Operations :


- Establish SLAs for crawl freshness, completeness, and accuracy.


- Design monitoring, alerting, and self-healing mechanisms for scraping pipelines.


- Optimise infrastructure cost (compute, proxies, storage) without compromising delivery KPIs.


Required Technical Skills Primary (Web Scraping) :


These are non-negotiable. The candidate must demonstrate deep, hands-on expertise in each of the following :


Scraping Frameworks & Tooling :


- Expert-level Python : Scrapy, BeautifulSoup, lxml, Requests, httpx, parsel.


- Headless browser automation : Playwright, Puppeteer, Selenium, Pyppeteer.


- Node.js scraping stack (where applicable) : Puppeteer, Cheerio, Crawlee.


Anti-Bot & Evasion Strategy :


- Proven experience bypassing Cloudflare, Akamai Bot Manager, DataDome, PerimeterX, Imperva, Kasada.


- CAPTCHA handling : reCAPTCHA v2/v3, hCaptcha, FunCaptcha, image / audio solvers; integration with 2Captcha, Anti-Captcha, CapSolver.


- TLS / JA3 / JA4 fingerprinting awareness; HTTP/2 fingerprint evasion; browser fingerprint spoofing.


- Stealth plugins, user-agent rotation, header normalisation, cookie / session management at scale.


Proxy & Network Infrastructure :


- Hands-on experience with rotating, residential, mobile, ISP, and datacenter proxies.


- Integration with providers : Bright Data, Oxylabs, Smartproxy, NetNut, IPRoyal, SOAX.


- Proxy pool design, health-checking, geo-targeting, sticky sessions, and cost optimisation.


Parsing & Extraction :


- XPath, CSS selectors, regex, JSON-LD, microdata, RDFa.


- Reverse-engineering of internal / mobile APIs, GraphQL endpoints, and XHR traffic.


- ML/LLM-assisted extraction for unstructured layouts (nice-to-have, increasingly expected).


Distributed Crawling & Orchestration :


- Scrapy-Redis, Scrapy Cluster, Frontera, Crawlee, or equivalent distributed crawling frameworks.


- Job scheduling and orchestration with Apache Airflow, Prefect, Dagster, or Celery.


- Queue-based architectures using Kafka, RabbitMQ, AWS SQS, GCP Pub/Sub.


Required Technical Skills Secondary (Data Engineering) :


Data Pipelines & Orchestration :


- ETL / ELT design patterns, idempotent pipelines, CDC (Change Data Capture), incremental loads.


- Apache Airflow, Prefect, Dagster, AWS Glue, Azure Data Factory, GCP Dataflow.


- Stream processing : Kafka Streams, Apache Flink, Spark Structured Streaming.


- Batch processing : Apache Spark (PySpark), Databricks, EMR, Dataproc.


Data Modelling & Storage :


- Dimensional modelling (Kimball), Data Vault 2.0, normalised vs. denormalised trade-offs.


- Data warehouses : Snowflake, BigQuery, Redshift, Synapse, Databricks SQL Warehouse.


- Data lakes / lakehouses : Delta Lake, Apache Iceberg, Apache Hudi on S3 / GCS / ADLS.


- OLTP databases : PostgreSQL, MySQL; NoSQL : MongoDB, DynamoDB, Cassandra, Elasticsearch, Redis.


- File formats : Parquet, Avro, ORC, JSON, CSV; partitioning, bucketing, compaction strategies.


Transformation & Quality :


- dbt (data build tool) for transformation, testing, and documentation.


- Data quality frameworks : Great Expectations, Soda, Deequ, custom validators.


- Data lineage and cataloguing : DataHub, OpenMetadata, Amundsen, Atlan, Collibra.


Cloud, DevOps & Observability :


- Strong on at least one of AWS, GCP, Azure (compute, storage, IAM, networking, serverless).


- Containerisation and orchestration : Docker, Kubernetes (EKS / GKE / AKS), ECS.


- Infrastructure as Code : Terraform, Pulumi, CloudFormation.


- CI/CD : GitHub Actions, GitLab CI, Jenkins, Argo CD.


- Observability : Prometheus, Grafana, ELK / OpenSearch, Datadog, Sentry, OpenTelemetry.


Experience & Qualifications :


- Bachelor's or Master's degree in Computer Science, Engineering, or related field.


- 10+ years of overall software engineering experience.


- 5+ years dedicated to large-scale web scraping / data extraction (primary).


- 3+ years of hands-on data engineering experience covering pipelines, warehouses / lakehouses, and orchestration (secondary).


- Proven track record of architecting scraping platforms processing millions of pages per day across diverse target sites.


- Prior experience as Solution Architect, Tech Lead, or Principal Engineer leading teams of 5+ engineers.


- Demonstrated experience in pre-sales / proposal authoring / RFP responses for scraping or data engineering engagements.


- Client-facing consulting experience strongly preferred.


Soft Skills :


- Excellent written and verbal communication; able to defend technical proposals to CXO-level audiences.


- Strong commercial acumen understands the cost / risk / quality trade-offs in estimation.


- Analytical, structured problem-solving mindset.


- Ownership-driven, comfortable being the single point of technical accountability.


- High standards for documentation, knowledge transfer, and reusability.


Expected Deliverables (RFP Scope) :


- Technical solution sections of all proposal documents.


- Target-site feasibility and risk assessment reports.


- Effort estimation models, assumption logs, and complexity matrices.


- Solution architecture diagrams and tech stack recommendations for proposals.


Delivery Phase :


- High-Level Design (HLD) and Low-Level Design (LLD) documents per initiative.


- Architecture Decision Records (ADRs).


- Reference implementations / PoCs for complex extraction scenarios.


- Code review records and engineering standards documentation.


- Operational runbooks, monitoring playbooks, and incident response procedures.


- Compliance and data governance documentation.


- Knowledge transfer sessions and final handover documentation.


Nice-to-Have :


- ML / NLP-based extraction, entity resolution, or LLM-assisted parsing experience.


- Exposure to GraphQL, gRPC, mobile API reverse-engineering, mitmproxy / Charles workflows.


- Domain experience : e-commerce price intelligence, market research, financial data, real estate, travel aggregation, hospitality.


- Open-source contributions to scraping or data engineering ecosystems.


- Familiarity with cross-jurisdictional legal frameworks for automated data collection.


- Experience with reverse ETL tools (Hightouch, Census) and feature stores.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...