Posted on: 08/09/2026
About the role :
We build and operate a large-scale data platform that collects, structures, and delivers web data at high volume. Millions of records flow through our pipelines every week, and the quality of that data is the product.
We're looking for a Data Analyst to own data quality end to end. This is a hands-on, investigative role : roughly 70% building automated checks, queries, scripts, and dashboards, and 30% manual validation and root-cause investigation.
Data pipelines at scale rarely fail loudly. They fail quietly - a source silently returns empty results, a field starts capturing the wrong value, a reported metric measures something different from what its label says. Your job is to find these issues before anyone downstream does, and then build the automation that keeps finding them.
What you'll do :
- Build automated data quality checks across our datasets - completeness, validity, deduplication, referential integrity, freshness - and run them on a schedule.
- Monitor volume and drift. Track day-over-day and week-over-week changes in record counts, field fill rates, and value distributions, and flag anomalies.
- Validate metrics and formulas. Verify every reported number is calculated the way its definition claims. Maintain a metric dictionary and reconcile the same metric across APIs, dashboards, and exports.
- Perform manual QA and ground-truth validation. Sample records against their original source, field by field, and maintain labelled reference datasets that automated checks are measured against.
- Investigate data issues to root cause. Use logs, metrics, pipeline run histories, and direct database queries to determine why a dataset looks wrong, and file clear, reproducible reports for engineering.
- Review new data sources before launch - confirm the right fields are captured, coverage is complete, and output meets quality standards.
- Evaluate automated and AI-assisted extraction output. Measure how often it's wrong, and in what way, and feed findings back into the pipeline.
- Build dashboards and reporting that answer coverage, quality, and freshness questions without manual effort, and publish a regular data quality report.
- Partner with engineering - validate fixes and backfills on real data before sign-off, and contribute checks and analysis scripts to the codebase.
What we're looking for :
- 1 - 3 years in a data role - data analyst, data quality, analytics engineering, or similar.
- Strong SQL, and comfort with (or willingness to quickly learn) NoSQL querying such as MongoDB aggregations.
- Python for data work - pandas or equivalent, and the ability to write a standalone script that runs on a schedule and reports a result.
- Working statistical literacy - sampling, distributions, drift, and a sense of how much evidence a claim actually needs.
- Investigative mindset. The core skill here is refusing to take a number at face value and tracing it back to prove what it really measures.
- Clear written communication - able to explain what's broken, show the evidence, and describe how to reproduce it.
Good to have :
- Experience with web-scraped, third-party, or externally ingested data and its typical failure modes.
- Elasticsearch / OpenSearch querying, or Grafana and Prometheus (PromQL).
- Familiarity with workflow orchestration (Airflow, Temporal, Celery) - enough to read a run history.
- Experience evaluating output of AI/LLM-based data extraction.
Did you find something suspicious?
Posted by
Posted in
Data Analytics & BI
Functional Area
Data Mining / Analysis
Job Code
1669611