Skip to content

Web data, collected and maintained.

I turn prices, listings, catalogues, and directories into data your team can use every day.

I handle collection, validation, delivery, and monitoring, with ongoing source maintenance agreed in the scope.

A production web-data platform

I led development of a data engine and worked across its backend, workers, storage, and AWS infrastructure.

Read the production case study
Collection
Scrapy · multi-source extraction
Processing
Django · Celery · Redis
Storage
PostgreSQL
Operations
AWS · queues · monitoring

Problems I help solve

  • Price and availability monitoring

    Track changing products, stock, offers, or market signals on a dependable schedule.

  • Listings and catalogue aggregation

    Bring fragmented public sources into one consistent schema for search, analysis, or a product.

  • Research data feeds

    Collect recurring datasets for market, operations, or commercial research without repeated manual work.

  • Existing scraper recovery

    Stabilise a brittle collection system and add the visibility needed to operate it with confidence.

What the work covers

Collect

Map sources, access patterns, fields, schedules, and technical constraints before choosing the collection approach.

  • Public websites and feeds
  • Scheduled and incremental runs
  • Pagination and source discovery
  • Rate and failure handling

Make it usable

Normalise source-specific output into a stable contract for the team or product that consumes it.

  • Validation and deduplication
  • Schema design
  • Files, databases, or APIs
  • Freshness and quality checks

Keep it running

Monitor runs and maintain extraction as sources change.

  • Run monitoring
  • Source-change maintenance
  • Failure diagnosis
  • Capacity and cost review

How the work starts

  1. Check the sources

    We check source access and sample data, then agree on fields, volume, and delivery. I flag risks that could change the scope.

    Result: Source assessment and delivery plan

  2. Set up the data feed

    I build and deploy collection, validation, and delivery for the agreed data.

    Result: Data feed running in production

  3. Maintain the feed

    I monitor runs and adapt collection when sources change. We agree on any changes to quality checks or capacity.

    Result: Monitoring and source maintenance

Is this a fit?

A good fit

  • The data supports an active product, analysis, or operating workflow.
  • The sources are public and can be assessed responsibly.
  • Freshness, consistency, and maintenance matter after launch.

Outside this service

  • A one-off export with no meaningful engineering or maintenance need.
  • Bypassing private access, authentication, or source protections.
  • A volume or legal requirement that cannot be validated during discovery.

Questions before we start

Can every website be scraped?

No. Access, source behaviour, data rights, volume, and maintenance risk must be assessed first. The feasibility step exists to make those boundaries explicit before a build is proposed.

How can the data be delivered?

Typical options are CSV or JSON files, a database, object storage, or an API. The right contract depends on who consumes the data and how fresh it needs to be.

What happens when a source changes?

A managed engagement can include run monitoring, source-change diagnosis, extraction updates, and validation checks. The exact response boundary is agreed before operations begin.

Will we own the implementation?

Ownership, hosting, credentials, documentation, and handover are defined in the proposal. The default goal is to avoid hidden platform lock-in.

Tell me which data you need.

Send me the source, fields, delivery frequency, and destination. I can assess feasibility and propose the next step.

Email me about your data