All case studies

Node.js Processing Performance: 5 Minutes → 15 Seconds

Removing unnecessary browser overhead and sequential execution from a scheduled ticket scraper.

Independent client work · Ticket scraper · 2023

Node.jsAWS LambdaAxiosCheerio
THE OUTCOME5 minutes → 15 seconds

Problem

A ticket scraper needed to collect information from three additional websites. The existing Puppeteer implementation ran sequentially, took about five minutes and often failed before finishing. It ran every two hours.

Context and constraints

The workload ran on AWS Lambda. A browser can be necessary when a page requires JavaScript execution or browser interaction, but that cost should be justified by the target site. The recorded solution used HTTP fetching and HTML parsing for this workload.

Existing architecture

Puppeteer-based scraping tasks ran one after another. Each task carried browser-automation overhead, while sequential execution added the duration of independent tasks together.

Investigation

I assessed the existing Node.js implementation and the information needed from the target websites. The key decisions were whether a browser was actually required and which scraping operations could run concurrently.

Root cause

The implementation combined a heavyweight extraction mechanism with sequential execution. That made the process slower than the underlying data-fetching problem required.

Solution

I replaced Puppeteer with Axios for HTTP requests and Cheerio for HTML parsing, and ran scraping work concurrently.

Architecture and implementation

The new path fetched pages directly and parsed the required information. Concurrency allowed independent work to overlap. The retained account does not specify a concurrency limit or retry policy, so neither is claimed as an implemented feature.

Conceptual architecture · simplified from the project account

Before

  1. Scheduled Lambda
  2. Sequential browser tasks
  3. Extract ticket information
  4. ~5 minutes

After

  1. Scheduled Lambda
  2. Concurrent Axios requests
  3. Cheerio extraction
  4. ~15 seconds

Result

Execution fell from approximately five minutes to 15 seconds—about a 20× improvement for this process. The client independently described the revised Node.js code as nearly 20× faster.

Tradeoffs

HTTP parsing does not replace browser execution for every website. Concurrency also needs to respect upstream limits, Lambda resources and failure isolation. Infrastructure savings and a quantified reliability improvement were not measured.

What I would carry forward

Start with the least expensive execution mechanism that meets the requirement. For a future extension, document which targets need a browser and retain bounded-concurrency and partial-failure tests.

“Muhammad did a great job. He improved on the previous NodeJS code base and made it run nearly 20X faster! He was readily available for any small changes we requested and understood immediately what we're asking for. Would recommend.”
Upwork client · AWS Lambda scraper · 5/5 · Feb–Mar 2023
View Upwork profile

More engineering evidence.

All case studies

Have a difficult engineering problem?

If you’re dealing with a backend bottleneck, production issue, unreliable integration or an AI prototype that needs production engineering, let’s discuss it.

Discuss the problem