Problem
A ticket scraper needed to collect information from three additional websites. The existing Puppeteer implementation ran sequentially, took about five minutes and often failed before finishing. It ran every two hours.
Context and constraints
The workload ran on AWS Lambda. A browser can be necessary when a page requires JavaScript execution or browser interaction, but that cost should be justified by the target site. The recorded solution used HTTP fetching and HTML parsing for this workload.
Existing architecture
Puppeteer-based scraping tasks ran one after another. Each task carried browser-automation overhead, while sequential execution added the duration of independent tasks together.
Investigation
I assessed the existing Node.js implementation and the information needed from the target websites. The key decisions were whether a browser was actually required and which scraping operations could run concurrently.
Root cause
The implementation combined a heavyweight extraction mechanism with sequential execution. That made the process slower than the underlying data-fetching problem required.
Solution
I replaced Puppeteer with Axios for HTTP requests and Cheerio for HTML parsing, and ran scraping work concurrently.
Architecture and implementation
The new path fetched pages directly and parsed the required information. Concurrency allowed independent work to overlap. The retained account does not specify a concurrency limit or retry policy, so neither is claimed as an implemented feature.
Before
- Scheduled Lambda
- Sequential browser tasks
- Extract ticket information
- ~5 minutes
After
- Scheduled Lambda
- Concurrent Axios requests
- Cheerio extraction
- ~15 seconds
Result
Execution fell from approximately five minutes to 15 seconds—about a 20× improvement for this process. The client independently described the revised Node.js code as nearly 20× faster.
Tradeoffs
HTTP parsing does not replace browser execution for every website. Concurrency also needs to respect upstream limits, Lambda resources and failure isolation. Infrastructure savings and a quantified reliability improvement were not measured.
What I would carry forward
Start with the least expensive execution mechanism that meets the requirement. For a future extension, document which targets need a browser and retain bounded-concurrency and partial-failure tests.
CLIENT EVIDENCE
“Muhammad did a great job. He improved on the previous NodeJS code base and made it run nearly 20X faster! He was readily available for any small changes we requested and understood immediately what we're asking for. Would recommend.”