cheerio vs domino vs jsdom vs puppeteer
Server-Side HTML Parsing and DOM Manipulation Strategies
cheeriodominojsdompuppeteerSimilar Packages:

Server-Side HTML Parsing and DOM Manipulation Strategies

cheerio, domino, jsdom, and puppeteer are essential tools for handling HTML and the Document Object Model (DOM) in Node.js environments, but they serve distinct architectural roles. cheerio is a fast, lightweight library that implements a subset of the jQuery core API for parsing and manipulating HTML strings without a real browser environment. domino provides a pure JavaScript implementation of the DOM and HTML5 standards, focusing on speed and low memory usage for server-side rendering tasks where a full browser is unnecessary. jsdom is a comprehensive implementation of the WHATWG DOM and HTML standards, designed to mimic a browser environment closely enough to run client-side JavaScript code and test frameworks. puppeteer is a Node.js library that provides a high-level API to control headless Chrome or Chromium browsers, enabling true end-to-end testing, screenshot generation, and interaction with complex, JavaScript-heavy web applications.

Npm Package Weekly Downloads Trend

3 Years

Github Stars Ranking

Stat Detail

Package
Downloads
Stars
Size
Issues
Publish
License
cheerio030,5001.01 MB608 months agoMIT
domino0791818 kB28a month agoBSD-2-Clause
jsdom021,6877.14 MB3022 days agoMIT
puppeteer095,61643.2 kB26510 hours agoApache-2.0

Server-Side HTML Parsing and DOM Manipulation: A Technical Deep Dive

When working with HTML in Node.js, developers often face a critical architectural decision: do we need a real browser, a simulated environment, or just a fast parser? The choices between cheerio, domino, jsdom, and puppeteer define the trade-offs between speed, fidelity, and resource consumption. Let's break down how each tool handles common engineering challenges.

โšก Execution Model: Static Parsing vs. Real Browser

cheerio works purely on HTML strings. It parses the markup into a lightweight data structure and lets you manipulate it using a jQuery-like syntax. It does not render CSS, execute JavaScript, or calculate layout.

// cheerio: Fast static parsing
import * as cheerio from 'cheerio';

const html = '<ul><li class="item">Apple</li><li class="item">Banana</li></ul>';
const $ = cheerio.load(html);

$('.item').each((i, el) => {
  console.log($(el).text());
});
// Output: Apple, then Banana

domino creates a real DOM tree in memory but still runs without a browser engine. It implements standard DOM APIs (like getElementById) but skips JavaScript execution and layout.

// domino: In-memory DOM without JS execution
import domino from 'domino';

const html = '<div id="app"><p>Hello</p></div>';
const window = domino.createWindow(html);
const document = window.document;

console.log(document.getElementById('app').textContent);
// Output: Hello

jsdom builds a full DOM environment that mimics a browser, including support for running simple JavaScript found within <script> tags. It bridges the gap between a static parser and a real browser.

// jsdom: DOM with script execution
import { JSDOM } from 'jsdom';

const html = `<div id="count">0</div><script>document.getElementById('count').innerText = '1';</script>`;
const dom = new JSDOM(html, { runScripts: "dangerously" });

console.log(dom.window.document.getElementById('count').textContent);
// Output: 1 (Script executed)

puppeteer launches a real headless Chrome instance. It loads the page exactly as a user would see it, executing all JavaScript, loading assets, and applying CSS.

// puppeteer: Real browser automation
import puppeteer from 'puppeteer';

(async () => {
  const browser = await puppeteer.launch();
  const page = await browser.newPage();
  await page.goto('https://example.com');
  
  const title = await page.title();
  console.log(title);
  
  await browser.close();
})();

๐Ÿ“‰ Performance and Resource Overhead

The cost of using these libraries varies wildly based on your needs.

  • cheerio is the fastest. It starts instantly and uses negligible memory because it only cares about the HTML structure.
  • domino is slightly heavier than cheerio but still extremely light. It builds a full DOM tree, which takes more CPU cycles than a simple parse, but it avoids the overhead of a JS engine.
  • jsdom is heavier still. It must initialize a JavaScript engine (usually V8) to run scripts and simulate browser quirks. Startup time and memory usage are significantly higher than cheerio or domino.
  • puppeteer is the most expensive. It spins up a full Chromium process, which can take seconds to start and consumes hundreds of megabytes of RAM. It is overkill for simple data extraction.

๐Ÿ’ก Tip: If you are processing thousands of HTML documents per second, cheerio or domino are your only viable options. puppeteer will bottleneck your system.

๐Ÿ› ๏ธ API Surface and Developer Experience

How you interact with the document changes depending on the tool.

cheerio uses a jQuery-style API. If you know jQuery, you know cheerio. It is concise and chainable.

// cheerio: jQuery syntax
const $ = cheerio.load('<h1 class="title">Header</h1>');
const text = $('.title').text().trim();

domino and jsdom use standard Web APIs. You interact with document, window, and Node objects just like in the browser.

// domino/jsdom: Standard DOM API
const doc = domino.createWindow('<h1 class="title">Header</h1>').document;
const text = doc.querySelector('.title').textContent.trim();

puppeteer uses an async API to control the browser remotely. You often need to wait for selectors to appear before interacting.

// puppeteer: Async browser control
await page.waitForSelector('.title');
const text = await page.$eval('.title', el => el.textContent.trim());

๐ŸŒ Handling Client-Side Rendering (CSR)

This is the most common failure point in scraping architectures.

  • cheerio and domino cannot handle CSR. If the HTML you fetch contains only <div id="app"></div> and the content is injected by React or Vue later, these tools will see an empty div. They do not run the JavaScript required to fill it.
  • jsdom can handle simple CSR if the scripts are inline and do not rely on complex browser features (like canvas or specific network policies). However, it often fails with modern frameworks that use advanced APIs or complex bundlers.
  • puppeteer handles CSR perfectly. Since it runs a real browser, it waits for React, Vue, or Angular to hydrate the page and render the final DOM.
// Scenario: Scraping a React app

// โŒ cheerio fails here
const $ = cheerio.load(htmlFromReactApp);
console.log($('#root').text()); // Empty string

// โœ… puppeteer succeeds here
await page.goto('https://react-app.com');
await page.waitForSelector('#root > div'); // Wait for render
const content = await page.content();

๐Ÿงช Testing and Simulation

When writing tests for frontend libraries, the choice depends on what you are testing.

  • Use jsdom for unit tests. It is fast enough to run thousands of tests and provides enough browser API coverage for most logic. Tools like Jest and Vitest use jsdom by default.
  • Use puppeteer for integration or end-to-end (E2E) tests. If you need to verify that a button click actually triggers a network request and updates the UI, you need a real browser.
  • Avoid cheerio and domino for testing interactive components, as they lack the event loop and rendering engine required to simulate user actions realistically.
// jsdom: Unit testing a component
import { render, screen } from '@testing-library/react';
// Runs in jsdom environment
render(<Button />);
fireEvent.click(screen.getByText('Click me'));

// puppeteer: E2E testing a flow
await page.click('#submit-btn');
await page.waitForNetworkIdle();
expect(await page.url()).toContain('/success');

๐Ÿ“Š Summary Table

Featurecheeriodominojsdompuppeteer
EngineCustom ParserPure JS DOMJS DOM + JS EngineHeadless Chrome
JS ExecutionโŒ NoโŒ Noโœ… Yes (Limited)โœ… Yes (Full)
CSS LayoutโŒ NoโŒ NoโŒ Noโœ… Yes
Speedโšก Very Fastโšก Fast๐Ÿข Moderate๐ŸŒ Slow
Memory๐Ÿ’พ Low๐Ÿ’พ Low๐Ÿ’พ Medium๐Ÿ’พ High
Best ForStatic ScrapingSSR ManipulationUnit TestingE2E / CSR Scraping

๐Ÿ’ก The Big Picture

cheerio is your scalpel ๐Ÿช’. Use it when you need to slice and dice static HTML with surgical precision and speed. It is the go-to for extracting data from blogs, documentation, or legacy servers.

domino is your lightweight framework ๐Ÿ—๏ธ. It shines in server-side rendering pipelines where you need a standard DOM to manipulate trees without the bloat of a full browser simulation.

jsdom is your sandbox ๐Ÿงช. It provides a safe, simulated browser environment for testing logic that depends on the DOM, striking a balance between realism and performance.

puppeteer is your robot operator ๐Ÿค–. When the job requires a real human-like experienceโ€”clicking buttons, waiting for animations, or scraping modern single-page appsโ€”it is the only tool that gets the job done.

Final Thought: Don't reach for puppeteer by default. The resource cost is high. Start with cheerio or jsdom, and only escalate to a full browser when the page's complexity demands it.

How to Choose: cheerio vs domino vs jsdom vs puppeteer

  • cheerio:

    Choose cheerio when you need to scrape static HTML or perform simple DOM manipulation on the server with maximum speed and minimal overhead. It is ideal for extracting data from well-formed HTML strings where you do not need to execute JavaScript or simulate a real browser environment. Avoid it if your target page relies heavily on client-side rendering to populate content.

  • domino:

    Choose domino if you are building a server-side rendering (SSR) engine and need a fast, standards-compliant DOM implementation that consumes very little memory. It is perfect for frameworks like Angular Universal or custom SSR solutions where you need to manipulate the DOM tree before sending it to the client. Do not use it if you need to run embedded scripts or require a complete browser API surface.

  • jsdom:

    Choose jsdom when you need to run unit tests for frontend code that interacts with the DOM or when you need to execute simple client-side scripts within a Node.js environment. It strikes a balance between performance and fidelity, making it suitable for testing libraries and tools that depend on browser APIs like window or document. Avoid it for heavy automation tasks or sites that require complex layout calculations.

  • puppeteer:

    Choose puppeteer for end-to-end testing, generating PDFs or screenshots, and scraping websites that rely heavily on JavaScript to render content. It is the only choice when you need to simulate real user interactions like clicking, typing, or scrolling in a genuine browser engine. Be aware that it requires more resources and startup time compared to the other libraries, so it is not suitable for high-volume, low-latency HTML parsing.

README for cheerio

cheerio

The fast, flexible, and elegant library for parsing and manipulating HTML and XML.

ไธญๆ–‡ๆ–‡ๆกฃ (Chinese Readme)

import * as cheerio from 'cheerio';
const $ = cheerio.load('<h2 class="title">Hello world</h2>');

$('h2.title').text('Hello there!');
$('h2').addClass('welcome');

$.html();
//=> <html><head></head><body><h2 class="title welcome">Hello there!</h2></body></html>

Installation

Install Cheerio using a package manager like npm, yarn, or bun.

npm install cheerio
# or
bun add cheerio

Features

โค Proven syntax: Cheerio implements a subset of core jQuery. Cheerio removes all the DOM inconsistencies and browser cruft from the jQuery library, revealing its truly gorgeous API.

ฯŸ Blazingly fast: Cheerio works with a very simple, consistent DOM model. As a result parsing, manipulating, and rendering are incredibly efficient.

โ Incredibly flexible: Cheerio wraps around parse5 for parsing HTML and can optionally use the forgiving htmlparser2. Cheerio can parse nearly any HTML or XML document. Cheerio works in both browser and server environments.

API

Loading

First you need to load in the HTML. This step in jQuery is implicit, since jQuery operates on the one, baked-in DOM. With Cheerio, we need to pass in the HTML document.

// ESM or TypeScript:
import * as cheerio from 'cheerio';

// In other environments:
const cheerio = require('cheerio');

const $ = cheerio.load('<ul id="fruits">...</ul>');

$.html();
//=> <html><head></head><body><ul id="fruits">...</ul></body></html>

Selectors

Once you've loaded the HTML, you can use jQuery-style selectors to find elements within the document.

$( selector, [context], [root] )

selector searches within the context scope which searches within the root scope. selector and context can be a string expression, DOM Element, array of DOM elements, or cheerio object. root, if provided, is typically the HTML document string.

This selector method is the starting point for traversing and manipulating the document. Like in jQuery, it's the primary method for selecting elements in the document.

$('.apple', '#fruits').text();
//=> Apple

$('ul .pear').attr('class');
//=> pear

$('li[class=orange]').html();
//=> Orange

Rendering

When you're ready to render the document, you can call the html method on the "root" selection:

$.root().html();
//=>  <html>
//      <head></head>
//      <body>
//        <ul id="fruits">
//          <li class="apple">Apple</li>
//          <li class="orange">Orange</li>
//          <li class="pear">Pear</li>
//        </ul>
//      </body>
//    </html>

If you want to render the outerHTML of a selection, you can use the outerHTML prop:

$('.pear').prop('outerHTML');
//=> <li class="pear">Pear</li>

You may also render the text content of a Cheerio object using the text method:

const $ = cheerio.load('This is <em>content</em>.');
$('body').text();
//=> This is content.

The "DOM Node" object

Cheerio collections are made up of objects that bear some resemblance to browser-based DOM nodes. You can expect them to define the following properties:

  • tagName
  • parentNode
  • previousSibling
  • nextSibling
  • nodeValue
  • firstChild
  • childNodes
  • lastChild

Screencasts

https://vimeo.com/31950192

This video tutorial is a follow-up to Nettut's "How to Scrape Web Pages with Node.js and jQuery", using cheerio instead of JSDOM + jQuery. This video shows how easy it is to use cheerio and how much faster cheerio is than JSDOM + jQuery.

Cheerio in the real world

Are you using cheerio in production? Add it to the wiki!

Sponsors

Does your company use Cheerio in production? Please consider sponsoring this project! Your help will allow maintainers to dedicate more time and resources to its development and support.

Headlining Sponsors

Tidelift Github AirBnB HasData

Other Sponsors

OnlineCasinosSpelen Nieuwe-Casinos.net

Backers

Become a backer to show your support for Cheerio and help us maintain and improve this open source project.

Vasy Kafidoff

License

MIT