cheerio vs dom-serializer vs domutils vs htmlparser2 vs jsdom vs parse5
HTML Parsing and DOM Manipulation Strategies in Node.js
cheeriodom-serializerdomutilshtmlparser2jsdomparse5Similar Packages:

HTML Parsing and DOM Manipulation Strategies in Node.js

These libraries form the backbone of HTML processing in the Node.js ecosystem, serving distinct roles from fast, jQuery-like scraping to full browser environment emulation. cheerio offers a familiar API for server-side DOM manipulation without the overhead of a real browser. jsdom provides a complete implementation of web standards, allowing code written for browsers to run in Node. The remaining packagesβ€”htmlparser2, parse5, domutils, and dom-serializerβ€”are lower-level primitives that handle specific parts of the parsing and serialization pipeline, often used internally by higher-level tools or for specialized high-performance tasks.

Npm Package Weekly Downloads Trend

3 Years

Github Stars Ranking

Stat Detail

Package
Downloads
Stars
Size
Issues
Publish
License
cheerio030,4741.01 MB547 months agoMIT
dom-serializer014738.2 kB24 months agoMIT
domutils0227119 kB116 months agoBSD-2-Clause
htmlparser204,789235 kB125 months agoMIT
jsdom021,6687.09 MB403a month agoMIT
parse503,928337 kB404 months agoMIT

HTML Parsing and DOM Manipulation Strategies in Node.js

Processing HTML in Node.js is not a one-size-fits-all task. The ecosystem offers a range of tools, from fast, lenient scrapers to full browser simulations. Choosing the right one depends on whether you need to simply extract text, manipulate a DOM tree, or execute JavaScript. Let's break down how these six packages tackle common engineering challenges.

⚑ Parsing Speed and Spec Compliance

The first decision is often between raw speed and strict adherence to web standards.

htmlparser2 prioritizes speed and forgiveness. It can parse broken HTML that would crash a browser. It uses a streaming approach, making it memory efficient for large files.

import * as htmlparser2 from 'htmlparser2';

const parser = new htmlparser2.Parser({
  onopentag(name, attribs) {
    if (name === "script") {
      console.log("Found script tag:", attribs);
    }
  },
  ontext(text) {
    console.log("Text content:", text);
  }
}, { decodeEntities: true });

parser.write("<div><script>alert('hi')</script></div>");
parser.end();

parse5 prioritizes correctness. It implements the official HTML5 parsing algorithm exactly as browsers do. If you need your tool to behave exactly like Chrome or Firefox, this is the choice.

import * as parse5 from 'parse5';

const document = parse5.parse('<div><script>alert("hi")</script></div>');

// Traverse the AST manually
function traverse(node) {
  if (node.tagName === 'script') {
    console.log('Found script tag in spec-compliant AST');
  }
  if (node.childNodes) {
    node.childNodes.forEach(traverse);
  }
}

traverse(document);

jsdom uses parse5 internally but wraps it in a full DOM implementation. It is slower because it builds a complete browser environment, but it guarantees that the resulting DOM matches what a user would see.

import { JSDOM } from 'jsdom';

const dom = new JSDOM('<div><script>alert("hi")</script></div>', { runScripts: 'dangerously' });
const doc = dom.window.document;

// Access via standard DOM API
const script = doc.querySelector('script');
console.log('Script found via querySelector:', script.textContent);

πŸ•΅οΈβ€β™€οΈ Data Extraction: jQuery Syntax vs Standard DOM

How you query the parsed HTML depends on the library's API surface.

cheerio loads HTML into a structure that mimics jQuery. This is incredibly familiar for frontend developers and concise for scraping tasks.

import * as cheerio from 'cheerio';

const html = '<ul><li class="item">Apple</li><li class="item">Banana</li></ul>';
const $ = cheerio.load(html);

// jQuery-style selection
const items = $('.item').map((i, el) => $(el).text()).get();
console.log(items); // ['Apple', 'Banana']

jsdom gives you the real browser document object. You use standard querySelector and getElementById methods. This is better if you are sharing code between client and server.

import { JSDOM } from 'jsdom';

const html = '<ul><li class="item">Apple</li><li class="item">Banana</li></ul>';
const dom = new JSDOM(html);
const doc = dom.window.document;

// Standard DOM selection
const items = Array.from(doc.querySelectorAll('.item')).map(el => el.textContent);
console.log(items); // ['Apple', 'Banana']

domutils provides low-level helper functions for traversing the DOM tree produced by htmlparser2. It lacks a CSS selector engine by itself but is powerful for custom logic.

import * as htmlparser2 from 'htmlparser2';
import * as domutils from 'domutils';

const html = '<ul><li class="item">Apple</li></ul>';
const dom = htmlparser2.parseDOM(html);

// Find elements using domutils helpers
const items = domutils.getElementsByTagName('li', dom);
const text = items.map(node => domutils.getText(node));
console.log(text); // ['Apple']

πŸ”„ Executing JavaScript and Handling Dynamic Content

Some sites render content via JavaScript. Static parsers cannot see this content.

jsdom is the only package in this list capable of executing scripts. You can configure it to run code just like a browser.

import { JSDOM } from 'jsdom';

const html = `<div id="app"></div><script>document.getElementById('app').innerText = 'Loaded';</script>`;
const dom = new JSDOM(html, { runScripts: 'dangerously' });

// Wait for script execution
setTimeout(() => {
  console.log(dom.window.document.getElementById('app').innerText); 
  // Output: 'Loaded'
}, 100);

cheerio, htmlparser2, parse5, domutils, and dom-serializer do not execute JavaScript. They treat <script> tags as plain text or ignore them entirely. If you use cheerio on a React app that renders client-side, you will only see the empty root div.

import * as cheerio from 'cheerio';

const html = `<div id="app"></div><script>document.getElementById('app').innerText = 'Loaded';</script>`;
const $ = cheerio.load(html);

// Script is not executed; content remains empty
console.log($('#app').text()); 
// Output: '' (Empty string)

πŸ› οΈ Serialization: Turning DOM Back into HTML

After manipulating the DOM, you often need to convert it back to a string.

dom-serializer is a dedicated tool for converting DOM nodes (specifically those from htmlparser2 or domhandler) back into HTML strings. It is often used internally but can be used directly.

import * as htmlparser2 from 'htmlparser2';
import serialize from 'dom-serializer';

const dom = htmlparser2.parseDOM('<div class="test">Hello</div>');
// Modify the DOM directly (e.g., change text)
dom[0].children[0].data = 'World';

const html = serialize(dom);
console.log(html); // <div class="test">World</div>

cheerio handles serialization automatically via the .html() or .root().html() methods.

import * as cheerio from 'cheerio';

const $ = cheerio.load('<div class="test">Hello</div>');
$('div').text('World');

console.log($.html()); // <div class="test">World</div>

jsdom uses the standard browser .innerHTML or .serialize() methods.

import { JSDOM } from 'jsdom';

const dom = new JSDOM('<div class="test">Hello</div>');
const div = dom.window.document.querySelector('div');
div.textContent = 'World';

console.log(div.outerHTML); // <div class="test">World</div>

πŸ“¦ Architecture: Building Your Own Pipeline

Sometimes you don't want a full library. You might want to assemble your own pipeline using primitives.

htmlparser2 + domutils + dom-serializer creates a lightweight, custom scraper stack. This gives you maximum control and minimal dependencies compared to pulling in cheerio or jsdom.

import * as htmlparser2 from 'htmlparser2';
import * as domutils from 'domutils';
import serialize from 'dom-serializer';

// 1. Parse
const dom = htmlparser2.parseDOM('<p>Original</p>');

// 2. Manipulate using domutils
const pTag = domutils.getElementsByTagName('p', dom)[0];
if (pTag) {
  pTag.children[0].data = 'Modified';
}

// 3. Serialize
const result = serialize(dom);
console.log(result); // <p>Modified</p>

parse5 is often used as the foundation for tools like jsdom or linting utilities where spec compliance is non-negotiable. It returns its own AST format, which requires specific traversal logic.

import * as parse5 from 'parse5';

// Parse into parse5's specific AST format
const ast = parse5.parse('<p>Original</p>');

// Note: Modifying parse5 ASTs directly is complex and verbose 
// compared to domutils/cheerio, usually requiring a custom walker.
// This highlights why parse5 is often a backend for other tools.

🌐 Real-World Scenarios

Scenario 1: Scraping a Static Blog

You need to extract titles and links from a static HTML site.

  • βœ… Best choice: cheerio
  • Why? The jQuery API is concise, and you don't need to execute scripts.
const $ = cheerio.load(html);
const posts = $('.post').map((i, el) => ({
  title: $(el).find('h2').text(),
  link: $(el).find('a').attr('href')
})).get();

Scenario 2: Testing a React Component

You need to render a component and test its behavior in a Node environment.

  • βœ… Best choice: jsdom
  • Why? You need window, document, and event simulation.
const dom = new JSDOM('<!DOCTYPE html><html><body></body></html>');
global.window = dom.window;
global.document = dom.window.document;
// Now you can mount React components

Scenario 3: Building an HTML Linter

You need to verify if HTML follows the exact HTML5 specification.

  • βœ… Best choice: parse5
  • Why? It catches errors that lenient parsers like htmlparser2 might ignore.
const errors = [];
const parser = new parse5.Parser({
  onParseError(error) {
    errors.push(error);
  }
});
parser.parse(html);

Scenario 4: Processing Massive Log Files

You need to parse gigabytes of HTML logs without crashing memory.

  • βœ… Best choice: htmlparser2 (Streaming mode)
  • Why? It supports streaming chunks of data, processing them as they arrive.
const parser = new htmlparser2.Parser({
  onopentag(name) { /* process tag */ }
});
stream.on('data', chunk => parser.write(chunk));
stream.on('end', () => parser.end());

πŸ“Š Summary: Key Differences

Featurecheeriojsdomhtmlparser2parse5domutils / dom-serializer
Primary GoaljQuery-like scrapingBrowser emulationFast, lenient parsingSpec-compliant parsingLow-level DOM utils
Executes JS❌ Noβœ… Yes❌ No❌ No❌ No
API StylejQuery ($)Standard DOMEvents / CallbacksAST TraversalHelper Functions
Speed⚑ Fast🐒 Slow (Heavy)⚑⚑ Very Fast⚑ Fast⚑ Very Fast
StrictnessLenientStrict (Browser-like)LenientStrict (Spec)N/A (Utility)

πŸ’‘ The Big Picture

cheerio is the daily driver for scraping. It strikes the best balance between ease of use and performance for static content.

jsdom is the heavy artillery. Use it when you absolutely need a browser environment, but be aware of the performance cost.

htmlparser2, parse5, domutils, and dom-serializer are the building blocks. Use them when you need to optimize for specific constraints like streaming, strict compliance, or minimal bundle size, or when you are building your own higher-level tools.

Final Thought: Don't reach for jsdom unless you need to run scripts. Don't build a custom parser pipeline unless cheerio doesn't fit your specific performance or compliance needs. Start with the high-level tool that matches your data source, and drop down to the primitives only when necessary.

How to Choose: cheerio vs dom-serializer vs domutils vs htmlparser2 vs jsdom vs parse5

  • cheerio:

    Choose cheerio when you need to scrape websites or manipulate HTML strings with a jQuery-like syntax in a lightweight server environment. It is ideal for extracting data from static HTML where executing JavaScript is not required. Avoid it if you need to handle complex, modern web standards or execute client-side scripts.

  • dom-serializer:

    Choose dom-serializer only if you are building a custom tooling pipeline and already have a DOM tree (compliant with htmlparser2 or domhandler) that needs to be converted back into an HTML string. It is rarely used directly by application developers; instead, it serves as a dependency for higher-level libraries.

  • domutils:

    Choose domutils when you need low-level helper functions to traverse, query, or manipulate DOM nodes generated by htmlparser2 or cheerio. It is best suited for library authors or advanced users who need fine-grained control over node structures without the overhead of a full query selector engine.

  • htmlparser2:

    Choose htmlparser2 when you need a fast, forgiving HTML parser that can handle broken markup often found in the wild. It is excellent for streaming large files or building custom scrapers where speed and tolerance for errors are more important than strict spec compliance.

  • jsdom:

    Choose jsdom when your application relies on browser-specific APIs like window, document, or needs to execute JavaScript embedded within the HTML. It is the go-to solution for testing frontend components in Node or scraping dynamic sites that rely on client-side rendering, despite its higher resource cost.

  • parse5:

    Choose parse5 when strict adherence to the official HTML5 parsing algorithm is critical, such as in linters, sanitizers, or tools that must match browser behavior exactly. It is the most spec-compliant parser available but does not include a DOM implementation or query engine out of the box.

README for cheerio

cheerio

The fast, flexible, and elegant library for parsing and manipulating HTML and XML.

δΈ­ζ–‡ζ–‡ζ‘£ (Chinese Readme)

import * as cheerio from 'cheerio';
const $ = cheerio.load('<h2 class="title">Hello world</h2>');

$('h2.title').text('Hello there!');
$('h2').addClass('welcome');

$.html();
//=> <html><head></head><body><h2 class="title welcome">Hello there!</h2></body></html>

Installation

Install Cheerio using a package manager like npm, yarn, or bun.

npm install cheerio
# or
bun add cheerio

Features

❀ Proven syntax: Cheerio implements a subset of core jQuery. Cheerio removes all the DOM inconsistencies and browser cruft from the jQuery library, revealing its truly gorgeous API.

ϟ Blazingly fast: Cheerio works with a very simple, consistent DOM model. As a result parsing, manipulating, and rendering are incredibly efficient.

❁ Incredibly flexible: Cheerio wraps around parse5 for parsing HTML and can optionally use the forgiving htmlparser2. Cheerio can parse nearly any HTML or XML document. Cheerio works in both browser and server environments.

API

Loading

First you need to load in the HTML. This step in jQuery is implicit, since jQuery operates on the one, baked-in DOM. With Cheerio, we need to pass in the HTML document.

// ESM or TypeScript:
import * as cheerio from 'cheerio';

// In other environments:
const cheerio = require('cheerio');

const $ = cheerio.load('<ul id="fruits">...</ul>');

$.html();
//=> <html><head></head><body><ul id="fruits">...</ul></body></html>

Selectors

Once you've loaded the HTML, you can use jQuery-style selectors to find elements within the document.

$( selector, [context], [root] )

selector searches within the context scope which searches within the root scope. selector and context can be a string expression, DOM Element, array of DOM elements, or cheerio object. root, if provided, is typically the HTML document string.

This selector method is the starting point for traversing and manipulating the document. Like in jQuery, it's the primary method for selecting elements in the document.

$('.apple', '#fruits').text();
//=> Apple

$('ul .pear').attr('class');
//=> pear

$('li[class=orange]').html();
//=> Orange

Rendering

When you're ready to render the document, you can call the html method on the "root" selection:

$.root().html();
//=>  <html>
//      <head></head>
//      <body>
//        <ul id="fruits">
//          <li class="apple">Apple</li>
//          <li class="orange">Orange</li>
//          <li class="pear">Pear</li>
//        </ul>
//      </body>
//    </html>

If you want to render the outerHTML of a selection, you can use the outerHTML prop:

$('.pear').prop('outerHTML');
//=> <li class="pear">Pear</li>

You may also render the text content of a Cheerio object using the text method:

const $ = cheerio.load('This is <em>content</em>.');
$('body').text();
//=> This is content.

The "DOM Node" object

Cheerio collections are made up of objects that bear some resemblance to browser-based DOM nodes. You can expect them to define the following properties:

  • tagName
  • parentNode
  • previousSibling
  • nextSibling
  • nodeValue
  • firstChild
  • childNodes
  • lastChild

Screencasts

https://vimeo.com/31950192

This video tutorial is a follow-up to Nettut's "How to Scrape Web Pages with Node.js and jQuery", using cheerio instead of JSDOM + jQuery. This video shows how easy it is to use cheerio and how much faster cheerio is than JSDOM + jQuery.

Cheerio in the real world

Are you using cheerio in production? Add it to the wiki!

Sponsors

Does your company use Cheerio in production? Please consider sponsoring this project! Your help will allow maintainers to dedicate more time and resources to its development and support.

Headlining Sponsors

Tidelift Github AirBnB HasData

Other Sponsors

OnlineCasinosSpelen Nieuwe-Casinos.net

Backers

Become a backer to show your support for Cheerio and help us maintain and improve this open source project.

Vasy Kafidoff

License

MIT