These libraries form the backbone of HTML processing in the Node.js ecosystem, serving distinct roles from fast, jQuery-like scraping to full browser environment emulation. cheerio offers a familiar API for server-side DOM manipulation without the overhead of a real browser. jsdom provides a complete implementation of web standards, allowing code written for browsers to run in Node. The remaining packagesβhtmlparser2, parse5, domutils, and dom-serializerβare lower-level primitives that handle specific parts of the parsing and serialization pipeline, often used internally by higher-level tools or for specialized high-performance tasks.
Processing HTML in Node.js is not a one-size-fits-all task. The ecosystem offers a range of tools, from fast, lenient scrapers to full browser simulations. Choosing the right one depends on whether you need to simply extract text, manipulate a DOM tree, or execute JavaScript. Let's break down how these six packages tackle common engineering challenges.
The first decision is often between raw speed and strict adherence to web standards.
htmlparser2 prioritizes speed and forgiveness. It can parse broken HTML that would crash a browser. It uses a streaming approach, making it memory efficient for large files.
import * as htmlparser2 from 'htmlparser2';
const parser = new htmlparser2.Parser({
onopentag(name, attribs) {
if (name === "script") {
console.log("Found script tag:", attribs);
}
},
ontext(text) {
console.log("Text content:", text);
}
}, { decodeEntities: true });
parser.write("<div><script>alert('hi')</script></div>");
parser.end();
parse5 prioritizes correctness. It implements the official HTML5 parsing algorithm exactly as browsers do. If you need your tool to behave exactly like Chrome or Firefox, this is the choice.
import * as parse5 from 'parse5';
const document = parse5.parse('<div><script>alert("hi")</script></div>');
// Traverse the AST manually
function traverse(node) {
if (node.tagName === 'script') {
console.log('Found script tag in spec-compliant AST');
}
if (node.childNodes) {
node.childNodes.forEach(traverse);
}
}
traverse(document);
jsdom uses parse5 internally but wraps it in a full DOM implementation. It is slower because it builds a complete browser environment, but it guarantees that the resulting DOM matches what a user would see.
import { JSDOM } from 'jsdom';
const dom = new JSDOM('<div><script>alert("hi")</script></div>', { runScripts: 'dangerously' });
const doc = dom.window.document;
// Access via standard DOM API
const script = doc.querySelector('script');
console.log('Script found via querySelector:', script.textContent);
How you query the parsed HTML depends on the library's API surface.
cheerio loads HTML into a structure that mimics jQuery. This is incredibly familiar for frontend developers and concise for scraping tasks.
import * as cheerio from 'cheerio';
const html = '<ul><li class="item">Apple</li><li class="item">Banana</li></ul>';
const $ = cheerio.load(html);
// jQuery-style selection
const items = $('.item').map((i, el) => $(el).text()).get();
console.log(items); // ['Apple', 'Banana']
jsdom gives you the real browser document object. You use standard querySelector and getElementById methods. This is better if you are sharing code between client and server.
import { JSDOM } from 'jsdom';
const html = '<ul><li class="item">Apple</li><li class="item">Banana</li></ul>';
const dom = new JSDOM(html);
const doc = dom.window.document;
// Standard DOM selection
const items = Array.from(doc.querySelectorAll('.item')).map(el => el.textContent);
console.log(items); // ['Apple', 'Banana']
domutils provides low-level helper functions for traversing the DOM tree produced by htmlparser2. It lacks a CSS selector engine by itself but is powerful for custom logic.
import * as htmlparser2 from 'htmlparser2';
import * as domutils from 'domutils';
const html = '<ul><li class="item">Apple</li></ul>';
const dom = htmlparser2.parseDOM(html);
// Find elements using domutils helpers
const items = domutils.getElementsByTagName('li', dom);
const text = items.map(node => domutils.getText(node));
console.log(text); // ['Apple']
Some sites render content via JavaScript. Static parsers cannot see this content.
jsdom is the only package in this list capable of executing scripts. You can configure it to run code just like a browser.
import { JSDOM } from 'jsdom';
const html = `<div id="app"></div><script>document.getElementById('app').innerText = 'Loaded';</script>`;
const dom = new JSDOM(html, { runScripts: 'dangerously' });
// Wait for script execution
setTimeout(() => {
console.log(dom.window.document.getElementById('app').innerText);
// Output: 'Loaded'
}, 100);
cheerio, htmlparser2, parse5, domutils, and dom-serializer do not execute JavaScript. They treat <script> tags as plain text or ignore them entirely. If you use cheerio on a React app that renders client-side, you will only see the empty root div.
import * as cheerio from 'cheerio';
const html = `<div id="app"></div><script>document.getElementById('app').innerText = 'Loaded';</script>`;
const $ = cheerio.load(html);
// Script is not executed; content remains empty
console.log($('#app').text());
// Output: '' (Empty string)
After manipulating the DOM, you often need to convert it back to a string.
dom-serializer is a dedicated tool for converting DOM nodes (specifically those from htmlparser2 or domhandler) back into HTML strings. It is often used internally but can be used directly.
import * as htmlparser2 from 'htmlparser2';
import serialize from 'dom-serializer';
const dom = htmlparser2.parseDOM('<div class="test">Hello</div>');
// Modify the DOM directly (e.g., change text)
dom[0].children[0].data = 'World';
const html = serialize(dom);
console.log(html); // <div class="test">World</div>
cheerio handles serialization automatically via the .html() or .root().html() methods.
import * as cheerio from 'cheerio';
const $ = cheerio.load('<div class="test">Hello</div>');
$('div').text('World');
console.log($.html()); // <div class="test">World</div>
jsdom uses the standard browser .innerHTML or .serialize() methods.
import { JSDOM } from 'jsdom';
const dom = new JSDOM('<div class="test">Hello</div>');
const div = dom.window.document.querySelector('div');
div.textContent = 'World';
console.log(div.outerHTML); // <div class="test">World</div>
Sometimes you don't want a full library. You might want to assemble your own pipeline using primitives.
htmlparser2 + domutils + dom-serializer creates a lightweight, custom scraper stack. This gives you maximum control and minimal dependencies compared to pulling in cheerio or jsdom.
import * as htmlparser2 from 'htmlparser2';
import * as domutils from 'domutils';
import serialize from 'dom-serializer';
// 1. Parse
const dom = htmlparser2.parseDOM('<p>Original</p>');
// 2. Manipulate using domutils
const pTag = domutils.getElementsByTagName('p', dom)[0];
if (pTag) {
pTag.children[0].data = 'Modified';
}
// 3. Serialize
const result = serialize(dom);
console.log(result); // <p>Modified</p>
parse5 is often used as the foundation for tools like jsdom or linting utilities where spec compliance is non-negotiable. It returns its own AST format, which requires specific traversal logic.
import * as parse5 from 'parse5';
// Parse into parse5's specific AST format
const ast = parse5.parse('<p>Original</p>');
// Note: Modifying parse5 ASTs directly is complex and verbose
// compared to domutils/cheerio, usually requiring a custom walker.
// This highlights why parse5 is often a backend for other tools.
You need to extract titles and links from a static HTML site.
cheerioconst $ = cheerio.load(html);
const posts = $('.post').map((i, el) => ({
title: $(el).find('h2').text(),
link: $(el).find('a').attr('href')
})).get();
You need to render a component and test its behavior in a Node environment.
jsdomwindow, document, and event simulation.const dom = new JSDOM('<!DOCTYPE html><html><body></body></html>');
global.window = dom.window;
global.document = dom.window.document;
// Now you can mount React components
You need to verify if HTML follows the exact HTML5 specification.
parse5htmlparser2 might ignore.const errors = [];
const parser = new parse5.Parser({
onParseError(error) {
errors.push(error);
}
});
parser.parse(html);
You need to parse gigabytes of HTML logs without crashing memory.
htmlparser2 (Streaming mode)const parser = new htmlparser2.Parser({
onopentag(name) { /* process tag */ }
});
stream.on('data', chunk => parser.write(chunk));
stream.on('end', () => parser.end());
| Feature | cheerio | jsdom | htmlparser2 | parse5 | domutils / dom-serializer |
|---|---|---|---|---|---|
| Primary Goal | jQuery-like scraping | Browser emulation | Fast, lenient parsing | Spec-compliant parsing | Low-level DOM utils |
| Executes JS | β No | β Yes | β No | β No | β No |
| API Style | jQuery ($) | Standard DOM | Events / Callbacks | AST Traversal | Helper Functions |
| Speed | β‘ Fast | π’ Slow (Heavy) | β‘β‘ Very Fast | β‘ Fast | β‘ Very Fast |
| Strictness | Lenient | Strict (Browser-like) | Lenient | Strict (Spec) | N/A (Utility) |
cheerio is the daily driver for scraping. It strikes the best balance between ease of use and performance for static content.
jsdom is the heavy artillery. Use it when you absolutely need a browser environment, but be aware of the performance cost.
htmlparser2, parse5, domutils, and dom-serializer are the building blocks. Use them when you need to optimize for specific constraints like streaming, strict compliance, or minimal bundle size, or when you are building your own higher-level tools.
Final Thought: Don't reach for jsdom unless you need to run scripts. Don't build a custom parser pipeline unless cheerio doesn't fit your specific performance or compliance needs. Start with the high-level tool that matches your data source, and drop down to the primitives only when necessary.
Choose cheerio when you need to scrape websites or manipulate HTML strings with a jQuery-like syntax in a lightweight server environment. It is ideal for extracting data from static HTML where executing JavaScript is not required. Avoid it if you need to handle complex, modern web standards or execute client-side scripts.
Choose dom-serializer only if you are building a custom tooling pipeline and already have a DOM tree (compliant with htmlparser2 or domhandler) that needs to be converted back into an HTML string. It is rarely used directly by application developers; instead, it serves as a dependency for higher-level libraries.
Choose domutils when you need low-level helper functions to traverse, query, or manipulate DOM nodes generated by htmlparser2 or cheerio. It is best suited for library authors or advanced users who need fine-grained control over node structures without the overhead of a full query selector engine.
Choose htmlparser2 when you need a fast, forgiving HTML parser that can handle broken markup often found in the wild. It is excellent for streaming large files or building custom scrapers where speed and tolerance for errors are more important than strict spec compliance.
Choose jsdom when your application relies on browser-specific APIs like window, document, or needs to execute JavaScript embedded within the HTML. It is the go-to solution for testing frontend components in Node or scraping dynamic sites that rely on client-side rendering, despite its higher resource cost.
Choose parse5 when strict adherence to the official HTML5 parsing algorithm is critical, such as in linters, sanitizers, or tools that must match browser behavior exactly. It is the most spec-compliant parser available but does not include a DOM implementation or query engine out of the box.
import * as cheerio from 'cheerio';
const $ = cheerio.load('<h2 class="title">Hello world</h2>');
$('h2.title').text('Hello there!');
$('h2').addClass('welcome');
$.html();
//=> <html><head></head><body><h2 class="title welcome">Hello there!</h2></body></html>
Install Cheerio using a package manager like npm, yarn, or bun.
npm install cheerio
# or
bun add cheerio
β€ Proven syntax: Cheerio implements a subset of core jQuery. Cheerio removes all the DOM inconsistencies and browser cruft from the jQuery library, revealing its truly gorgeous API.
Ο Blazingly fast: Cheerio works with a very simple, consistent DOM model. As a result parsing, manipulating, and rendering are incredibly efficient.
β Incredibly flexible: Cheerio wraps around parse5 for parsing HTML and can optionally use the forgiving htmlparser2. Cheerio can parse nearly any HTML or XML document. Cheerio works in both browser and server environments.
First you need to load in the HTML. This step in jQuery is implicit, since jQuery operates on the one, baked-in DOM. With Cheerio, we need to pass in the HTML document.
// ESM or TypeScript:
import * as cheerio from 'cheerio';
// In other environments:
const cheerio = require('cheerio');
const $ = cheerio.load('<ul id="fruits">...</ul>');
$.html();
//=> <html><head></head><body><ul id="fruits">...</ul></body></html>
Once you've loaded the HTML, you can use jQuery-style selectors to find elements within the document.
selector searches within the context scope which searches within the root
scope. selector and context can be a string expression, DOM Element, array
of DOM elements, or cheerio object. root, if provided, is typically the HTML
document string.
This selector method is the starting point for traversing and manipulating the document. Like in jQuery, it's the primary method for selecting elements in the document.
$('.apple', '#fruits').text();
//=> Apple
$('ul .pear').attr('class');
//=> pear
$('li[class=orange]').html();
//=> Orange
When you're ready to render the document, you can call the html method on the
"root" selection:
$.root().html();
//=> <html>
// <head></head>
// <body>
// <ul id="fruits">
// <li class="apple">Apple</li>
// <li class="orange">Orange</li>
// <li class="pear">Pear</li>
// </ul>
// </body>
// </html>
If you want to render the
outerHTML
of a selection, you can use the outerHTML prop:
$('.pear').prop('outerHTML');
//=> <li class="pear">Pear</li>
You may also render the text content of a Cheerio object using the text
method:
const $ = cheerio.load('This is <em>content</em>.');
$('body').text();
//=> This is content.
Cheerio collections are made up of objects that bear some resemblance to browser-based DOM nodes. You can expect them to define the following properties:
tagNameparentNodepreviousSiblingnextSiblingnodeValuefirstChildchildNodeslastChildThis video tutorial is a follow-up to Nettut's "How to Scrape Web Pages with Node.js and jQuery", using cheerio instead of JSDOM + jQuery. This video shows how easy it is to use cheerio and how much faster cheerio is than JSDOM + jQuery.
Are you using cheerio in production? Add it to the wiki!
Does your company use Cheerio in production? Please consider sponsoring this project! Your help will allow maintainers to dedicate more time and resources to its development and support.
Headlining Sponsors
Other Sponsors
Become a backer to show your support for Cheerio and help us maintain and improve this open source project.
MIT