Markdownee is a web scraper that crawls websites and saves their content as Markdown, HTML or plain text for LLMs, retrieval pipelines, and research datasets.

Choose which links to follow, set page and depth limits, and select how much page content to keep. Control tables, links, images, and user comments separately.

- Boilerplate removal uses Trafilatura Core, our open-source extraction engine based on Trafilatura, available in TypeScript and Python. Markdownee’s TypeScript and Python versions each use the matching Core implementation. - The Core in its name means it is reduced to one task: extracting main content by removing boilerplate. Other packages handle output conversion, including Markdown. - Core's TypeScript engine ports the original Python Trafilatura, with go-trafilatura as a DOM translation aid. Its native Python engine translates that port. - Compared with Mozilla Readability, Trafilatura and Trafilatura Core use layered structural and content heuristics with fallback and recall escalation, rather than centering extraction on the candidate scoring inherited from Arc90’s original readability.js article extractor; Trafilatura Core also offers configuration options for boilerplate removal.

- Save image files with the optional image downloading mode

Try Markdownee with one npx command — no browser install or API key is needed. Use Node.js 22.22.2+ on the 22.x line, 24.15.0+ on 24.x, or 26+:

npx --package=@markdownee/markdownee markdownee fetch \

https://en.wikipedia.org/wiki/Web_scraping \

--crawler-type cheerio --save markdown-file -o page.md

Open page.md to inspect the extracted article. --crawler-type cheerio fetches over plain HTTP. For a whole-site crawl or specific formats, use the playground to build a command visually, then copy it.