Best open source web scrapers for LLMs

For LLM and RAG pipelines, Crawl4AI and Firecrawl are the strongest open source picks because they return clean markdown and render JavaScript. Scrapy remains best for large, structured crawls, Playwright for interactive pages, and Trafilatura for fast article text extraction without a browser.
What makes a web scraper “LLM-ready”?
A classic scraper returns HTML or structured fields you define with selectors. An LLM-ready scraper goes one step further: it strips navigation, cookie banners and footers, then returns readable markdown or plain text that fits into a context window without wasting tokens.
The second difference is rendering. Many modern sites load content with JavaScript, so a scraper that only fetches raw HTML sees an empty shell. Tools built for LLM work usually drive a headless browser or offer it as an option.
The third is shape of output. RAG pipelines want chunks with metadata such as source URL, title and headings, so the retriever can cite where an answer came from.
- Clean markdown or text output, not raw HTML
- Optional JavaScript rendering through a headless browser
- Crawl controls: depth, URL filters, rate limits, robots.txt handling
- Metadata per page for citations and deduplication
- A way to run it on your own infrastructure
Which open source scrapers are worth using?
The field splits into LLM-native crawlers, general crawling frameworks and extraction libraries. You often combine one from each group rather than picking a single winner.
| Tool | Language | JS rendering | Output for LLMs | Best for | Trade-off |
|---|---|---|---|---|---|
| Crawl4AI | Python | Yes, via Playwright | Markdown, structured extraction | Self-hosted RAG ingestion | Browser runtime adds memory and setup |
| Firecrawl | TypeScript | Yes | Markdown, JSON, crawl API | Teams wanting an API-shaped service | Self-hosted version can lag the hosted one; AGPL-style copyleft, check the licence file |
| Scrapy | Python | Only with plugins | Whatever your pipeline emits | Large, structured, repeatable crawls | You write the cleaning step yourself |
| Playwright | TS, Python, Java, .NET | Yes, full browser | Raw DOM, you convert | Logins, clicks, infinite scroll | Not a crawler; no queue or dedup built in |
| Trafilatura | Python | No | Main text, markdown, XML | News and article pages at speed | Misses content that needs JavaScript |
| Crawlee | TypeScript, Python | Yes, optional | Raw data, you convert | Robust crawling with retries and proxies | More framework to learn |
When should you choose Crawl4AI?
Crawl4AI is a Python library designed around the LLM use case. It crawls with a headless browser, converts pages to markdown and supports extraction strategies, including CSS-based schemas and LLM-based extraction for messier pages.
Pick it when you control a Python ingestion pipeline and want everything in-process: fetch, clean, chunk, embed. It fits well next to LlamaIndex, LangChain or a hand-written loader.
The cost is operational. A browser per worker is heavier than an HTTP client, so plan memory and concurrency before you point it at thousands of URLs.
When should you choose Firecrawl?
Firecrawl exposes scraping as an API: send a URL, get markdown or structured JSON back, or start a crawl job that follows links. That shape suits multi-language teams and agents that call tools over HTTP.
It can be self-hosted, which matters if you scrape internal sites or cannot send URLs to a third party. Expect to run several services, typically an API, workers, a queue and a browser service.
Before building a commercial product on the self-hosted code, read the licence file carefully. Copyleft terms can affect how you offer a modified version as a network service.
Is Scrapy still relevant for AI projects?
Yes. Scrapy is mature, fast and built for crawling at scale, with middlewares, pipelines, throttling and resumable jobs. If you need millions of product pages or a nightly refresh of a documentation site, it is still one of the most dependable options.
What it lacks is LLM-specific cleaning. The common pattern is Scrapy for fetching and scheduling, then Trafilatura or a markdown converter in an item pipeline, then your chunker.
How to choose a scraper for your LLM pipeline
- Check whether target pages need JavaScript: view source and search for the text you want. If it is missing, you need a browser-based tool.
- Estimate volume. Tens of pages a day works with anything; hundreds of thousands favours Scrapy or Crawlee with a proper queue.
- Decide where it runs. Library inside your app (Crawl4AI, Trafilatura) or separate service with an API (Firecrawl).
- Test output quality on ten real pages and read the markdown yourself. Boilerplate that slips through becomes noise in retrieval.
- Confirm the licence fits your distribution model, especially for hosted products.
- Plan refresh: store content hashes so you only re-embed pages that changed.
Common mistakes when scraping for RAG
- Embedding navigation and footers, which makes every chunk look similar to the retriever.
- Ignoring robots.txt, terms of service and rate limits. Being blocked is the mild outcome; legal trouble is the serious one.
- Chunking by fixed character count and splitting tables or code blocks in half.
- Dropping the source URL, so the model cannot cite and you cannot debug bad answers.
- Running a browser for every page when most of the site is static HTML.
- Letting an agent scrape arbitrary URLs from user input without an allow list, which opens you to server-side request forgery against internal addresses.
A practical stack that works
For most teams a hybrid holds up: a fast HTTP fetch plus Trafilatura for static pages, with a fallback to Crawl4AI or Playwright when the extracted text is suspiciously short. Store raw HTML, cleaned markdown and a hash side by side, so you can re-clean without re-crawling.
If you want to compare these projects by licence, language and difficulty before committing, RepoLoot’s catalog tags scraping and data-ingestion repos with exactly those attributes.
How do you turn scraped pages into good RAG chunks?
Scraping is only half the job. The way you split pages decides what the retriever can find. Split on headings first, so each chunk carries one topic, and fall back to paragraph boundaries when a section is long.
Keep tables and code blocks whole even if a chunk runs over your target size. A half table is worse than a long chunk, because the model loses the column headers that give numbers meaning.
Prepend the page title and the heading path to each chunk before embedding. A chunk that reads “Pricing > Enterprise > Limits” is far easier to retrieve than the same text without context.
- Store: URL, title, heading path, crawl date, content hash
- Normalise whitespace and remove repeated boilerplate lines across pages
- Deduplicate near-identical pages such as print versions and tracking-parameter URLs
- Record the HTTP status and final URL after redirects
How much does it cost to run a scraper for LLMs?
The software is free; the costs are compute, bandwidth and, for some sites, proxies. Plain HTTP fetching is cheap and runs happily on a small server. Headless browsers are the expensive part, because each page loads scripts, images and fonts.
You can cut browser cost by blocking images, media and fonts, reusing browser contexts, and only rendering pages that failed a cheap first pass. Scheduling matters too: re-crawl changelogs daily, but stable reference pages weekly or monthly.
LLM-based extraction adds token costs on top. Use CSS or schema-based extraction for predictable layouts and reserve model calls for pages that genuinely vary.
Should agents scrape the web live?
Agents that browse on demand are useful for fresh questions, but they are slower and less predictable than a pre-built index. A good compromise is an index for your known sources plus a tightly scoped live fetch tool for anything newer.
Treat every live page as untrusted input. Web content can contain hidden instructions aimed at the model, so keep scraped text separated from system instructions and never let it trigger actions without checks.
Frequently asked questions
- What is the best free web scraper for LLMs?
- Crawl4AI is a strong free choice if you work in Python, because it renders JavaScript and outputs markdown directly. Firecrawl is the better fit if you want an HTTP API that any language or agent can call, and Trafilatura is ideal for static article pages.
- Can I self-host Firecrawl?
- Yes, Firecrawl publishes a self-hostable version. It runs as several services, including workers, a queue and a browser component, so it is heavier than a library. Some features of the hosted product may not be available, and you should read the licence file before commercial use.
- Do I need a headless browser to scrape for RAG?
- Only for pages that build content with JavaScript. Many documentation sites, blogs and news pages serve full HTML, and a plain HTTP fetch plus an extractor is faster and cheaper. Use a browser as a fallback when extracted text comes back empty or truncated.
- Is web scraping for AI training or RAG legal?
- It depends on jurisdiction, the site’s terms, the type of data and how you use it. Respect robots.txt, avoid personal data, honour rate limits and get legal advice for commercial projects. Scraping for internal retrieval is usually lower risk than redistributing content.