Web Scraping & Crawling

Frameworks and tools for web scraping, crawling, and data extraction

Web Scraping & Crawling — comparison of firecrawl, crawl4ai, Scrapling, scrapy, crawlee
The context API to search, scrape, and interact with the web at scale. 🔥
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Scrapy, a fast high-level web crawling & scraping framework for Python.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Popularity
Stars170,26478,82775,62063,97925,453
Global Rank#48#207#226#313#1564
Weekly Activity(Aug 14 – Aug 20)
New Stars+14+1+800
Pushes50001
Issues Closed00000
Community
Forks9,4758,1697,55611,9171,632
Contributors1709131743142
Open Issues5421674464153
Project Info
OwnerfirecrawlunclecodeD4Vinciscrapyapify
LicenseAGPL-3.0Apache-2.0BSD-3-ClauseBSD-3-ClauseApache-2.0
LanguageTypeScriptPythonPythonPythonTypeScript
CreatedApr 2024May 2024Oct 2024Feb 2010Aug 2016