Web-Scraping

Frameworks und Tools zum Crawlen von Websites, Headless-Browsing und Datenextraktion.

Repositories

Node.js-Browserautomatisierungsbibliothek zur Steuerung von Chrome und Firefox über DevTools Protocol oder WebDriver BiDi. Standardmäßig headless, ideal für Web-Scraping, automatisierte Tests, Screenshots, PDF-Erstellung und Browserautomatisierung.

TypeScript
95.5k
16 hours ago

Open-Source-Webcrawler für LLMs optimiert, wandelt Webinhalte in sauberes Markdown für KI-Anwendungen um. Bietet asynchrone Verarbeitung, Browserautomatisierung und strukturierte Datenextraktion.

Python
77.9k
7 hours ago
D4Vinci/Scrapling

Scrapling ist ein adaptives Python-Web-Scraping-Framework, das alles von einzelnen Anfragen bis hin zu großangelegtem Crawling abdeckt. Der intelligente Parser findet Elemente nach Website-Änderungen automatisch wieder, integrierte Fetcher umgehen Anti-Bot-Systeme wie Cloudflare, und das Spider-Framework unterstützt konkurrentes Crawling mit Pause/Fortsetzen, Proxy-Rotation und KI-Integration über einen MCP-Server.

Python
73.6k
a day ago
scrapy/scrapy

Scrapy ist ein leistungsstarkes Python-Framework für Web-Crawling und -Scraping, das einen vollständigen Werkzeugkasten zur effizienten und skalierbaren Extraktion strukturierter Daten von Websites bereitstellt.

Python
63.8k
18 hours ago

Ein Multiplattform-Social-Media-Crawler für Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba und Zhihu. Basierend auf Playwright mit CDP-Modus bietet er Stichwortsuche, Extraktion von Beiträgen und verschachtelten Kommentaren, Erstellerprofil-Crawling, IP-Proxy-Pool, Login-Caching, Kommentar-Wortwolke, WebUI sowie mehrere Speicherformate wie CSV, JSON, Excel, SQLite und MySQL.

Python
62.0k
13 hours ago

The fast, flexible, and elegant library for parsing and manipulating HTML and XML.

TypeScript
30.4k
3 days ago

Elegant Scraper and Crawler Framework for Golang

Go
25.4k
2 months ago

⬛️ CLI tool and library for saving complete web pages as a single HTML file

Rust
15.4k
3 months ago