Package profile
trafilatura
- Summary: Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.
- Author: Adrien Barbaresi
- License: Apache-2.0
- Homepage: https://trafilatura.readthedocs.io
- Source: https://github.com/adbar/trafilatura
- Number of releases: 53
- First release: 0.0.1 on 2019-07-17
- Latest release: 2.3.0 on 2026-10-02
- Latest release size: 153.9 KB (pure Python wheel)
Dependencies
Trafilatura has 39 dependencies, 32 of which optional.Dependent packages
| Package | Optional | Group |
|---|---|---|
| agno | true | trafilatura |
| gnews | true | fulltext |
| instructor | true | trafilatura |
Similar packages
- livereloadPython LiveReload is an awesome tool for web developers
- importlib-metadataRead metadata from Python packages
- urlextractCollects and extracts URLs from given text.
- html-testrunnerA Test Runner in python, for Human Readable HTML Reports
- readability-lxmlfast html to text parser (article readability tool) with python 3 support
- scrapyA high-level Web Crawling and Web Scraping framework
- panelThe powerful data exploration & web app framework for Python.