Package profile
htmldate
- Summary: Fast and robust extraction of original and updated publication dates from URLs and web pages.
- Author: Adrien Barbaresi <adrien.barbaresi@gmail.com>
- License: Apache-2.0
- Homepage: https://htmldate.readthedocs.io
- Source: https://github.com/adbar/htmldate
- Number of releases: 61
- First release: 0.1.0 on 2017-08-25
- Latest release: 1.11.0 on 2026-10-02
- Latest release size: 30.8 KB (pure Python wheel)
Dependencies
Htmldate has 13 dependencies, 8 of which optional.Dependent packages
| Package | Optional | Group |
|---|---|---|
| trafilatura | false |
Similar packages
- dateparserDate parsing library designed to parse dates from HTML pages
- urlextractCollects and extracts URLs from given text.
- newspaper3kSimplified python article discovery & extraction.
- dateformatParse and format dates quickly
- scrapyA high-level Web Crawling and Web Scraping framework
- datefinderExtract datetime objects from natural language text
- trafilaturaPython & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.