Package profile
lxml-html-clean
- Summary: HTML cleaner from lxml project
- Author: Lumír Balhar
- License: BSD-3-Clause
- Homepage: https://github.com/fedora-python/lxml_html_clean/
- Documentation: https://lxml-html-clean.readthedocs.io/
- Source: https://github.com/fedora-python/lxml_html_clean
- Number of releases: 13
- First release: 0.1.0 on 2024-02-26
- Latest release: 0.4.5 on 2026-05-20
- Latest release size: 14.2 KB (pure Python wheel)
4 known vulnerabilitiesLatest: GHSA-4jhm-jv67-739f — `lxml_html_clean.Cleaner` does not strip `javascript:` URLs from namespaced URL attributes View all →
Dependencies
Lxml-html-clean has one dependency (non-optional).View all 1 dependencies →Dependent packages
| Package | Optional | Group |
|---|---|---|
| extruct | false | |
| html-text | false | |
| llama-index-readers-web | false | |
| news-please | false | |
| readability-lxml | false |
Similar packages
- prettierfierIntelligently pretty-print HTML/XML with inline tags.
- markuppyAn HTML/XML generator
- elementpathXPath 1.0/2.0/3.0/3.1 parsers and selectors for ElementTree and lxml
- bleachAn easy safelist-based HTML-sanitizing tool.
- mammothConvert Word documents from docx to simple and clean HTML and Markdown
- parselParsel is a library to extract data from HTML and XML using XPath and CSS selectors
- htmldocxConvert html to docx