Summary: To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
Weekly downloads over the last yearSeptemberOctoberNovemberDecember2026FebruaryMarchAprilMayJuneJulyAugustDate0102030405060708090100 thousand downloads per week
Source: ClickPy
Dependencies
Llmlingua has 12 dependencies, 6 of which optional.