Originally created by: aakashH242
Our frozen benchmark currently evaluates:
- LLMLingua GPT-2 token-pruning adapter
- LongLLMLingua GPT-2 single-context adapter
- RECOMP NQ extractive sentence adapter
- Trimwise Lexical and Hybrid
The results are fully disclosed and useful, but the two LLMLingua adapters are intentionally scoped: GPT-2 is a lightweight scorer, and LongLLMLingua currently receives one monolithic source string. Its context-level selection is more
meaningfully exercised with a source-backed list of passages.
We would welcome a reproducible stronger-baseline contribution.
## Requested work
- Add a separately identified stronger LLMLingua configuration, preferably an officially documented scorer such as microsoft/phi-2.
- Add a LongLLMLingua configuration that receives source-backed paragraph or section segments as a list.
- Use the same stronger scorer for both LLMLingua and LongLLMLingua where possible.
- Run a small VRAM/thermal preflight on representative long cases before attempting the full benchmark.
- Preserve the existing frozen GPT-2 rows and reports. New runs must use new method/config identities and separate manifests.
- Regenerate the v1.2 strict source-span-survival summary from the saved outputs.
## Contribution requirements
Please include:
- exact model ID and revision;
- adapter parameters and segmentation rule;
- hardware, CUDA, package, and temperature record;
- raw compression rows and result hashes;
- v1.2 normalized contiguous source-span results, plus local ordered 80% and 90% sensitivities;
- a short note describing any failures, OOMs, thermal pauses, or budget violations.
This issue is about a fairer and stronger comparison—not changing the existing result or overwriting historical artifacts.