Menu ▾ ▴

#15 Help wanted: benchmark stronger LLMLingua and segmented LongLLMLingua adapters

open
nobody
help wanted (2)
2026-07-28
2026-07-28
Anonymous
No

Originally created by: aakashH242

Our frozen benchmark currently evaluates:

  • LLMLingua GPT-2 token-pruning adapter
  • LongLLMLingua GPT-2 single-context adapter
  • RECOMP NQ extractive sentence adapter
  • Trimwise Lexical and Hybrid

The results are fully disclosed and useful, but the two LLMLingua adapters are intentionally scoped: GPT-2 is a lightweight scorer, and LongLLMLingua currently receives one monolithic source string. Its context-level selection is more
meaningfully exercised with a source-backed list of passages.

We would welcome a reproducible stronger-baseline contribution.

## Requested work

  • Add a separately identified stronger LLMLingua configuration, preferably an officially documented scorer such as microsoft/phi-2.
  • Add a LongLLMLingua configuration that receives source-backed paragraph or section segments as a list.
  • Use the same stronger scorer for both LLMLingua and LongLLMLingua where possible.
  • Run a small VRAM/thermal preflight on representative long cases before attempting the full benchmark.
  • Preserve the existing frozen GPT-2 rows and reports. New runs must use new method/config identities and separate manifests.
  • Regenerate the v1.2 strict source-span-survival summary from the saved outputs.

## Contribution requirements

Please include:

  • exact model ID and revision;
  • adapter parameters and segmentation rule;
  • hardware, CUDA, package, and temperature record;
  • raw compression rows and result hashes;
  • v1.2 normalized contiguous source-span results, plus local ordered 80% and 90% sensitivities;
  • a short note describing any failures, OOMs, thermal pauses, or budget violations.

This issue is about a fairer and stronger comparison—not changing the existing result or overwriting historical artifacts.

Discussion


Log in to post a comment.