Skip to content

EthenEthenEthen

Open Source Model Profile · microsoft

llmlingua-2-bert-base-multilingual-cased-meetingbank

Microsoft's LLMLingua-2 MeetingBank model is a 177M-parameter multilingual BERT classifier for prompt compression. According to the model card, it scores each token's preservation probability to compress prompts.

Publisher
microsoft
Task
token-classification
Model type
bert
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

llmlingua-2-bert-base-multilingual-cased-meetingbank is published by Microsoft as a token-classification model. The captured configuration identifies BertForTokenClassification with model type bert. According to the model card, it is a multilingual cased BERT base model fine-tuned for task-agnostic prompt compression, introduced in the LLMLingua-2 paper (Pan et al., 2024).

Recorded capabilities

Token-level compression scoring

According to the model card, the model performs token classification where each token's preservation probability serves as the compression metric.

MeetingBank-seeded training

According to the model card, training uses an extractive compression dataset constructed with the LLMLingua-2 methodology from MeetingBank seed examples.

Documented downstream evaluation

According to the model card, compressed meeting transcripts can be evaluated on question answering and summarization with the linked dataset.

PromptCompressor integration

According to the model card, the llmlingua PromptCompressor exposes compression rate, forced tokens, chunk-end tokens, and word-label outputs.

Use cases in the source record

  • Compressing long meeting-transcript prompts at a chosen rate while preserving forced tokens such as sentence boundaries.
  • Evaluating question answering and summarization over compressed meeting transcripts with the linked dataset.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Training and methodology details come from the publisher model card and were not independently verified.

Source and provenance

Source: microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank

Captured: Unknown. Processed: 2026-09-07T19:34:51.973991+00:00.

LLMLingua-2-Bert-base-Multilingual-Cased-MeetingBank This model was introduced in the paper LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression (Pan et al, 2024) . It is a BERT multilingual base model (cased) finetuned to perform token classification for task agnostic prompt compression. The probability $p_{preserve}$ of each token $x_i$ is used as the metric for compression. This model is trained on the extractive text compression dataset constructed with the methodology proposed in the LLMLingua-2 , using training examples from MeetingBank (Hu et al, 2023) as the seed data. You can evaluate t…

F001F002F003F004F005F006F007F010F011F012F014F015