Skip to content

EthenEthenEthen

Open Source Model Profile · aubmindlab

bert-base-arabertv2

bert-base-arabertv2 is a 135.85M-parameter Arabic BERT fill-mask model from aubmindlab with Farasa pre-segmentation.

Publisher
aubmindlab
Task
fill-mask
Model type
bert
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

bert-base-arabertv2 is published by aubmindlab as a BERT fill-mask model for Arabic. The captured configuration identifies BertForMaskedLM and Safetensors metadata reports 135,851,522 parameters. According to the model card, it is the pre-segmented v2-base AraBERT variant at about 136M parameters, trained on about 200M sentences.

Recorded capabilities

Arabic BERT-Base fill-mask

According to the model card, AraBERT is an Arabic pretrained model on Google's BERT architecture using the same BERT-Base config.

Farasa pre-segmentation

The model card says AraBERTv1 and v2 use Farasa-segmented text, with v2 adding revised preprocessing and a new wordpiece vocabulary.

Documented downstream scope

According to the model card, the publisher evaluated AraBERT on sentiment analysis, ANERcorp named-entity recognition, and Arabic question answering, with v2-base training reported on TPUv3-8.

Use cases in the source record

  • Arabic fill-mask and language-understanding experiments using the publisher's documented preprocessing function.
  • Arabic sentiment analysis, named-entity recognition, and question-answering research in the task areas the model card names.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No license value was extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: aubmindlab/bert-base-arabertv2

Captured: Unknown. Processed: 2026-09-07T19:34:40.505128+00:00.

AraBERT v1 & v2 : Pre-training BERT for Arabic Language Understanding AraBERT is an Arabic pretrained lanaguage model based on Google's BERT architechture . AraBERT uses the same BERT-Base config. More details are available in the AraBERT Paper and in the AraBERT Meetup There are two versions of the model, AraBERTv0.1 and AraBERTv1, with the difference being that AraBERTv1 uses pre-segmented text where prefixes and suffixes were splitted using the Farasa Segmenter . We evalaute AraBERT models on different downstream tasks and compare them to mBERT , and other state of the art models ( To the extent of our knowledge ). The Tasks were…

F001F002F003F004F005F006F008F009F010F011F012F013F015