Skip to content

EthenEthenEthen

Open Source Model Profile · nazimali

Mistral-Nemo-Kurdish

Mistral-Nemo-Kurdish is a 12.25B-parameter Mistral text-generation model from nazimali. According to the model card, it continued pre-training for Kurdish language understanding.

Publisher
nazimali
Task
text-generation
Model type
mistral
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Mistral-Nemo-Kurdish is published by nazimali as a text-generation model. The captured configuration identifies MistralForCausalLM, and Safetensors metadata reports about 12.25B parameters. According to the model card, it continued pre-training Mistral-Nemo-Instruct-2407 on Kurdish Wikipedia with Unsloth.

Recorded capabilities

Mistral text-generation base

The captured configuration identifies MistralForCausalLM with model type mistral and Transformers support.

Kurdish continued pre-training

According to the model card, the model continued pre-training on Mistral-Nemo-Instruct-2407 with Kurdish Wikipedia articles.

Memory-efficient quantization note

According to the model card, bitsandbytes quantization is used so the model uses less memory.

Explicit evaluation gap

According to the model card, no standard Kurdish metric was found for evaluation, with future Kurmanji and Sorani work planned.

Published training log

According to the model card, the run reports Transformers 4.44.2, one A100 80GB, loss near 1.21, and full step and runtime figures.

Use cases in the source record

  • Kurdish-language understanding experiments building on the documented continued pre-training.
  • Further fine-tuning for Kurmanji and Sorani coverage as suggested by the publisher.

Limitations and unknowns

  • According to the model card, no standard Kurdish evaluation metric was available, so no verified performance is stated here.
  • According to the model card, further fine-tuning is recommended before treating this as a finished task model.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: nazimali/Mistral-Nemo-Kurdish

Captured: Unknown. Processed: 2026-09-07T19:34:52.912743+00:00.

Continued pre-training on mistralai/Mistral-Nemo-Instruct-2407 using the Kurdish wiki dataset with unsloth . This model should be further fine-tuned since the pre-training was to improve Kurdish language understanding. It's a quantized model using bitsandbytes so that it uses less memory. See bitsandbytes documentation . There isn't a standard or even a good Kurdish metric to evaluate the model (that I could find). Will make it my next project to create an evaluation so that there's a reproducible baseline for Kurdish. Will look into a multi-GPU training setup so don't have to wait all day for results. Would like to train it with bo…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F017