Skip to content

EthenEthenEthen

Open Source Model Profile · Grogros

dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1

dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1 is a 1.24B-parameter Llama fine-tune from Grogros. The model card lists meta-llama/Llama-3.2-1B-Instruct and openwebtext.

Publisher
Grogros
Task
text-generation
Model type
llama
License
llama3.2
Library
transformers
Publication status
Accepted · not indexed

Model overview

dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1 is published by Grogros as a Llama text-generation model. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 1,235,816,448 parameters. According to the model card, it fine-tunes meta-llama/Llama-3.2-1B-Instruct on openwebtext.

Recorded capabilities

Llama 3.2 1B openwebtext fine-tune

Captured configuration records LlamaForCausalLM with about 1.24B parameters, and the model card describes a fine-tune of meta-llama/Llama-3.2-1B-Instruct on openwebtext.

Documented Adafactor training setup

According to the model card, training used Adafactor, learning rate 2e-05, cosine scheduling with 0.1 warmup ratio, and 5,000 training steps.

Llama3.2 licensing record

Card data records llama3.2 for this model.

Use cases in the source record

  • Conversational text-generation experiments using the documented Llama-3.2-1B-Instruct plus openwebtext setup.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • Fine-tuning purpose and training details are publisher claims from the model card and were not independently verified.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: Grogros/dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1

Captured: Unknown. Processed: 2026-09-07T19:35:05.980935+00:00.

dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1 This model is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the openwebtext dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 8 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 64 optimizer: Use OptimizerNames.ADAFACTOR and the args are: No additional optimizer arguments lr_schedul…

F001F002F003F004F005F006F007F009F010F011F012