Skip to content

EthenEthenEthen

Open Source Model Profile · Grogros

dmWM-llama-3.2-1B-Instruct-HarmData-Al4-OWT-d4-a0.25

dmWM-llama-3.2-1B-Instruct-HarmData-Al4-OWT-d4-a0.25 is a 1.24B-parameter Llama text-generation fine-tune from Grogros. According to the model card, it derives from Llama-3.2-1B-Instruct.

Publisher
Grogros
Task
text-generation
Model type
llama
License
llama3.2
Library
transformers
Publication status
Accepted · not indexed

Model overview

dmWM-llama-3.2-1B-Instruct-HarmData-Al4-OWT-d4-a0.25 is published by Grogros as a Transformers text-generation fine-tune. The captured configuration identifies LlamaForCausalLM with model type llama and about 1.24B Safetensors parameters. According to the model card, it fine-tunes meta-llama/Llama-3.2-1B-Instruct.

Recorded capabilities

Llama 3.2 1B lineage

According to the model card and hub tags, this is a fine-tune of meta-llama/Llama-3.2-1B-Instruct.

1.24B Llama build

Captured configuration identifies LlamaForCausalLM with 1235814400 Safetensors parameters.

Documented hyperparameters

The model card documents Adafactor optimization, cosine scheduling, 2500 training steps, and Transformers 4.46.3 with PyTorch 2.5.1 and related framework versions.

Use cases in the source record

  • Small-model conversational text-generation experiments built on Llama-3.2-1B-Instruct.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • The model card lists the training dataset as None and marks intended uses, limitations, and training data as needing more information.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: Grogros/dmWM-llama-3.2-1B-Instruct-HarmData-Al4-OWT-d4-a0.25

Captured: Unknown. Processed: 2026-09-07T19:35:05.860911+00:00.

dmWM-llama-3.2-1B-Instruct-HarmData-Al4-OWT-d4-a0.25 This model is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the None dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 4 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 32 optimizer: Use OptimizerNames.ADAFACTOR and the args are: No additional optimizer arguments lr_scheduler_type: cosine lr…

F001F002F003F004F005F006F007F009F010F011F012