Skip to content

EthenEthenEthen

Open Source Model Profile · Grogros

dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-Al4-wmToken-d4-APP

This Grogros model is a 1.24B-parameter Llama text-generation fine-tune. Its model card identifies a Llama-3.2-1B-Instruct base with only hyperparameter detail.

Publisher
Grogros
Task
text-generation
Model type
llama
License
llama3.2
Library
transformers
Publication status
Accepted · not indexed

Model overview

This Grogros model is published as a text-generation fine-tune of Meta-Llama-3.2-1B-Instruct. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 1,235,816,448 parameters. According to the model card, the training dataset is recorded as None, with llama3.2 recorded as the license.

Recorded capabilities

Documented Llama-3.2-1B-Instruct base

Hub tags and the model card both identify Meta-Llama-3.2-1B-Instruct as the base and fine-tune source.

Llama 1.24B configuration

Captured config identifies LlamaForCausalLM and llama, with Safetensors metadata reporting 1,235,816,448 parameters.

Reported hyperparameter record

According to the model card, training lists learning rate 2e-05, batch size 8, Adafactor optimizer, cosine scheduling, and 2,500 training steps.

Use cases in the source record

  • Conversational text-generation experiments using the Transformers stack recorded for this Llama model.

Limitations and unknowns

  • The model card states More information needed for intended uses, limitations, and training and evaluation data.
  • No evaluation results, dataset detail, context-window value, or hardware requirement was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: Grogros/dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-Al4-wmToken-d4-APP

Captured: Unknown. Processed: 2026-09-07T19:35:05.811843+00:00.

dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-Al4-wmToken-d4-APP This model is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the None dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 8 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 64 optimizer: Use OptimizerNames.ADAFACTOR and the args are: No additional optimizer arguments lr_scheduler_t…

F001F002F003F004F005F006F007F008F009F010F011F012