Documented Llama-3.2-1B-Instruct base
Hub tags and the model card both identify Meta-Llama-3.2-1B-Instruct as the base and fine-tune source.
Open Source Model Profile · Grogros
This Grogros model is a 1.24B-parameter Llama text-generation fine-tune. Its model card identifies a Llama-3.2-1B-Instruct base with only hyperparameter detail.
This Grogros model is published as a text-generation fine-tune of Meta-Llama-3.2-1B-Instruct. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 1,235,816,448 parameters. According to the model card, the training dataset is recorded as None, with llama3.2 recorded as the license.
Hub tags and the model card both identify Meta-Llama-3.2-1B-Instruct as the base and fine-tune source.
Captured config identifies LlamaForCausalLM and llama, with Safetensors metadata reporting 1,235,816,448 parameters.
According to the model card, training lists learning rate 2e-05, batch size 8, Adafactor optimizer, cosine scheduling, and 2,500 training steps.
Source: Grogros/dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-Al4-wmToken-d4-APP
Captured: Unknown. Processed: 2026-09-07T19:35:05.811843+00:00.
dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-Al4-wmToken-d4-APP This model is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the None dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 8 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 64 optimizer: Use OptimizerNames.ADAFACTOR and the args are: No additional optimizer arguments lr_scheduler_t…
F001F002F003F004F005F006F007F008F009F010F011F012