Llama 3.2 1B openwebtext fine-tune
Captured configuration records LlamaForCausalLM with about 1.24B parameters, and the model card describes a fine-tune of meta-llama/Llama-3.2-1B-Instruct on openwebtext.
Open Source Model Profile · Grogros
dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1 is a 1.24B-parameter Llama fine-tune from Grogros. The model card lists meta-llama/Llama-3.2-1B-Instruct and openwebtext.
dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1 is published by Grogros as a Llama text-generation model. The captured configuration identifies LlamaForCausalLM with model type llama, and Safetensors metadata reports 1,235,816,448 parameters. According to the model card, it fine-tunes meta-llama/Llama-3.2-1B-Instruct on openwebtext.
Captured configuration records LlamaForCausalLM with about 1.24B parameters, and the model card describes a fine-tune of meta-llama/Llama-3.2-1B-Instruct on openwebtext.
According to the model card, training used Adafactor, learning rate 2e-05, cosine scheduling with 0.1 warmup ratio, and 5,000 training steps.
Card data records llama3.2 for this model.
Source: Grogros/dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1
Captured: Unknown. Processed: 2026-09-07T19:35:05.980935+00:00.
dmWM-llama-3.2-1B-Instruct-OWTWM-DistillationWM-wmToken-d4-a0.1 This model is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the openwebtext dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 8 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 64 optimizer: Use OptimizerNames.ADAFACTOR and the args are: No additional optimizer arguments lr_schedul…
F001F002F003F004F005F006F007F009F010F011F012