Skip to content

EthenEthenEthen

Open Source Model Profile · Grogros

dmWM-llama-3.2-1B-Instruct-KGWB-OWT_WMBoundary-OWT2-WB-v4

dmWM-llama-3.2-1B-Instruct-KGWB-OWT_WMBoundary-OWT2-WB-v4 is a 1.24B-parameter Llama text-generation fine-tune from Grogros. According to the model card, it fine-tunes meta-llama/Llama-3.2-1B-Instruct on openwebtext.

Publisher
Grogros
Task
text-generation
Model type
llama
License
llama3.2
Library
transformers
Publication status
Accepted · not indexed

Model overview

dmWM-llama-3.2-1B-Instruct-KGWB-OWT_WMBoundary-OWT2-WB-v4 is published by Grogros as a Llama text-generation model. The captured configuration identifies LlamaForCausalLM and Safetensors metadata reports 1,235,818,496 parameters. According to the model card, it is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the openwebtext dataset.

Recorded capabilities

1.24B Llama with Transformers

The captured configuration reports LlamaForCausalLM with model type llama and 1,235,818,496 parameters, with Transformers library support.

Llama-3.2-1B-Instruct base

According to the model card, the model fine-tunes meta-llama/Llama-3.2-1B-Instruct; hub tags list the same base.

Openwebtext dataset

According to the model card and hub tags, training uses the openwebtext dataset.

Published hyperparameters

According to the model card, training used learning rate 2e-05, batch sizes 8 and 8 with accumulation 8, seed 42, cosine schedule with 0.1 warmup, and 5,000 steps.

Use cases in the source record

  • Conversational text-generation experiments on the documented Llama-3.2-1B-Instruct lineage.
  • Reproduction-style study of the published Adafactor, cosine-schedule, 5,000-step training setup.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted.
  • According to the model card, intended uses, limitations, and training and evaluation data need more information.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: Grogros/dmWM-llama-3.2-1B-Instruct-KGWB-OWT_WMBoundary-OWT2-WB-v4

Captured: Unknown. Processed: 2026-09-07T19:35:05.951193+00:00.

dmWM-llama-3.2-1B-Instruct-KGWB-OWT_WMBoundary-OWT2-WB-v4 This model is a fine-tuned version of meta-llama/Llama-3.2-1B-Instruct on the openwebtext dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 8 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 64 optimizer: Use OptimizerNames.ADAFACTOR and the args are: No additional optimizer arguments lr_scheduler_typ…

F001F002F003F004F005F006F007F008F009F010F011F012