Skip to content

EthenEthenEthen

Open Source Model Profile · secmlr

VD-DS-Clean-8k_VD-DS-Clean-16k_Qwen2.5-7B-Instruct_full_sft_1e-5

This secmlr release is a 7.62B-parameter Qwen2 fine-tune of Qwen2.5-7B-Instruct on VD-DS-Clean-8k and VD-DS-Clean-16k.

Publisher
secmlr
Task
text-generation
Model type
qwen2
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

This secmlr release is a Qwen2-family text-generation fine-tune. The captured configuration identifies Qwen2ForCausalLM and Safetensors metadata reports 7,615,616,512 parameters. According to the model card, it is a full fine-tune of Qwen/Qwen2.5-7B-Instruct on the VD-DS-Clean-8k and VD-DS-Clean-16k datasets.

Recorded capabilities

7.62B Qwen2 record

Captured configuration records Qwen2ForCausalLM and Safetensors metadata reports 7,615,616,512 parameters.

Named Qwen2.5-7B-Instruct lineage

According to the model card and hub tags, the release is a full fine-tune of Qwen/Qwen2.5-7B-Instruct on VD-DS-Clean-8k and VD-DS-Clean-16k.

Documented 3-epoch schedule

According to the model card, training used learning rate 1e-05, 3.0 epochs, cosine scheduling, warmup ratio 0.1, and 48-sample total train batch size.

Apache-2.0 record

Card data records apache-2.0.

Use cases in the source record

  • Conversational text-generation experimentation on the named Qwen2.5-7B-Instruct lineage using the Transformers stack.
  • Fine-tune comparison work that references the documented VD-DS-Clean datasets and 3-epoch schedule as publisher claims.

Limitations and unknowns

  • According to the model card, intended uses, limitations, and training and evaluation data sections are recorded as More information needed.
  • No evaluation results, context-window value, or hardware requirements were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: secmlr/VD-DS-Clean-8k_VD-DS-Clean-16k_Qwen2.5-7B-Instruct_full_sft_1e-5

Captured: Unknown. Processed: 2026-09-07T19:35:30.058700+00:00.

VD-DS-Clean-8k_VD-DS-Clean-16k_Qwen2.5-7B-Instruct_full_sft_1e-5 This model is a fine-tuned version of Qwen/Qwen2.5-7B-Instruct on the VD-DS-Clean-8k and the VD-DS-Clean-16k datasets. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 1e-05 train_batch_size: 1 eval_batch_size: 8 seed: 42 distributed_type: multi-GPU num_devices: 4 gradient_accumulation_steps: 12 total_train_batch_size: 48 total_eval_batch_size: 32 optimizer:…

F001F002F003F004F005F006F007F008F009F010F011F012