Skip to content

EthenEthenEthen

Open Source Model Profile · penfever

glm46-ling-coder-sft-sandboxes-1-maxeps-131k

glm46-ling-coder-sft-sandboxes-1-maxeps-131k is a penfever text-generation release with a captured Qwen3 configuration and unknown training data.

Publisher
penfever
Task
text-generation
Model type
qwen3
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

glm46-ling-coder-sft-sandboxes-1-maxeps-131k is published by penfever as a text-generation model. The captured configuration identifies Qwen3ForCausalLM and Safetensors metadata reports 308,224 parameters. The model card says it was trained from scratch on an unknown dataset.

Recorded capabilities

Captured Qwen3 configuration

The captured configuration identifies Qwen3ForCausalLM with 308,224 parameters in Safetensors metadata.

From-scratch training claim

The model card says the model was trained from scratch on an unknown dataset.

Documented hyperparameters

According to the model card, training used 7 epochs, learning rate 4e-05, 8 GPUs, cosine scheduling, and Transformers 4.57.3 with PyTorch 2.9.0.

Use cases in the source record

  • Hyperparameter and trainer-output inspection for Llama-Factory text-generation runs, using the documented 7-epoch setup.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No license value was extracted from this record.
  • No context-window value was extracted from this record.
  • The training dataset is stated as unknown in the model card.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: penfever/glm46-ling-coder-sft-sandboxes-1-maxeps-131k

Captured: Unknown. Processed: 2026-09-07T19:35:59.844433+00:00.

glm46-ling-coder-sft-sandboxes-1-maxeps-131k This model was trained from scratch on an unknown dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 4e-05 train_batch_size: 1 eval_batch_size: 8 seed: 42 distributed_type: multi-GPU num_devices: 8 gradient_accumulation_steps: 2 total_train_batch_size: 16 total_eval_batch_size: 64 optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.98) and epsilon=1e-08 and…

F001F002F003F004F005F006F009F010F011