Skip to content

EthenEthenEthen

Open Source Model Profile · shanchen

ds-limo-te-500

ds-limo-te-500 is a 7.62B-parameter Qwen2 text-generation fine-tune from shanchen built on DeepSeek-R1-Distill-Qwen-7B.

Publisher
shanchen
Task
text-generation
Model type
qwen2
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

ds-limo-te-500 is published by shanchen as a qwen2 text-generation fine-tune. The captured configuration identifies Qwen2ForCausalLM and Safetensors metadata reports 7,615,616,512 parameters. According to the model card, it fine-tunes deepseek-ai/DeepSeek-R1-Distill-Qwen-7B with SFT using TRL.

Recorded capabilities

DeepSeek-R1 distill fine-tune

According to the model card, the model fine-tunes deepseek-ai/DeepSeek-R1-Distill-Qwen-7B.

TRL supervised fine-tuning

According to the model card, training used SFT with TRL.

7.6B Qwen2 text-generation build

The captured configuration identifies Qwen2ForCausalLM with model type qwen2 and Safetensors metadata reports about 7.62B parameters.

Use cases in the source record

  • Conversational text-generation experiments using the publisher's documented Transformers pipeline quick-start.
  • TRL supervised fine-tuning studies building on the documented DeepSeek-R1 distill lineage.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • No license value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: shanchen/ds-limo-te-500

Captured: Unknown. Processed: 2026-09-07T19:36:01.715436+00:00.

Model Card for ds-limo-te-500 This model is a fine-tuned version of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B . It has been trained using TRL . Quick start from transformers import pipeline question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?" generator = pipeline( "text-generation" , model= "shanchen/ds-limo-te-500" , device= "cuda" ) output = generator([{ "role" : "user" , "content" : question}], max_new_tokens= 128 , return_full_text= False )[ 0 ] print (output[ "generated_text" ]) Training procedure This model was trained with SFT. Framework versi…

F001F002F003F004F005F006F008F009F010F011F012F013