Skip to content

EthenEthenEthen

Open Source Model Profile · shanchen

ds-limo-ja-500

ds-limo-ja-500 is a 7.62B-parameter Qwen2-family text-generation fine-tune from shanchen. According to the model card, it fine-tunes deepseek-ai/DeepSeek-R1-Distill-Qwen-7B.

Publisher
shanchen
Task
text-generation
Model type
qwen2
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

ds-limo-ja-500 is published by shanchen as a text-generation model. The captured configuration identifies Qwen2ForCausalLM with model type qwen2, and Safetensors metadata reports 7,615,616,512 parameters. According to the model card, it is a fine-tuned version of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B, and hub tags record TRL, SFT, and conversational markers.

Recorded capabilities

DeepSeek-R1-Distill-Qwen-7B base

Hub tags and the model card identify deepseek-ai/DeepSeek-R1-Distill-Qwen-7B as the fine-tune base.

7.62B parameter record

Captured Safetensors metadata reports 7,615,616,512 parameters, or about 7.62B.

TRL with SFT

According to the model card, the model was trained using TRL, and the training procedure is described as SFT.

Transformers pipeline example

According to the model card, a transformers text-generation pipeline quick-start example is provided.

Use cases in the source record

  • Conversational text-generation workflows matching the record's text-generation pipeline tag and conversational markers.
  • Transformers pipeline experiments following the model card's text-generation quick-start example.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No license value was extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: shanchen/ds-limo-ja-500

Captured: Unknown. Processed: 2026-09-07T19:36:01.680055+00:00.

Model Card for ds-limo-ja-500 This model is a fine-tuned version of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B . It has been trained using TRL . Quick start from transformers import pipeline question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?" generator = pipeline( "text-generation" , model= "shanchen/ds-limo-ja-500" , device= "cuda" ) output = generator([{ "role" : "user" , "content" : question}], max_new_tokens= 128 , return_full_text= False )[ 0 ] print (output[ "generated_text" ]) Training procedure This model was trained with SFT. Framework versi…

F001F002F003F004F005F006F008F009F010F011F012F013