Skip to content

EthenEthenEthen

Open Source Model Profile · arcee-ai

Virtuoso-Small-v2

Virtuoso-Small-v2 is a 14.77B-parameter Qwen2 text-generation fine-tune from arcee-ai. Its model card documents DeepSeek-V3 distillation on a Qwen-2.5-14B base.

Publisher
arcee-ai
Task
text-generation
Model type
qwen2
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Virtuoso-Small-v2 is published by arcee-ai as a Qwen2-family text-generation fine-tune. The captured configuration identifies Qwen2ForCausalLM and Safetensors metadata reports 14,765,947,904 parameters, or about 14.77B. The model card describes it as a 14B model built on Qwen-2.5-14B and distilled from DeepSeek-V3.

Recorded capabilities

DeepSeek-V3 logit distillation

The model card describes distillation from DeepSeek-V3 with full logit-level replication, mentioning an expanded set of 5B+ tokens worth of logits and about 1.1B tokens/logits from DeepSeek-V3 training data.

Qwen-2.5-14B base with tokenizer surgery

The model card names Qwen-2.5-14B as the architecture base and says final alignment uses the Qwen tokenizer with specialized tokenizer surgery after initial DeepSeek-V3 tokenizer use for logit extraction.

128k context with June 2024 cutoff note

According to the model card, context length is 128k tokens and training data may not reflect developments beyond June 2024.

Use cases in the source record

  • Conversational text-generation workflows consistent with the record's text-generation and conversational tags.
  • Experiments in technical and scientific queries, code generation, and mathematical problem-solving, which the model card names as the intended transference areas for the distillation.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • No VRAM, quantization, pricing, or benchmark figures were extracted; local requirements and performance are unknown.
  • Distillation-data wording varies within the model card between 5B+ tokens worth of logits and about 1.1B tokens/logits; the publisher text was preserved without reconciling the figures.

Source and provenance

Source: arcee-ai/Virtuoso-Small-v2

Captured: Unknown. Processed: 2026-09-07T19:35:18.951763+00:00.

Virtuoso-Small-v2 (14B) is our next-generation, 14-billion-parameter language model that builds upon the original Virtuoso-Small architecture. This version is distilled from Deepseek-v3, leveraging an expanded dataset of 5B+ tokens worth of logits. Model Details Architecture Base: Qwen-2.5-14B Parameter Count: 14B Tokenizer: Initially integrated with Deepseek-v3 tokenizer for logit extraction. Final alignment uses the Qwen tokenizer, using specialized “tokenizer surgery” for cross-architecture compatibility. Distillation Data: ~1.1B tokens/logits from Deepseek-v3’s training data. Logit-level distillation using a proprietary “fusion…

F001F002F003F004F005F006F007F009F010F011F012F013F014F015F017F018