Skip to content

EthenEthenEthen

Open Source Model Profile · codefuse-ai

F2LLM-v2-330M

F2LLM-v2-330M is a 334M-parameter Qwen3-family multilingual embedding model from codefuse-ai. The model card describes 200-language support and instruction-style retrieval.

Publisher
codefuse-ai
Task
feature-extraction
Model type
qwen3
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

F2LLM-v2-330M is published by codefuse-ai as a feature-extraction embedding model. The captured configuration identifies Qwen3Model with a qwen3 model type and Safetensors metadata reports 334349184 parameters. According to the model card, it is the 330M member of a multilingual F2LLM-v2 family trained on 60 million curated data.

Recorded capabilities

Multilingual family position

According to the model card, F2LLM-v2 spans 80M to 14B, and this 330M instruct model was pruned and trained from the 0.6B base.

Instruction-style retrieval

The model card documents Sentence Transformers and Transformers encoding with Instruct plus Query prompts for queries and unprompted passages.

MRL truncation

According to the model card, Matryoshka training supports truncating embeddings to fewer dimensions to reduce storage and speed vector search.

Apache-2.0 licensing

Captured metadata records an apache-2.0 license for this repository.

Use cases in the source record

  • Instruction-guided retrieval and semantic search using query prompts with unprompted passages.
  • Multilingual embedding for symmetric STS and clustering tasks, with optional prompts and Matryoshka truncation.

Limitations and unknowns

  • No independent Ethen evaluation results were extracted; MTEB leadership language comes from the publisher model card.
  • No context-window, VRAM, latency, or pricing values were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Family training and openness claims come from the publisher model card and have not been independently verified by Ethen.

Source and provenance

Source: codefuse-ai/F2LLM-v2-330M

Captured: Unknown. Processed: 2026-09-07T19:35:55.004940+00:00.

F2LLM-v2-330M F2LLM-v2 is a family of general-purpose, multilingual embedding models in 8 distinct sizes ranging from 80M to 14B. Trained on a curated composite of 60 million publicly available high-quality data, F2LLM-v2 supports more than 200 languages, with a particular emphasis on previously underserved mid- and low-resource languages. F2LLM-v2 is fully open. We release base models in 5 sizes, instruct models in 8 sizes, the training data, the training code, and intermediate checkpoints. The three smallest instruct models are pruned and trained from the 0.6B base model. Model Base Instruct 80M 🤗F2LLM-v2-80M 160M 🤗F2LLM-v2-160M…

F001F002F003F004F005F006F007F010F011F012F015F016F018F019