Skip to content

EthenEthenEthen

Open Source Model Profile · codefuse-ai

F2LLM-v2-160M

F2LLM-v2-160M is a 159M-parameter multilingual embedding model from codefuse-ai. According to the model card, it belongs to the eight-size F2LLM-v2 family supporting more than 200 languages.

Publisher
codefuse-ai
Task
feature-extraction
Model type
qwen3
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

F2LLM-v2-160M is published by codefuse-ai as a feature-extraction model. The captured configuration identifies Qwen3Model with model type qwen3, and Safetensors metadata reports about 159M parameters under apache-2.0. According to the model card, it is the 160M instruct member of a multilingual embedding family trained on 60 million curated records.

Recorded capabilities

160M multilingual embedding variant

According to the model card, F2LLM-v2 spans 80M to 14B, and the three smallest instruct models were pruned and trained from the 0.6B base model.

Retrieval-ready encoding

According to the model card, Sentence Transformers and Transformers examples encode queries separately from documents and compare them with cosine similarity.

Documented prompt format

According to the model card, retrieval queries use the prompt for queries but not for passages, while STS, clustering, and bitext mining can encode either way.

Matryoshka truncation

According to the model card, Matryoshka Representation Learning allows keeping only the first dimensions, with a smallest trained dimension of 8.

Use cases in the source record

  • Multilingual retrieval, semantic search, and classification using the documented query-document encoding and cosine-similarity path.
  • Storage- and speed-sensitive vector search using Matryoshka truncation and instruction-conditioned retrieval prompts.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • No structured evaluation scores were extracted; family MTEB leadership statements are publisher claims referring to the external leaderboard.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • State-of-the-art and language-coverage statements are publisher claims and were not independently verified.

Source and provenance

Source: codefuse-ai/F2LLM-v2-160M

Captured: Unknown. Processed: 2026-09-07T19:35:54.984804+00:00.

F2LLM-v2-160M F2LLM-v2 is a family of general-purpose, multilingual embedding models in 8 distinct sizes ranging from 80M to 14B. Trained on a curated composite of 60 million publicly available high-quality data, F2LLM-v2 supports more than 200 languages, with a particular emphasis on previously underserved mid- and low-resource languages. F2LLM-v2 is fully open. We release base models in 5 sizes, instruct models in 8 sizes, the training data, the training code, and intermediate checkpoints. The three smallest instruct models are pruned and trained from the 0.6B base model. Model Base Instruct 80M 🤗F2LLM-v2-80M 160M 🤗F2LLM-v2-160M…

F001F002F003F004F005F006F007F009F010F011F012F013F015F016F018F019