Skip to content

EthenEthenEthen

Open Source Model Profile · rubenroy

Geneva-12B-GCv2-5m

Geneva-12B-GCv2-5m is a 12.25B-parameter Mistral NeMo text-generation fine-tune from rubenroy. According to the model card, it fine-tunes Mistral Nemo Instruct 2407 on GammaCorpus v2 5m with Unsloth.

Publisher
rubenroy
Task
text-generation
Model type
mistral
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Geneva-12B-GCv2-5m is published by rubenroy as a mistral-based text-generation fine-tune. The captured configuration identifies MistralForCausalLM and Safetensors metadata reports 12247782400 parameters. According to the model card, it fine-tunes mistralai/Mistral-Nemo-Instruct-2407 on the 5m split of GammaCorpus v2.

Recorded capabilities

Documented Mistral NeMo Instruct lineage

According to the model card and hub tags, the source is mistralai/Mistral-Nemo-Instruct-2407, with the card listing 12B parameters, 40 layers, and SwiGLU activation.

GammaCorpus v2 5m dataset

According to the model card, training used GammaCorpus v2 5m, described as structured and filtered multi-turn conversations; hub tags record dataset rubenroy/GammaCorpus-v2-5m.

Short Unsloth fine-tune

According to the model card, fine-tuning used the Unsloth framework on one A100 GPU for about 70 minutes over 60 epochs.

Mistral 12.25B Transformers record

Captured config identifies MistralForCausalLM and mistral, with Safetensors metadata reporting 12247782400 parameters and transformers library support.

Apache-2.0 licensing

Card data records apache-2.0, and the model card states the model is released under the Apache 2.0 License.

Use cases in the source record

  • Conversational text-generation workflows using the captured transformers configuration with text-generation-inference compatibility.
  • Multi-turn conversation experiments building on the documented GammaCorpus structured and filtered conversation data.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Performance characterizations such as designed to outperform similarly sized models are publisher claims and were not independently verified by Ethen.

Source and provenance

Source: rubenroy/Geneva-12B-GCv2-5m

Captured: Unknown. Processed: 2026-09-07T19:35:29.791984+00:00.

Geneva 12B GammaCorpus v2-5m A Mistral NeMo model fine-tuned on the GammaCorpus dataset Overview Geneva 12B GammaCorpus v2-5m is a fine-tune of Mistral's Mistral Nemo Instruct 2407 model. Geneva is designed to outperform other models that have a similar size while also showcasing GammaCorpus v2-5m . Model Details Base Model: mistralai/Mistral-Nemo-Instruct-2407 Parameters: 12B Layers: 40 Dim: 5,120 Head dim: 128 Hidden dim: 14,336 Activation Function: SwiGLU Number of heads: 32 Number of kv-heads: 8 (GQA) Vocabulary size: 2**17 ~= 128k Rotary embeddings (theta = 1M) Training Details Geneva-12B-GCv2-5m underwent fine-tuning with 1 A1…

F001F002F003F004F005F006F007F008F009F010F011F012F013F017F018F020F022