Documented Mistral NeMo Instruct lineage
According to the model card and hub tags, the source is mistralai/Mistral-Nemo-Instruct-2407, with the card listing 12B parameters, 40 layers, and SwiGLU activation.
Open Source Model Profile · rubenroy
Geneva-12B-GCv2-5m is a 12.25B-parameter Mistral NeMo text-generation fine-tune from rubenroy. According to the model card, it fine-tunes Mistral Nemo Instruct 2407 on GammaCorpus v2 5m with Unsloth.
Geneva-12B-GCv2-5m is published by rubenroy as a mistral-based text-generation fine-tune. The captured configuration identifies MistralForCausalLM and Safetensors metadata reports 12247782400 parameters. According to the model card, it fine-tunes mistralai/Mistral-Nemo-Instruct-2407 on the 5m split of GammaCorpus v2.
According to the model card and hub tags, the source is mistralai/Mistral-Nemo-Instruct-2407, with the card listing 12B parameters, 40 layers, and SwiGLU activation.
According to the model card, training used GammaCorpus v2 5m, described as structured and filtered multi-turn conversations; hub tags record dataset rubenroy/GammaCorpus-v2-5m.
According to the model card, fine-tuning used the Unsloth framework on one A100 GPU for about 70 minutes over 60 epochs.
Captured config identifies MistralForCausalLM and mistral, with Safetensors metadata reporting 12247782400 parameters and transformers library support.
Card data records apache-2.0, and the model card states the model is released under the Apache 2.0 License.
Source: rubenroy/Geneva-12B-GCv2-5m
Captured: Unknown. Processed: 2026-09-07T19:35:29.791984+00:00.
Geneva 12B GammaCorpus v2-5m A Mistral NeMo model fine-tuned on the GammaCorpus dataset Overview Geneva 12B GammaCorpus v2-5m is a fine-tune of Mistral's Mistral Nemo Instruct 2407 model. Geneva is designed to outperform other models that have a similar size while also showcasing GammaCorpus v2-5m . Model Details Base Model: mistralai/Mistral-Nemo-Instruct-2407 Parameters: 12B Layers: 40 Dim: 5,120 Head dim: 128 Hidden dim: 14,336 Activation Function: SwiGLU Number of heads: 32 Number of kv-heads: 8 (GQA) Vocabulary size: 2**17 ~= 128k Rotary embeddings (theta = 1M) Training Details Geneva-12B-GCv2-5m underwent fine-tuning with 1 A1…
F001F002F003F004F005F006F007F008F009F010F011F012F013F017F018F020F022