Skip to content

EthenEthenEthen

Open Source Model Profile · M4-ai

Hercules-5.0-Qwen2-1.5B

Hercules-5.0-Qwen2-1.5B is a 1.54B-parameter Qwen2-family text-generation fine-tune from M4-ai. According to the model card, it fine-tunes qwen2-1.5B and uses ChatML.

Publisher
M4-ai
Task
text-generation
Model type
qwen2
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

Hercules-5.0-Qwen2-1.5B is published by M4-ai as a text-generation model. The captured configuration identifies Qwen2ForCausalLM with model type qwen2, and Safetensors metadata reports about 1.54B parameters. According to the model card, it is a qwen2-1.5B fine-tune on a high-quality mix for general-purpose assistants using ChatML.

Recorded capabilities

Qwen2 text-generation architecture

The captured configuration identifies Qwen2ForCausalLM with model type qwen2 and Transformers support.

1.54B Qwen2-1.5B fine-tune

Safetensors metadata reports 1,543,714,304 parameters; according to the model card, the model is fine-tuned from qwen2-1.5B.

Documented ChatML format

According to the model card, the publisher uses the ChatML prompt format.

General-assistant positioning

According to the model card, the publisher positions it for general-purpose assistance, question answering, and chain-of-thought use, with English and possibly Chinese.

Disclosed TPU training setup

According to the model card, training used eight Kaggle TPUs with bf16 non-mixed precision, batch size 256, and sequence length 1536.

Use cases in the source record

  • General-purpose assistant workflows such as question answering and chain-of-thought prompting, as described by the publisher.
  • ChatML-based text-generation experiments consistent with the publisher-documented prompt format and math, coding, and writing focus.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • According to the model card, the publisher includes an anecdotal coding-achievement note; it is a publisher claim, not an Ethen-verified benchmark.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: M4-ai/Hercules-5.0-Qwen2-1.5B

Captured: Unknown. Processed: 2026-09-07T19:34:33.656402+00:00.

Hercules-5.0-Qwen2-1.5B We fine-tuned qwen2-1.5B on a high quality mix for general-purpose assistants. A DPO version of this will be released soon. We use the ChatML prompt format. Model Details Model Description This model has capabilities in math, coding, writing, and more. We fine-tuned it using a high quality mix for general-purpose assistants. Developed by: M4-ai Language(s) (NLP): English and maybe Chinese License: apache-2.0 Finetuned from model: qwen2-1.5B Uses General purpose assistant, question answering, chain-of-thought, etc.. This language model made an impressive achievement, and correctly implemented a Multi Head Atte…

F001F002F003F004F005F006F007F008F010F011F012F013F014F015