Instruction-tuned Qwen2.5 derivative
According to the model card, Athena-1 3B fine-tunes Qwen/Qwen2.5-3B-Instruct for precise adherence to user prompts.
Open Source Model Profile · aayanmishra-ml
Athena-1-3B is a 3.09B-parameter Qwen2 text-generation fine-tune from aayanmishra-ml. According to the model card, it derives from Qwen2.5-3B-Instruct for instruction following.
Athena-1-3B is published by aayanmishra-ml as a text-generation model. The captured configuration identifies Qwen2ForCausalLM, and Safetensors metadata reports 3,085,938,688 parameters. According to the model card, it is an instruction-following fine-tune of Qwen/Qwen2.5-3B-Instruct.
According to the model card, Athena-1 3B fine-tunes Qwen/Qwen2.5-3B-Instruct for precise adherence to user prompts.
According to the model card, the 3.09B-parameter model targets lightweight applications, conversational AI, and structured data tasks.
According to the model card, the publisher describes the model as proficient in coding challenges and mathematical tasks.
According to the model card, the model supports up to 8K tokens of output on a Transformers stack with RoPE, SwiGLU, RMSNorm, QKV bias, and tied embeddings.
Source: aayanmishra-ml/Athena-1-3B
Captured: Unknown. Processed: 2026-09-07T19:34:39.245189+00:00.
Athena-1 3B: Athena-1 3B is a fine-tuned, instruction-following large language model derived from Qwen/Qwen2.5-3B-Instruct . It is designed to provide efficient, high-quality text generation while maintaining a compact size. Athena 3B is optimized for lightweight applications, conversational AI, and structured data tasks, making it ideal for real-world use cases where performance and resource efficiency are critical. Key Features ⚡ Lightweight and Efficient Compact Size : At just 3.09 billion parameters , Athena-1 3B offers excellent performance with reduced computational requirements. Instruction Following : Fine-tuned for precise…
F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016