Skip to content

EthenEthenEthen

Open Source Model Profile · aayanmishra-ml

Athena-1-3B

Athena-1-3B is a 3.09B-parameter Qwen2 text-generation fine-tune from aayanmishra-ml. According to the model card, it derives from Qwen2.5-3B-Instruct for instruction following.

Publisher
aayanmishra-ml
Task
text-generation
Model type
qwen2
License
qwen-research
Library
transformers
Publication status
Approved for indexing

Model overview

Athena-1-3B is published by aayanmishra-ml as a text-generation model. The captured configuration identifies Qwen2ForCausalLM, and Safetensors metadata reports 3,085,938,688 parameters. According to the model card, it is an instruction-following fine-tune of Qwen/Qwen2.5-3B-Instruct.

Recorded capabilities

Instruction-tuned Qwen2.5 derivative

According to the model card, Athena-1 3B fine-tunes Qwen/Qwen2.5-3B-Instruct for precise adherence to user prompts.

Compact conversational design

According to the model card, the 3.09B-parameter model targets lightweight applications, conversational AI, and structured data tasks.

Publisher-described code and math coverage

According to the model card, the publisher describes the model as proficient in coding challenges and mathematical tasks.

Documented output and architecture notes

According to the model card, the model supports up to 8K tokens of output on a Transformers stack with RoPE, SwiGLU, RMSNorm, QKV bias, and tied embeddings.

Use cases in the source record

  • Lightweight conversational text generation following the card's instruction-following fine-tune description.
  • Structured-data, coding, and mathematical text tasks within the card's stated 8K-token output scope, as publisher-described uses.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record; only the card's 8K-token output statement is available.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Capability claims for instruction following, coding, mathematics, and efficiency come from the publisher model card and were not independently verified.

Source and provenance

Source: aayanmishra-ml/Athena-1-3B

Captured: Unknown. Processed: 2026-09-07T19:34:39.245189+00:00.

Athena-1 3B: Athena-1 3B is a fine-tuned, instruction-following large language model derived from Qwen/Qwen2.5-3B-Instruct . It is designed to provide efficient, high-quality text generation while maintaining a compact size. Athena 3B is optimized for lightweight applications, conversational AI, and structured data tasks, making it ideal for real-world use cases where performance and resource efficiency are critical. Key Features ⚡ Lightweight and Efficient Compact Size : At just 3.09 billion parameters , Athena-1 3B offers excellent performance with reduced computational requirements. Instruction Following : Fine-tuned for precise…

F001F002F003F004F005F006F007F009F010F011F012F013F014F015F016